Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Way With Words | Speech Collection Datasets

Domain:

natural language processing

Record type:

company

We create speech datasets including transcripts for machine learning purposes. Our service is used for technologies looking to create or improve existing automatic speech recognition models (ASR) using natural language processing (NLP) for select languages and various domains.

Each dataset can be created according to dialect, demographics, domain or any other required conditions.

Speech datasets for select languages and industries are available, or bespoke speech collection projects available on request.

Visit

waywithwords.net

Tasks

text to speechautomatic speech recognitionspeech processingspeech translation

Languages

AfrikaansSotho, SouthernZulu

Tags

waywithwords

Licenses

License belongs to WayWithWords Limited ©

Similar

Roboflow Transportation Datasets CollectionAfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African LanguagesEnhancing text pre-processing for Swahili language: Datasets for common Swahili stop-words, slangs and typos with equivalent proper wordsSpeech-to-text-data-collection/STT-data-collectionAfri Code Datasets (A collection of datasets for code generation in African languages)Collection of study datasets for the HSU group

Roboflow Transportation Datasets Collection

Autonomous vehicles, traffic monitoring, smart cities, infrastructure inspection, AI-based logistics

AfriHate: A Multilingual Collection of Hate Speech and Abusive Language Datasets for African Languages

Hate speech and abusive language are global phenomena that need socio-cultural background knowledge

Enhancing text pre-processing for Swahili language: Datasets for common Swahili stop-words, slangs and typos with equivalent proper words

Natural Language Processing requires data to be pre-processed to guarantee quality models in different machine learning tasks. However, Swahili language have been disadvantaged and is classified as low resource language because of inadequate data for NLP especially

Speech-to-text-data-collection/STT-data-collection

A data engineering pipeline that allows recording millions of Amharic and Swahili speakers reading d

Afri Code Datasets (A collection of datasets for code generation in African languages)

Training and evaluating Large Language Models (LLMs) for code generation, building AI-powered coding

Collection of study datasets for the HSU group

This dataset contains data for various studies undertaken by the Health Services Unit (HSU) of the K