Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Adja Speech Dataset for ASR and TTS

Domain:

natural language processing

Record type:

dataset
Creator:
Jos
Host:
This is the canonical public Adja speech dataset for the May 2026 thesis release. It is intended for automatic speech recognition, text-to-speech, and speech pipeline experiments. This component duplicates the Orpheus speech source from JosueG/adja-tts-orpheus. The source repository is treated as read-only provenance; this dataset repo is the public canonical release surface.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processingtext to speech

Languages

AdioukrouAjaAja

Licenses

other

Similar

Common Voice Scripted Speech 26.0 - AdjaRamsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTSYoruba TTS (text-to-speech) training datasetAfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASRText Selection scripts for ASR/TTSEhugbo TTS: biblical text to speech dataset in Ehugbo Language

Common Voice Scripted Speech 26.0 - Adja

A collection of read speech recordings in Adja (ajg).

Ramsa: A Large Sociolinguistically Rich Emirati Arabic Speech Corpus for ASR and TTS

Ramsa is a developing 41-hour speech corpus of Emirati Arabic designed to support sociolinguistic re

Yoruba TTS (text-to-speech) training dataset

Textbook audio archive size: total 36M archive created: 8 July 2011 mp3 file size ======== ==== 01-

AfriSpeech-200: Pan-African Accented Speech Dataset for Clinical and General Domain ASR

Africa has a very poor doctor-to-patient ratio. At very busy clinics, doctors could see 30+ patients

Text Selection scripts for ASR/TTS

Scripts for text selection of phonetically balanced sentences for ASR/TTS corpora. Based on phonetis

Ehugbo TTS: biblical text to speech dataset in Ehugbo Language

This dataset contains audio recordings of Bible verses in Ehugbo, a dialect of Igbo (a Niger-Congo language spoken in Nigeria). It contains 312 audio recordings of biblical text-to-speech data comprising 1 hour and 30 seconds of speech data.

This dataset