Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

galsenai/WaxalNLP

Domain:

natural language processing

Record type:

dataset
Creator:
gal
Host:
Google introduced WAXAL, a new open dataset for 21 African languages, to tackle data scarcity and build inclusive speech technology. However, the Wolof language has experienced alignment issues between the audio files and their transcriptions, making the dataset unusable. We therefore propose to correct this using a simple and effective approach: For each audio clip, we generated a transcription using Google Gemini ASR.

Visit

huggingface.co

Tasks

automatic speech recognitionspeech processing

Languages

Wolof

Similar

Tonic/WaxalNLPgalsenai/centralized_wolof_french_translation_datagalsenai/waxal_datasetgalsenai/llama2-7B_wolofgalsenai/wolof-ttshausa-tts-24khz-waxalnlp-5

Tonic/WaxalNLP

Google introduced WAXAL, a new open dataset for 21 African languages, to tackle data scarcity and bu

galsenai/centralized_wolof_french_translation_data

galsenai/waxal_dataset

Keyword spotting refers to the task of learning to detect spoken keywords. It interfaces all modern

galsenai/llama2-7B_wolof

galsenai/wolof-tts

hausa-tts-24khz-waxalnlp-5

Single-speaker Hausa TTS dataset from WaxalNLP (google/WaxalNLP hau_tts), speaker 5. Audio is stored