This dataset contains synthetic text generated using large language models for ten African languages. It is intended to support research and evaluation in automatic speech recognition (ASR), natural language processing (NLP), and related fields for low-resource languages.
Data Generation and Licensing