This dataset contains 13,488 synthetic sentences across 10 African languages (Bambara, Chichewa, Hausa, Kanuri, Luo, Nande, Somali, Twi, Wolof, Yoruba) generated using large language models (GPT-4o, GPT-4.5, Claude 3.5 Sonnet, Claude 3.7 Sonnet). Each sentence has been evaluated by human linguists on readability and naturalness (1-7 scale), translation adequacy and accuracy (1-7 scale), grammatical correctness, word validity, and presence of notable errors. Corrected versions are provided where applicable. The dataset was created by Dimagi to support ASR, NLP research, and evaluation for low-resource African languages. See: DeRenzi et al. (2025), "Synthetic Voice Data for Automatic Speech Recognition in African Languages", arXiv:2507.17578.