This paper analyzes conversational self-repair, which refers to reconstructing problematic portions
A large-scale, unified collection of speech data for text-to-speech (TTS) and speech research. This
Amharic Pretraining Corpus is a large-scale dataset (~103M) for general amharic language pretraining
Cover title
This is IsiZulu Pretraining Dataset. The dataset was used to pre-train BafoGPT-3B Books: Zulu-Englis