This dataset, curated by Digital Divide Data (DDD), provides high-quality audio recordings and corresponding text transcriptions for the Somali (som) language. The collection includes thousands of unique utterances per language to support diverse acoustic modeling. All transcriptions have undergone a manual verification process to ensure high linguistic accuracy. Recordings feature a balanced mix of genders and various age groups to minimize bias in downstream AI models. This data is specifically designed for training Automatic Speech Recognition (ASR) systems, Text-to-Speech (TTS) synthesis, and general linguistic research for underrepresented African languages.