# Project Title
# Yoruba Learner Speech Corpus (YLSC)
_The first phonemically and tonally annotated speech corpus of Yoruba learner pronunciation designed for linguistic research and CAPT development._
The Yoruba learner Speech Corpus is a pioneering, phonemically and tonally annotated dataset designed to support research in Second Language Acquisition (SLA), Automatic Speech Recognition (ASR), and Computer-Assisted Pronunciation Training (CAPT). By capturing authentic learner speech with expert phonetic diagnostics, this project addresses the significant gap in resources for tonal African languages.
**Project lead: Kaosarat Aina**, **Indiana University Bloomington**
**Contact: kbaina@iu.edu**
# Overview
The Yoruba Learner Speech Corpus (YLSC) is a developing dataset of speech produced by second-language and heritage learners of Yoruba.
Learning a second language (L2) is a complex process where speaking proficiency often serves as the foundation for learner confidence. However, learners frequently lack immediate corrective feedback during self-study. While commercial CAPT tools exist, their proprietary nature limits research and modification. This corpus provides an open-source foundational resource for developing robust, diagnostic feedback systems.
# Research Motivation
* **Tonal Complexity**: Yoruba's three-way tonal contrast (High, Mid, and Low) presents a steep challenge for L2 learners, particularly those from non-tonal L1 backgrounds.
* **Segmental Challenges**: Beyond tones, learners often struggle with unique Yoruba phonemes, such as co-articulated labial-velar stops ([ɡb], [kp]) and the Advanced Tongue Root (±ATR) vowel harmony system.
* **Native vs. Learner Data**: Existing Yoruba corpora are designed for native speech, lacking the specific phonological errors—such as vowel harmony violations or tonal misassignments—inherent to learners.
* **Diagnostic Gaps**: Current research requires error-annotated corpora to move beyond simple "pass/fail" error …