Logo Lanfrica

Izere01/tekana-kinyarwanda-tts

Domain:

natural language processing

Record type:

modelsoftware
Creator:
Ize
Host:
Kinyarwanda Text-to-Speech system for Tekana health support line using MMS-TTS with CPU optimization and latency evaluation. # Tekana Project: Kinyarwanda Text-to-Speech This repository contains a Kinyarwanda TTS system built for the Tekana health support line assessment. ## Model - Base Model: facebook/mms-tts-kin - Fine-tuned on 14-hour Kinyarwanda dataset - Model size: ~118 MB - CPU-optimized inference ## Performance - Average CPU latency (10-word sentence): 1094 ms - Best-case latency: 978 ms - Real-Time Factor (RTF): ~0.5 - Model size under 200MB constraint ## Run Inference ```bash venv\Scripts\activate python -m src.inference.infer ``` Audio outputs are saved to: ```bash outputs/ ``` ## Production Recommendation GPU-backed deployment or ONNX optimization would reduce latency below the 800 ms constraint. --- # Evaluation Summary ## Required Sentences 1. Muraho, nagufasha gute uyu munsi? 2. Niba ufite ibibazo bijyanye n’ubuzima bwawe, twagufasha. 3. Ni ngombwa ko ubonana umuganga vuba. 4. Twabanye nawe kandi tuzakomeza kukwitaho. 5. Ushobora kuduhamagara igihe cyose ukeneye ubufasha. Audio files are located in `outputs/`. --- ## Latency Results (CPU Only) | Run | Latency (ms) | |-----|-------------| | 1 | 1059 | | 2 | 978 | | 3 | 1245 | | **Average** | **1094 ms** | Real-Time Factor (RTF): ~0.5 --- ## Subjective Evaluation - Speech intelligibility: High - Accent naturalness: Acceptable for Rwandan listeners - No robotic artifacts observed - Clear pronunciation of Kinyarwanda phonemes --- ## Constraint Analysis Target latency: <800 ms Observed latency: ~1094 ms (CPU) RTF < 1 indicates generation faster than real-time. GPU-backed deployment or ONNX optimization is expected to meet the 800 ms requirement.