Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

buumba641/Low-Resource-Multilingual-ASR-for-Zambian-Languages-Nyanja-Tonga-and-Bemba

Domaine:

natural language processing

Type de record:

project
Créateur:
buu
Hôte:
# Low-Resource ASR for Zambian Languages (Bemba, Nyanja, Tonga) **Final Year Research Project (UNZA — Department of Computing and Informatics, 2026)** This repository contains **Jupyter notebooks** for a final year research project titled: > **Low-Resource Automatic Speech Recognition for Zambian Languages: A Comparative Analysis of Pre-Trained Models on Bemba, Nyanja, and Tonga** The work focuses on **monolingual Automatic Speech Recognition (ASR)** for three low-resource Zambian languages—**Bemba, Nyanja, and Tonga**—using the **Zambezi Voice** dataset (UNZA Speech Lab). The main goal is to **fine-tune and benchmark pre-trained speech models** in extremely low-resource conditions and evaluate performance using **Word Error Rate (WER)**. --- ## Project Information - **Student:** Buumba Chinjila - **Institution:** The University of Zambia (UNZA), School of Natural and Applied Sciences - **Academic Year:** 2026 - **Proposal Submission Date:** March 20, 2026 --- ## Motivation Many Zambian communities primarily communicate orally, yet most digital services are English-first. ASR for local languages can improve: - accessibility (voice interfaces, transcription, captioning) - digital record keeping (meetings, consultations, reporting) - inclusion for users with limited English literacy --- ## Problem Statement ASR development for Zambian languages faces: - **Limited labeled data** (~22–24 hours per language in Zambezi Voice) - **Limited compute**, making large-scale training difficult --- ## Aim To implement and evaluate **monolingual ASR pipelines** for Bemba, Nyanja, and Tonga by **fine-tuning and comparing open-source pre-trained models** on Zambezi Voice subsets. --- ## Objectives 1. **Model Benchmarking:** Fine-tune and compare multiple pre-trained models (e.g., **XLS-R, Whisper, MMS, HuBERT**). 2. **Data Augmentation:** Use techniques such as **speed perturbation** and **SpecAugment** where appropriate. 3. **Performance Evaluation:** Evaluate using **W …

Visit

github.com

Languages

BembaChichewaTonga

Similaires

buumba641/nyanja-asr-whisper-tinybuumba641/nyanja-asr-whisper-basebuumba641/bemba-asr-mms-300mbuumba641/bemba-asr-whisper-tinybuumba641/bemba-asr-whisper-smallbuumba641/bemba-asr-xlsr-300m

buumba641/nyanja-asr-whisper-tiny

buumba641/nyanja-asr-whisper-base

buumba641/bemba-asr-mms-300m

buumba641/bemba-asr-whisper-tiny

buumba641/bemba-asr-whisper-small

buumba641/bemba-asr-xlsr-300m