Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Dealing with the Hard Facts of Low-Resource African NLP

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
DiaCouKamTal
Hôte:avatar
Creating speech datasets, models, and evaluation frameworks for low-resource languages remains challenging given the lack of a broad base of pertinent experience to draw from. This paper reports on the field collection of 612 hours of spontaneous speech in Bambara, a low-resource West African language; the semi-automated annotation of that dataset with transcriptions; the creation of several monolingual ultra-compact and small models using the dataset; and the automatic and human evaluation of their output. We offer practical suggestions for data collection protocols, annotation, and model design, as well as evidence for the importance of performing human evaluation. In addition to the main dataset, multiple evaluation datasets, models, and code are made publicly available. 10 pages, 4 figures

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Languages

BamanankanLame

Tags

Computation and Language

Similaires

oyinkanchekwas/low-resource-nlp-toolkitToluClassics/Low-Resource-NLP-Tutorialsellis-nlp/low-resource-MTDecolonizing NLP for “Low-resource Languages”MariaYasmeen/NLP-Low-Resource-Paraphrase-detectionSaifWiyar/low-resource-nlp-ml-research

oyinkanchekwas/low-resource-nlp-toolkit

Selective language routing and code-switch audits, evaluated on a reproducible AfriSenti benchmark.

ToluClassics/Low-Resource-NLP-Tutorials

Getting started in NLP for low resource languages # Low-Resource-NLP The goal of this repository i

ellis-nlp/low-resource-MT

Work on machine translation in low resource scenarios and for minority and under-resourced languages

Decolonizing NLP for “Low-resource Languages”

Today African languages are spoken by more than a billion people, yet in the world of machine transl

MariaYasmeen/NLP-Low-Resource-Paraphrase-detection

# Monolingual Paraphrase Detection - Low Resource Sindhi Lang at Sentence Level This project focuses

SaifWiyar/low-resource-nlp-ml-research

Research portfolio for low-resource language NLP, dataset creation, annotation, machine learning, de