A Low-Resource Bété–French Parallel Corpus for NLP
# OpenBété: A Low-Resource Bété–French Parallel Corpus
-orange.svg)
> *"A language dies when its last speaker dies. Digital preservation begins today."*
---
## Overview
**OpenBété** is an open-source, research-grade parallel corpus for the **Bété language** (ISO 639-3: `bev`), specifically the **Zédígbo dialect** spoken in the Gagnoa region of Côte d'Ivoire. The corpus provides **302 parallel sentence pairs** in French and Bété, annotated with semantic categories, pragmatic intent labels, grammatical tense, and difficulty levels.
This project addresses a critical gap in African NLP: Bété, spoken by an estimated 300,000–450,000 speakers, has **no existing publicly available machine-readable corpus**, no trained translation models, and minimal computational linguistic resources. OpenBété is a foundational step toward reversing this.
---
## Motivation
Approximately **2,000 of Africa's ~3,000 languages** are considered low-resource or endangered in the context of natural language processing. The consequences are profound:
- Speakers of these languages are **excluded from AI-driven tools** (translation, speech recognition, education)
- Cultural and oral knowledge encoded in these languages risks **permanent digital erasure**
- Computational linguistics research continues to concentrate on **less than 1% of the world's languages**
Bété is a **Kru language** of the Niger-Congo family, with complex tonal systems, agglutinative morphology, and oral literary traditions of significant cultural value. It deserves computational representation.
OpenBété is both a **research dataset** and a **cultural preservation initiative**.
---
## Academic Significance
| Dimension | Contribution |
|-----------|--------------|
| **NLP** | First public French–Bété parallel corpus; baseline for MT systems |
| **Linguistics** | Documented Zédígbo dialect with phonological notation |
| **Digital Humanities** | Machine-readable record of oral language |
| **African AI** | Contrib …