Logo Lanfrica

aidyai/Ibibio-NMT-Corpus-Analysis

Domain:

natural language processing

Record type:

dataset
Creator:
aid
Host:
Reproducible corpus-linguistic analysis of a 205,000-sentence English dataset for English–Ibibio Neural Machine Translation | NLLB fine-tuning | Low-resource NLP | Akwa Ibom, Nigeria # Ibibio NMT Corpus Analysis Reproducible corpus-linguistic analysis of a 205,000-sentence English dataset built for fine-tuning NLLB on English-Ibibio neural machine translation. Ibibio is spoken by approximately 6.3 million first-language speakers and over 10 million people in total when second-language speakers are included (Ethnologue, 2020). Speakers are concentrated in Akwa Ibom, Cross River, and Abia States in southeastern Nigeria, with diaspora communities in Ghana, Cameroon, Equatorial Guinea, the United States, and the United Kingdom. Despite this scale, it has virtually no existing parallel data for machine translation. This project documents the construction and analysis of the English source side of a corpus intended to change that. **Author:** Idara Samuel Osu, ML Engineer --- ## What this is Before translating 205,000 sentences into Ibibio, I needed to know what was actually inside the corpus. This repository contains the Python scripts that analyse the corpus, the plots those scripts produce, and a short research paper that explains everything. Every number in the paper comes from running these scripts. You can verify all of it yourself. --- ## Dataset at a glance | | | |---|---| | Sentences | 205,000 | | Tokens | 2,570,100 | | Vocabulary | 18,385 word types | | Median sentence length | 11 words | | Questions | 44,555 (21.7%) | | Dialogues | 7,999 multi-turn units | | Duplicates | 0 | | Target language | Ibibio, Akwa Ibom State, Nigeria | | Intended model | NLLB fine-tuning | The corpus was assembled from three components: a large synthetic Nigerian-English collection, a topic-structured pedagogical collection covering 70 themes, and a targeted supplement that filled gaps in health, government, legal, and technology domains. All three were cross-deduplicated before the final merge. --- ## Analysis results ### Sentence length The median sentence is 11 words. 69.5% of sentences fall in the 6-15 word range. Sentences above 30 words make …