This repository hosts the first large-scale, quality-benchmarked Chinese-Hausa parallel corpus. It contains one million sentence pairs designed to advance machine translation for low-resource languages. The work provides a reproducible framework to foster research and cross-cultural communication.
# Quality-Benchmarked Hausa-Chinese Parallel Corpus
---
This repository hosts the **first large-scale, quality-benchmarked parallel corpus** for the Chinese-Hausa language pair.
It is designed to address the critical resource gap for low-resource machine translation (MT) and to catalyze research and development in this area.
The corpus contains approximately **one million sentence pairs**, constructed through a rigorous methodology detailed in our research paper.
---
## 📖 Project Overview
Machine translation for low-resource languages like Hausa is severely hindered by the scarcity of high-quality parallel data.
This project tackles this challenge by providing a **foundational dataset** for Chinese-Hausa MT, a linguistically complex and previously under-resourced pair.
Our work introduces a reproducible framework for dataset construction, including **novel quality metrics** to ensure the corpus is robust and suitable for training state-of-the-art models.
By making this resource public, we aim to:
- Enhance information accessibility for millions of Hausa speakers
- Support linguistic diversity
- Foster cross-cultural communication
---
### **Key Features**
- **Massive Scale:**
Contains ~1M sentence pairs, with over 700K unique pairs — one of the most significant resources for any low-resource African language.
- **High-Quality & Diverse:**
Sourced from literature, news, and web-crawled text. Covers 20+ domains, including healthcare, education, and politics. Data has undergone meticulous cleaning.
- **Advanced Augmentation:**
Two-step back-translation process + weak supervision for size and linguistic diversity, improving model adaptability.
- **Linguistically-Aware:**
Specialized tokenization strategies — SentencePiece for Hausa, BPEmb for Chinese — to respect morphological and tonal features.
- **Novel Quality Benchmarks:**
Three new metrics proposed: **Coverage**, **Unique Semantic Alignment Rate**, and **Speciality Preservation**.
---
## 📂 Repo …