Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

kumshey/hausa-chinese-corpus

Domain:

natural language processing

Record type:

dataset
Creator:
kum
Host:
This repository hosts the first large-scale, quality-benchmarked Chinese-Hausa parallel corpus. It contains one million sentence pairs designed to advance machine translation for low-resource languages. The work provides a reproducible framework to foster research and cross-cultural communication. # Quality-Benchmarked Hausa-Chinese Parallel Corpus --- This repository hosts the **first large-scale, quality-benchmarked parallel corpus** for the Chinese-Hausa language pair. It is designed to address the critical resource gap for low-resource machine translation (MT) and to catalyze research and development in this area. The corpus contains approximately **one million sentence pairs**, constructed through a rigorous methodology detailed in our research paper. --- ## 📖 Project Overview Machine translation for low-resource languages like Hausa is severely hindered by the scarcity of high-quality parallel data. This project tackles this challenge by providing a **foundational dataset** for Chinese-Hausa MT, a linguistically complex and previously under-resourced pair. Our work introduces a reproducible framework for dataset construction, including **novel quality metrics** to ensure the corpus is robust and suitable for training state-of-the-art models. By making this resource public, we aim to: - Enhance information accessibility for millions of Hausa speakers - Support linguistic diversity - Foster cross-cultural communication --- ### **Key Features** - **Massive Scale:** Contains ~1M sentence pairs, with over 700K unique pairs — one of the most significant resources for any low-resource African language. - **High-Quality & Diverse:** Sourced from literature, news, and web-crawled text. Covers 20+ domains, including healthcare, education, and politics. Data has undergone meticulous cleaning. - **Advanced Augmentation:** Two-step back-translation process + weak supervision for size and linguistic diversity, improving model adaptability. - **Linguistically-Aware:** Specialized tokenization strategies — SentencePiece for Hausa, BPEmb for Chinese — to respect morphological and tonal features. - **Novel Quality Benchmarks:** Three new metrics proposed: **Coverage**, **Unique Semantic Alignment Rate**, and **Speciality Preservation**. --- ## 📂 Repo …

Visit

github.com

Tasks

machine translation

Languages

Hausa

Licenses

Apache-2.0

Similar

Kumshe/DeepSeek-Hausa-Chinese-TranslationKumshe/testing_mT5-hausa-to-ChineseHausa Corpushausa-corpusmradermacher/DeepSeek-Hausa-Chinese-Translation-GGUFKumshe/mt5-small_ne_-hausa-to-chinese

Kumshe/DeepSeek-Hausa-Chinese-Translation

Kumshe/testing_mT5-hausa-to-Chinese

Hausa Corpus

Hausa datasets with stopwords

hausa-corpus

A collection of textual datasets in Hausa language and the corresponding translation in English language.

mradermacher/DeepSeek-Hausa-Chinese-Translation-GGUF

Kumshe/mt5-small_ne_-hausa-to-chinese