# NLP Resources for Hausa and Fongbe Languages
A comprehensive catalog of publicly available text and speech resources for Hausa and Fongbe, two low-resource African languages. This repository accompanies our survey paper on NLP resources for these languages.
---
## Table of Contents
- Overview
- Hausa Resources
- Text Corpora
- Parallel Corpora & Machine Translation
- Named Entity Recognition
- Sentiment Analysis
- Speech & ASR
- Question Answering & Reasoning
- Other Resources
- Fongbe Resources
- Text Corpora
- Parallel Corpora & Machine Translation
- Speech & ASR
- Other Resources
- Multilingual Resources
- How to Contribute
- Citation
- License
---
## Overview
| Language | ISO 639-3 | Speakers | Region |
|----------|-----------|----------|--------|
| Hausa | hau | ~80 million | Nigeria, Niger, Ghana, Cameroon |
| Fongbe | fon | ~2 million | Benin, Togo, Nigeria |
This repository catalogs **60 datasets for Hausa** and **18 datasets for Fongbe** spanning various NLP tasks including machine translation, named entity recognition, sentiment analysis, speech recognition, and more.
---
## Hausa Resources
### Hausa Text Corpora
| Resource | Description | Size | License | Links |
|----------|-------------|------|---------|-------|
| **CC-100 (Hausa)** | Monolingual data from Common Crawl for language modeling | Large | - | Download |
| **AfriBERTa Corpus** | Multilingual corpus including Hausa from BBC news and Common Crawl | - | Apache 2.0 | GitHub | HuggingFace |
| **Naijaweb Dataset** | 270k+ documents from Nigerian web sources | ~230M tokens | MIT | HuggingFace |
| **BloomLM** | Bloom Library data for language modeling (~400 languages) | Varies | CC-BY variants | HuggingFace |
| **African Storybook** | Multilingual children's stories | ~6,700 stories | - | GitHub |
### Hausa Parallel Corpora & Machine Translation
| Resource | Description | Size | License | Links |
|----------|-------------|------|---------|-------|
| **MAFAND-MT** | News domain pa …