Tone-marked linguistic data for low-resource NLP.
# π³π¬ Tone-Marked Igbo Proverbs Parallel Corpus
This repository hosts the documentation, research context and linguistic frameworks for the **Igbo Proverbs Dataset**. This project serves as a highly specialised, low-resource language resource designed to bridge the gap between traditional African oral literature and modern Natural Language Processing (NLP).
## Project Overview
Rather than focusing on sheer volume, this dataset acts as a deeply curated **linguistic micro-corpus**. It provides a granular, multidimensional look at structural variations, tonal inflections, and contextual applications within localised Igbo dialects. It establishes a high quality data baseline meant for fine-tuning linguistic models, testing AI translation systems, and preserving indigenous knowledge systems.
## Dataset Access
To protect the integrity of this specialised data and track its academic or commercial application, the primary dataset file is securely hosted and gated on Hugging Face.
π **Download the Dataset on Hugging Face**
*Note: Access is granted automatically. You simply need to log in with a Hugging Face account and quickly fill out the short form stating your intended project or research use case.*
## π Dataset Specifications
- **Core Language:** Igbo (`ig`)
- **Dataset Volume:** 50 highly curated entries (Ongoing compilation)
- **Format Type:** Structured CSV
- **License Framework:** CC-BY-4.0 (Attribution Required)
### Rich 10-Unit Data Schema
To ensure maximum analytical depth, every single proverb row is mapped across 10 precise linguistic and contextual dimensions:
1. **Proverb Code:** Unique identifier for systematic indexing and cross-referencing.
2. **Raw Dialectal Proverb:** The original proverb preserved with native, localised pronunciation and tone markings.
3. **Standard Orthography:** The proverb structurally mapped to central standard Igbo orthography with full tone markings.
4. **Literal Meaning:** Word-for-word translation to assist comparative sem β¦