DAtaset Hassaniya
# DAH (DAtaset Hassaniya)
## Project Overview
**DAH (DAtaset Hassaniya)** is the first publicly available, open-source dataset for Hassaniya Arabic. Hassaniya is a dialect of Arabic spoken in Mauritania. Despite its widespread use, the language is not well-represented in digital resources.
The primary goal of the DAH project is to bridge this gap. By creating a high-quality dataset, we aim to facilitate:
1. **Machine Translation:** The development of accurate translation tools between English and Hassaniya. This is crucial for improving technologies like online translators and other AI-driven language models.
2. **Linguistic Research:** Providing researchers with a structured dataset to study the grammar, vocabulary, and nuances of the Hassaniya dialect.
3. **Language Preservation:** Documenting Hassaniya in a digital format that is accessible to both the Hassaniya-speaking community and a global audience.
***
## Dataset Details
The DAH dataset is structured as a parallel corpus, meaning that each entry consists of a sentence in English paired with its translation in Hassaniya.
| Component | Description |
| :--- | :--- |
| `english` | The source sentence in English. |
| `hassaniya-ar` | The Hassaniya translation written in the traditional Arabic script. |
| `hassaniya-en` | The Hassaniya translation written in a Latin-based script, often referred to as "Arabizi." |
* **Source:** The sentences are adapted from the Tatoeba Project (
tatoeba.org) , a large, community-based collection of sentences and translations.
* **Quality:** All entries have been reviewed by the founders of the Hassan-IA community to ensure accuracy.
* **License:** The dataset is released under the CC-BY-4.0 license, which allows for open use and adaptation, provided that appropriate credit is given.
***
## Creators
- Ahlam Abdelkader
- Emani Babe
- Oumoukelthoum Sidenna
***
## The Hassaniya Latin Script
To maintain consistency, a standardized transliteration system …