This repository contains open-source datasets designed for training and fine-tuning Large Language Models (LLMs) on the Hassaniya dialect of Mauritania.
The repository currently hosts two primary datasets in JSONL format, ready for LLM training:
Content: Parallel corpus of Standard Arabic to Hassaniya translation pairs.
Size: ~4,430 pairs.