Logo Lanfrica

ombabiker89-collab/sudanese-dialect-corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
omb
Hôte:
Open-source project to build a high-quality Sudanese Arabic dialect dataset and resources for training AI models to speak natural Sudanese. Focus on conversational data, vocabulary, and fine-tuning examples. # Sudanese Dialect Corpus for AI Open-source project to build a high-quality Sudanese Arabic dialect dataset and resources for training AI models to speak natural Sudanese. ## Goals - Collect authentic conversational Sudanese Arabic - Build vocabulary, phrases, and examples - Create instruction-tuning data for LLMs - Normalize spelling and regional variations (Khartoum, Darfur, etc.) ## Structure - `data/`: Raw and cleaned datasets (JSONL) - `vocab/`: Sudanese-specific words and expressions - `examples/`: Sample conversations - `scripts/`: Cleaning and processing tools ## How to contribute 1. Fork the repo 2. Add your Sudanese phrases or conversations 3. Open a PR Inspired by projects like Sudaverse and AnwarCS/Sudanese-Arabic-LLM. License: MIT