Open-source project to build a high-quality Sudanese Arabic dialect dataset and resources for training AI models to speak natural Sudanese. Focus on conversational data, vocabulary, and fine-tuning examples.
# Sudanese Dialect Corpus for AI
Open-source project to build a high-quality Sudanese Arabic dialect dataset and resources for training AI models to speak natural Sudanese.
## Goals
- Collect authentic conversational Sudanese Arabic
- Build vocabulary, phrases, and examples
- Create instruction-tuning data for LLMs
- Normalize spelling and regional variations (Khartoum, Darfur, etc.)
## Structure
- `data/`: Raw and cleaned datasets (JSONL)
- `vocab/`: Sudanese-specific words and expressions
- `examples/`: Sample conversations
- `scripts/`: Cleaning and processing tools
## How to contribute
1. Fork the repo
2. Add your Sudanese phrases or conversations
3. Open a PR
Inspired by projects like Sudaverse and AnwarCS/Sudanese-Arabic-LLM.
License: MIT