Logo Lanfrica

lemneya/hdrp-data-collection-protocol

Domain:

natural language processing

Record type:

project
Creator:
lem
Host:
HDRP: Hassaniya Data Refinery Pipeline - Master Data Collection Protocol for LLM Training # HDRP: Hassaniya Data Refinery Pipeline **Master Data Collection Protocol for Hassaniya LLM Training** ## Overview HDRP is a comprehensive data collection and processing protocol designed to build high-quality training datasets for Hassaniya Arabic language models. The protocol emphasizes **chat-first understanding** while maintaining dialect authenticity and preventing MSA/Darija contamination. ## Core Principles ### 1. Chat-First Philosophy The model should think like a WhatsApp/Facebook user first: short turns, messy spelling, code-switch, emojis, and intent-heavy phrasing. ### 2. Dialect Preservation - **Prioritize authentic Hassaniya** expressions - **Avoid Darija mixing** (Moroccan Arabic contamination) - **Exclude Mauritania news** as a source (non-representative) - **Include Parliament sessions** (exception - rich Hassaniya debate since 2007) ### 3. Balanced Source Mix Prevent any single source from dominating the training signal. ## Data Collection Buckets | Bucket | DAPT Weight | SFT Weight | Sources | |--------|-------------|------------|---------| | **everyday_chat** | 65% | 85% | WhatsApp, Facebook DMs | | **public_comments** | 17% | 10% | Website comments, Facebook | | **marketplace_qa** | 10% | 3% | Marketplace listings | | **tv_discussion** | 5% | 2% | TV/Radio, YouTube | | **culture_story_poetry** | 2% | 0% | Film, Poetry, Azawan | | **monologue_specialist** | 1% | 0% | Parliament, Religion | ## Global Caps (Hard Limits) | Cap | Value | Purpose | |-----|-------|---------| | Max single source | 3% | Prevent style imprinting | | Max single thread | 1% | Prevent topic bias | | Max monologue | 10% | Keep chat rhythm | | Min dialogue | 60% | Ensure conversational focus | ## Heat Policy (Content Moderation) | Level | Name | DAPT | SFT | Notes | |-------|------|------|-----|-------| | 0 | Normal | ✅ | ✅ | Standard content | | 1 | Mild | ✅ | ✅ | Light disagreement | | 2 | Heated | ✅ | ❌ | Strong debate | | 3 | Toxic | ❌ | ❌ | Eval only | ## D …