Logo Lanfrica

Enoch208/Foundry-Y

Domaine:

natural language processing

Type de record:

datasetmodel
Créateur:
Eno
Hôte:
An execution-driven, zero-contamination data pipeline designed to fine-tune a dual-capability Yoruba-English LLM, governed by strict human-in-the-loop quality gates to guarantee compliance and maximize evaluation performance. # Foundry-Y ### A contamination-proof, 100% human Yorùbá corpus — and the 3B model it taught to speak. Most Yorùbá "AI data" is either scraped from the web (broken diacritics, unnatural language) or generated by an LLM (the same failure, at scale). Foundry-Y answers a harder question: **can you build a Yorùbá training corpus that is provably clean, provably human, provably not the eval set — and actually move a model with it?** A reproducible pipeline turns three approved human sources into a bucket-balanced corpus, native speakers gate its quality, Adaption's AutoScientist trains a LoRA on it, and a live demo shows the result base-vs-fine-tuned, side by side. Every number here is measured. Nothing is generated. > **Human data only. Train splits only. Reproducible by SHA-256. The model learns Yorùbá — including the tone marks — without a single synthetic word.** *Clean if it's human. Rejected if it's contaminated or the wrong license. Verified by tests on every push.* ** The pipeline ↗ **  ·  ** The diacritics trick ↗ **  ·  ** Results ↗ **  ·  ** Run it locally ↗ ** *Dataset, weights, and live demo publish at Gate 3 (after native validation + sign-off).* --- ## Table of contents - The problem - What I built - Architecture - The pipeline, stage by stage - The mixture — and why it's the hard part - The one clever part: teaching the tone marks - The rules I held - Contamination-proofing & reproducibility - Native validation — humans gate the quality - Results — measured honestly - The demo - Licensing — the part most people skip - What's real vs. honest limitations - Quality gates & tests - Tech stack - Project layout - Run it locally - Roadmap --- ## The problem Yorùbá is spoken by **~45 million people**, and it is **tonal**: `ẹ` vs `e`, `ọ` vs `o`, `ṣ` vs `s`, plus grave/acute tone marks — these change the *meaning* of words. Most language models mangle them or just answer in English. The usual "fix" makes it worse …