Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MaDi-12/linguistic-validation-low-resource-datasets

Domain:

natural language processing
Creator:
MaD
Host:
Linguistic audit of the Wolof-labeled subset in the MURI-IT dataset using hierarchical language detection and model triangulation. # Linguistic Audit of the Wolof-Labeled Subset (MURI-IT) ## Overview This repository presents a methodological linguistic audit of the subset labeled as “Wolof” in the MURI-IT dataset. The objective is to evaluate whether the linguistic labeling is consistent with the actual textual content using independent language identification models. ## Methodology A hierarchical detection strategy was applied: 1. **Global multilingual detection** using fastText (176 languages) 2. Confidence-based filtering (τ = 0.7) 3. Specialized African language detection using AfroLID 4. Comparative triangulation analysis ## Key Findings - 0% of instances were detected as Wolof with strong confidence using fastText. - AfroLID classified 50.33% of ambiguous cases as Wolof, but with extremely low confidence scores (~0.0017). - No robust Wolof signal was detected across models. The results support the hypothesis of linguistic discordance between the labeling and actual content. ## Related Work AfroLID selection was informed by prior evaluation: MaDi-12/wolof-language-dete… ## Author Marème Diop AI & Big Data Engineering

Visit

github.com

Tasks

language identification

Languages

BariMaMa’diWolof