Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-Lingual Morpheme Networks (CLMN) for Endangered Language Preservation and Disinformation Detection via Topological Wave Interference

Domaine:

natural language processing

Type de record:

model
Créateur:
Kab
Éditeur:
Zenodo
Hôte:avatar
CLMN (Cross-Lingual Morpheme Network) is a novelty physics-inspired framework that models language contact, code-switching, and morphological dynamics as acoustic wave interference phenomena. By treating morpheme boundaries as probability fields governed by reaction-diffusion equations and analyzing their topological persistence, CLMN enables three critical applications: (1) ultra-low-resource endangered language preservation requiring only 10 hours of audio, (2) disinformation detection achieving 99.2% accuracy in identifying manipulated political speech, and (3) cross-lingual translation for language pairs with fewer than 1,000 parallel sentences. Our architecture integrates wave-based morpheme boundary detection with topological data analysis, achieving state-of-the-art performance while maintaining extreme computational efficiency (52,847 parameters, ~0.2 MB). We validate CLMN on three Kenyan language pairs (Swahili-English, Kikuyu-Swahili, Luo-English) and demonstrate successful reconstruction of Yaaku, an endangered Kenyan language with fewer than 50 native speakers. This work establishes the first computational framework connecting wave physics, topology, and linguistics for practical language technology applications in low-resource settings. Keywords: Morpheme networks, endangered languages, disinformation detection, topological data analysis, wave interference, code-switching, low-resource NLP

Visit

doi.orgzenodo.org

Tasks

code switching

Languages

GikuyuSwahiliYaaku

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2025 Kabbey, F.http://rightsstatements.org/vocab/InC/1.0/

Similaires

Automated Cross-Lingual Semantic Alignment via Hierarchical Generative Adversarial Networks for Endangered Language DocumentationSynthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice CloningCross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuningXLM-R Performance Enhancement via Intermediate Task Fine-Tuning for Cross-Lingual Euphemism DetectionEffectively Prompting Small-sized Language Models for Cross-lingual Tasks via Winning TicketsIntermediate Language Fine-Tuning for Cross-Lingual Euphemism Detection in XLM-R

Automated Cross-Lingual Semantic Alignment via Hierarchical Generative Adversarial Networks for Endangered Language Documentation

**Abstract:** Low-resource languages face an urgent threat of extinction, largely due to limited doc

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen

Cross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuning

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

XLM-R Performance Enhancement via Intermediate Task Fine-Tuning for Cross-Lingual Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Effectively Prompting Small-sized Language Models for Cross-lingual Tasks via Winning Tickets

Current soft prompt methods yield limited performance when applied to small-sized models (fewer than

Intermediate Language Fine-Tuning for Cross-Lingual Euphemism Detection in XLM-R

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec