Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Cross-Lingual Morpheme Networks (CLMN) for Endangered Language Preservation and Disinformation Detection via Topological Wave Interference

Domain:

natural language processing

Record type:

model
Creator:
Kab
Publisher:
Zenodo
Host:avatar
CLMN (Cross-Lingual Morpheme Network) is a novelty physics-inspired framework that models language contact, code-switching, and morphological dynamics as acoustic wave interference phenomena. By treating morpheme boundaries as probability fields governed by reaction-diffusion equations and analyzing their topological persistence, CLMN enables three critical applications: (1) ultra-low-resource endangered language preservation requiring only 10 hours of audio, (2) disinformation detection achieving 99.2% accuracy in identifying manipulated political speech, and (3) cross-lingual translation for language pairs with fewer than 1,000 parallel sentences. Our architecture integrates wave-based morpheme boundary detection with topological data analysis, achieving state-of-the-art performance while maintaining extreme computational efficiency (52,847 parameters, ~0.2 MB). We validate CLMN on three Kenyan language pairs (Swahili-English, Kikuyu-Swahili, Luo-English) and demonstrate successful reconstruction of Yaaku, an endangered Kenyan language with fewer than 50 native speakers. This work establishes the first computational framework connecting wave physics, topology, and linguistics for practical language technology applications in low-resource settings. Keywords: Morpheme networks, endangered languages, disinformation detection, topological data analysis, wave interference, code-switching, low-resource NLP

Visit

doi.orgzenodo.org

Tasks

code switching

Languages

GikuyuSwahiliYaaku

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcodeCopyright (C) 2025 Kabbey, F.http://rightsstatements.org/vocab/InC/1.0/

Similar

Automated Cross-Lingual Semantic Alignment via Hierarchical Generative Adversarial Networks for Endangered Language DocumentationSynthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice CloningCross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuningXLM-R Performance Enhancement via Intermediate Task Fine-Tuning for Cross-Lingual Euphemism DetectionEffectively Prompting Small-sized Language Models for Cross-lingual Tasks via Winning TicketsIntermediate Language Fine-Tuning for Cross-Lingual Euphemism Detection in XLM-R

Automated Cross-Lingual Semantic Alignment via Hierarchical Generative Adversarial Networks for Endangered Language Documentation

**Abstract:** Low-resource languages face an urgent threat of extinction, largely due to limited doc

Synthetic Speech, Real Signal: Paralinguistic Preservation and Cross-Lingual Augmentation via Voice Cloning

Synthetic data augmentation in speech is common practice for linguistic tasks like ASR, but has seen

Cross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuning

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

XLM-R Performance Enhancement via Intermediate Task Fine-Tuning for Cross-Lingual Euphemism Detection

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec

Effectively Prompting Small-sized Language Models for Cross-lingual Tasks via Winning Tickets

Current soft prompt methods yield limited performance when applied to small-sized models (fewer than

Intermediate Language Fine-Tuning for Cross-Lingual Euphemism Detection in XLM-R

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec