Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Neural Recovery of Historical Lexical Structure in Bantu Languages from Modern Data

Domain:

natural language processing

Record type:

papersoftwaredataset
Creator:
MutMug
Host:avatar
We investigate whether neural models trained exclusively on modern morphological data can recover cross-lingual lexical structure consistent with historical reconstruction. Using BantuMorph v7, a transformer over Bantu morphological paradigms, we analyze 14 Eastern and Southern Bantu languages, extract encoder embeddings for their noun and verb lemmas, and identify 728 noun and 1,525 verb cognate candidates shared across 5+ languages. Evaluating these candidates against established historical resources-the Bantu Lexical Reconstructions database (BLR3; 4,786 reconstructed Proto-Bantu forms) and the ASJP basic vocabulary-we confirm 10 of the top 11 noun candidates (90.9%) align with previously reconstructed Proto-Bantu forms, including *-ntU 'person' (8 languages), *gombe 'cow' (9 languages), and *mUn (9 languages). Extending to verbs, 12 verb cognates align with reconstructed Proto-Bantu roots, including *-bon- 'see' and *-jIm- 'stand', each attested across wide geographic ranges. Cross-model validation using an independent translation model (NLLB-600M) confirms these patterns: both models recover cognate clusters and phylogenetic groupings consistent with established Guthrie-zone classifications (p < 0.01). Cross-lingual noun class analysis reveals that all 13 productive classes maintain >0.83 cosine similarity across languages (within-class > between-class, p < 10^-9). Our dataset is restricted to Eastern and Southern Bantu, so we interpret these results as recovering shared Bantu lexical structure consistent with Proto-Bantu rather than definitively distinguishing Proto-Bantu retentions from later regional innovations.

Visit

arxiv.org

Languages

Fulfulde, AdamawaSena

Tags

Machine LearningComputation and Language

Similar

Lexical cluster analysis of 10 Bantu A80 languagesAn Outline of the Grammatical Structure of Central Bantu LanguagesOptimised Recovery of Critical Minerals from Historical Pegmatite Tailings in RwandaTargeted Lexical Injection for Cross-Lingual Alignment in Underrepresented Bantu Languages on XCOPABantu sentence structureBantu Lexical Loans in Hadza: An Introduction

Lexical cluster analysis of 10 Bantu A80 languages

International audience In this presentation, I will show the results of a small-scale

An Outline of the Grammatical Structure of Central Bantu Languages

Draft of a grammatical outline of the Bantu languages spoken in Zambia by Thomas Givon.

Optimised Recovery of Critical Minerals from Historical Pegmatite Tailings in Rwanda

ABSTRACT Historical pegmatite mining tailings in Rwanda represent a significant but largely unquant

Targeted Lexical Injection for Cross-Lingual Alignment in Underrepresented Bantu Languages on XCOPA

Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their performance in low

Bantu sentence structure

Bantu Lexical Loans in Hadza: An Introduction

Hadza is a language rich in lexical loans from many different languages. The study of the H