Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

The Geometry of Low-Resource Language Representations

Domain:

natural language processing

Record type:

paper
Creator:
MeyBuys, Jan
Publisher:
arXiv
Host:avatar
The performance gap between low- and high-resource languages in LLMs is widely known, but it remains unclear which internal model factors drive these disparities. In this paper, we characterise this gap through the lens of representational geometry. Comparing the geometric properties of hidden representations across 30 languages reveals that LLM geometry is systematically related to language data availability. The most consistent effect is in final layers, where low-resource languages exhibit representational degeneration. To counter this, we investigate the effectiveness of regularisation terms to penalise degeneration during continued pretraining (CPT). Experiments monolingually adapting 9 base LLMs to 10 African languages show that geometric regularisation successfully reduces representational degeneration during CPT. For larger models, cosine similarity-based regularisation marginally improves performance over vanilla CPT, with more consistent gains on the most challenging tasks. We establish that the representational geometry of low- and high-resource languages in LLMs is measurably distinct, and that targeted geometric intervention is a viable strategy for improving CPT for low-resource languages.

Visit

doi.org

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Low-Resource Parsing with Crosslingual Contextualized RepresentationsCombining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource LanguagesBitext Mining Using Distilled Sentence Representations for Low-Resource LanguagesMultilingual Representations for Low Resource Speech Recognition and Keyword SearchMultilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingImproving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations

Low-Resource Parsing with Crosslingual Contextualized Representations

Despite advances in dependency parsing, languages with small treebanks still present challenges. We

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we exp

Bitext Mining Using Distilled Sentence Representations for Low-Resource Languages

Scaling multilingual representation learning beyond the hundred most frequent languages is challengi

Multilingual Representations for Low Resource Speech Recognition and Keyword Search

This paper examines the impact of multilingual (ML) acoustic representations on Automatic Speech Rec

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations

In this paper, we propose to boost low-resource cross-lingual document retrieval performance with de