Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Contrastive Pretraining Objectives and Cross-Lingual Retrieval Accuracy in XTREME Low-Resource Language Pairs

Domaine:

natural language processing

Type de record:

paper
Créateur:
SOV
Éditeur:
Zenodo
Hôte:avatar
Large Language Models (LLMs) have demonstrated remarkable capabilities, yet their performance in low-resource languages (LRLs), such as Swahili, often lags due to data scarcity and underrepresentation in pre-training. A key challenge is achieving robust cross-lingual lexical alignment, crucial for tasks like translation and cross-lingual information retrieval. This paper introduces Targeted Lexical Injection (TLI), a novel and efficient fine-tuning approach. We first demonstrate that Lugha-Llama-8B-wura, a Swahili-centric LLM, exhibits strong, near-perfect lexical alignment for Swahili-English Research goal: How does contrastive pretraining objective selection impact cross-lingual retrieval accuracy for low-resource language pairs in the XTREME benchmark? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.6/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.6/10.

Visit

doi.orgzenodo.org

Languages

Swahili

Tags

contrastivepretrainingobjectiveselectionimpactcross-lingualretrievalaccuracy

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Noise Impact on Zero-Shot Cross-Lingual Retrieval Accuracy in Low-Resource Language PairsImpact of Hybrid Training on Zero-Shot Cross-Lingual Retrieval Accuracy in Low-Resource Language PairsPretraining Dense Retrieval Models on WebFAQ for Zero-Shot Cross-Lingual Recall in Low-Resource XTREME SubsetsZero-shot Cross-lingual Retrieval Accuracy of XLM-R and mBART on Code-switched Low-resource Language PairsImpact of Monolingual vs. Cross-Lingual Training Data Proportions on Zero-Shot Retrieval Accuracy in Low-Resource Language PairsSimultaneous Monolingual and Cross-Lingual Optimization for mXLS Retrieval in Low-Resource Language Pairs

Noise Impact on Zero-Shot Cross-Lingual Retrieval Accuracy in Low-Resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Impact of Hybrid Training on Zero-Shot Cross-Lingual Retrieval Accuracy in Low-Resource Language Pairs

Information retrieval across different languages is an increasingly important challenge in natural l

Pretraining Dense Retrieval Models on WebFAQ for Zero-Shot Cross-Lingual Recall in Low-Resource XTREME Subsets

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Zero-shot Cross-lingual Retrieval Accuracy of XLM-R and mBART on Code-switched Low-resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Impact of Monolingual vs. Cross-Lingual Training Data Proportions on Zero-Shot Retrieval Accuracy in Low-Resource Language Pairs

Information retrieval across different languages is an increasingly important challenge in natural l

Simultaneous Monolingual and Cross-Lingual Optimization for mXLS Retrieval in Low-Resource Language Pairs

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi