Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Diversity in Training Languages for Dense Retrieval and Zero-Shot Accuracy in Low-Resource Benchmarks

Domain:

natural language processing

Record type:

paper
Creator:
SOV
Publisher:
Zenodo
Host:avatar
Accuracy of English-language Question Answering (QA) systems has improved significantly in recent years with the advent of Transformer-based models (e.g., BERT). These models are pre-trained in a self-supervised fashion with a large English text corpus and further fine-tuned with a massive English QA dataset (e.g., SQuAD). However, QA datasets on such a scale are not available for most of the other languages. Multi-lingual BERT-based models (mBERT) are often used to transfer knowledge from high-resource languages to low-resource languages. Since these models are pre-trained with huge text corp Research goal: How does increasing the diversity of training languages in dense retrieval models affect zero-shot accuracy on low-resource non-English benchmarks compared to high-resource language performance? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 8.2/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.2/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

increasingdiversitytraininglanguagesdenseretrievalmodelsaffect

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Hybrid Batch Training and Zero-Shot Retrieval Accuracy in Low-Resource BEIR LanguagesHybrid Batch Training Ratios and Zero-Shot Retrieval Accuracy in Low-Resource LanguagesXLM-R and mBERT Zero-Shot Retrieval Accuracy in Low-Resource Language BenchmarksSynthetic Cross-Lingual Query Volume and Zero-Shot Dense Retrieval Accuracy in XNLI Low-Resource LanguagesHybrid Batch Training Effects on Zero-Shot Retrieval Accuracy in Low-Resource LanguagesScaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Hybrid Batch Training and Zero-Shot Retrieval Accuracy in Low-Resource BEIR Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Hybrid Batch Training Ratios and Zero-Shot Retrieval Accuracy in Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

XLM-R and mBERT Zero-Shot Retrieval Accuracy in Low-Resource Language Benchmarks

Information retrieval across different languages is an increasingly important challenge in natural l

Synthetic Cross-Lingual Query Volume and Zero-Shot Dense Retrieval Accuracy in XNLI Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) perform strongly in cross-lingual transfer. We propo

Hybrid Batch Training Effects on Zero-Shot Retrieval Accuracy in Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Scaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to