Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MIRACL Retrieval Accuracy via Combined Monolingual, Cross-Lingual, and Multilingual Data Augmentation for Low-Resource Languages

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Information retrieval across different languages is an increasingly important challenge in natural language processing. Recent approaches based on multilingual pre-trained language models have achieved remarkable success, yet they often optimize for either monolingual, cross-lingual, or multilingual retrieval performance at the expense of others. This paper proposes a novel hybrid batch training strategy to simultaneously improve zero-shot retrieval performance across monolingual, cross-lingual, and multilingual settings while mitigating language bias. The approach fine-tunes multilingual lang Research goal: What is the impact of combining monolingual, cross-lingual, and multilingual data augmentation techniques during training on the retrieval accuracy of the MIRACL dataset for low-resource languages, compared to using only one augmentation strategy? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

impactcombiningmonolingualcross-lingualmultilingualdataaugmentationtechniques

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Query Augmentation for Cross-Lingual Dense Retrieval in Low-Resource LanguagesSynergistic Optimization of Monolingual and Cross-Lingual Objectives for Long-Tail Low-Resource Retrieval in MIRACLPerformance Degradation in Monolingual Retrieval Accuracy for Low-Resource Languages with Hybrid Batch Training on MIRACLCross-lingual to monolingual sample ratio effects on zero-shot retrieval accuracy in low-resource languagesOptimal Transport Distillation for Robust Cross-Lingual Retrieval in Low-Resource Languages on MIRACLImpact of Monolingual vs. Cross-Lingual Training Proportions on XLM-R Zero-Shot Retrieval Accuracy for Low-Resource Languages

Query Augmentation for Cross-Lingual Dense Retrieval in Low-Resource Languages

Effective cross-lingual dense retrieval methods that rely on multilingual pre-trained language model

Synergistic Optimization of Monolingual and Cross-Lingual Objectives for Long-Tail Low-Resource Retrieval in MIRACL

Information retrieval across different languages is an increasingly important challenge in natural l

Performance Degradation in Monolingual Retrieval Accuracy for Low-Resource Languages with Hybrid Batch Training on MIRACL

Information retrieval across different languages is an increasingly important challenge in natural l

Cross-lingual to monolingual sample ratio effects on zero-shot retrieval accuracy in low-resource languages

Information retrieval across different languages is an increasingly important challenge in natural l

Optimal Transport Distillation for Robust Cross-Lingual Retrieval in Low-Resource Languages on MIRACL

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Impact of Monolingual vs. Cross-Lingual Training Proportions on XLM-R Zero-Shot Retrieval Accuracy for Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l