Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

M2M-100 Zero-Shot Cross-Lingual Retrieval with Language-Family Data Augmentation

Domain:

natural language processing

Record type:

paper
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Information retrieval across different languages is an increasingly important challenge in natural language processing. Recent approaches based on multilingual pre-trained language models have achieved remarkable success, yet they often optimize for either monolingual, cross-lingual, or multilingual retrieval performance at the expense of others. This paper proposes a novel hybrid batch training strategy to simultaneously improve zero-shot retrieval performance across monolingual, cross-lingual, and multilingual settings while mitigating language bias. The approach fine-tunes multilingual lang Research goal: What is the impact of incorporating language-family-specific data augmentation during hybrid batch training on the zero-shot cross-lingual retrieval performance of M2M-100, measured by nDCG@10 and MRR on BEIR for both mid- and low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.0/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.0/10.

Visit

doi.org

Tasks

information retrieval

Tags

impactincorporatinglanguage-family-specificdataaugmentationduringhybridbatch

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode