Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

M2M-100 Zero-Shot Cross-Lingual Retrieval with Language-Family Data Augmentation

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Information retrieval across different languages is an increasingly important challenge in natural language processing. Recent approaches based on multilingual pre-trained language models have achieved remarkable success, yet they often optimize for either monolingual, cross-lingual, or multilingual retrieval performance at the expense of others. This paper proposes a novel hybrid batch training strategy to simultaneously improve zero-shot retrieval performance across monolingual, cross-lingual, and multilingual settings while mitigating language bias. The approach fine-tunes multilingual lang Research goal: What is the impact of incorporating language-family-specific data augmentation during hybrid batch training on the zero-shot cross-lingual retrieval performance of M2M-100, measured by nDCG@10 and MRR on BEIR for both mid- and low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.0/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.0/10.

Visit

doi.org

Tasks

information retrieval

Tags

impactincorporatinglanguage-family-specificdataaugmentationduringhybridbatch

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode