Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Robustness of mT5 and XLM-R versus Monolingual Code-Switched Models in Zero-Shot Low-Resource Retrieval

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: How does the robustness of mT5 and XLM-R in zero-shot cross-lingual retrieval compare to monolingual models trained on code-switched data when evaluated on low-resource language pairs using nDCG@10? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.3/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.3/10.

Visit

doi.orgzenodo.org

Tasks

code switchinginformation retrieval

Tags

robustnessmT5XLM-Rzero-shotcross-lingualretrievalmonolingualmodels

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Robustness of Zero-Shot Cross-Lingual Retrieval Models via Code-Switched Pre-Training in Low-Resource LanguagesZero-shot Cross-lingual Retrieval Accuracy of XLM-R and mBART on Code-switched Low-resource Language PairsmT5 Code-Switching Training for Low-Resource Zero-Shot RetrievalMultilingual BERT versus XLM-R for Zero-Shot Cross-Lingual Retrieval on Low-Resource Pairs with Artificial Code-SwitchingZero-Shot Cross-Lingual Retrieval Robustness Under Varying Low-Resource Language Proportions in Code-Switched Training DataComparative Analysis of Hybrid Batch Training Against XLM-R and mT5 for Zero-Shot Cross-Lingual Retrieval in Low-Resource

Robustness of Zero-Shot Cross-Lingual Retrieval Models via Code-Switched Pre-Training in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Zero-shot Cross-lingual Retrieval Accuracy of XLM-R and mBART on Code-switched Low-resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

mT5 Code-Switching Training for Low-Resource Zero-Shot Retrieval

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Multilingual BERT versus XLM-R for Zero-Shot Cross-Lingual Retrieval on Low-Resource Pairs with Artificial Code-Switching

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Zero-Shot Cross-Lingual Retrieval Robustness Under Varying Low-Resource Language Proportions in Code-Switched Training Data

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Comparative Analysis of Hybrid Batch Training Against XLM-R and mT5 for Zero-Shot Cross-Lingual Retrieval in Low-Resource

Information retrieval across different languages is an increasingly important challenge in natural l