Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Effectiveness of Zero-Shot Cross-Lingual Retrieval with Synthetic Code-Switched Data

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: How does the effectiveness of zero-shot cross-lingual retrieval models trained on synthetic code-switched data scale with the size of the bilingual lexicons used for data generation, as evaluated by precision@k on low-resource language pairs in XGLUE? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.6/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.6/10.

Visit

doi.orgzenodo.org

Tasks

code switchinginformation retrieval

Tags

effectivenesszero-shotcross-lingualretrievalmodelstrainedsyntheticcode-switched

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Scaling Performance of Zero-Shot Cross-Lingual Retrievers Trained on Synthetic Code-Switched Data on XTREMEZero-shot cross-lingual retrieval effectiveness in low-resource languages with code-switched trainingScaling Code-Switched Data Ratios in XLM-R Training for Zero-Shot Cross-Lingual Retrieval PerformanceScaling Zero-Shot Cross-Lingual Retrieval Models Trained on Code-Switched Data for Low-Resource LanguagesPerformance of Zero-Shot Cross-Lingual Retrieval Models with Artificially Code-Switched Training DataScaling Model Size for Zero-Shot Cross-Lingual Retrieval on Code-Switched Data in Low-Resource Languages

Scaling Performance of Zero-Shot Cross-Lingual Retrievers Trained on Synthetic Code-Switched Data on XTREME

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Zero-shot cross-lingual retrieval effectiveness in low-resource languages with code-switched training

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Code-Switched Data Ratios in XLM-R Training for Zero-Shot Cross-Lingual Retrieval Performance

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Zero-Shot Cross-Lingual Retrieval Models Trained on Code-Switched Data for Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Performance of Zero-Shot Cross-Lingual Retrieval Models with Artificially Code-Switched Training Data

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Model Size for Zero-Shot Cross-Lingual Retrieval on Code-Switched Data in Low-Resource Languages

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi