Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Bilingual Lexicon Size for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: To what extent does scaling the size of the bilingual lexicon used to generate artificial code-switching data improve the robustness of zero-shot cross-lingual retrieval models on low-resource languages in the XQA benchmark? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.6/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.6/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

extentscalingsizebilinguallexiconusedgenerateartificial

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesBilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesScaling Model Size for Zero-Shot Cross-Lingual Retrieval on Code-Switched Data in Low-Resource LanguagesHybrid Batch Training for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesMultilingual Contrastive Learning for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesScaling Intermediate-Task Dataset Size for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Model Size for Zero-Shot Cross-Lingual Retrieval on Code-Switched Data in Low-Resource Languages

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Hybrid Batch Training for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Multilingual Contrastive Learning for Robust Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Scaling Intermediate-Task Dataset Size for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni