Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: What is the impact of scaling the volume of artificially code-switched training data on the zero-shot nDCG@10 performance of multilingual dense retrievers across low-resource languages? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.orgzenodo.org

Tasks

code switchinginformation retrieval

Tags

impactscalingvolumeartificiallycode-switchedtrainingdatazero-shot

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Scaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic LanguagesArtificially Code-Switched Training for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesImpact of Artificially Code-Switched Training Data Volume on Zero-Shot Retrieval for Low-Resource Languages in MIRACLScaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Low-Resource Data in MIRACLCross-Lingual Embeddings for Zero-Shot Retrieval on Artificially Code-Switched Low-Resource DataScaling Behavior of Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Versus Native Data Across Low-Resource

Scaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificially Code-Switched Training for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Impact of Artificially Code-Switched Training Data Volume on Zero-Shot Retrieval for Low-Resource Languages in MIRACL

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Low-Resource Data in MIRACL

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Cross-Lingual Embeddings for Zero-Shot Retrieval on Artificially Code-Switched Low-Resource Data

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Behavior of Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Versus Native Data Across Low-Resource

Transferring information retrieval (IR) models from a high-resource language (typically English) to