Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Scalability of Artificially Code-Switched Data for Low-Resource Language Rankers in MTOP

Domaine:

natural language processing
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: Does the scalability of artificially code-switched data for training rankers hold when applied to low-resource languages in the MTOP dataset, and how does it compare to zero-shot cross-lingual transfer from high-resource languages in terms of retrieval performance metrics like nDCG and MAP? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

scalabilityartificiallycode-switcheddatatrainingrankersholdapplied

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language PairsCross-Lingual Embeddings for Zero-Shot Retrieval on Artificially Code-Switched Low-Resource DataScaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic LanguagesArtificially Code-Switched Training Data for Robust Zero-Shot Cross-Lingual Ranking in Low-Resource SettingsScaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource LanguagesArtificially Code-Switched Data for Cross-Lingual Embedding Alignment in Low-Resource Zero-Shot XNLI Settings

Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language Pairs

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Cross-Lingual Embeddings for Zero-Shot Retrieval on Artificially Code-Switched Low-Resource Data

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificially Code-Switched Training Data for Robust Zero-Shot Cross-Lingual Ranking in Low-Resource Settings

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificially Code-Switched Data for Cross-Lingual Embedding Alignment in Low-Resource Zero-Shot XNLI Settings

Transferring information retrieval (IR) models from a high-resource language (typically English) to