Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Scaling Performance of Models Trained on Artificially Code-Switched Data for Unseen Low-Resource Language Pairs

Domain:

natural language processing
Creator:
Ass
Publisher:
Zenodo
Host:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: How does the performance of models trained on artificially code-switched data scale when evaluated on low-resource language pairs not covered by the bilingual lexicons used during training? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

performancemodelstrainedartificiallycode-switcheddatascaleevaluated

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Performance of Zero-Shot Cross-Lingual Retrieval Models Trained on Artificial Code-Switched Data on Unseen Low-Resource LanguagesScaling Zero-Shot Cross-Lingual Retrieval Models Trained on Code-Switched Data for Low-Resource LanguagesScaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic LanguagesScalability of Artificially Code-Switched Data for Low-Resource Language Rankers in MTOPScaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Low-Resource Data in MIRACLScaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Performance of Zero-Shot Cross-Lingual Retrieval Models Trained on Artificial Code-Switched Data on Unseen Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Zero-Shot Cross-Lingual Retrieval Models Trained on Code-Switched Data for Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Artificially Code-Switched Data for Zero-Shot Retrieval in Low-Resource Afro-Asiatic Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scalability of Artificially Code-Switched Data for Low-Resource Language Rankers in MTOP

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval on Artificially Code-Switched Low-Resource Data in MIRACL

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Artificially Code-Switched Training Data for Zero-Shot Multilingual Dense Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to