Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Impact of Synthetic Code-Switched Training on T5 Zero-Shot Retrieval Accuracy in Adversarial Multilingual Benchmarks

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: How does training T5 models on synthetic code-switched data from low-resource language pairs impact zero-shot retrieval accuracy on adversarial multilingual benchmarks compared to monolingual training? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.7/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.7/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

trainingmodelssyntheticcode-switcheddatalow-resourcelanguagepairs

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Code-Switched Token Ratios in Synthetic Data and Multilingual Encoder Performance on Zero-Shot BEIR RetrievalArtificially Code-Switched Training for Zero-Shot Cross-Lingual Retrieval Robustness Against Adversarial PerturbationsPerformance of Zero-Shot Cross-Lingual Retrieval Models on Low-Resource Languages with Adversarial Code-Switched TrainingCode-switched Training vs Multilingual Fine-tuning for Zero-shot Cross-lingual RetrievalScaling Effectiveness of Zero-Shot Cross-Lingual Retrieval with Synthetic Code-Switched DataArtificially Code-Switched Training Data Volume and Zero-Shot MIRACL Retrieval Accuracy Across Language Resource Levels

Code-Switched Token Ratios in Synthetic Data and Multilingual Encoder Performance on Zero-Shot BEIR Retrieval

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificially Code-Switched Training for Zero-Shot Cross-Lingual Retrieval Robustness Against Adversarial Perturbations

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Performance of Zero-Shot Cross-Lingual Retrieval Models on Low-Resource Languages with Adversarial Code-Switched Training

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Code-switched Training vs Multilingual Fine-tuning for Zero-shot Cross-lingual Retrieval

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling Effectiveness of Zero-Shot Cross-Lingual Retrieval with Synthetic Code-Switched Data

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Artificially Code-Switched Training Data Volume and Zero-Shot MIRACL Retrieval Accuracy Across Language Resource Levels

Transferring information retrieval (IR) models from a high-resource language (typically English) to