Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Performance comparison of cross-lingual retrieval models on code-switched vs. M2C2 conversational query data

Domaine:

natural language processing

Type de record:

paper
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
Transferring information retrieval (IR) models from a high-resource language (typically English) to other languages in a zero-shot fashion has become a widely adopted approach. In this work, we show that the effectiveness of zero-shot rankers diminishes when queries and documents are present in different languages. Motivated by this, we propose to train ranking models on artificially code-switched data instead, which we generate by utilizing bilingual lexicons. To this end, we experiment with lexicons induced from (1) cross-lingual word embeddings and (2) parallel Wikipedia page titles. We use Research goal: How does the performance of cross-lingual retrieval models trained on artificially code-switched data compare to models fine-tuned on M2C2's multilingual conversational query dataset, measured by mean average precision (MAP) across low-resource language pairs? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 8.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.5/10.

Visit

doi.org

Tasks

code switchinginformation retrieval

Tags

performancecross-lingualretrievalmodelstrainedartificiallycode-switcheddata

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode