Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improving Low-Resource Cross-lingual Document Retrieval by Reranking with Deep Bilingual Representations

Domaine:

natural language processing

Type de record:

paper
Créateur:
Zhang, RuiWesShiBin
Hôte:avatar
In this paper, we propose to boost low-resource cross-lingual document retrieval performance with deep bilingual query-document representations. We match queries and documents in both source and target languages with four components, each of which is implemented as a term interaction-based deep neural network with cross-lingual word embeddings as input. By including query likelihood scores as extra features, our model effectively learns to rerank the retrieved documents by using a small number of relevance labels for low-resource language pairs. Due to the shared cross-lingual word embedding space, the model can also be directly applied to another language pair without any training label. Experimental results on the MATERIAL dataset show that our model outperforms the competitive translation-based baselines on English-Swahili, English-Tagalog, and English-Somali cross-lingual information retrieval tasks. ACL 2019, short paper

Visit

arxiv.org

Tasks

information retrieval

Languages

SomaliSwahili

Tags

Computation and Language

Similaires

Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesZero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource LanguagesScaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource LanguagesScaling of Artificially Code-Switched Data Effectiveness with Bilingual Lexicon Size for Low-Resource Cross-Lingual RetrievalCONCRETE: Improving Cross-lingual Fact-checking with Cross-lingual RetrievalImproving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation

Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Zero-Shot Cross-Lingual Reranking with Large Language Models for Low-Resource Languages

Large language models (LLMs) have shown impressive zero-shot capabilities in various document rerank

Scaling Bilingual Lexicons for Zero-Shot Cross-Lingual Retrieval in Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Scaling of Artificially Code-Switched Data Effectiveness with Bilingual Lexicon Size for Low-Resource Cross-Lingual Retrieval

Transferring information retrieval (IR) models from a high-resource language (typically English) to

CONCRETE: Improving Cross-lingual Fact-checking with Cross-lingual Retrieval

Fact-checking has gained increasing attention due to the widespread of falsified information. Most f

Improving Cross-lingual Information Retrieval on Low-Resource Languages via Optimal Transport Distillation

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi