Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Performance of Multilingual Dense Retrieval Models on WebFAQ Benchmarks with Cross-Lingual Contrastive Learning

Domain:

natural language processing

Record type:

dataset
Creator:
SOV
Publisher:
Zenodo
Host:avatar
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49\%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multil Research goal: How do state-of-the-art multilingual dense retrieval models compare to monolingual models on WebFAQ benchmarks when fine-tuned with cross-lingual contrastive learning, as evaluated by performance on zero-shot and few-shot retrieval tasks in low-resource languages? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 8.1/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 8.1/10.

Visit

doi.orgzenodo.org

Tasks

information retrievalquestion answering

Tags

state-of-the-artmultilingualdenseretrievalmodelsmonolingualWebFAQbenchmarks

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Multilingual Contrastive Learning for Zero-Shot Cross-Lingual Retrieval on XNLIScaling Performance of Zero-Shot Cross-Lingual Retrieval Models with Model Size on Adversarial BenchmarksPretraining Dense Retrieval Models on WebFAQ for Zero-Shot Cross-Lingual Recall in Low-Resource XTREME SubsetsUnsupervised Dense Information Retrieval with Contrastive LearningPerformance Comparison of Cross-Lingual Retrieval Models via Optimal Transport Distillation and Contrastive Learning on XTR-TRECDegradation of Cross-Lingual Dense Retrieval MRR on Low-Resource WebFAQ Language Families Versus Monolingual Baselines

Multilingual Contrastive Learning for Zero-Shot Cross-Lingual Retrieval on XNLI

Information retrieval across different languages is an increasingly important challenge in natural l

Scaling Performance of Zero-Shot Cross-Lingual Retrieval Models with Model Size on Adversarial Benchmarks

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Pretraining Dense Retrieval Models on WebFAQ for Zero-Shot Cross-Lingual Recall in Low-Resource XTREME Subsets

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Unsupervised Dense Information Retrieval with Contrastive Learning

Recently, information retrieval has seen the emergence of dense retrievers, using neural networks, a

Performance Comparison of Cross-Lingual Retrieval Models via Optimal Transport Distillation and Contrastive Learning on XTR-TREC

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Degradation of Cross-Lingual Dense Retrieval MRR on Low-Resource WebFAQ Language Families Versus Monolingual Baselines

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from