Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Pretraining Dense Retrieval Models on WebFAQ for Zero-Shot Cross-Lingual Recall in Low-Resource XTREME Subsets

Domaine:

natural language processing

Type de record:

dataset
Créateur:
SOV
Éditeur:
Zenodo
Hôte:avatar
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49\%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multil Research goal: Does pretraining dense retrieval models on WebFAQ's 47 million non-English samples improve zero-shot cross-lingual recall on low-resource subsets of XTREME compared to English-only baselines? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 9.0/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 9.0/10.

Visit

doi.orgzenodo.org

Tasks

information retrieval

Tags

pretrainingdenseretrievalmodelsWebFAQmillionnon-Englishsamples

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Fine-Tuning Dense Retrievers on WebFAQ Low-Resource Subsets for Zero-Shot Cross-Lingual XQuAD AccuracyCross-lingual Zero-shot Retrieval Accuracy of WebFAQ vs. TyDi QA Models on Low-resource XQuAD SubsetsImpact of Non-English WebFAQ Pretraining on Zero-Shot Cross-Lingual Retrieval AccuracySynthetic Query Augmentation for Zero-Shot Cross-Lingual Dense Retrieval in Low-Resource LanguagesEffectiveness of Zero-Shot Cross-Lingual Retrieval Models on Low-Resource LanguagesComparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on Low-Resource XQuAD Subsets

Fine-Tuning Dense Retrievers on WebFAQ Low-Resource Subsets for Zero-Shot Cross-Lingual XQuAD Accuracy

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Cross-lingual Zero-shot Retrieval Accuracy of WebFAQ vs. TyDi QA Models on Low-resource XQuAD Subsets

Benefiting from transformer-based pre-trained language models, neural ranking models have made signi

Impact of Non-English WebFAQ Pretraining on Zero-Shot Cross-Lingual Retrieval Accuracy

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Synthetic Query Augmentation for Zero-Shot Cross-Lingual Dense Retrieval in Low-Resource Languages

Multilingual Pretrained Language Models (MPLMs) perform strongly in cross-lingual transfer. We propo

Effectiveness of Zero-Shot Cross-Lingual Retrieval Models on Low-Resource Languages

Transferring information retrieval (IR) models from a high-resource language (typically English) to

Comparative Analysis of Hybrid Batch Training for Zero-Shot Cross-Lingual Retrieval on Low-Resource XQuAD Subsets

Information retrieval across different languages is an increasingly important challenge in natural l