Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Impact of Multilingual Dense Retriever Model Scaling on Low-Resource WebFAQ 2.0 Performance

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ass
Éditeur:
Zenodo
Hôte:avatar
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49\%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multil Research goal: What is the impact of scaling the multilingual dense retriever model size (e.g., small vs. large) on retrieval performance across low-resource languages in WebFAQ 2.0, as evaluated by recall@k and precision@k on the XGLUE retrieval benchmark? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.5/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.5/10.

Visit

doi.orgzenodo.org

Tasks

information retrievalquestion answering

Tags

impactscalingmultilingualdenseretrievermodelsizesmall

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Scaling WebFAQ 2.0 Dataset Size and Its Impact on MTEB Retrievers for Low-Resource LanguagesMultilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ BenchmarkPerformance comparison of dense retrieval models trained on WebFAQ versus Wikipedia-based datasets for low-resource languagePerformance of Multilingual Dense Retrieval Models on WebFAQ Benchmarks with Cross-Lingual Contrastive LearningImpact of WebFAQ Fine-Tuning on Cross-Lingual NLI Performance in Low-Resource LanguagesDiminishing Returns of Scaling Bilingual QA Pairs in WebFAQ 2.0 for Cross-Lingual Alignment in Low-Resource Languages

Scaling WebFAQ 2.0 Dataset Size and Its Impact on MTEB Retrievers for Low-Resource Languages

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Multilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ Benchmark

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Performance comparison of dense retrieval models trained on WebFAQ versus Wikipedia-based datasets for low-resource language

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Performance of Multilingual Dense Retrieval Models on WebFAQ Benchmarks with Cross-Lingual Contrastive Learning

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Impact of WebFAQ Fine-Tuning on Cross-Lingual NLI Performance in Low-Resource Languages

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Diminishing Returns of Scaling Bilingual QA Pairs in WebFAQ 2.0 for Cross-Lingual Alignment in Low-Resource Languages

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from