Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Scaling WebFAQ 2.0 Dataset Size and Its Impact on MTEB Retrievers for Low-Resource Languages

Domain:

natural language processing

Record type:

dataset
Creator:
SOV
Publisher:
Zenodo
Host:avatar
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49\%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multil Research goal: How does the scaling of WebFAQ 2.0's dataset size (198M vs. smaller subsets) influence the trade-off between MTEB retrieval scores and inference efficiency of dense retrievers in low-resource languages? Autonomous synthesis report generated by SOVEREIGN Research Kernel. Tribunal consensus score: 7.7/10. This report was generated autonomously by SOVEREIGN Research Kernel, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.7/10.

Visit

doi.orgzenodo.org

Tasks

information retrievalquestion answering

Tags

scalingWebFAQdatasetsizesmallersubsetsinfluencetrade-off

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Impact of Multilingual Dense Retriever Model Scaling on Low-Resource WebFAQ 2.0 PerformanceDiminishing Returns of Scaling Bilingual QA Pairs in WebFAQ 2.0 for Cross-Lingual Alignment in Low-Resource LanguagesFine-Tuning Dense Retrievers on WebFAQ Low-Resource Subsets for Zero-Shot Cross-Lingual XQuAD AccuracyScaling Intermediate-Task Dataset Size for Zero-Shot Cross-Lingual Transfer in Low-Resource LanguagesImpact of WebFAQ Fine-Tuning on Cross-Lingual NLI Performance in Low-Resource LanguagesMultilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ Benchmark

Impact of Multilingual Dense Retriever Model Scaling on Low-Resource WebFAQ 2.0 Performance

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Diminishing Returns of Scaling Bilingual QA Pairs in WebFAQ 2.0 for Cross-Lingual Alignment in Low-Resource Languages

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Fine-Tuning Dense Retrievers on WebFAQ Low-Resource Subsets for Zero-Shot Cross-Lingual XQuAD Accuracy

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Scaling Intermediate-Task Dataset Size for Zero-Shot Cross-Lingual Transfer in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Impact of WebFAQ Fine-Tuning on Cross-Lingual NLI Performance in Low-Resource Languages

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Multilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ Benchmark

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from