Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ Benchmark

Domain:

natural language processing

Record type:

dataset
Creator:
Ass
Publisher:
Zenodo
Host:avatar
We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from FAQ-style schema.org annotations. In total, the data collection consists of 96 million natural question-answer (QA) pairs across 75 languages, including 47 million (49\%) non-English samples. WebFAQ further serves as the foundation for 20 monolingual retrieval benchmarks with a total size of 11.2 million QA pairs (5.9 million non-English). These datasets are carefully curated through refined filtering and near-duplicate detection, yielding high-quality resources for training and evaluating multil Research goal: Does increasing the scale of multilingual pretraining data improve robustness against domain shift in low-resource language retrieval tasks within the WebFAQ benchmark? Autonomous synthesis report generated by Assignee Research. Tribunal consensus score: 7.7/10. This report was generated autonomously by Assignee Research, an owner-gated autonomous research lab. The content synthesizes findings from peer-reviewed papers. Tribunal score: 7.7/10.

Visit

doi.orgzenodo.org

Tasks

information retrievalquestion answering

Tags

increasingscalemultilingualpretrainingdataimproverobustnessagainst

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode