Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MegaWika

Domain:

natural language processing

Record type:

dataset
Creator:
hlt
Host:
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided. Furthermore, nearly 130 million English question/answer pairs were extracted from the passages, and FrameNet events occurring in the passages are detected using the LOME FrameNet parser.

Visit

huggingface.co

Languages

AfrikaansXhosa

Licenses

cc-by-sa-4.0

Similar

MegaWika-Report-Generation

MegaWika-Report-Generation

MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with the