Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MegaWika-Report-Generation

Domain:

natural language processing

Record type:

dataset
Creator:
hlt
Host:
MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with their scraped and cleaned web citations. The passages span 50 Wikipedias in 50 languages, and the articles in which the passages were originally embedded are included for convenience. Where a Wikipedia passage is in a non-English language, an automated English translation is provided.

Visit

huggingface.co

Languages

AfrikaansXhosa

Licenses

cc-by-sa-4.0

Similar

MegaWikaRole-Adapted Clinical Report Generation for Ultrasound Measurements in Low-Resource SettingsAutomatic spike train analysis and report generation. An implementation with R, R2HTML and STAR.

MegaWika

MegaWika is a multi- and crosslingual text dataset containing 30 million Wikipedia passages with the

Role-Adapted Clinical Report Generation for Ultrasound Measurements in Low-Resource Settings

International audience Obstetric ultrasound is critical for monitoring fetal growth,

Automatic spike train analysis and report generation. An implementation with R, R2HTML and STAR.

International audience Multi-electrode arrays (MEA) allow experimentalists to record