Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

MassiveSumm: a very large-scale, very multilingual, news summarisation dataset

Domain:

natural language processing

Record type:

paper
Current research in automatic summarisation is unapologetically anglo-centered–a persistent state-of-affairs, which also predates neural net approaches. High-quality automatic summarisation datasets are notoriously expensive to create, posing a challenge for any language. However, with digitalisation, archiving, and social media advertising of newswire articles, recent work has shown how, with careful methodology application, large-scale datasets can now be simply gathered instead of written. In this paper, we present a large-scale multilingual summarisation dataset containing articles in 92 languages, spread across 28.8 million articles, in more than 35 writing scripts. This is both the largest, most inclusive, existing automatic summarisation dataset, as well as one of the largest, most inclusive, ever published datasets for any NLP task. We present the first investigation on the efficacy of resource building from news platforms in the low-resource language setting. Finally, we provide some first insight on how low-resource language settings impact state-of-the-art automatic summarisation system performance.

Visit

aclanthology.orgopenreview.net

Connected records

dataset

Tasks

summarizationnatural language generation

Languages

AfrikaansAmharicBamanankanDanFulaHausaIgboJiruKaanKinyarwanda+16

Tags

MassiveSumm

Similar

COLD-CI: A large-scale very high-resolution label polygon dataset for cocoa and non-cocoa classification in Cote d'IvoireCOLD-CI: A large-scale very high-resolution label polygon dataset for cocoa and non-cocoa classification in Côte d'IvoirePest24: A large-scale very small object data set of agricultural pests for multi-target detectionA hybrid approach to very small scale electrical demand forecastingA Statistical Model for Large and Very Large Hail: Development, Global Climate Applications and Use in ForecastingDevelopment of a clinical feeding assessment scale for very young infants in South Africa

COLD-CI: A large-scale very high-resolution label polygon dataset for cocoa and non-cocoa classification in Cote d'Ivoire

Spatially explicit information on cocoa cultivation is essential for land-use planning, deforestatio

COLD-CI: A large-scale very high-resolution label polygon dataset for cocoa and non-cocoa classification in Côte d'Ivoire

COLD-CI consists of 123,736 vector polygons corresponding to a total labelled area of 5,996 km², inc

Pest24: A large-scale very small object data set of agricultural pests for multi-target detection

A hybrid approach to very small scale electrical demand forecasting

Microgrid management and scheduling can considerably benefit from day-ahead demand forecasting. Unti

A Statistical Model for Large and Very Large Hail: Development, Global Climate Applications and Use in Forecasting

We have developed additive regression convective hazard models (AR-CHaMo) for predicting the occurre

Development of a clinical feeding assessment scale for very young infants in South Africa

Background: There is a need for validated neonatal feeding assessment instruments in South Africa. A