Logo Lanfrica
en
Home
Atlas
Insights
Docs
Sign in
Back
MaLA-LM/MassiveSumm_long
Domain:
natural language processing
Record type:
dataset
Creator:
MaL
Host:
Visit
Actions
Share
Report an issue
This dataset is a subset of MassiveSumm by subsampling and setting a maximum sequence length of 7.5k tokens. Links to reproduce the whole set of MassiveSumm via Common Crawl and the Wayback Machine are provided in the repository of MassiveSumm.
Visit
huggingface.co
Tasks
natural language generation
summarization
Languages
Afrikaans
Amharic
Bamanankan
Fula
Hausa
Igbo
Kinyarwanda
Lingala
Malagasy
Ndebele
+8
View more
Similar
MaLA-LM/MassiveSumm_short
MaLA-LM/MassiveSumm_short
This dataset is a subset of MassiveSumm by subsampling and setting a maximum sequence length of 1.5k