Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Multilingual Open Text Release 1: Public Domain News in 44 Languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
PalKimLig
Host:avatar
We present Multilingual Open Text (MOT), a new multilingual corpus containing text in 44 languages, many of which have limited existing text resources for natural language processing. The first release of the corpus contains over 2.8 million news articles and an additional 1 million short snippets (photo captions, video descriptions, etc.) published between 2001--2022 and collected from Voice of America's news websites. We describe our process for collecting, filtering, and processing the data. The source material is in the public domain, our collection is licensed using a creative commons license (CC BY 4.0), and all software used to create the corpus is released under the MIT License. The corpus will be regularly updated as additional documents are published. Submitted to LREC 2022

Visit

arxiv.org

Tags

Computation and Language

Similar

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 LanguagesMassively Multilingual Text Translation For Low-Resource LanguagesMultilingual Parallel Text Corpora for East African LanguagesTaxi1500: A Dataset for Multilingual Text Classification in 1500 LanguagesMassively Multilingual ASR: 50 Languages, 1 Model, 1 Billion ParametersOmnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

XL-Sum: Large-Scale Multilingual Abstractive Summarization for 44 Languages

Contemporary works on abstractive text summarization have focused primarily on highresource languages like English, mostly due to the limited availability of datasets for low/midresource ones. In this work, we present XLSum, a comprehensive and diverse dataset comp

Massively Multilingual Text Translation For Low-Resource Languages

Translation into severely low-resource languages has both the cultural goal of saving and reviving t

Multilingual Parallel Text Corpora for East African Languages

This is a partial multilingual parallel corpora of 5 East African languages. The dataset contains an

Taxi1500: A Dataset for Multilingual Text Classification in 1500 Languages

While broad-coverage multilingual natural language processing tools have been developed, a significa

Massively Multilingual ASR: 50 Languages, 1 Model, 1 Billion Parameters

We study training a single acoustic model for multiple languages with the aim of improving automatic

Omnilingual ASR: Open-Source Multilingual Speech Recognition for 1600+ Languages

Automatic speech recognition (ASR) has advanced in high-resource languages, but most of the world's