This dataset contains news articles, stories, and cultural reports scraped from Global Voices Malagasy. It is intended to support Natural Language Processing (NLP) tasks for the Malagasy language, such as language modeling and text generation.
The data is automatically scraped and updated once a month to ensure freshness.
The dataset is formatted in JSONL (JSON Lines). Each entry represents a single article.