This dataset consists of clean, structured sentences extracted via Optical Character Recognition (OCR) from approximately 1GB of Malagasy thesis documents. These documents were collected based on educational, cultural, and linguistic themes.
The dataset is saved in CSV format, and is particularly useful for NLP tasks involving sentence-level modeling in Malagasy — a low-resource language.
Language: Malagasy