CeyNews: A Trilingual Sri Lankan News Corpus
This Zenodo release contains the CeyNews dataset, a large-scale trilingual news corpus for Sri Lanka covering Sinhala, Tamil, and English.
The dataset includes over 1.02 million news articles collected from three popular Sri Lankan news outlets. The articles span the period from 2013 to 2026 and contain approximately 150 million tokens in total.
Each record includes the news article text together with available metadata, including the news source, publication timestamp, headline, category, URL, and language.
The dataset has been preprocessed to improve consistency and usability. Preprocessing includes deduplication and Unicode/script normalisation for Sinhala and Tamil text.
The dataset is intended to support research in multilingual NLP, low-resource language processing, cross-lingual transfer learning, LLM adaptation, Sri Lankan and South Asian digital humanities, and comparative multilingual news analysis.
The dataset can also be used to construct downstream NLP tasks such as news source identification, news category classification, and headline generation.