Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CeyNews: A Sri Lankan Trilingual News Corpus

Domain:

natural language processing

Record type:

dataset
Creator:
AjiSarRan
Editor:
KenTha
Publisher:
Zenodo
Host:avatar
CeyNews: A Trilingual Sri Lankan News Corpus This Zenodo release contains the CeyNews dataset, a large-scale trilingual news corpus for Sri Lanka covering Sinhala, Tamil, and English. The dataset includes over 1.02 million news articles collected from three popular Sri Lankan news outlets. The articles span the period from 2013 to 2026 and contain approximately 150 million tokens in total. Each record includes the news article text together with available metadata, including the news source, publication timestamp, headline, category, URL, and language. The dataset has been preprocessed to improve consistency and usability. Preprocessing includes deduplication and Unicode/script normalisation for Sinhala and Tamil text. The dataset is intended to support research in multilingual NLP, low-resource language processing, cross-lingual transfer learning, LLM adaptation, Sri Lankan and South Asian digital humanities, and comparative multilingual news analysis. The dataset can also be used to construct downstream NLP tasks such as news source identification, news category classification, and headline generation.

Visit

doi.orgzenodo.org

Tags

News CorpusLow-resource languagesSri Lankan Trilingual Corpus

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

LKValues: Aligning Large Language Models with Sri Lankan Societal ValuesAfrican News CorpusSwahili News CorpusAmharic News CorpusSAE Radio News Speech CorpusAfrikaans Radio News Speech Corpus

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Wester

African News Corpus

This consist of a monolingual news corpus for 19 languages from various sources like VOA, B

Swahili News Corpus

Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes

Amharic News Corpus

Amharic news text/category csv, stop words. Scraped/compiled from other sources.

SAE Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co

Afrikaans Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co