Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

CeyNews: A Sri Lankan Trilingual News Corpus

Domaine:

natural language processing

Type de record:

dataset
Créateur:
AjiSarRan
Éditeur:
KenTha
Éditeur:
Zenodo
Hôte:avatar
CeyNews: A Trilingual Sri Lankan News Corpus This Zenodo release contains the CeyNews dataset, a large-scale trilingual news corpus for Sri Lanka covering Sinhala, Tamil, and English. The dataset includes over 1.02 million news articles collected from three popular Sri Lankan news outlets. The articles span the period from 2013 to 2026 and contain approximately 150 million tokens in total. Each record includes the news article text together with available metadata, including the news source, publication timestamp, headline, category, URL, and language. The dataset has been preprocessed to improve consistency and usability. Preprocessing includes deduplication and Unicode/script normalisation for Sinhala and Tamil text. The dataset is intended to support research in multilingual NLP, low-resource language processing, cross-lingual transfer learning, LLM adaptation, Sri Lankan and South Asian digital humanities, and comparative multilingual news analysis. The dataset can also be used to construct downstream NLP tasks such as news source identification, news category classification, and headline generation.

Visit

doi.orgzenodo.org

Tags

News CorpusLow-resource languagesSri Lankan Trilingual Corpus

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

LKValues: Aligning Large Language Models with Sri Lankan Societal ValuesAfrican News CorpusSwahili News CorpusAmharic News CorpusSAE Radio News Speech CorpusAfrikaans Radio News Speech Corpus

LKValues: Aligning Large Language Models with Sri Lankan Societal Values

Value alignment of Large Language Models (LLMs) has been shown to be culturally biased toward Wester

African News Corpus

This consist of a monolingual news corpus for 19 languages from various sources like VOA, B

Swahili News Corpus

Language modeling, topic classification, AI training for Swahili NLP, digital literacy tools Notes

Amharic News Corpus

Amharic news text/category csv, stop words. Scraped/compiled from other sources.

SAE Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co

Afrikaans Radio News Speech Corpus

News bulletins purchased from the SABC. Data to be used for the development of a large vocabulary co