Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

A 10-Million-Row Sinhala Narrative Corpus for Low-Resource NLP: Dataset Construction, Statistical Characterisation, and Ethical Considerations

Domaine:

natural language processing

Type de record:

dataset
Créateur:
SamEka
Éditeur:
Zenodo
Hôte:avatar
sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 rows (~1.95 GB) in Apache Parquet format, collected through a public Streamlit submission portal integrated with the Hugging Face Hub. Each submission was validated for Sinhala script content, length, language-identification confidence, and heuristic spam/duplicate checks before being merged into the corpus via an automated CI/CD pipeline. Content warning: a substantial portion of the corpus consists of explicit adult narrative content, submitted with minimal moderation beyond character-level validation. Anyone training generative models on this data should apply content filtering and safety alignment before deployment. See the accompanying data note for full ethical considerations. Dataset: huggingface.co app source: github.com Corpus contains a significant proportion of explicit adult content; no automated or human content moderation was applied beyond character-level validation at submission time.

Visit

doi.org

Tasks

language modeling

Tags

Sinhala NLPlow-resource languagetext corpusnarrative datasetsmall language modelstokeniser fertilitydata papercrowdsourced dataset

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia AssistanceA Sheng Phishing Corpus for Low-Resource Cybersecurity NLPUnveiling Swahili Verb Conjugations: A Comprehensive Dataset for Low-Resource NLPAkan POS Tagging Dataset - 10 Million SentencesPashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource LanguageRIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

A Low-Resource Speech-Driven NLP Pipeline for Sinhala Dyslexia Assistance

Dyslexia in adults remains an under-researched and under-served area, particularly in non-English-sp

A Sheng Phishing Corpus for Low-Resource Cybersecurity NLP

 

This dataset, the Sheng-English

Unveiling Swahili Verb Conjugations: A Comprehensive Dataset for Low-Resource NLP

Akan POS Tagging Dataset - 10 Million Sentences

The dataset includes sentences and part-of-speech tags and is aimed at supporting the development of

Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language

We present the Pashto Common Voice corpus -- the first large-scale, openly licensed speech resource

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples develop