sinhala_stories is a crowdsourced corpus of Sinhala-language narrative text comprising 10,949,004 rows (~1.95 GB) in Apache Parquet format, collected through a public Streamlit submission portal integrated with the Hugging Face Hub. Each submission was validated for Sinhala script content, length, language-identification confidence, and heuristic spam/duplicate checks before being merged into the corpus via an automated CI/CD pipeline.
Content warning: a substantial portion of the corpus consists of explicit adult narrative content, submitted with minimal moderation beyond character-level validation. Anyone training generative models on this data should apply content filtering and safety alignment before deployment. See the accompanying data note for full ethical considerations.
Dataset:
huggingface.co app source:
github.com Corpus contains a significant proportion of explicit adult content; no automated or human content moderation was applied beyond character-level validation at submission time.