This dataset, titled ChattogramSent, is a novel and large-scale multilingual resource developed specifically for sentiment classification in the Chattogram regional language (a low-resource language spoken by approximately 13–16 million people). The dataset follows a parallel structure across three languages: Chattogram dialect, standard Bangla, and English.
Dataset Specifications:
Total Instances: 7,053 unique entries.
Languages: Chattogram (Regional), Bangla (Standard), and English (Global).
Sentiment Classes: Balanced across Positive, Negative, and Neutral categories.
Data Sources: Scraped from social media (Facebook, Twitter/X), transcripts of regional dramas (Natoks), and public comments on news portals.
Validation: All regional translations and sentiment labels have been manually verified by native speakers to ensure linguistic accuracy.
This dataset is designed to support research in Natural Language Processing (NLP) for low-resource languages, machine translation, and benchmarking multilingual transformer models.