Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

ChattogramSent: A Comprehensive Multilingual Sentiment Dataset for Chattogram Dialect, Bangla, and English

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MouMomDr.Md.
Éditeur:
Zenodo
Hôte:avatar
This dataset, titled ChattogramSent, is a novel and large-scale multilingual resource developed specifically for sentiment classification in the Chattogram regional language (a low-resource language spoken by approximately 13–16 million people). The dataset follows a parallel structure across three languages: Chattogram dialect, standard Bangla, and English. Dataset Specifications: Total Instances: 7,053 unique entries. Languages: Chattogram (Regional), Bangla (Standard), and English (Global). Sentiment Classes: Balanced across Positive, Negative, and Neutral categories. Data Sources: Scraped from social media (Facebook, Twitter/X), transcripts of regional dramas (Natoks), and public comments on news portals. Validation: All regional translations and sentiment labels have been manually verified by native speakers to ensure linguistic accuracy. This dataset is designed to support research in Natural Language Processing (NLP) for low-resource languages, machine translation, and benchmarking multilingual transformer models.

Visit

doi.orgzenodo.org

Tasks

machine translationsentiment analysistext classification

Languages

Samba Leko

Tags

Chattogram Language, Sentiment Analysis, Multilingual Dataset, Low-resource NLP, Bangla NLP, Dataset for Machine Learning.

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

TroubleQA-BE: A Bilingual Troubleshooting Question Answering Dataset for Bangla and EnglishSentiment Analysis of Multilingual Dataset of Bahraini Dialects, Arabic, and EnglishAdvancements in Sentiment Analysis for the Algerian Dialect: A Comprehensive ReviewBanglaLekha-Isolated: A multi-purpose comprehensive dataset of Handwritten Bangla Isolated charactersSentiment dataset of Algerian dialectFine-tuning Multilingual Transformers for Hausa-English Sentiment Analysis

TroubleQA-BE: A Bilingual Troubleshooting Question Answering Dataset for Bangla and English

BTQA is a bilingual troubleshooting question answering dataset created to facilitate research and de

Sentiment Analysis of Multilingual Dataset of Bahraini Dialects, Arabic, and English

This dataset was generated using two cascading stages of translation—a machine translation followed

Advancements in Sentiment Analysis for the Algerian Dialect: A Comprehensive Review

BanglaLekha-Isolated: A multi-purpose comprehensive dataset of Handwritten Bangla Isolated characters

Sentiment dataset of Algerian dialect

* This sentiment dataset of Algerian dialect consists of 11760 comments (6111 positive/ 5649 negativ

Fine-tuning Multilingual Transformers for Hausa-English Sentiment Analysis