Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Preparing Bengali-English Code-Mixed Corpus for Sentiment Analysis of Indian Languages

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ManMahDas
Hôte:avatar
Analysis of informative contents and sentiments of social users has been attempted quite intensively in the recent past. Most of the systems are usable only for monolingual data and fails or gives poor results when used on data with code-mixing property. To gather attention and encourage researchers to work on this crisis, we prepared gold standard Bengali-English code-mixed data with language and polarity tag for sentiment analysis purposes. In this paper, we discuss the systems we prepared to collect and filter raw Twitter data. In order to reduce manual work while annotation, hybrid systems combining rule based and supervised models were developed for both language and sentiment tagging. The final corpus was annotated by a group of annotators following a few guidelines. The gold standard corpus thus obtained has impressive inter-annotator agreement obtained in terms of Kappa values. Various metrics like Code-Mixed Index (CMI), Code-Mixed Factor (CF) along with various aspects (language and emotion) also qualitatively polled the code-mixed and sentiment properties of the corpus. The 13th Workshop on Asian Language Resources (ALR), collocated with LREC 2018

Visit

arxiv.org

Tasks

code switchingsentiment analysistext classification

Tags

Computation and Language

Similaires

Sentiment Analysis for Amharic-English Code-Mixed Sociopolitical Posts Using Deep LearningDeveloping a Code-Mixed Sentiment Analysis Dataset of Xitsonga-English Music ReviewsA Corpus of English-Hindi Code-Mixed Tweets for Sarcasm DetectionBengali Political Sentiment Analysis DatasetBengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI SystemsA Performance-efficiency Analysis of Transformer Models for Code-mixed Hausa Sentiment Data

Sentiment Analysis for Amharic-English Code-Mixed Sociopolitical Posts Using Deep Learning

Abstract Sentiment analysis is crucial in natural language processing for identifying emot

Developing a Code-Mixed Sentiment Analysis Dataset of Xitsonga-English Music Reviews

A Corpus of English-Hindi Code-Mixed Tweets for Sarcasm Detection

Social media platforms like twitter and facebook have be- come two of the largest mediums used by pe

Bengali Political Sentiment Analysis Dataset

This dataset comprises 3,290 Bengali political comments sourced from social media platforms, news co

BengVoice: A Stratified Dataset of Code-Mixed Bengali-English Voice Commands for Intent Classification in Conversational AI Systems

This dataset presents a meticulously curated benchmark collection of 1,200 Bengali voice assistant u

A Performance-efficiency Analysis of Transformer Models for Code-mixed Hausa Sentiment Data

Abstract The deployment of large language models (LLMs) in African contexts offers