Logo Lanfrica

Sindhi Sentiment Analysis Dataset

Domain:

natural language processing

Record type:

dataset
Creator:
Ali
Host:
The Sindhi Sentiment Analysis Dataset is an open NLP resource containing 4,420 manually labeled sentences in the Sindhi language, annotated for three sentiment classes: positive, negative, and neutral. Every sentence has been hand-labeled to ensure linguistic accuracy and cultural relevance. Sindhi is a low-resource language spoken by over 30 million people in Pakistan and India, yet it remains severely underrepresented in NLP research and digital language tools. This dataset addresses that gap by providing a structured, high-quality text corpus for sentiment analysis and text classification tasks. The dataset is also available on Kaggle (kaggle.com) and HuggingFace (huggingface.co). It is intended to serve researchers, linguists, and developers working on Sindhi language technology and low-resource NLP.