Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

An Annotated Corpus of Uzbek Business Reviews for Aspect-Based Sentiment Analysis

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Mat
Éditeur:
Zenodo
Hôte:avatar
This dataset contains 5,038 annotated business reviews designed for Aspect-Based Sentiment Analysis (ABSA). The reviews were scraped from Commeta Sharh, a publicly accessible business review platform in Uzbekistan. The dataset captures the natural linguistic diversity of the region, featuring mixed-language text (including Russian and Uzbek) alongside Uzbek-language metadata categories. The corpus spans 630 unique businesses across 23 domains (e.g., Education/Ta'lim). It serves as a valuable resource for evaluating low-resource and code-switched NLP models, specifically for extracting specific business aspects and their associated sentiment polarities. Dataset Characteristics Total Reviews: 5,038 (filtered from a larger pool, excluding entries with fewer than five words). Businesses Covered: 630 Business Domains: 23 Task: Aspect-Based Sentiment Analysis (ABSA) – Aspect Term Extraction (ATE) and Aspect Polarity Classification (APC). Data Structure The dataset is provided in JSON format. Each entry represents a single user review and contains the following fields: review_id: A unique identifier for the review. text: The raw text of the user review. business_name: The name of the reviewed business. business_category: The domain/industry of the business (e.g., "Ta'lim" for Education). user_rating: The numerical rating given by the user (typically 1-5). aspects: A list of extracted aspects, where each aspect contains: term: The specific word or phrase from the text representing the aspect. category: The broader category of the aspect (e.g., "xizmat" for service, "boshqalar" for others). polarity: The sentiment expressed toward the aspect (positive, negative, or neutral). num_aspects: The total count of aspects identified in the text. annotation_source: The model used for the automated annotation pipeline (e.g., qwen2.5-7b-finetuned). parse_success: A boolean indicating if the model output was successfully parsed into the JSON structure. raw_output: The raw JSON string generated by the fine-tuned LLM before parsing.

Visit

doi.orgzenodo.org

Tasks

sentiment analysistext classification

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Human Annotated Arabic Dataset of Book Reviews for Aspect Based Sentiment AnalysisAn Uzbek Medical-Domain Dataset for Aspect-Based Sentiment AnalysisAspect-Based Sentiment Analysis for Afaan Oromoo Movie Reviews Using Machine Learning TechniquesHauBERT: A Transformer Model for Aspect-Based Sentiment Analysis of Hausa-Language Movie ReviewsAn Annotated Huge Dataset for Standard and Colloquial Arabic Reviews for Subjective Sentiment AnalysisAspect-Based Sentiment Analysis of Arabic Restaurants Customers' Reviews Using a Hybrid Approach

Human Annotated Arabic Dataset of Book Reviews for Aspect Based Sentiment Analysis

An Uzbek Medical-Domain Dataset for Aspect-Based Sentiment Analysis

UzMedSentiment is a manually annotated Uzbek medical-domain dataset designed for sentiment classific

Aspect-Based Sentiment Analysis for Afaan Oromoo Movie Reviews Using Machine Learning Techniques

Aspect-based sentiment analysis (ABSA) is the subfield of natural language processing that deals wit

HauBERT: A Transformer Model for Aspect-Based Sentiment Analysis of Hausa-Language Movie Reviews

An Annotated Huge Dataset for Standard and Colloquial Arabic Reviews for Subjective Sentiment Analysis

Aspect-Based Sentiment Analysis of Arabic Restaurants Customers' Reviews Using a Hybrid Approach