Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Do language families matter? Evaluating LLMs for sentiment analysis through a hierarchical cross-lingual lens

Domain:

natural language processing

Record type:

paper
Creator:
MuhAbdAliZen
Publisher:
Fro
Host:
Social media sentiment analysis has become one of the most significant instruments for understanding the opinion of the population in the spheres of healthcare, politics, and education. Yet, large language models (LLMs) remain unevenly distributed in their linguistic coverage, failing to adequately serve a large portion of the world's languages. This study evaluates five state-of-the-art LLMs: GPT-4o, Gemini 2.0 Flash, DeepSeek-V3, Mistral Large, and Claude 3.7 Sonnet on three-class sentiment classification across 36 datasets spanning 36 languages, with emphasis on low- and medium-resource settings, using zero-shot and few-shot prompting without task-specific fine-tuning. In addition to the traditional measures of performance per language, the study presents a hierarchical analysis of languages based on a genealogical tree of Indo-European, Afro-Asiatic, Niger-Congo, Turkic, Austronesian, and English Creole language families, so that it is possible to identify the systematic patterns of performance superiority and inferiority among the language families. The findings show that few-shot prompting improves the results of a vast majority of languages, with several models approaching or surpassing the performance of the state-of-the-art benchmark of task-specific models. The GPT-4o and Claude achieved the highest performance in the high-resource and medium-resource settings, and Gemini is a competent trade-off that allows balancing the performance and the computational cost. Although it has lower zero-shot performance, Mistral benefits the most from few-shot prompting and becomes highly competitive in the few-shot setting. Despite these developments, the level of performance on low-resource languages, such as Oromo, Xitsonga, Azerbaijani, and Twi, remains significantly lower, underscoring that progress in multilingual LLMs requires moving beyond English-centric evaluation toward genuinely representative and globally inclusive benchmarks.

Visit

doi.org

Tasks

sentiment analysistext classification

Languages

OromoOromo, Borana-Arsi-GujiTsonga

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence ModelsAbelAdissu/Cross-Lingual-Question-Answering-for-Amharic-Language-Using-Pretrained-LLMs- Cross-Lingual and Low-Resource Sentiment AnalysisCross-Lingual Auto Evaluation for Assessing Multilingual LLMsDeep Persian sentiment analysis: Cross-lingual training for low-resource languagesnossamchakri05/Cross-Lingual-Transfer-Learning-Based-Sentiment-Analysis-for-Low-Resource-Languages

Advancing Cross-lingual Aspect-Based Sentiment Analysis with LLMs and Constrained Decoding for Sequence-to-Sequence Models

Aspect-based sentiment analysis (ABSA) has made significant strides, yet challenges remain for low-r

AbelAdissu/Cross-Lingual-Question-Answering-for-Amharic-Language-Using-Pretrained-LLMs-

## **INTRODUCTION** 📖 Welcome to the Amharic Text Generation project, a journey into the realm of na

Cross-Lingual and Low-Resource Sentiment Analysis

Identifying sentiment in a low-resource language is essential for understanding opinions internation

Cross-Lingual Auto Evaluation for Assessing Multilingual LLMs

Evaluating machine-generated text remains a significant challenge in NLP, especially for non-English

Deep Persian sentiment analysis: Cross-lingual training for low-resource languages

With the advent of deep neural models in natural language processing tasks, having a large amount of

nossamchakri05/Cross-Lingual-Transfer-Learning-Based-Sentiment-Analysis-for-Low-Resource-Languages

# Cross-Lingual Transfer Learning-Based Sentiment Analysis for Low-Resource Languages ## 📋 Overview