Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Beyond Metrics: Evaluating LLMs' Effectiveness in Culturally Nuanced, Low-Resource Real-World Scenarios

Domain:

natural language processing

Record type:

paper
Creator:
OchGumSitWan
Host:avatar
The deployment of Large Language Models (LLMs) in real-world applications presents both opportunities and challenges, particularly in multilingual and code-mixed communication settings. This research evaluates the performance of seven leading LLMs in sentiment analysis on a dataset derived from multilingual and code-mixed WhatsApp chats, including Swahili, English and Sheng. Our evaluation includes both quantitative analysis using metrics like F1 score and qualitative assessment of LLMs' explanations for their predictions. We find that, while Mistral-7b and Mixtral-8x7b achieved high F1 scores, they and other LLMs such as GPT-3.5-Turbo, Llama-2-70b, and Gemma-7b struggled with understanding linguistic and contextual nuances, as well as lack of transparency in their decision-making process as observed from their explanations. In contrast, GPT-4 and GPT-4-Turbo excelled in grasping diverse linguistic inputs and managing various contextual information, demonstrating high consistency with human alignment and transparency in their decision-making process. The LLMs however, encountered difficulties in incorporating cultural nuance especially in non-English settings with GPT-4s doing so inconsistently. The findings emphasize the necessity of continuous improvement of LLMs to effectively tackle the challenges of culturally nuanced, low-resource real-world settings and the need for developing evaluation benchmarks for capturing these issues.

Visit

arxiv.org

Tasks

code switchingsentiment analysistext classification

Languages

Swahili

Tags

Computation and Language

Similar

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced ContextsLLM Probe: Evaluating LLMs for Low-Resource LanguagesCulturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and SundaneseLeveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languagesBeyond English: Evaluating LLMs for Arabic Grammatical Error CorrectionEvaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Reasoning Beyond Labels: Measuring LLM Sentiment in Low-Resource, Culturally Nuanced Contexts

Sentiment analysis in low-resource, culturally nuanced contexts challenges conventional NLP approach

LLM Probe: Evaluating LLMs for Low-Resource Languages

Despite rapid advances in large language models (LLMs), their linguistic abilities in low-resource a

Culturally-Nuanced Story Generation for Reasoning in Low-Resource Languages: The Case of Javanese and Sundanese

Culturally grounded commonsense reasoning is underexplored in low-resource languages due to scarce d

Leveraging LLMs for MT in Crisis Scenarios: a blueprint for low-resource languages

In an evolving landscape of crisis communication, the need for robust and adaptable Machine Translat

Beyond English: Evaluating LLMs for Arabic Grammatical Error Correction

Large language models (LLMs) finetuned to follow human instruction have recently exhibited significa

Evaluating Multimodal LLMs for Inpatient Diagnosis: Real-World Performance, Safety, and Cost Across Ten Frontier Models

Background: Large language models (LLMs) are increasingly proposed for diagnostic support, but few e