Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

How Good Are Large Language Models at Arithmetic Reasoning in Low-Resource Language Settings?—A Study on Yorùbá Numerical Probes with Minimal Contamination

Domaine:

natural language processing

Type de record:

dataset
Créateur:
FiyPet
Éditeur:
MDP
Hôte:
We study the performance of large language models (LLMs) in natural language understanding and natural language reasoning tasks in a low-resourced-language (LRL) setting. Using Yorùbá, an LRL, we curated a set of numerical probes with minimal contamination. The probes comprise three sets of questions—the first covers basic arithmetic, the second covers date and time (calendar system), and the last focuses on numerals and counting systems. Assessed in a zero-shot setup, three LLMs (ChatGPT, Gemini, and PaLM) were evaluated based on several metrics. The best-performing model, ChatGPT, generated some correct answers, showing logical steps in attaining the answers in Yorùbá (with an accuracy of 56% in set one, and 44% in set two). The second-best model (with an accuracy of 56% in set one, and 32% in set two) is Gemini. PaLM (with an accuracy of 16% in set one, and 8% in set two) showed the answers without logic. The three models performed poorly on the Yorùbá numerals question set (ChatGPT scored 8%, and Gemini and PaLM each had 0% accuracy). The study also revealed that there is significant room for improvement in the state of the art of LLMs when it comes to Yorùbá numerals.

Visit

doi.org

Languages

Yoruba

Licenses

https://creativecommons.org/licenses/by/4.0/

Similaires

How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings – A Benchmarking Study & DatasetHow Good are Commercial Large Language Models on African Languages?AfroBench: How Good are Large Language Models on African Languages?Yorùbá Arithmetic Reasoning DatasetLarge language models for frontline healthcare support in low-resource settingsIn-Context Learning for Low-Resource Machine Translation: A Study on Tarifit with Large Language Models

How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings – A Benchmarking Study & Dataset

Abstract Large language models (LLMs) have demonstrated strong performance in medi

How Good are Commercial Large Language Models on African Languages?

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretr

AfroBench: How Good are Large Language Models on African Languages?

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African language

Yorùbá Arithmetic Reasoning Dataset

#YORUBA ARITHMETIC REASONING DATASET This dataset consists of arithmetic and numerical reasoning que

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str

In-Context Learning for Low-Resource Machine Translation: A Study on Tarifit with Large Language Models

This study presents the first systematic evaluation of in-context learning for Tarifit machine trans