Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

How Good Are Large Language Models at Arithmetic Reasoning in Low-Resource Language Settings?—A Study on Yorùbá Numerical Probes with Minimal Contamination

Domain:

natural language processing

Record type:

dataset
Creator:
FiyPet
Publisher:
MDP
Host:
We study the performance of large language models (LLMs) in natural language understanding and natural language reasoning tasks in a low-resourced-language (LRL) setting. Using Yorùbá, an LRL, we curated a set of numerical probes with minimal contamination. The probes comprise three sets of questions—the first covers basic arithmetic, the second covers date and time (calendar system), and the last focuses on numerals and counting systems. Assessed in a zero-shot setup, three LLMs (ChatGPT, Gemini, and PaLM) were evaluated based on several metrics. The best-performing model, ChatGPT, generated some correct answers, showing logical steps in attaining the answers in Yorùbá (with an accuracy of 56% in set one, and 44% in set two). The second-best model (with an accuracy of 56% in set one, and 32% in set two) is Gemini. PaLM (with an accuracy of 16% in set one, and 8% in set two) showed the answers without logic. The three models performed poorly on the Yorùbá numerals question set (ChatGPT scored 8%, and Gemini and PaLM each had 0% accuracy). The study also revealed that there is significant room for improvement in the state of the art of LLMs when it comes to Yorùbá numerals.

Visit

doi.org

Languages

Yoruba

Licenses

https://creativecommons.org/licenses/by/4.0/

Similar

How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings – A Benchmarking Study & DatasetHow Good are Commercial Large Language Models on African Languages?AfroBench: How Good are Large Language Models on African Languages?Yorùbá Arithmetic Reasoning DatasetLarge language models for frontline healthcare support in low-resource settingsIn-Context Learning for Low-Resource Machine Translation: A Study on Tarifit with Large Language Models

How Good Are Large Language Models at Supporting Frontline Healthcare Workers in Low-Resource Settings – A Benchmarking Study & Dataset

Abstract Large language models (LLMs) have demonstrated strong performance in medi

How Good are Commercial Large Language Models on African Languages?

Recent advancements in Natural Language Processing (NLP) has led to the proliferation of large pretr

AfroBench: How Good are Large Language Models on African Languages?

Large-scale multilingual evaluations, such as MEGA, often include only a handful of African language

Yorùbá Arithmetic Reasoning Dataset

#YORUBA ARITHMETIC REASONING DATASET This dataset consists of arithmetic and numerical reasoning que

Large language models for frontline healthcare support in low-resource settings

Abstract Large language models (LLMs) have demonstrated str

In-Context Learning for Low-Resource Machine Translation: A Study on Tarifit with Large Language Models

This study presents the first systematic evaluation of in-context learning for Tarifit machine trans