Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Evaluating artificial intelligence large language models’ performances in a South African high school chemistry exam

Domaine:

educationnatural language processing

Type de record:

paper
Créateur:
Sam
Éditeur:
Mod
Hôte:
Gemini, ChatGPT Plus, and Claude 3.5 Sonnet are artificial intelligence (AI) chatbots with potential in education. Their capabilities, such as acting as virtual teaching assistants, offering personalized responses to learners’ queries, and summarizing content, make them versatile tools with the potential to assist learners. The chemistry section of physical sciences in South Africa is often considered challenging, and learners could benefit from virtual teaching assistants to supplement traditional instruction. However, little is known about AI chatbots’ abilities in solving high school chemistry problems. This descriptive case study examined the capabilities of Gemini, Claude 3.5 Sonnet, and ChatGPT Plus in accurately answering questions from the final grade 12 physical sciences chemistry exam in South Africa. The conceptual framework that guided the study was Bloom’s taxonomy of educational objectives. The responses were rigorously evaluated using the same criteria and rubrics applied to the candidates that year, ensuring a fair and robust comparison. The findings were that ChatGPT Plus performed at 47%, Gemini at 51% and Claude 3.5 Sonnet at 65%. All chatbots performed above the average performance of the candidates who sat for the paper that year, which was 46%. This has significant implications for policymakers, teachers, and learners regarding integrating large language models in teaching physical sciences and exam preparation.

Visit

doi.org

Tasks

question answering

Similaires

Evaluating the Usage of African-American Vernacular English in Large Language ModelsAfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language ModelsPREPAREDNESS OF HIGH SCHOOL STUDENTS FOR ARTIFICIAL INTELLIGENCE IN ACCOUNTING: INSIGHTS FROM SCHOOLS IN SOUTH AFRICAEvaluating Metalinguistic Knowledge in Large Language Models across the World's LanguagesEvaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"Imagining a South African Climate Change Adaptation-aligned school curriculum using Generative Artificial Intelligence

Evaluating the Usage of African-American Vernacular English in Large Language Models

In AI, most evaluations of natural language understanding tasks are conducted in standardized dialec

AfriStereo: A Culturally Grounded Dataset for Evaluating Stereotypical Bias in Large Language Models

Existing AI bias evaluation benchmarks largely reflect Western perspectives, leaving African context

PREPAREDNESS OF HIGH SCHOOL STUDENTS FOR ARTIFICIAL INTELLIGENCE IN ACCOUNTING: INSIGHTS FROM SCHOOLS IN SOUTH AFRICA

This paper examines how prepared South African high school students are to interact with Artificial

Evaluating Metalinguistic Knowledge in Large Language Models across the World's Languages

LLMs are routinely evaluated on language use, yet their explicit knowledge about linguistic structur

Evaluating Racial Bias in Large Language Models: The Necessity for "SMOKY"

This paper evaluates the understanding and biases of large language models (LLMs) regarding

Imagining a South African Climate Change Adaptation-aligned school curriculum using Generative Artificial Intelligence