Logo Lanfrica

MadaFactBench: A Benchmark for Evaluating Factual Reliability and Hallucination in Large Language Models for Malagasy

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Rat
Éditeur:
Zenodo
Hôte:avatar
MadaFactBench is a benchmark designed to evaluate the factual reliability and hallucination behavior of Large Language Models (LLMs) when answering factual questions in Malagasy. The benchmark contains 500 factual questions written in Malagasy, comprising 175 general-knowledge questions (35%) and 325 questions focusing on Madagascar-specific knowledge (65%). Questions cover multiple domains and difficulty levels and are accompanied by reference answers, accepted answer variants where applicable, and supporting sources. Five publicly accessible conversational AI systems were evaluated on the complete benchmark: GPT-5.6 Luna (ChatGPT), Gemini, Claude Sonnet 5, Meta AI powered by Muse Spark 1.1, and Mistral Large 2 (Le Chat). The experiment produced 2,500 model responses, which were manually evaluated using four behavioral categories: Correct (C), Partially Correct (PC), Hallucination/Factual Error (H), and Abstention (A). The dataset also uses NA for rare technical missing responses caused by collection or alignment issues; these cases are not considered model behavior. The repository provides both the reusable benchmark dataset and the corresponding model responses with their human evaluation labels. MadaFactBench is intended to support research on factuality, hallucination, multilingual LLM evaluation, low-resource languages, and Malagasy Natural Language Processing.

Similaires