Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SA-AEO-Bench v1: An Open Benchmark for Measuring LLM Citation Behavior on South African Queries

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Ban
Éditeur:
Cen
Éditeur:
OSF
Hôte:avatar
The first publicly pre-registered benchmark measuring how three frontier Large Language Models (OpenAI GPT-5, Anthropic Claude Sonnet 4.5, Google Gemini 2.5 Pro) cite source material when answering South Africa–specific queries. Conducted by Cited Brands (citedbrands.co.za). 5,500 prompts across 10 South African consumer industries (banking, telecom, grocery retail, medical aid, short-term insurance, automotive/EV, e-commerce, restaurants, streaming, real estate). 100 brands. 1,100 unique prompts × 5 replications = 5,500 prompts × 3 LLMs = 16,500 API calls. Pre-registered hypotheses (H1–H7) include: SA-domain citation share, model-specific citation patterns, Latin Square position bias, reputation polarity gap, multilingual coverage in Afrikaans and isiZulu, Reddit citation asymmetry, and search-budget standardization validity. Methodology: blind prompts, Latin Square counterbalancing for comparison questions, Bradley-Terry pairwise MLE for brand strength, bootstrap confidence intervals at prompt level, Cohen's kappa for inter-rater URL classification reliability. Public release: protocol, code, prompts, classification rubric, and aggregate citation dataset under CC-BY-4.0. Raw response text retained by Cited Brands for commercial use (disclosed). Total study cost: ~USD 1,450 (~ZAR 26,500). API Cost only no development cost included Re-runnable by any researcher with API access for the same cost.

Visit

doi.orgosf.io

Languages

AfrikaansZulu

Tags

BusinessCommunicationMass CommunicationMarketingFOS: Economics and businessPhysical Sciences and MathematicsInformation LiteracyLibrary and Information ScienceComputer SciencesSocial and Behavioral Sciences+16

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

Bias Amplification Cascade: LLM Benchmark for African Health EquityGgboykxz/gabon-llm-v1Kambaata LLM Cultural Benchmarkrifaasa/hassaniya-llm-benchmarkKrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural AdvisoryMulti-Hall-SA: A Cross-lingual Benchmark for Multi-Type Hallucination Detection in Low-Resource South African Languages

Bias Amplification Cascade: LLM Benchmark for African Health Equity

Analysis code, data, and results for "The Bias Amplification Cascade" (Communic

Ggboykxz/gabon-llm-v1

Kambaata LLM Cultural Benchmark

Maintenance release for Zenodo archival of the Kambaata LLM Cultural Benchmark. This release contain

rifaasa/hassaniya-llm-benchmark

benchmark of modern LLMs on Hassaniya Arabic dialect » # hassaniya-llm-benchmark Code and evaluati

KrishokChat: A Citation-Grounded Dataset and Benchmark for Bengali Agricultural Advisory

We present KrishokChat, the first citation-grounded Bengali agricultural instruction-tuning dataset

Multi-Hall-SA: A Cross-lingual Benchmark for Multi-Type Hallucination Detection in Low-Resource South African Languages