Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Lost in Simulation: LLM-Simulated Users are Unreliable Proxies for Human Users in Agentic Evaluations

Domaine:

natural language processing

Type de record:

paper
Créateur:
SesCahOduSin
Hôte:avatar
Agentic benchmarks increasingly rely on LLM-simulated users to scalably evaluate agent performance, yet the robustness, validity, and fairness of this approach remain unexamined. Through a user study with participants across the United States, India, Kenya, and Nigeria, we investigate whether LLM-simulated users serve as reliable proxies for real human users in evaluating agents on τ-Bench retail tasks. We find that user simulation lacks robustness, with agent success rates varying up to 9 percentage points across different user LLMs. Furthermore, evaluations using simulated users exhibit systematic miscalibration, underestimating agent performance on challenging tasks and overestimating it on moderately difficult ones. African American Vernacular English (AAVE) speakers experience consistently worse success rates and calibration errors than Standard American English (SAE) speakers, with disparities compounding significantly with age. We also find simulated users to be a differentially effective proxy for different populations, performing worst for AAVE and Indian English speakers. Additionally, simulated users introduce conversational artifacts and surface different failure patterns than human users. These findings demonstrate that current evaluation practices risk misrepresenting agent capabilities across diverse user populations and may obscure real-world deployment challenges.

Visit

arxiv.org

Tags

Human-Computer InteractionArtificial IntelligenceComputers and SocietyMachine Learning

Similaires

Swahili Terminological Modernization in Tanzania. What Are the Register Users’ Views?Adoption of Virtual Assistants for Human-Computer Interaction among Smartphone Users in Lagos, NigeriaImproving Methodologies for LLM Evaluations Across Global LanguagesMaternal and Perinatal Outcomes among Maternity Waiting Home Users and Non-Users in Rural RwandaUsers' Traces for Enhancing Arabic Facebook SearchLIS Education for 21st Century Information Users

Swahili Terminological Modernization in Tanzania. What Are the Register Users’ Views?

Adoption of Virtual Assistants for Human-Computer Interaction among Smartphone Users in Lagos, Nigeria

Steady advancements in digital technologies are facilitating human-machine interaction to rival huma

Improving Methodologies for LLM Evaluations Across Global Languages

As frontier AI models are deployed globally, it is essential that their behaviour remains safe and r

Maternal and Perinatal Outcomes among Maternity Waiting Home Users and Non-Users in Rural Rwanda

Most maternal and perinatal deaths could be prevented through timely access to skilled birth attenda

Users' Traces for Enhancing Arabic Facebook Search

International audience This paper proposes an approach on Facebook search in Arabic,

LIS Education for 21st Century Information Users

This chapter is on library and information science education for the 21st century users. It aims at