Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Towards Reliable Evaluation of Large Language Models for Multilingual and Multimodal E-Commerce Applications

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
XieLieZhaZha
Hôte:avatar
Large Language Models (LLMs) excel on general-purpose NLP benchmarks, yet their capabilities in specialized domains remain underexplored. In e-commerce, existing evaluations-such as EcomInstruct, ChineseEcomQA, eCeLLM, and Shopping MMLU-suffer from limited task diversity (e.g., lacking product guidance and after-sales issues), limited task modalities (e.g., absence of multimodal data), synthetic or curated data, and a narrow focus on English and Chinese, leaving practitioners without reliable tools to assess models on complex, real-world shopping scenarios. We introduce EcomEval, a comprehensive multilingual and multimodal benchmark for evaluating LLMs in e-commerce. EcomEval covers six categories and 37 tasks (including 8 multimodal tasks), sourced primarily from authentic customer queries and transaction logs, reflecting the noisy and heterogeneous nature of real business interactions. To ensure both quality and scalability of reference answers, we adopt a semi-automatic pipeline in which large models draft candidate responses subsequently reviewed and modified by over 50 expert annotators with strong e-commerce and multilingual expertise. We define difficulty levels for each question and task category by averaging evaluation scores across models with different sizes and capabilities, enabling challenge-oriented and fine-grained assessment. EcomEval also spans seven languages-including five low-resource Southeast Asian languages-offering a multilingual perspective absent from prior work.

Visit

arxiv.org

Tasks

question answering

Tags

Artificial Intelligence

Similaires

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language ModelsTowards Multimodal Cultural Context Modeling for African Languages in Large Language ModelsHealMed: Multilingual Evaluation of Large Language Models in MedicineGlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language ModelsEthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task EvaluationGlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

M3Exam: A Multilingual, Multimodal, Multilevel Benchmark for Examining Large Language Models

Despite the existence of various benchmarks for evaluating natural language processing models, we ar

Towards Multimodal Cultural Context Modeling for African Languages in Large Language Models

This preliminary work addresses the critical gap in multimodal Large Language Models (LLMs) for Afri

HealMed: Multilingual Evaluation of Large Language Models in Medicine

We present HealMed, an expert-reviewed benchmark for multilingual evaluation of large language model

GlotEval: A Test Suite for Massively Multilingual Evaluation of Large Language Models

Large language models (LLMs) are advancing at an unprecedented pace globally, with regions increasin

EthioLLM: Multilingual Large Language Models for Ethiopian Languages with Task Evaluation

Large language models (LLMs) have gained popularity recently due to their outstanding performance in

GlobeAudio: A Multilingual Multicultural Benchmark for Naturalistic Evaluation of Large Audio-Language Models

Large Audio-Language Models (LALMs) integrate audio perception and language understanding within a u