Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

VisionArena: 230K Real World User-VLM Conversations with Preference Labels

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
ChoDunMasMan
Hôte:avatar
With the growing adoption and capabilities of vision-language models (VLMs) comes the need for benchmarks that capture authentic user-VLM interactions. In response, we create VisionArena, a dataset of 230K real-world conversations between users and VLMs. Collected from Chatbot Arena - an open-source platform where users interact with VLMs and submit preference votes - VisionArena spans 73K unique users, 45 VLMs, and 138 languages. Our dataset contains three subsets: VisionArena-Chat, 200k single and multi-turn conversations between a user and a VLM; VisionArena-Battle, 30K conversations comparing two anonymous VLMs with user preference votes; and VisionArena-Bench, an automatic benchmark of 500 diverse user prompts that efficiently approximate the live Chatbot Arena model rankings. Additionally, we highlight the types of question asked by users, the influence of response style on preference, and areas where models often fail. We find open-ended tasks like captioning and humor are highly style-dependent, and current VLMs struggle with spatial reasoning and planning tasks. Lastly, we show finetuning the same base model on VisionArena-Chat outperforms Llava-Instruct-158K, with a 17-point gain on MMMU and a 46-point gain on the WildVision benchmark. Dataset at huggingface.co updated for CVPR Camera Ready

Visit

arxiv.org

Tags

Computer Vision and Pattern Recognition

Similaires

FewShotCropNet: Real-Time Detection of Emerging Crop Diseases with Limited Labels Using Spectral-Temporal Attention Prototypical NetworksReal to H-space Autoencoders for Theme Identification in Telephone ConversationsMulti-lingual Learning with Limited Labelsafrica-ai/vlmCulTex-VLM/EC-VCRHumachine/egypt-constitution-vlm

FewShotCropNet: Real-Time Detection of Emerging Crop Diseases with Limited Labels Using Spectral-Temporal Attention Prototypical Networks

The acquisition of labelled data for new or emerging plant diseases is difficult due to the high cos

Real to H-space Autoencoders for Theme Identification in Telephone Conversations

International audience Machine learning (ML) and deep learning with deep neural netwo

Multi-lingual Learning with Limited Labels

Recent advances in natural language processing with neural networks have largely benefited high-reso

africa-ai/vlm

# Kalenjin Dictionary OCR & Extraction A clean and efficient pipeline to extract dictionary entries

CulTex-VLM/EC-VCR

EC-VCR (Egyptian Culture Visual Commonsense Reasoning) is a multimodal benchmark designed to evaluat

Humachine/egypt-constitution-vlm