Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Compar:IA conversations

Domaine:

natural language processing

Type de record:

dataset
Créateur:
Com
Hôte:
The compar:IA dataset is a large-scale collection of real user conversations from compar:IA, a public chatbot arena run by the French Ministry of Culture. Users chat with two anonymous AI models side by side and say which answer they prefer, in a blind setting. The platform's goals are both educational, helping users understand how different models behave, and technical, contributing open alignment and evaluation data with a strong focus on French-language use. The dataset contains over 675,000 paired responses and around 208,000 human preference judgments across more than 115 conversational AI models, both open-source and proprietary. Each row is a single turn: the two models' answers to the same user message, the preference given on that turn (if any), and the full conversation each answer belongs to. Most interactions are in French and reflect unconstrained, real-world uses of conversational AI across writing, programming, administration, creative tasks, and everyday questions. Prompts are not curated or engineered, making the data representative of actual user behavior rather than benchmark-style evaluations. Each entry links the model pair, identifies the two models, and records the dialogue structure. Additional metadata fields provide automatically generated summaries, thematic categories, detected languages, output token counts, generation duration and latency, time to vote, and estimated electricity consumption, enabling analysis of performance and efficiency trade-offs. About 208,000 turns carry an explicit preference; the remaining turns are unrated and serve as raw French conversation data. User consent is collected through the platform's terms of use. Automated detection is used to identify and anonymize personally identifiable information, but no filtering is applied to remove potentially toxic or sensitive content, in order to support research on safety and real-world risks. The dataset is released under the open Etalab 2.0 and CC-BY-4.0 licenses and is intended for research and development in conversational model alignment, evaluation methods, human-AI interaction, and AI safety, particularly for French and other under-resourced languages.

Visit

mozilladatacollective.com

Tags

mdcmozilla data collectiveNLGPARQUET

Licenses

Etalab 2.0

Similaires

Youth Conversations DatasetFacebook and LinkedIn conversationsReza2kn/adhdfree-conversations-v1Conversations in Northern PomoConversations in Kirika CONV_01Conversations about an unknown topic

Youth Conversations Dataset

This dataset contains synthetic transcripts of conversations between a counsellor and a young person

Facebook and LinkedIn conversations

This dataset consists of conversations collected from interactions generated by Algerian users in Fa

Reza2kn/adhdfree-conversations-v1

3,378 multilingual ADHD-support conversations in OpenAI chat format (messages array of user and assi

Conversations in Northern Pomo

No English glosses. http://cla.berkeley.edu/item/19275

Conversations in Kirika CONV_01

A narration by Anthonia Alali (AA) and Elder Zedekiah Frank Opunye of an incident in the village. Ma

Conversations about an unknown topic

Mostly Central Pomo with some English. http://cla.berkeley.edu/item/18980