Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Finding Diamonds in Conversation Haystacks: A Benchmark for Conversational Data Retrieval

Domain:

natural language processing

Record type:

paperdataset
Creator:
LeeSonKim
Host:avatar
We present the Conversational Data Retrieval (CDR) benchmark, the first comprehensive test set for evaluating systems that retrieve conversation data for product insights. With 1.6k queries across five analytical tasks and 9.1k conversations, our benchmark provides a reliable standard for measuring conversational data retrieval performance. Our evaluation of 16 popular embedding models shows that even the best models reach only around NDCG@10 of 0.51, revealing a substantial gap between document and conversational data retrieval capabilities. Our work identifies unique challenges in conversational data retrieval (implicit state recognition, turn dynamics, contextual references) while providing practical query templates and detailed error analysis across different task categories. The benchmark dataset and code are available at github.com. Accepted by EMNLP 2025 Industry Track

Visit

arxiv.org

Tasks

information retrieval

Tags

Computation and Language

Similar

Conversational Implicature in Hausa: A Study of Doctor-Patient ConversationMultilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ BenchmarkLeveraging Retrieval-Augmented Generation for Swahili Language Conversation SystemsCan LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacksCan LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacksPerformance comparison of cross-lingual retrieval models on code-switched vs. M2C2 conversational query data

Conversational Implicature in Hausa: A Study of Doctor-Patient Conversation

The paper explores the conversational implicature with the view to determine the conversational maxi

Multilingual Pretraining Data Scaling for Robust Low-Resource Retrieval in the WebFAQ Benchmark

We present WebFAQ, a large-scale collection of open-domain question answering datasets derived from

Leveraging Retrieval-Augmented Generation for Swahili Language Conversation Systems

A conversational system is an artificial intelligence application designed to interact with users in

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval and haystacks

Existing multilingual long-context benchmarks, often based on the popular needle-in-a-haystack test,

Can LLMs reason over extended multilingual contexts? Towards long-context evaluation beyond retrieval over haystacks

Existing multilingual long-context benchmarks, often based on the popular needle-in-a-haystack test,

Performance comparison of cross-lingual retrieval models on code-switched vs. M2C2 conversational query data

Transferring information retrieval (IR) models from a high-resource language (typically English) to