Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

CaLMQA: Exploring culturally specific long-form question answering across 23 languages

Domain:

natural language processing

Record type:

paperdataset
Creator:
AroKarCheBha
Host:avatar
Despite rising global usage of large language models (LLMs), their ability to generate long-form answers to culturally specific questions remains unexplored in many languages. To fill this gap, we perform the first study of textual multilingual long-form QA by creating CaLMQA, a dataset of 51.7K culturally specific questions across 23 different languages. We define culturally specific questions as those that refer to concepts unique to one or a few cultures, or have different answers depending on the cultural or regional context. We obtain these questions by crawling naturally-occurring questions from community web forums in high-resource languages, and by hiring native speakers to write questions in under-resourced, rarely-studied languages such as Fijian and Kirundi. Our data collection methodologies are translation-free, enabling the collection of culturally unique questions like "Kuber iki umwami wa mbere w'uburundi yitwa Ntare?" (Kirundi; English translation: "Why was the first king of Burundi called Ntare (Lion)?"). We evaluate factuality, relevance and surface-level quality of LLM-generated long-form answers, finding that (1) for many languages, even the best models make critical surface-level errors (e.g., answering in the wrong language, repetition), especially for low-resource languages; and (2) answers to culturally specific questions contain more factual errors than answers to culturally agnostic questions -- questions that have consistent meaning and answer across many cultures. We release CaLMQA to facilitate future research in cultural and multilingual long-form QA. 46 pages, 26 figures. Accepted as a main conference paper at ACL 2025. Code and data available at github.com . Dataset expanded to 51.7K questions

Visit

arxiv.org

Tasks

question answering

Languages

MbereRundi

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

Exploring Open Domain Question Answering in FrenchBhavik-Ardeshna/Question-Answering-for-Low-Resource-LanguagesDomain-specific Embeddings for Question-Answering Systems: FAQs for Health CoachingAfri-MCQA: Multimodal Cultural Question Answering for African LanguagesAfri-MCQA: Multimodal Cultural Question Answering for African LanguagesBerhanu948/Question-Answering

Exploring Open Domain Question Answering in French

Exploring Open Domain Question Answering in French

Poster presented at the Deep Learning Indaba 2022 by Espoir Murhabazi

Bhavik-Ardeshna/Question-Answering-for-Low-Resource-Languages

Cascading Adaptors to Leverage English Data to Improve Performance ofQuestion Answering for Low-Reso

Domain-specific Embeddings for Question-Answering Systems: FAQs for Health Coaching

FAQs are widely used to respond to users’ knowledge needs within knowledge domains. While LLM might

Afri-MCQA: Multimodal Cultural Question Answering for African Languages

Paper Afri-MCQA is the first multilingual cultural question-answering benchmark covering 8k Q&A pai

Afri-MCQA: Multimodal Cultural Question Answering for African Languages

Africa is home to over one-third of the world's languages, yet remains underrepresented in AI resear

Berhanu948/Question-Answering

This is an Amharic question answering for Healthcare