Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

Domain:

natural language processing

Record type:

paper
Creator:
HosMatBru
Publisher:
arXiv
Host:avatar
Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully recorded speech, limiting the reach of speech-LLM methods in low-resource settings. This paper investigates whether text-to-speech (TTS) can provide task-specific training data for Luxembourgish SQA without requiring a large human-recorded QA corpus. Starting from existing text-based QA resources, we translate questions into Luxembourgish, synthesize spoken questions with multiple TTS systems, and pair them with textual answers. We train a parameter-efficient SLAM-style architecture that connects a frozen Whisper encoder to frozen multilingual LLM backends through a learned projector and LoRA adapters. We compare MMS-TTS, Qwen3-TTS, and OmniVoice variants, including single-source corpora of about 48k questions and a 4TTS multi-source mix of approximately 230k questions. Evaluation on LLAMA-LB-Test with two real Luxembourgish speaker conditions shows that multi-source and voice-design-based synthetic training configurations yield the strongest SQA performance. The results also show that no-reference TTS quality scores do not monotonically predict downstream QA performance, indicating that synthetic speech must be evaluated as task-specific training data rather than only as natural-sounding audio. 7 pages, under review

Visit

doi.org

Tasks

text to speechquestion answeringspeech processing

Tags

Computation and Language (cs.CL)FOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Spoken Dialectal Question Answering for the Real WorldDUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question AnsweringImproving Amharic Legal Question Answering with Retrieval-Augmented Generation and Locally-Sourced DataSD-QA: Spoken Dialectal Question Answering for the Real WorldCPIQA: Climate Paper Image Question Answering Dataset for Retrieval-Augmented Generation with Context-based Query ExpansionBerhanu948/Question-Answering

Spoken Dialectal Question Answering for the Real World

Question answering (QA) systems are now available through numerous commercial applications for a wid

DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering

Spoken Question Answering (SQA) is to find the answer from a spoken document given a question, which

Improving Amharic Legal Question Answering with Retrieval-Augmented Generation and Locally-Sourced Data

SD-QA: Spoken Dialectal Question Answering for the Real World

Question answering (QA) systems are now available through numerous commercial applications for a wid

CPIQA: Climate Paper Image Question Answering Dataset for Retrieval-Augmented Generation with Context-based Query Expansion

CPIQA is a large scale QA dataset focused on figured extracted from scientific research papers from

Berhanu948/Question-Answering

This is an Amharic question answering for Healthcare