Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

DUAL: Discrete Spoken Unit Adaptive Learning for Textless Spoken Question Answering

Domaine:

natural language processing

Type de record:

paperdatasetsoftware
Créateur:
LinChuChuYan
Hôte:avatar
Spoken Question Answering (SQA) is to find the answer from a spoken document given a question, which is crucial for personal assistants when replying to the queries from the users. Existing SQA methods all rely on Automatic Speech Recognition (ASR) transcripts. Not only does ASR need to be trained with massive annotated data that are time and cost-prohibitive to collect for low-resourced languages, but more importantly, very often the answers to the questions include name entities or out-of-vocabulary words that cannot be recognized correctly. Also, ASR aims to minimize recognition errors equally over all words, including many function words irrelevant to the SQA task. Therefore, SQA without ASR transcripts (textless) is always highly desired, although known to be very difficult. This work proposes Discrete Spoken Unit Adaptive Learning (DUAL), leveraging unlabeled data for pre-training and fine-tuned by the SQA downstream task. The time intervals of spoken answers can be directly predicted from spoken documents. We also release a new SQA benchmark corpus, NMSQA, for data with more realistic scenarios. We empirically showed that DUAL yields results comparable to those obtained by cascading ASR and text QA model and robust to real-world data. Our code and model will be open-sourced. Accepted by Interspeech 2022

Visit

arxiv.org

Tasks

question answering

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

Spoken Dialectal Question Answering for the Real WorldSD-QA: Spoken Dialectal Question Answering for the Real WorldLuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question AnsweringDeep learning-based approach for Arabic open domain question answeringA Transfer Learning Approach For Identifying Spoken Maghrebi DialectsBerhanu948/Question-Answering

Spoken Dialectal Question Answering for the Real World

Question answering (QA) systems are now available through numerous commercial applications for a wid

SD-QA: Spoken Dialectal Question Answering for the Real World

Question answering (QA) systems are now available through numerous commercial applications for a wid

LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering

Spoken Question Answering (SQA) remains largely focused on high-resource languages and carefully rec

Deep learning-based approach for Arabic open domain question answering

Open-domain question answering (OpenQA) is one of the most challenging yet widely investigated probl

A Transfer Learning Approach For Identifying Spoken Maghrebi Dialects

This paper investigates a transfer learning approach to solve the spoken dialects identification pro

Berhanu948/Question-Answering

This is an Amharic question answering for Healthcare