Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Cross-lingual topic prediction for speech using translations

Domaine:

natural language processing

Type de record:

papermodel
Créateur:
BanKamLopGol
Hôte:avatar
Given a large amount of unannotated speech in a low-resource language, can we classify the speech utterances by topic? We consider this question in the setting where a small amount of speech in the low-resource language is paired with text translations in a high-resource language. We develop an effective cross-lingual topic classifier by training on just 20 hours of translated speech, using a recent model for direct speech-to-text translation. While the translations are poor, they are still good enough to correctly classify the topic of 1-minute speech segments over 70% of the time - a 20% improvement over a majority-class baseline. Such a system could be useful for humanitarian applications like crisis response, where incoming speech in a foreign low-resource language must be quickly assessed for further action. Accepted to ICASSP 2020

Visit

arxiv.org

Tasks

speech processingtext classificationtopic classification

Tags

Computation and Language

Similaires

Unsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabelsangeet2020/Cross-lingual-topic-identification-in-low-resource-scenariosExploiting Adapters for Cross-lingual Low-resource Speech RecognitionPOTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text TranslationTesting Cross-Lingual Text Comprehension In LLMs Using Next Sentence PredictionRunyankore-Rukiga Simulated Speech Corpus for Cross-Lingual Computational Speech Research (Version 1.0)

Unsupervised Cross-Lingual Speech Emotion Recognition Using Pseudo Multilabel

Speech Emotion Recognition (SER) in a single language has achieved remarkable results through deep l

sangeet2020/Cross-lingual-topic-identification-in-low-resource-scenarios

Topic prediction for low-resource language # Cross-lingual topic identification in low resource sce

Exploiting Adapters for Cross-lingual Low-resource Speech Recognition

Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource langu

POTSA: A Cross-Lingual Speech Alignment Framework for Speech-to-Text Translation

Speech Large Language Models have achieved breakthroughs in multilingual speech-to-text translation.

Testing Cross-Lingual Text Comprehension In LLMs Using Next Sentence Prediction

While large language models are trained on massive datasets, this data is heavily skewed towards Eng

Runyankore-Rukiga Simulated Speech Corpus for Cross-Lingual Computational Speech Research (Version 1.0)

The Runyankore-Rukiga Simulated Speech Corpus (RR-SC) is a structured speech dataset develo