Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

On the Contributions of Visual and Textual Supervision in Low-Resource Semantic Speech Retrieval

Domaine:

natural language processing

Type de record:

paper
Créateur:
PasShiKamLiv
Hôte:avatar
Recent work has shown that speech paired with images can be used to learn semantically meaningful speech representations even without any textual supervision. In real-world low-resource settings, however, we often have access to some transcribed speech. We study whether and how visual grounding is useful in the presence of varying amounts of textual supervision. In particular, we consider the task of semantic speech retrieval in a low-resource setting. We use a previously studied data set and task, where models are trained on images with spoken captions and evaluated on human judgments of semantic relevance. We propose a multitask learning approach to leverage both visual and textual modalities, with visual supervision in the form of keyword probabilities from an external tagger. We find that visual grounding is helpful even in the presence of textual supervision, and we analyze this effect over a range of sizes of transcribed data sets. With ~5 hours of transcribed speech, we obtain 23% higher average precision when also using visual supervision.

Visit

arxiv.org

Tags

Computation and LanguageSoundAudio and Speech Processing

Similaires

Semantic Alignment Impact on Zero-Shot Retrieval Performance in Low-Resource LanguagesImproving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource LanguagesReusable Component Retrieval: A Semantic Search Approach for Low-Resource LanguagesMulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and FusionPerformance of XLM-R Models on Zero-Shot Cross-Lingual Semantic Textual Similarity in XTREME-R for Low-Resource African LanguagesA Resource-Light Method for Cross-Lingual Semantic Textual Similarity

Semantic Alignment Impact on Zero-Shot Retrieval Performance in Low-Resource Languages

Information retrieval across different languages is an increasingly important challenge in natural l

Improving Multilingual Semantic Textual Similarity with Shared Sentence Encoder for Low-resource Languages

Measuring the semantic similarity between two sentences (or Semantic Textual Similarity - STS) is fu

Reusable Component Retrieval: A Semantic Search Approach for Low-Resource Languages

A common practice among programmers is to reuse existing code, accomplished by performing natural la

MulMoSenT: Multimodal Sentiment Analysis for a Low-Resource Language Using Textual-Visual Cross-Attention and Fusion

First-ever Bengali Multimodal Sentiment Analysis (BMSA) corpus and details in https://www.sciencedir

Performance of XLM-R Models on Zero-Shot Cross-Lingual Semantic Textual Similarity in XTREME-R for Low-Resource African Languages

Pre-trained multilingual language encoders, such as multilingual BERT and XLM-R, show great potentia

A Resource-Light Method for Cross-Lingual Semantic Textual Similarity

Recognizing semantically similar sentences or paragraphs across languages is beneficial for many tas