Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?

Domain:

natural language processing

Record type:

papermodel
Creator:
OsaKin
Host:avatar
Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models, are widely used, especially where there are limited data for the downstream task, such as for a low-resource language. Typically, discretization of speech into a sequence of symbols is achieved by unsupervised clustering of the latents from an SSL model. Our study evaluates whether discrete symbols - found using k-means - adequately capture tone in two example languages, Mandarin and Yoruba. We compare latent vectors with discrete symbols, obtained from HuBERT base, MandarinHuBERT, or XLS-R, for vowel and tone classification. We find that using discrete symbols leads to a substantial loss of tone information, even for language-specialised SSL models. We suggest that discretization needs to be task-aware, particularly for tone-dependent downstream tasks. Submitted to ICASSP 2025

Visit

arxiv.org

Tasks

speech processing

Languages

Yoruba

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Self-supervised Speech Representations Still Struggle with African American Vernacular EnglishFine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched SpeechMultilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingDo self-supervised speech models develop human-like perception biases?Non-Contrastive Self-Supervised Speech Representations vs. Wav2Vec 2.0 in Low-Resource LanguagesA comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings

Self-supervised Speech Representations Still Struggle with African American Vernacular English

Underperformance of ASR systems for speakers of African American Vernacular English (AAVE) and other

Fine-Tuned Self-Supervised Speech Representations for Language Diarization in Multilingual Code-Switched Speech

Annotating a multilingual code-switched corpus is a painstaking process requiring specialist linguis

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

Do self-supervised speech models develop human-like perception biases?

Self-supervised models for speech processing form representational spaces without using any external

Non-Contrastive Self-Supervised Speech Representations vs. Wav2Vec 2.0 in Low-Resource Languages

This report synthesises findings from 13 peer-reviewed papers addressing the following research ques

A comparison of self-supervised speech representations as input features for unsupervised acoustic word embeddings

Many speech processing tasks involve measuring the acoustic similarity between speech segments. Acou