Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Self-supervised Speech Representations Still Struggle with African American Vernacular English

Domain:

natural language processing

Record type:

paper
Creator:
ChaChoShiChe
Host:avatar
Underperformance of ASR systems for speakers of African American Vernacular English (AAVE) and other marginalized language varieties is a well-documented phenomenon, and one that reinforces the stigmatization of these varieties. We investigate whether or not the recent wave of Self-Supervised Learning (SSL) speech models can close the gap in ASR performance between AAVE and Mainstream American English (MAE). We evaluate four SSL models (wav2vec 2.0, HuBERT, WavLM, and XLS-R) on zero-shot Automatic Speech Recognition (ASR) for these two varieties and find that these models perpetuate the bias in performance against AAVE. Additionally, the models have higher word error rates on utterances with more phonological and morphosyntactic features of AAVE. Despite the success of SSL speech models in improving ASR for low resource varieties, SSL pre-training alone may not bridge the gap between AAVE and MAE. Our code is publicly available at github.com. INTERSPEECH 2024

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Computation and LanguageSoundAudio and Speech Processing

Similar

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitchingEnglish in South Africa: parallels with African American vernacular EnglishDo Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?African American Vernacular English in CaliforniaDialect-Specific Models for Automatic Speech Recognition of African American Vernacular EnglishThe Origins of African American Vernacular English

Multilingual self-supervised speech representations improve the speech recognition of low-resource African languages with codeswitching

While many speakers of low-resource languages regularly code-switch between their languages and othe

English in South Africa: parallels with African American vernacular English

A comparison between Black English usage in South Africa and the United States There has been a lon

Do Discrete Self-Supervised Representations of Speech Capture Tone Distinctions?

Discrete representations of speech, obtained from Self-Supervised Learning (SSL) foundation models,

African American Vernacular English in California

Dialect-Specific Models for Automatic Speech Recognition of African American Vernacular English

African American Vernacular English (AAVE) is a widely-spoken dialect of English, yet it is under-represented in major speech corpora. As a result, speakers of this dialect are often misunderstood by NLP applications. This study explores the effect on transcription

The Origins of African American Vernacular English