Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Compiling a corpus of African American Language from oral histories

Domaine:

natural language processing

Type de record:

dataset
Créateur:
SarAleWilMic
Éditeur:
Res
Hôte:
African American Language (AAL) is a marginalized variety of American English that has been understudied due to a lack of accessible data. This lack of data has made it difficult to research language in African American communities and has been shown to cause emerging technologies such as Automatic Speech Recognition (ASR) to perform worse for African American speakers. To address this gap, the Joel Buchanan Archive of African American Oral History (JBA) at the University of Florida is being compiled into a time-aligned and linguistically annotated corpus. Through Natural Language Processing (NLP) techniques, this project will automatically time-align spoken data with transcripts and automatically tag AAL features. Transcription and time-alignment challenges have arisen as we ensure accuracy in depicting AAL morphosyntactic and phonetic structure. Two linguistic studies illustrate how the African American Corpus from Oral Histories betters our understanding of this lesser-studied variety.

Visit

doi.org

Tasks

automatic speech recognitionspeech processing

Similaires

ALFAH -- A Toolkit for Annotating Linguistic Features in African American Oral HistoriesOn compiling a corpus of South African EnglishCompiling The Oxford Dictionary of African American English : A Progress ReportThe Corpus of Regional African American Language Corpus of Regional African American Language (CORAAL)Streaming audio from African‐American oral history collections

ALFAH -- A Toolkit for Annotating Linguistic Features in African American Oral Histories

ALFAH is a toolkit designed for annotating linguistic features in African American oral histories. I

On compiling a corpus of South African English

Compiling The Oxford Dictionary of African American English : A Progress Report

ABSTRACT: This article provides an overview of the progress on the forthcoming Oxford Dictionary of

The Corpus of Regional African American Language

Corpus of Regional African American Language (CORAAL)

This dataset comprises more than 150 socio-linguistic interviews with African-American English speak

Streaming audio from African‐American oral history collections

Through the use of OCLC/DiMeMa’s CONTENTdm software and a RealSystems Server 8, this article outline