Logo Lanfrica

Variation across Regions and Demographics in African American Language Morphosyntax: Evidence from Large-Scale Twitter Data

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
TesChlLisTay
Éditeur:
Duk
Hôte:
Abstract Social media data and computational tools have become increasingly powerful alternatives to traditional methods of data collection and annotation for sociolinguistic research. While not without their own drawbacks, these resources can help address critical problems faced by researchers, such as sparse data, the observer's paradox, and annotating corpora at scale. This article presents a Twitter dataset of 227 million conversational messages across the U.S. and uses deep learning natural language processing (NLP) methods to analyze how 18 morphosyntactic features used by speakers of African American Language (AAL) vary geographically and across 12 demographic factors. Results demonstrate more frequent usage of AAL features in the rural South and in Mexican American communities, both of which are underrepresented in the literature. This work constitutes the first national-level description and analysis of overall morphosyntactic variation in AAL, and demonstrates how NLP tools enable the study of large-scale data to gain a more representative understanding of speech in marginalized communities.

Similaires