Logo Lanfrica

Extracting "non-standard" data from the Twitter API

Domain:

natural language processing

Record type:

paper
Creator:
Baxter, Kimberley
Publisher:
Language Science Press
Host:avatar

The present paper examines methodology in the use of Twitter in the corpus-based
analysis of African American English (AAE) syntax. I discuss the extraction and
geospatial mapping of indices of use of the perfective marker done (hereafter
perfective done), which, alongside a simple past form, indicates that the action described
in the past form has been completed. Widespread use of AAE on archived social
media posts creates a living database of timed, dated, and geotagged utterances from
which corpora may be built. The Academic Twitter API (ACTW) allows access to
their full database of tweets, which is a much larger and more accessible dataset
than its large social media contemporaries. I discuss two methods of extracting
perfective done from the ACTW: a front-end approach which aims to isolate uses
of perfective done by eliminating non-perfective uses of done from the search prior
to running the query, and a back-end approach which first extracts a set of all uses
of done from 2012–2015 and aims to isolate uses of perfective done afterwards. I
discuss the results of this method, as well as the implications therein and directions for
future research. I conclude that while both methods are effective at extracting per-
fective done from the ACTW, the back-end approach is better suited to geospatial
mapping.