Perception Verb Database
This dataset accompanies the following publication: Norcliffe, E & Majid, A. 2024. Verbs of Perception: A quantitative typological study. Language, 100 (1), 81-123. 10.1353/lan.2024.a922000
The database is stored as a .CSV file in long-table format, in which each row corresponds to a unique datapoint (a form-meaning association) within the set of basic perception verbs in the 100 language sample and each cell specifies a single type of data. Columns specify information pertaining to the following fields:
AREA: macro-area
FAMILY: language family
LANGUAGE: language name
GLOTTOCODE: glottocode
LAT: latitude of language
LONG: longitude of language
MEANING_SOURCE: meaning of the form as listed in the source
MEANING_STANDARDISED: standardised meaning of the form. In most cases these are identical to the source meanings, but in certain cases we recode closely related synonyms between languages as a single meaning (see Supplementary Materials, S4).
FORM: form of the word as provided in source. Note that the diacritics and some characters have not always been properly rendered, so the original source should always be consulted to check for accuracy.
COMPLEXITY: whether the verb form is morphologically simple or complex. Complex forms include multi-word and morphologically derived expressions.
SOURCE: author and date of the original data source. For the full publication reference, consult the bibliography of the article.
SOURCE_ADDITIONAL: author and date of any additional data source consulted. Entries for columns 1-6 are taken from/standardised according to Glottolog (
glottolog.org)