Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Significant Sequences in World Englishes: A Data-driven Approach to Variationist Models

Domain:

natural language processing

Record type:

paper
Creator:
Koc
Editor:
KreAng
Publisher:
Phi
Host:avatar
The present study attempts a strictly data-driven/bottom-up evaluation of models of World Englishes. Methodologically, diverging degrees of association within lexical and grammatical n-grams are chosen as the linguistic basis on which to estimate similarities and differences between 15 varieties of World Englishes as represented by the International Corpus of English. To this end, sequences of both dynamic and statics lengths are generated from homogenized components of the corpus in both its regular lexical format as well as a POS-annotated version. Analysis of collocational preference is carried out by applying five association measures of both traditional (MI-score, t-score, log-likelihood) as well as more innovative designs (lexical gravity, Delta P) to the respective datasets. On the basis of these association patterns within the different datasets, groups of varieties exhibiting similar association profiles can be established through the application of various clustering techniques. In essence, these methods identify binary pairs of varieties which display the least amount of difference, and consecutively merge these. This process is repeated until all varieties are accounted for within the cluster structure. Since clustering methods, however, have a tendency of discovering patterns even in random data, results from various clustering techniques (hierarchical clustering, k-means, phylogenetic clustering) are triangulated for the present study, and segmentations within the hierarchical structures are empirically substantiated through the application of random resampling of the data. The variety clusters thus obtained are in turn contrasted to expectations derived from extra-linguistic assessments informed by major language-externally grounded models. This particularly concerns three types of models for the description of World Englishes: 1) traditional, tripartite distinctions into English as a native/second/foreign language, 2) models of regional standardization and epicentral effects (Hundt 2013), as well as 3) evolutionary models of language and identity formation in postcolonial settings (Schneider 2007, 2014). Each of these models would suggest different patterns of similarity within the data, and are thus in turn contrasted against the empirical findings based on variety-specific association profiles. Results of the study support an interpretation along regional criteria most strongly. In particular, the African varieties are commonly found to differentiate clearly from the remaining data, while exhibiting internal regional separation. While there is some separation between traditional ENL/ESL varieties within the spoken data, a regional explanation emerges most strongly from the more comprehensive written dataset. The Asian data as a whole least support this interpretation and commonly fragment into smaller and more fluid groups, but on a more fine-grained level, pairs based on regional proximity still recur frequently. Support for groups based on Schneider’s dynamic model is generally low: Some clusters emerging from the data match those expected on the basis of the model, but convergence is generally lower than within an analysis based on proximity. In the latter case, the data frequently mirror not only large-scale but also more fine-grained patterns, while several groups of varieties based on similar degrees of exo- or endonormative normative stabilization fail to reliably emerge from the data. Thus, the analysis concludes by favoring regional and cultural proximity over other explanative approaches for the description of association patterns in World Englishes.

Visit

doi.orgarchiv.ub.uni-marburg.de

Tags

New Englishesinstitutionalized L2 varietiesClusteranalysen-gramsinstitutionalisierte L2-Varietätencluster analysisLanguage, LinguisticsSprachwissenschaft, LinguistikphraseologyMehrworteinheiten+6

Licenses

Creative Commons Attribution Non Commercial No Derivatives 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-nd/4.0/legalcode

Similar

The comparative correlative construction in World Englishes a usage-based construction grammar approachData-Driven Approach to Capitation Reform in RwandaSIGNBASE: A DATA-DRIVEN APPROACH TO ABSTRACT SIGNS IN THE PALEOLITHICEnglishes around the WorldData Mining Methods to Compare EnglishesHybridity, globalisation and models of Englishes

The comparative correlative construction in World Englishes a usage-based construction grammar approach

Employing a two-pronged approach to analyzing language, this study uses evidence obtained from a lar

Data-Driven Approach to Capitation Reform in Rwanda

As part of Rwanda's transition toward universal health coverage, the national Community-Based Health

SIGNBASE: A DATA-DRIVEN APPROACH TO ABSTRACT SIGNS IN THE PALEOLITHIC

In the Paleolithic around 100,000 to 10,000 years ago, abstract motives also referred to as signs, p

Englishes around the World

The two volumes of Englishes around the World present high-quality original research papers written

Data Mining Methods to Compare Englishes

The paper presents the results of the corpus-based research of noun cryptotypes in 20 varieties of E

Hybridity, globalisation and models of Englishes

Abstract Current models of Englishes face empirical challenges, such as multilingualism, hybrid va