Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Corpus linguistics and non‐native varieties of English

Domain:

natural language processing

Record type:

dataset
Creator:
JOS
Publisher:
WILEY
Host:
ABSTRACT: This article derives from the internal discussions of a project that has just been launched and which may provide a useful example of modern comparative linguistics: the International Corpus of English (ICE). It concentrates on the problems which arise when the principles of corpus compilation, which were developed in native communities (ENL corpora) in the pre‐sociolinguistic age, are applied to non‐native communities (ESL corpora) such as Africa. In my opinion this reveals a crucial difficulty in corpus compilation that has been neglected in most corpus‐linguistic work: the contrast and relationship between variation according to use and that according to user, or between stylistic sampling categories based on text types and sociolinguistic ones based on speaker/writer identity. Examples of such problems will be derived from the second‐language corpus I am primarily concerned with, the Corpus of East African English, but the principles of socio‐stylistic variation in native and non‐native varieties of English go far beyond this immediate context. They aim at combining two modern quantitatively oriented linguistic subdisciplines to their mutual benefit. After a brief introduction to the ICE project the following points are dealt with: first, the uses of computer‐readable corpora for modern grammars and dictionaries in general (Section 2) and for applied (Section 3) and theoretical (Section 4) research on non‐native varieties of English in particular, then the text type approach applied in ENL corpora so far (Section 5) and the sociolinguistic dimension with its relationship to stylistic variation (Section 6), followed by practical considerations for Third World Englishes (Section 7), and finally a multidimensional approach to socio‐stylistic variation (Section 8) which may be necessary for transferring the ENL‐based methodology of corpus compilation to ESL varieties.

Visit

doi.org

Licenses

http://onlinelibrary.wiley.com/termsAndConditions#vor

Similar

Corpus-based Study on African English VarietiesThe timing of English words by non-native speakersTowards Better Inclusivity: A Diverse Tweet Corpus of English VarietiesBeyond Translating French into English: Experiences of a Non-Native TranslatorTHE PRONUNCIATION OF NON-NATIVE PHONEMES BY EDUCATED HAUSA SPEAKERS OF ENGLISHMispronunciation Detection in Non-native (L2) English with Uncertainty Modeling

Corpus-based Study on African English Varieties

Corpus-based research is more and more used in linguistics. English varieties are used a lot in dail

The timing of English words by non-native speakers

Difficulties with rhythm may be a major cause of lack of intelligibility or naturalness of non-nativ

Towards Better Inclusivity: A Diverse Tweet Corpus of English Varieties

The prevalence of social media presents a growing opportunity to collect and analyse examples of Eng

Beyond Translating French into English: Experiences of a Non-Native Translator

This paper documents a non-native translator’s experience in an academic setting, focusing on the ch

THE PRONUNCIATION OF NON-NATIVE PHONEMES BY EDUCATED HAUSA SPEAKERS OF ENGLISH

This study examines the pronunciation of non-native phonemes by educated Hausa speakers of English,

Mispronunciation Detection in Non-native (L2) English with Uncertainty Modeling

A common approach to the automatic detection of mispronunciation in language learning is to recogniz