Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A corpus‐based method for identifying word class in an English lexified extended pidgin

Domain:

natural language processing

Record type:

paper
Creator:
Sar
Publisher:
WILEY
Host:
Abstract This paper outlines an innovative corpus‐based method for establishing word class using as a case study the verbal status of the locative copula in Cameroon Pidgin English. The corpus‐based methodology outlined here uses a combination of distributional, sociolinguistic, and frequency analysis to establish that the Cameroon Pidgin English locative copula, deiy , is a verb and in doing so outlines a blueprint for establishing word class in under described languages. Pidgin creole languages are often under described and identifying the word class of multifunctional words can be particularly challenging. The methodology presented here can resolve such ambiguity by basing the classification on the language in which the word occurs rather than on typological generalisations. This corpus‐based methodology has wide ranging applicability and could be especially useful in establishing word class in lesser described world English varieties and pidgin creole languages.

Visit

doi.org

Languages

Ghanaian Pidgin English

Licenses

http://onlinelibrary.wiley.com/termsAndConditions#vor

Similar

Fanakalo, a Bantu-lexified pidginA spoken corpus of Cameroon Pidgin EnglishExtended Parallel Corpus for Amharic-English Machine TranslationInformation structure in a spoken corpus of Cameroon Pidgin EnglishA Spoken Corpus of Cameroon Pidgin English: pilot studyA word-class tagset for Setswana

Fanakalo, a Bantu-lexified pidgin

Abstract This chapter provides an overview of Fanakalo pidgin, which is worthy o

A spoken corpus of Cameroon Pidgin English

This article reports on the construction of a 240,000-word pilot corpus of spoken Cameroon Pidgin En

Extended Parallel Corpus for Amharic-English Machine Translation

This paper describes the acquisition, preprocessing, segmentation, and alignment of an Amharic-English parallel corpus. It will be helpful for machine translation of a low-resource language, Amharic. We freely released the corpus for research purposes. Furthermore,

Information structure in a spoken corpus of Cameroon Pidgin English

We explore information structure in a spoken corpus of Cameroon Pidgin English, addressing the broad

A Spoken Corpus of Cameroon Pidgin English: pilot study

This resource is a 240,000-word corpus of spoken Cameroon Pidgin English (CPE), a widely-used yet st

A word-class tagset for Setswana