Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Corpus-Guided Contrast Sets for Morphosyntactic Feature Detection in Low-Resource English Varieties

Domaine:

natural language processing

Type de record:

modeldataset
Créateur:
MasNeaGreen, LisaO'C
Éditeur:
arXiv
Hôte:avatar
The study of language variation examines how language varies between and within different groups of speakers, shedding light on how we use language to construct identities and how social contexts affect language use. A common method is to identify instances of a certain linguistic feature - say, the zero copula construction - in a corpus, and analyze the feature's distribution across speakers, topics, and other variables, to either gain a qualitative understanding of the feature's function or systematically measure variation. In this paper, we explore the challenging task of automatic morphosyntactic feature detection in low-resource English varieties. We present a human-in-the-loop approach to generate and filter effective contrast sets via corpus-guided edits. We show that our approach improves feature detection for both Indian English and African American English, demonstrate how it can assist linguistic research, and release our fine-tuned models for use by other researchers. Field Matters Workshop at COLING 2022

Visit

doi.orgarxiv.org

Tags

Computation and Language (cs.CL)Artificial Intelligence (cs.AI)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution Non Commercial Share Alike 4.0 Internationalhttps://creativecommons.org/licenses/by-nc-sa/4.0/legalcode

Similaires

MOSAIKS Feature Sets for ZambiaFeature sets from ZI dataset.A morphosyntactic approach to language contact in African varieties of English. Studia NeophilologicaLow resource Twi-English parallel corpus for machine translation in multiple domains (Twi-2-ENG)DATASHI: A Parallel English-Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language ProcessingMorphosyntactic borrowing in closely related varieties

MOSAIKS Feature Sets for Zambia

These two features sets cover the Survey Enumeration Areas and the entire country of Zambia. They we

Feature sets from ZI dataset.

Urbanization and industrialization have led to a significant increase in air pollution, posi

A morphosyntactic approach to language contact in African varieties of English. Studia Neophilologica

This study aims to widen our knowledge of substrate influence and language transfer in World English

Low resource Twi-English parallel corpus for machine translation in multiple domains (Twi-2-ENG)

Abstract Although Ghana does not have one unique language for its citizens, the Twi dialect stands

DATASHI: A Parallel English-Tashlhiyt Corpus for Orthography Normalization and Low-Resource Language Processing

DATASHI is a new parallel English-Tashlhiyt corpus that fills a critical gap in computational resour

Morphosyntactic borrowing in closely related varieties

Abstract The paper examines contact-induced morphosyntactic change in Swahili, w