Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Semi-automatic discourse annotation in a low-resource language: Developing a connective lexicon for Nigerian Pidgin

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AssDemMarSch
Publisher:
Und
Host:avatar
Cross-linguistic research on discourse structure and coherence marking requires discourse-annotated corpora and connective lexicons in a large number of languages. However, the availability of such resources is limited, especially for languages for which linguistic resources are scarce in general, such as Nigerian Pidgin. In this study, we demonstrate how a semi-automatic approach can be used to source connectives and their relation senses and develop a discourse-annotated corpus in a low-resource language. Connectives and their relation senses were extracted from a parallel corpus combining automatic (PDTB end-to-end parser) and manual annotations. This resulted in Naija-Lex, a lexicon of discourse connectives in Nigerian Pidgin with English translations. The lexicon shows that the majority of Nigerian Pidgin connectives are borrowed from its English lexifier, but that there are also some connectives that are unique to Nigerian Pidgin.

Visit

doi.orgunderline.io

Tasks

information extraction

Tags

Natural Language ProcessingMachine LearningMachine Learning and Data MiningComputational Linguistics

Similar

Iammteo/Sentiment-analysis-model-for-Nigerian-pidgin-Low-resource-language-AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech TechnologiesLow-Resource Cross-Lingual Adaptive Training for Nigerian Pidginazziimm7/Low-Resource-Language-Translation-Arabic-and-Nigerian-Pidgin-in-the-SpotlightFostering Digital Inclusion for Low-Resource Nigerian Languages: A Case Study of Igbo and Nigerian PidginSemi-automatic news video annotation framework for Arabic text

Iammteo/Sentiment-analysis-model-for-Nigerian-pidgin-Low-resource-language-

This is a machine learning project of a Nigerian Pidgin sentiment analysis model specifically design

AnnoTheia: A Semi-Automatic Annotation Toolkit for Audio-Visual Speech Technologies

More than 7,000 known languages are spoken around the world. However, due to the lack of annotated r

Low-Resource Cross-Lingual Adaptive Training for Nigerian Pidgin

Developing effective spoken language processing systems for low-resource languages poses several cha

azziimm7/Low-Resource-Language-Translation-Arabic-and-Nigerian-Pidgin-in-the-Spotlight

Low-Resource Language Translation: Arabic and Nigerian Pidgin in the Spotlight Project title :Low-R

Fostering Digital Inclusion for Low-Resource Nigerian Languages: A Case Study of Igbo and Nigerian Pidgin

Semi-automatic news video annotation framework for Arabic text