Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

GlossAssist -- A Tool to Simplify Corpus Creation and Study the Effect of NLP Models in Low-Resource Documentation Settings

Domaine:

natural language processing

Type de record:

papersoftwaredataset
Créateur:
ShaBucPal
Hôte:avatar
Interlinear glossed text (IGT) is the standard format for linguistic annotation in language documentation. Producing it manually, however, is often slow and costly. Automated glossing systems have improved substantially in recent years, but adoption among field linguists remains limited. Existing tools are designed to be evaluated rather than used, offering no interpretable path for correction or the incorporation of linguistic expertise back into model behavior. We present GlossAssist, a glossing tool built around the retrieval-based architecture of CWoMP (Contrastive Word-Morpheme Pre-training), which grounds predictions in a mutable lexicon of learned morpheme representations. In conjunction with CWoMP, our system treats each correction by an annotator as part of an active learning setting, which expands the lexicon and improves future predictions without having to retrain the model. In this paper, we present our interface and argue that this feedback loop should be treated as a design requirement for NLP tools aimed at documentary linguists. 6 pages, 3 figures

Visit

arxiv.org

Tags

Computation and LanguageHuman-Computer Interaction

Similaires

Improving Resource Creation for Low-Resource Languages using NLP MethodsA Sheng Phishing Corpus for Low-Resource Cybersecurity NLPCorpus Voice Dataset Creation in Low-Resource Contexts: A systematic reviewThe Malawi Developmental Assessment Tool (MDAT): The Creation, Validation, and Reliability of a Tool to Assess Child Development in Rural African SettingsEmpirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-resource SettingssaltPAD: A New Analytical Tool for Monitoring Salt Iodization in Low Resource Settings

Improving Resource Creation for Low-Resource Languages using NLP Methods

To digitize existing high-quality text belonging to a certain low-resource language, we are often fa

A Sheng Phishing Corpus for Low-Resource Cybersecurity NLP

 

This dataset, the Sheng-English

Corpus Voice Dataset Creation in Low-Resource Contexts: A systematic review

Voice corpora are fundamental resources for developing speech technologies, such as automatic speech

The Malawi Developmental Assessment Tool (MDAT): The Creation, Validation, and Reliability of a Tool to Assess Child Development in Rural African Settings

Background

Although 80% of children with disabilities live in developing countries,

Empirical Evaluation of Sequence-to-Sequence Models for Word Discovery in Low-resource Settings

Since Bahdanau et al. [1] first introduced attention for neural machine translation, most sequence-t

saltPAD: A New Analytical Tool for Monitoring Salt Iodization in Low Resource Settings

We created a paper test card that measures a common iodizing agent, iodate, in salt. To test the ana