Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Grounded Unsupervised Universal Part-of-Speech Tagger for Low-Resource Languages

Domain:

natural language processing

Record type:

papersoftware
Creator:
CarLinJi,May
Host:avatar
Unsupervised part of speech (POS) tagging is often framed as a clustering problem, but practical taggers need to \textit{ground} their clusters as well. Grounding generally requires reference labeled data, a luxury a low-resource language might not have. In this work, we describe an approach for low-resource unsupervised POS tagging that yields fully grounded output and requires no labeled training data. We find the classic method of Brown et al. (1992) clusters well in our use case and employ a decipherment-based approach to grounding. This approach presumes a sequence of cluster IDs is a `ciphertext' and seeks a POS tag-to-cluster ID mapping that will reveal the POS sequence. We show intrinsically that, despite the difficulty of the task, we obtain reasonable performance across a variety of languages. We also show extrinsically that incorporating our POS tagger into a name tagger leads to state-of-the-art tagging performance in Sinhalese and Kinyarwanda, two languages with nearly no labeled POS data available. We further demonstrate our tagger's utility by incorporating it into a true `zero-resource' variant of the Malopa (Ammar et al., 2016) dependency parser model that removes the current reliance on multilingual resources and gold POS tags for new languages. Experiments show that including our tagger makes up much of the accuracy lost when gold POS tags are unavailable. NAACL-HLT 2019, 12 pages, code available at github.com

Visit

arxiv.org

Tasks

information extractionnamed entity recognitionpart of speech tagging

Languages

Hamer-BannaKinyarwanda

Tags

Computation and LanguageArtificial IntelligenceMachine Learning

Similar

Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource ScenariosVisually Grounded Speech Models for Low-resource Languages and Cognitive ModellingAn Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource LanguagesSepedi Part of Speech TaggerCopticScriptorium/tagger-part-of-speechABISOLAP/Part-of-Speech-Tagger-

Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios

With the high cost of manually labeling data and the increasing interest in low-resource languages,

Visually Grounded Speech Models for Low-resource Languages and Cognitive Modelling

This dissertation examines visually grounded speech (VGS) models that learn from unlabelled speech p

An Unsupervised Probability Model for Speech-to-Translation Alignment of Low-Resource Languages

For many low-resource languages, spoken language resources are more likely to be annotated with tran

Sepedi Part of Speech Tagger

Sesotho sa Leboa part of speech statistical tagger compiled with stochastic tagger of Helmut Scmidt

CopticScriptorium/tagger-part-of-speech

Part of speech tagger for Sahidic Coptic TreeTagger part-of-speech tagging models for Sahidic Copti

ABISOLAP/Part-of-Speech-Tagger-

Part of Speech Tagger For Yoruba Language using Artificial Neural Network Algorithm # Part-of-Speec