Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Unsupervised Cross-Lingual Part-of-Speech Tagging with Monolingual Corpora Only

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
Zhe
Hôte:avatar
Due to the scarcity of part-of-speech annotated data, existing studies on low-resource languages typically adopt unsupervised approaches for POS tagging. Among these, POS tag projection with word alignment method transfers POS tags from a high-resource source language to a low-resource target language based on parallel corpora, making it particularly suitable for low-resource language settings. However, this approach relies heavily on parallel corpora, which are often unavailable for many low-resource languages. To overcome this limitation, we propose a fully unsupervised cross-lingual part-of-speech(POS) tagging framework that relies solely on monolingual corpora by leveraging unsupervised neural machine translation(UNMT) system. This UNMT system first translates sentences from a high-resource language into a low-resource one, thereby constructing pseudo-parallel sentence pairs. Then, we train a POS tagger for the target language following the standard projection procedure based on word alignments. Moreover, we propose a multi-source projection technique to calibrate the projected POS tags on the target side, enhancing to train a more effective POS tagger. We evaluate our framework on 28 language pairs, covering four source languages (English, German, Spanish and French) and seven target languages (Afrikaans, Basque, Finnis, Indonesian, Lithuanian, Portuguese and Turkish). Experimental results show that our method can achieve performance comparable to the baseline cross-lingual POS tagger with parallel sentence pairs, and even exceeds it for certain target languages. Furthermore, our proposed multi-source projection technique further boosts performance, yielding an average improvement of 1.3% over previous methods. 16 pages, 6 figures, 7 tables, under review

Visit

arxiv.org

Tasks

part of speech tagging

Languages

Afrikaans

Tags

Computation and Language

Similaires

Zero Resource Cross-Lingual Part Of Speech Tagging Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource ScenariosSetswana Part of Speech TaggingMouhamedkhlifi/Part-of-speech-taggingnaftalindeapo/Part-of-speech-tagging-with-hidden-Markov-modelsPart of Speech Tagging for Amharic

Zero Resource Cross-Lingual Part Of Speech Tagging

Part of speech tagging in zero-resource settings can be an effective approach for low-resource langu

Unsupervised Morphological Segmentation and Part-of-Speech Tagging for Low-Resource Scenarios

With the high cost of manually labeling data and the increasing interest in low-resource languages,

Setswana Part of Speech Tagging

Mouhamedkhlifi/Part-of-speech-tagging

Training xlmroberta model on 20 typologically diverse african languages to classify 14 parts of spee

naftalindeapo/Part-of-speech-tagging-with-hidden-Markov-models

This project aims to develop a part-of-speech tagger for Afrikaans, a South African language, using

Part of Speech Tagging for Amharic

International audience