Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Self-Training for Unsupervised Parsing with PRPN

Domaine:

natural language processing

Type de record:

paper
Créateur:
MohKanBow
Hôte:avatar
Neural unsupervised parsing (UP) models learn to parse without access to syntactic annotations, while being optimized for another task like language modeling. In this work, we propose self-training for neural UP models: we leverage aggregated annotations predicted by copies of our model as supervision for future copies. To be able to use our model's predictions during training, we extend a recent neural UP architecture, the PRPN (Shen et al., 2018a) such that it can be trained in a semi-supervised fashion. We then add examples with parses predicted by our model to our unlabeled UP training data. Our self-trained model outperforms the PRPN by 8.1% F1 and the previous state of the art by 1.6% F1. In addition, we show that our architecture can also be helpful for semi-supervised parsing in ultra-low-resource settings. Accepted for publication at the 16th International Conference on Parsing Technologies (IWPT), 2020

Visit

arxiv.org

Tasks

parsing

Tags

Computation and Language

Similaires

Multilingual Constituency Parsing with Self-Attention and Pre-TrainingA Hierarchical Transformer for Unsupervised ParsingSpeaker Diarization With Unsupervised Training FrameworkUnsupervised Pidgin Text Generation By Pivoting English Data and Self-TrainingJoint Unsupervised and Supervised Training for Multilingual ASRComparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

Multilingual Constituency Parsing with Self-Attention and Pre-Training

We show that constituency parsing benefits from unsupervised pre-training across a variety of langua

A Hierarchical Transformer for Unsupervised Parsing

The underlying structure of natural language is hierarchical; words combine into phrases, which in turn form clauses. An awareness of this hierarchical structure can aid machine learning models in performing many linguistic tasks. However, most such models just pro

Speaker Diarization With Unsupervised Training Framework

International audience This paper investigates single and cross-show diarization base

Unsupervised Pidgin Text Generation By Pivoting English Data and Self-Training

West African Pidgin English is a language that is significantly spoken in West Africa, consisting of at least 75 million speakers. Nevertheless, proper machine translation systems and relevant NLP datasets for pidgin English are virtually absent. In this work, we d

Joint Unsupervised and Supervised Training for Multilingual ASR

Self-supervised training has shown promising gains in pretraining models and facilitating the downst

Comparing Self-Supervised Pre-Training and Semi-Supervised Training for Speech Recognition in Languages with Weak Language Models

International audience This paper investigates the potential of improving a hybrid au