Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Towards Better UD Parsing: Deep Contextualized Word Embeddings, Ensemble, and Treebank Concatenation

Domain:

natural language processing

Record type:

paper
Creator:
CheLiuWanZhe
Host:avatar
This paper describes our system (HIT-SCIR) submitted to the CoNLL 2018 shared task on Multilingual Parsing from Raw Text to Universal Dependencies. We base our submission on Stanford's winning system for the CoNLL 2017 shared task and make two effective extensions: 1) incorporating deep contextualized word embeddings into both the part of speech tagger and parser; 2) ensembling parsers trained with different initialization. We also explore different ways of concatenating treebanks for further improvements. Experimental results on the development data show the effectiveness of our methods. In the final evaluation, our system was ranked first according to LAS (75.84%) and outperformed the other systems by a large margin. System description paper of our system (HIT-SCIR) for the CoNLL 2018 shared task on Universal Dependency parsing, which was ranked first in the LAS evaluation. Fix typos and grammar errors. Add the results of parser without ensemble

Visit

arxiv.org

Tasks

dependency parsingparsing

Tags

Computation and Language

Similar

A UD Treebank for Bohairic CopticA Monolingual Approach to Contextualized Word Embeddings for Mid-Resource LanguagesA Surface-Syntactic UD Treebank for NaijaDetecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like BiasesLow-Resource Parsing with Crosslingual Contextualized RepresentationsTopic Modelling Swahili Using LDA and Contextualized Embeddings

A UD Treebank for Bohairic Coptic

Despite recent advances in digital resources for other Coptic dialects, especially Sahidic, Bohairic

A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages

We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages. We then compare the performance of OSCAR-based and Wik

A Surface-Syntactic UD Treebank for Naija

International audience This paper presents a syntactic treebank for spoken Naija, an

Detecting Emergent Intersectional Biases: Contextualized Word Embeddings Contain a Distribution of Human-like Biases

With the starting point that implicit human biases are reflected in the statistical regularities of

Low-Resource Parsing with Crosslingual Contextualized Representations

Despite advances in dependency parsing, languages with small treebanks still present challenges. We

Topic Modelling Swahili Using LDA and Contextualized Embeddings