Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

On the Importance of Subword Information for Morphological Tasks in Truly Low-Resource Languages

Domaine:

natural language processing

Type de record:

paper
Créateur:
ZhuHeiVulStr
Hôte:avatar
Recent work has validated the importance of subword information for word representation learning. Since subwords increase parameter sharing ability in neural models, their value should be even more pronounced in low-data regimes. In this work, we therefore provide a comprehensive analysis focused on the usefulness of subwords for word representation learning in truly low-resource scenarios and for three representative morphological tasks: fine-grained entity typing, morphological tagging, and named entity recognition. We conduct a systematic study that spans several dimensions of comparison: 1) type of data scarcity which can stem from the lack of task-specific training data, or even from the lack of unannotated data required to train word embeddings, or both; 2) language type by working with a sample of 16 typologically diverse languages including some truly low-resource ones (e.g. Rusyn, Buryat, and Zulu); 3) the choice of the subword-informed word representation method. Our main results show that subword-informed models are universally useful across all language types, with large gains over subword-agnostic embeddings. They also suggest that the effective use of subwords largely depends on the language (type) and the task at hand, as well as on the amount of available data for training the embeddings and task-based models, where having sufficient in-task data is a more critical requirement. CONLL2019

Visit

arxiv.org

Tags

Computation and Language

Similaires

Multilingual Projection for Parsing Truly Low-Resource LanguagesCombining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource LanguagesCross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource LanguagesTaxonomic Loss for Morphological Glossing of Low-Resource LanguagesMorphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource LanguagesWhat is the impact of intermediate-task training on English tasks for cross-lingual transfer in low-resource languages

Multilingual Projection for Parsing Truly Low-Resource Languages

We propose a novel approach to cross-lingual part-of-speech tagging and dependency parsing for truly

Combining Pretrained High-Resource Embeddings and Subword Representations for Low-Resource Languages

The contrast between the need for large amounts of data for current Natural Language Processing (NLP) techniques, and the lack thereof, is accentuated in the case of African languages, most of which are considered low-resource. To help circumvent this issue, we exp

Cross-Lingual Dependency Parsing with Late Decoding for Truly Low-Resource Languages

In cross-lingual dependency annotation projection, information is often lost during transfer because

Taxonomic Loss for Morphological Glossing of Low-Resource Languages

Morpheme glossing is a critical task in automated language documentation and can benefit other downs

Morphology-aware Subword Segmentation for Zero-shot Cross-lingual Transfer in Low-resource Languages

Multilingual modelling can improve machine translation for low-resource languages, partly through sh

What is the impact of intermediate-task training on English tasks for cross-lingual transfer in low-resource languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni