Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

SIGMORPHON 2020 Shared Task 0: Typologically Diverse Morphological Inflection

Domaine:

natural language processing

Type de record:

paperdataset
Créateur:
VylWhiSalMie
Éditeur:
arXiv
Hôte:avatar
A broad goal in natural language processing (NLP) is to develop a system that has the capacity to process any natural language. Most systems, however, are developed using data from just one language such as English. The SIGMORPHON 2020 shared task on morphological reinflection aims to investigate systems' ability to generalize across typologically distinct languages, many of which are low resource. Systems were developed using data from 45 languages and just 5 language families, fine-tuned with data from an additional 45 languages and 10 language families (13 in total), and evaluated on all 90 languages. A total of 22 systems (19 neural) from 10 teams were submitted to the task. All four winning systems were neural (two monolingual transformers and two massively multilingual RNN-based models with gated attention). Most teams demonstrate utility of data hallucination and augmentation, ensembles, and multilingual training for low-resource languages. Non-neural learners and manually designed grammars showed competitive and even superior performance on some languages (such as Ingrian, Tajik, Tagalog, Zarma, Lingala), especially with very limited data. Some language families (Afro-Asiatic, Niger-Congo, Turkic) were relatively easy for most systems and achieved over 90% mean accuracy while others were more challenging. 39 pages, SIGMORPHON

Visit

doi.orgarxiv.org

Languages

LingalaZarma

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesFOS: Computer and information sciences

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

CMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in MorphologyExploring Linguistic Probes for Morphological InflectionIntermediate-Task Training on Typologically Diverse Sources for Zero-Shot Accuracy in Low-Resource Languages on XTREME-RImproving Low-Resource Morphological Inflection via Self-Supervised ObjectivesMasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African LanguagesCross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuning

CMU-01 at the SIGMORPHON 2019 Shared Task on Crosslinguality and Context in Morphology

This paper presents the submission by the CMU-01 team to the SIGMORPHON 2019 task 2 of Morphological

Exploring Linguistic Probes for Morphological Inflection

Modern work on the cross-linguistic computational modeling of morphological inflection has typically

Intermediate-Task Training on Typologically Diverse Sources for Zero-Shot Accuracy in Low-Resource Languages on XTREME-R

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Improving Low-Resource Morphological Inflection via Self-Supervised Objectives

Self-supervised objectives have driven major advances in NLP by leveraging large-scale unlabeled dat

MasakhaPOS: Part-of-Speech Tagging for Typologically Diverse African Languages

In this paper, we present MasakhaPOS, the largest part-of-speech (POS) dataset for 20 typologically

Cross-lingual Euphemism Detection via Typologically Diverse Intermediate Fine-tuning

Euphemisms are culturally variable and often ambiguous, posing challenges for language models, espec