Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Heidelberg-Boston @ SIGTYP 2024 Shared Task: Enhancing Low-Resource Language Analysis With Character-Aware Hierarchical Transformers

Domain:

natural language processing

Record type:

papersoftwaremodel
Creator:
RieKra
Host:avatar
Historical languages present unique challenges to the NLP community, with one prominent hurdle being the limited resources available in their closed corpora. This work describes our submission to the constrained subtask of the SIGTYP 2024 shared task, focusing on PoS tagging, morphological tagging, and lemmatization for 13 historical languages. For PoS and morphological tagging we adapt a hierarchical tokenization method from Sun et al. (2023) and combine it with the advantages of the DeBERTa-V3 architecture, enabling our models to efficiently learn from every character in the training data. We also demonstrate the effectiveness of character-level T5 models on the lemmatization task. Pre-trained from scratch with limited data, our models achieved first place in the constrained subtask, nearly reaching the performance levels of the unconstrained task's winner. Our code is available at github.com Accepted for publication at the 6th Workshop on Research in Computational Linguistic Typology and Multilingual NLP (SIGTYP-WS) 2024; 11 pages, 1 figure, 9 tables

Visit

arxiv.org

Tasks

parsingpart of speech tagging

Tags

Computation and LanguageI.2.7

Similar

Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?PRiSM: Enhancing Low-Resource Document-Level Relation Extraction with Relation-Aware Score CalibrationAFRISENTI-SEMEVAL SHARED TASK 12: SENTIMENT ANALYSIS FOR 15 LOW-RESOURCE AFRICAN LANGUAGES USING TWITTER DATASETGMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared TaskLIA and ELYADATA systems for the IWSLT 2025 low-resource speech translation shared task

Can Character-based Language Models Improve Downstream Task Performances In Low-Resource And Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural language models, have been mostly demonstrated on at most a couple dozen high- resource languages. Building language mod- els and, more generally, NLP systems for non- standard

Can Character-based Language Models Improve Downstream Task Performance in Low-Resource and Noisy Language Scenarios?

Recent impressive improvements in NLP, largely based on the success of contextual neural la

PRiSM: Enhancing Low-Resource Document-Level Relation Extraction with Relation-Aware Score Calibration

Document-level relation extraction (DocRE) aims to extract relations of all entity pairs in a docume

AFRISENTI-SEMEVAL SHARED TASK 12: SENTIMENT ANALYSIS FOR 15 LOW-RESOURCE AFRICAN LANGUAGES USING TWITTER DATASET

AFRISENTI-SEMEVAL SHARED TASK 12: SENTIMENT ANALYSIS FOR 15 LOW-RESOURCE AFRICAN LANGUAGES USING TWITTER DATASET

Poster presented at the Deep Learning Indaba 2022 by Shamsuddeen Muhammad

GMU Systems for the IWSLT 2025 Low-Resource Speech Translation Shared Task

This paper describes the GMU systems for the IWSLT 2025 low-resource speech translation shared task.

LIA and ELYADATA systems for the IWSLT 2025 low-resource speech translation shared task

International audience

In this paper, we present the approach and system setu