Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Combining Automated and Manual Data for Effective Downstream Fine-Tuning of Transformers for Low-Resource Language Applications

Domaine:

natural language processing

Type de record:

paper
Créateur:
AssIsa
Éditeur:
Und
Hôte:avatar
This paper addresses the constraints of down-stream applications of pre-trained language models (PLMs) for low-resource languages. These constraints are pre-train data deficiency preventing a low-resource language from being well represented in a PLM and inaccessibility of high-quality task-specific data annotation that limits task learning. We propose to use automatically labeled texts combined with manually annotated data in a two-stage task fine-tuning approach. The experiments revealed that utilizing such methodology combined with vocabulary adaptation may compensate for the absence of a targeted PLM or the deficiency of manually annotated data. The methodology is validated on the morphological tagging task for the Udmurt language. We publish our best model that achieved 93.25% token accuracy on HuggingFace Hub along with the training code1.

Visit

doi.orgunderline.io

Tasks

part of speech tagging

Tags

Computational LinguisticsLanguage ModelsNatural Language ProcessingArtificial Intelligence

Similaires

Comparison of Intermediate-Task Fine-Tuning and Multilingual Fine-Tuning for Zero-Shot Low-Resource Language AccuracyWhispering in Amharic: Fine-tuning Whisper for Low-resource LanguageFine Tuning Methods for Low-resource LanguagesFine-Tuning SLAM-ASR for Low-Resource Language Speech Recognition with High-Resource Alignmentshruthi0429/BLOOMZ-and-mT5-Fine-Tuning-Optimizing-Large-Language-Models-for-a-Low-Resource-LanguageFine-tuning Multilingual Transformers for Hausa-English Sentiment Analysis

Comparison of Intermediate-Task Fine-Tuning and Multilingual Fine-Tuning for Zero-Shot Low-Resource Language Accuracy

Accuracy of English-language Question Answering (QA) systems has improved significantly in recent ye

Whispering in Amharic: Fine-tuning Whisper for Low-resource Language

This work explores fine-tuning OpenAI's Whisper automatic speech recognition (ASR) model for Amharic

Fine Tuning Methods for Low-resource Languages

The rise of Large Language Models has not been inclusive of all cultures. The models are mostly trai

Fine-Tuning SLAM-ASR for Low-Resource Language Speech Recognition with High-Resource Alignment

Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource

shruthi0429/BLOOMZ-and-mT5-Fine-Tuning-Optimizing-Large-Language-Models-for-a-Low-Resource-Language

# BLOOMZ and mT5 Fine-Tuning: Optimizing Large Language Models for a Low-Resource Language This rep

Fine-tuning Multilingual Transformers for Hausa-English Sentiment Analysis