Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Filling the Gap for Uzbek: Creating Translation Resources for Southern Uzbek

Domaine:

natural language processing

Type de record:

paperdatasetmodel
Créateur:
MamAraShoIno
Hôte:avatar
Southern Uzbek (uzs) is a Turkic language variety spoken by around 5 million people in Afghanistan and differs significantly from Northern Uzbek (uzn) in phonology, lexicon, and orthography. Despite the large number of speakers, Southern Uzbek is underrepresented in natural language processing. We present new resources for Southern Uzbek machine translation, including a 997-sentence FLORES+ dev set, 39,994 parallel sentences from dictionary, literary, and web sources, and a fine-tuned NLLB-200 model (lutfiy). We also propose a post-processing method for restoring Arabic-script half-space characters, which improves handling of morphological boundaries. All datasets, models, and tools are released publicly to support future work on Southern Uzbek and other low-resource languages.

Visit

arxiv.org

Tasks

machine translation

Tags

Computation and Language

Similaires

KITAABI FOR UZBEK LANGUAGE LEARNERSUzbekTagger: The rule-based POS tagger for Uzbek languagePRINCIPLES OF USING LINGUISTIC RESOURCES IN THE SENTIMENT ANALYSIS PROCESS OF UZBEK TEXTSAn Uzbek Medical-Domain Dataset for Aspect-Based Sentiment AnalysisAn Annotated Corpus of Uzbek Business Reviews for Aspect-Based Sentiment AnalysisUzbek text summarization based on TF-IDF

KITAABI FOR UZBEK LANGUAGE LEARNERS

Kitaabi is an innovative online platform designed to facilitate the language-learning process, offer

UzbekTagger: The rule-based POS tagger for Uzbek language

This research paper presents a part-of-speech (POS) annotated dataset and tagger tool for the low-re

PRINCIPLES OF USING LINGUISTIC RESOURCES IN THE SENTIMENT ANALYSIS PROCESS OF UZBEK TEXTS

This article deeply explores the principles of utilizing linguistic resources in the sentiment analy

An Uzbek Medical-Domain Dataset for Aspect-Based Sentiment Analysis

UzMedSentiment is a manually annotated Uzbek medical-domain dataset designed for sentiment classific

An Annotated Corpus of Uzbek Business Reviews for Aspect-Based Sentiment Analysis

This dataset contains 5,038 annotated business reviews designed for Aspect-Based Sentiment Analysis

Uzbek text summarization based on TF-IDF

The volume of information is increasing at an incredible rate with the rapid development of the Inte