Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

158 - ViDia2Std: A Parallel Corpus and Methods for Low-Resource Vietnamese Dialect-to-Standard Translation

Domain:

natural language processing

Record type:

datasetpaper
Creator:
AssNguTa,Van
Publisher:
Und
Host:avatar
Vietnamese exhibits extensive dialectal variation, posing challenges for NLP systems trained predominantly on standard Vietnamese. Such systems often underperform on dialectal inputs, especially from underrepresented Central and Southern regions. Previous work on dialect normalization has focused narrowly on Central-to-Northern dialect transfer using synthetic data and limited dialectal diversity. These efforts exclude Southern varieties and intra-regional variants within the North. We introduce ViDia2Std, the first manually annotated parallel corpus for dialect-to-standard Vietnamese translation covering all 63 provinces. Unlike prior datasets, ViDia2Std includes diverse dialects from Central, Southern, and non-standard Northern regions often absent from existing resources, making it the most dialectally inclusive corpus to date. The dataset consists of over 13,000 sentence pairs sourced from real-world Facebook comments and annotated by native speakers across all three dialect regions. To assess annotation consistency, we define a semantic mapping agreement metric that accounts for synonymous standard mappings across annotators. Based on this criterion, we report agreement rates of 86% (North), 82% (Central), and 85% (South). We benchmark several sequence-to-sequence models on ViDia2Std. mBART-large-50 achieves the best results (BLEU 0.8166, ROUGE-L 0.9384, METEOR 0.8925), while ViT5-base offers competitive performance with fewer parameters. ViDia2Std demonstrates that dialect normalization substantially improves downstream tasks, highlighting the need for dialect-aware resources in building robust Vietnamese NLP systems.

Visit

doi.orgunderline.io

Tasks

machine translationtext normalization

Similar

Improving neural machine translation for low resource languages through non-parallel corpora: a case study of Egyptian dialect to modern standard Arabic translationOpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource LanguagesIWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel CorpusFiltered Pseudo-parallel Corpus Improves Low-resource Neural Machine TranslationA Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation BenchmarksMachine Translation Experiments on PADIC: A Parallel Arabic DIalect Corpus

Improving neural machine translation for low resource languages through non-parallel corpora: a case study of Egyptian dialect to modern standard Arabic translation

Abstract Machine translation for low-resource languages poses significant challenges, primarily due

OpenWHO: A Document-Level Parallel Corpus for Health Translation in Low-Resource Languages

In machine translation (MT), health is a high-stakes domain characterised by widespread deployment a

IWSLT2022 - Low-resource Speech Translation Track: Tamasheq-French Parallel Corpus

Repository for sharing the data in the Tamasheq language, one of the languages for the low-resource speech translation track at IWSLT 2022.

Filtered Pseudo-parallel Corpus Improves Low-resource Neural Machine Translation

Large-scale parallel corpora are essential for training high-quality machine translation systems; ho

A Low-Resource English–Hassaniya Parallel Corpus with Neural Machine Translation Benchmarks

Machine Translation Experiments on PADIC: A Parallel Arabic DIalect Corpus

We present in this paper PADIC, a Parallel Arabic DIalect Corpus we built from scratch, then we conducted experiments on cross-dialect Arabic machine translation. PADIC is composed of dialects from both the Maghreb and the Middle-East. Each dialect has been aligned