Logo Lanfrica

<div class="page" title="Page 1"> <div class="layoutArea"> <div class="column"> Exploring Embedding Visualization Tools for Low-Resource Machine Translation </div> </div> </div>

Domaine:

natural language processing

Type de record:

paper
Créateur:
VarKev
Éditeur:
MDP
Hôte:
We investigate the potential of embedding visualization tools, specifically Latent Lab, for supporting machine translation tasks with a focus on low-resource languages. Through empirical analysis of bilingual datasets comprising high-resource (Spanish, German, French) and low-resource (Dzongkha) languages, we explore how embedding visualization can facilitate the identification of cross-lingual token correspondences and dataset quality assessment. Our findings suggest that while embedding visualization tools offer promising capabilities for dataset preparation and quality control in machine translation pipelines, several technical modifications are necessary to fully realize their potential. We identify some challenges and propose concrete modifications to both the datasets and visualization tools to address these limitations.

Similaires