Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Improving Cross-Lingual Phonetic Representation of Low-Resource Languages Through Language Similarity Analysis

Domaine:

natural language processing

Type de record:

paper
Créateur:
KimJanKim
Hôte:avatar
This paper examines how linguistic similarity affects cross-lingual phonetic representation in speech processing for low-resource languages, emphasizing effective source language selection. Previous cross-lingual research has used various source languages to enhance performance for the target low-resource language without thorough consideration of selection. Our study stands out by providing an in-depth analysis of language selection, supported by a practical approach to assess phonetic proximity among multiple language families. We investigate how within-family similarity impacts performance in multilingual training, which aids in understanding language dynamics. We also evaluate the effect of using phonologically similar languages, regardless of family. For the phoneme recognition task, utilizing phonologically similar languages consistently achieves a relative improvement of 55.6% over monolingual training, even surpassing the performance of a large-scale self-supervised learning model. Multilingual training within the same language family demonstrates that higher phonological similarity enhances performance, while lower similarity results in degraded performance compared to monolingual training. 10 pages, 5 figures, accepted to ICASSP 2025

Visit

arxiv.org

Tasks

automatic speech recognitionspeech processing

Tags

Audio and Speech ProcessingComputation and LanguageSoundI.2.7; J.5; H.5.5; I.5.4

Similaires

Impact of Language Similarity Metrics on Cross-Lingual NER Performance in Low-Resource LanguagesCross-lingual Transfer Accuracy and Task Similarity in Low-Resource LanguagesCross-lingual Transfer Accuracy in Low-Resource Languages via Task SimilarityLinguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource LanguagesDomain Similarity Impact on Zero-Shot Cross-Lingual Transfer Performance in Low-Resource LanguagesNamed Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Impact of Language Similarity Metrics on Cross-Lingual NER Performance in Low-Resource Languages

Cross-lingual Named Entity Recognition (NER) leverages knowledge transfer between languages to ident

Cross-lingual Transfer Accuracy and Task Similarity in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Cross-lingual Transfer Accuracy in Low-Resource Languages via Task Similarity

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Linguistic Entity Masking to Improve Cross-Lingual Representation of Multilingual Language Models for Low-Resource Languages

Multilingual Pre-trained Language models (multiPLMs), trained on the Masked Language Modelling (MLM)

Domain Similarity Impact on Zero-Shot Cross-Lingual Transfer Performance in Low-Resource Languages

Intermediate-task training---fine-tuning a pretrained model on an intermediate task before fine-tuni

Named Entity Recognition in Low-resource Languages using Cross-lingual distributional word representation

Named Entity Recognition (NER) is a fundamental task in many NLP applications that seek to identify