Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

DonorRank: Donor Language Selection for Low-Resource Cross-Lingual Speech Recognition

Domain:

natural language processing

Record type:

paper
Creator:
DhaSriChi
Publisher:
arXiv
Host:avatar
Low-resource automatic speech recognition (ASR) commonly relies on cross-lingual transfer, where models are adapted from higher-resource donor languages. However, selecting donors remains challenging for spontaneous speech from under-resourced language communities, due to linguistic variation, evolving orthographic conventions, and uneven resource availability. We present DonorRank, a learning-to-rank framework for predicting effective donor languages for zero-shot ASR. We evaluate DonorRank on two multilingual speech corpora of Indic and African language families. It accurately predicts donor language rankings and improves donor selection over common heuristics based on genetic similarity or high-resource languages. Beyond improving transfer, we show how DonorRank is a general framework for analyzing donor language selection itself. Our analyses show that the composition of the donor set determines which linguistic cues are useful in predicting successful transfer. We also identify transfer patterns that provide practical guidance for multilingual ASR in low-resource settings. 11 pages, 4 figures, 12 tables

Visit

doi.org

Tasks

automatic speech recognitiontransfer learningspeech processing

Tags

Computation and Language (cs.CL)FOS: Computer and information sciencesI.2.7

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Exploiting Adapters for Cross-lingual Low-resource Speech RecognitionCross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech RecognitionCross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech RecognitionXLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech RecognitionDiscrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource LanguagesTypological Features in Source Language Selection for Cross-Lingual NER on Low-Resource African Languages

Exploiting Adapters for Cross-lingual Low-resource Speech Recognition

Cross-lingual speech adaptation aims to solve the problem of leveraging multiple rich-resource langu

Cross-lingual Data Selection Using Clip-level Acoustic Similarity for Enhancing Low-resource Automatic Speech Recognition

This paper presents a novel donor data selection method to enhance low-resource automatic speech rec

Cross-lingual Embedding Clustering for Hierarchical Softmax in Low-Resource Multilingual Speech Recognition

We present a novel approach centered on the decoding stage of Automatic Speech Recognition (ASR) tha

XLST: Cross-lingual Self-training to Learn Multilingual Representation for Low Resource Speech Recognition

In this paper, we propose a weakly supervised multilingual representation learning framework, called

Discrete Audio Tokens Enhance Cross-Lingual Speech Recognition in Low-Resource Languages

This report synthesises findings from 13 peer-reviewed papers addressing the following research ques

Typological Features in Source Language Selection for Cross-Lingual NER on Low-Resource African Languages

Cross-lingual transfer learning enables NLP for low-resource languages by leveraging labeled data fr