Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Investigating Data Sharing in Speech Recognition for an Under-Resourced Language: The Case of Algerian Dialect

Domaine:

natural language processing

Type de record:

paper
Créateur:
MohKam
Éditeur:
AIR
Hôte:
The Arabic language has many varieties, including its standard form, Modern Standard Arabic (MSA), and its spoken forms, namely the dialects. Those dialects are representative examples of under-resourced languages for which automatic speech recognition is considered as an unresolved issue. To address this issue, we recorded several hours of spoken Algerian dialect and used them to train a baseline model. This model was boosted afterwards by taking advantage of other languages that impact this dialect by integrating their data in one large corpus and by investigating three approaches: multilingual training, multitask learning and transfer learning. The best performance was achieved using a limited and balanced amount of acoustic data from each additional language, as compared to the data size of the studied dialect. This approach led to an improvement of 3.8% in terms of word error rate in comparison to the baseline system trained only on the dialect data.

Visit

doi.org

Tasks

automatic speech recognitionspeech processingtransfer learning

Languages

Arabic, Algerian Spoken

Similaires

Speech recognition for under-resourced languages: Data sharing in hidden Markov model systemsAutomatic speech recognition for an under-resourced language - amharicSub-word Based End-to-End Speech Recognition for an Under-Resourced Language: AmharicImproving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data AugmentationEnhancement of spoken digits recognition for under-resourced languages: case of Algerian and Moroccan dialectsAutomatic speech recognition for under-resourced languages: A survey

Speech recognition for under-resourced languages: Data sharing in hidden Markov model systems

For purposes of automated speech recognition in under-resourced environments, t

Automatic speech recognition for an under-resourced language - amharic

Sub-word Based End-to-End Speech Recognition for an Under-Resourced Language: Amharic

In this work, we focused on end-to-end speech recognition for less-resourced language, Amharic. The result can be integrated with other tasks such as spoken content retrieval. We explored three models, which consist of Convolutional Neural Networks, Recurrent Neura

Improving Speech Recognition for Under-resourced Languages Utilizing Audio-codecs for Data Augmentation

Presenter: Nirayo Hailu Gebreegziabher, Ingo Siegert, Andreas Nürnberger, MMSP 2020, Virtual Event,

Enhancement of spoken digits recognition for under-resourced languages: case of Algerian and Moroccan dialects

Automatic speech recognition for under-resourced languages: A survey

(Impact-F 1.28 estim. in 2012) International audience no abstract