Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Convolutional Neural Networks and Language Embeddings for End-to-End Dialect Recognition

Domain:

natural language processing

Record type:

paper
Creator:
ShoAli, AhmedGla
Host:avatar
Dialect identification (DID) is a special case of general language identification (LID), but a more challenging problem due to the linguistic similarity between dialects. In this paper, we propose an end-to-end DID system and a Siamese neural network to extract language embeddings. We use both acoustic and linguistic features for the DID task on the Arabic dialectal speech dataset: Multi-Genre Broadcast 3 (MGB-3). The end-to-end DID system was trained using three kinds of acoustic features: Mel-Frequency Cepstral Coefficients (MFCCs), log Mel-scale Filter Bank energies (FBANK) and spectrogram energies. We also investigated a dataset augmentation approach to achieve robust performance with limited data resources. Our linguistic feature research focused on learning similarities and dissimilarities between dialects using the Siamese network, so that we can reduce feature dimensionality as well as improve DID performance. The best system using a single feature set achieves 73% accuracy, while a fusion system using multiple features yields 78% on the MGB-3 dialect test set consisting of 5 dialects. The experimental results indicate that FBANK features achieve slightly better results than MFCCs. Dataset augmentation via speed perturbation appears to add significant robustness to the system. Although the Siamese network with language embeddings did not achieve as good a result as the end-to-end DID system, the two approaches had good synergy when combined together in a fused system. Speaker Odyssey 2018, The Speaker and Language Recognition Workshop

Visit

arxiv.org

Tasks

language identificationspeech processing

Tags

SoundAudio and Speech Processing

Similar

Graph Neural Networks for end-to-end information extraction from handwritten documentsEnd-to-End Continuous Ethiopia Sign Language RecognitionEnd-To-End Continuous Ethiopian Sign Language RecognitionBi-directional Recurrent End-to-End Neural Network Classifier for Spoken Arab Digit RecognitionA Streaming End-to-End Speech Recognition Approach Based on WeNet for Tibetan Amdo DialectOkwuGbé: End-to-End Speech Recognition for Fon and Igbo

Graph Neural Networks for end-to-end information extraction from handwritten documents

Graph Neural Networks for end-to-end information extraction from handwritten documents

Poster presented at the Deep Learning Indaba 2023 by yessine khanfir

End-to-End Continuous Ethiopia Sign Language Recognition

This paper tackles the challenge of continuous Ethiopian Sign Language (EthSL) recogn

End-To-End Continuous Ethiopian Sign Language Recognition

Bi-directional Recurrent End-to-End Neural Network Classifier for Spoken Arab Digit Recognition

International audience —Automatic Speech Recognition can be considered as a transcrip

A Streaming End-to-End Speech Recognition Approach Based on WeNet for Tibetan Amdo Dialect

OkwuGbé: End-to-End Speech Recognition for Fon and Igbo

Language is inherent and compulsory for human communication. Whether expressed in a written or spoken way, it ensures understanding between people of the same and different regions. With the growing awareness and effort to include more low-resourced languages in NL