Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

MIT-QCRI Arabic Dialect Identification System for the 2017 Multi-Genre Broadcast Challenge

Domaine:

natural language processing

Type de record:

paper
Créateur:
ShoAli, AhmedGla
Hôte:avatar
In order to successfully annotate the Arabic speech con- tent found in open-domain media broadcasts, it is essential to be able to process a diverse set of Arabic dialects. For the 2017 Multi-Genre Broadcast challenge (MGB-3) there were two possible tasks: Arabic speech recognition, and Arabic Dialect Identification (ADI). In this paper, we describe our efforts to create an ADI system for the MGB-3 challenge, with the goal of distinguishing amongst four major Arabic dialects, as well as Modern Standard Arabic. Our research fo- cused on dialect variability and domain mismatches between the training and test domain. In order to achieve a robust ADI system, we explored both Siamese neural network models to learn similarity and dissimilarities among Arabic dialects, as well as i-vector post-processing to adapt domain mismatches. Both Acoustic and linguistic features were used for the final MGB-3 submissions, with the best primary system achieving 75% accuracy on the official 10hr test set. Submitted to the 2017 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU 2017)

Visit

arxiv.org

Tasks

language identificationspeech processing

Tags

Computation and LanguageMachine LearningSound

Similaires

LSTM-TDNN with convolutional front-end for Dialect Identification in the 2019 Multi-Genre Broadcast ChallengeThe MGB-2 Challenge: Arabic Multi-Dialect Broadcast Media RecognitionMulti-view Dimensionality Reduction for Dialect Identification of Arabic Broadcast SpeechQCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual FeaturesMulti-Dialect Arabic BERT for Country-Level Dialect IdentificationQCRI Arabic Dialects Identification (QADI) Corpus

LSTM-TDNN with convolutional front-end for Dialect Identification in the 2019 Multi-Genre Broadcast Challenge

This paper presents a novel Dialect Identification (DID) system developed for the Fifth Edition of t

The MGB-2 Challenge: Arabic Multi-Dialect Broadcast Media Recognition

This paper describes the Arabic Multi-Genre Broadcast (MGB-2) Challenge for SLT-2016. Unlike last ye

Multi-view Dimensionality Reduction for Dialect Identification of Arabic Broadcast Speech

In this work, we present a new Vector Space Model (VSM) of speech utterances for the task of spoken

QCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual Features

The paper describes the QCRI submissions to the task of automatic Arabic dialect classification into 5 Arabic variants, namely Egyptian, Gulf, Levantine, North-African, and Modern Standard Arabic (MSA). The training data is relatively small and is automatically gen

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

Arabic dialect identification is a complex problem for a number of inherent properties of the langua

QCRI Arabic Dialects Identification (QADI) Corpus

QCRI Arabic Dialects Identification (QADI) is a Country-level Arabic dialects identification (DI) dataset. It provides a collection for benchmarking DI task.The dataset contains 540,590 tweets from 18 Arab countries.