Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

Weighted combination of BERT and N-GRAM features for Nuanced Arabic Dialect Identification

Domaine:

natural language processing

Type de record:

paper
Créateur:
Int
Éditeur:
Und
Hôte:avatar
Around the Arab world, different Arabic dialects are spoken by more than 300M persons and are increasingly popular in social media texts. However, Arabic dialects are considered to be low-resource languages, limiting the development of machine-learning-based systems for these dialects. In this paper, we investigate the Arabic dialect identification task, from two perspectives: country-level dialect identification from 21 Arab countries, and province-level dialect identification from 100 provinces. We introduce an unified pipeline of state-of-the-art models, that can handle the two subtasks. Our experimental studies applied to the NADI shared task under the team name BERT-NGRAMS, show promising results both at the country-level (F1-score of 25.99%) and the province-level (F1-score of 6.39%), and thus allow us to be ranked 2nd for the country-level subtask and 1st in the province-level subtask.

Visit

doi.orgunderline.io

Tasks

language identification

Tags

Natural Language Processing

Similaires

Multi-Dialect Arabic BERT for Country-Level Dialect IdentificationOptimizing n‑gram Order of an n‑gram Based Language Identification Algorithm for 68 Written LanguagesQCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual FeaturesN-Gram Based HASSANIYA Dialect ClassificationMohammed2311/Arabic-Dialect-Identificationbashar-talafha/multi-dialect-bert-base-arabic

Multi-Dialect Arabic BERT for Country-Level Dialect Identification

Arabic dialect identification is a complex problem for a number of inherent properties of the langua

Optimizing n‑gram Order of an n‑gram Based Language Identification Algorithm for 68 Written Languages

Language identification technology is widely used in the domains of machine learning and text mining

QCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual Features

The paper describes the QCRI submissions to the task of automatic Arabic dialect classification into 5 Arabic variants, namely Egyptian, Gulf, Levantine, North-African, and Modern Standard Arabic (MSA). The training data is relatively small and is automatically gen

N-Gram Based HASSANIYA Dialect Classification

Mohammed2311/Arabic-Dialect-Identification

This project aims to identify the dialect of Arabic tweets among five dialects: Egypt (EG), Lebanon

bashar-talafha/multi-dialect-bert-base-arabic