Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

A Deep Learning Approach for the Romanized Tunisian Dialect Identification

Domain:

natural language processing

Record type:

paper
Creator:
JihHadEmnAhm
Publisher:
Zar
Host:
Language identification is an important task in natural language processing that consists in determining the language of a given text. It has increasingly picked the interest of researchers for the past few years, especially for code-switching informal textual content. In this paper, we focus on the identification of the Romanized user-generated Tunisian dialect on the social web. We segment and annotate a corpus extracted from social media and propose a deep learning approach for the identification task. We use a Bidirectional Long Short-Term Memory neural network with Conditional Random Fields decoding (BLSTM-CRF). For word embeddings, we combine word-character BLSTM vector representation and Fast Text embeddings that takes into consideration character n-gram features. The overall accuracy obtained is 98.65%.

Visit

doi.org

Tasks

language identification

Languages

Arabic, Tunisian Spoken

Similar

Word-Level Identification of Romanized Tunisian DialectSocial Media Sentiment Classification for Tunisian Dialect: A Deep Learning ApproachArabic Transliteration of Romanized Tunisian Dialect Text: A Preliminary InvestigationRomanized Berber and Romanized Arabic Automatic Language Identification Using Machine LearningA Medical Chatbot for Tunisian Dialect using a Rule-Based and Machine Learning ApproachDeep learning approach for Tunisian hate Speech detection on Facebook

Word-Level Identification of Romanized Tunisian Dialect

Social Media Sentiment Classification for Tunisian Dialect: A Deep Learning Approach

Arabic Transliteration of Romanized Tunisian Dialect Text: A Preliminary Investigation

Romanized Berber and Romanized Arabic Automatic Language Identification Using Machine Learning

The identification of the language of text/speech input is the first step to be able to properly do any language-dependent natural language processing. The task is called Automatic Language Identification (ALI). Being a well-studied field since early 1960{'}s, vari

A Medical Chatbot for Tunisian Dialect using a Rule-Based and Machine Learning Approach

Deep learning approach for Tunisian hate Speech detection on Facebook