Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Computational Approaches to Arabic-English Code-Switching

Domain:

natural language processing

Record type:

paperdataset
Creator:
Sab
Host:avatar
Natural Language Processing (NLP) is a vital computational method for addressing language processing, analysis, and generation. NLP tasks form the core of many daily applications, from automatic text correction to speech recognition. While significant research has focused on NLP tasks for the English language, less attention has been given to Modern Standard Arabic and Dialectal Arabic. Globalization has also contributed to the rise of Code-Switching (CS), where speakers mix languages within conversations and even within individual words (intra-word CS). This is especially common in Arab countries, where people often switch between dialects or between dialects and a foreign language they master. CS between Arabic and English is frequent in Egypt, especially on social media. Consequently, a significant amount of code-switched content can be found online. Such code-switched data needs to be investigated and analyzed for several NLP tasks to tackle the challenges of this multilingual phenomenon and Arabic language challenges. No work has been done before for several integral NLP tasks on Arabic-English CS data. In this work, we focus on the Named Entity Recognition (NER) task and other tasks that help propose a solution for the NER task on CS data, e.g., Language Identification. This work addresses this gap by proposing and applying state-of-the-art techniques for Modern Standard Arabic and Arabic-English NER. We have created the first annotated CS Arabic-English corpus for the NER task. Also, we apply two enhancement techniques to improve the NER tagger on CS data using CS contextual embeddings and data augmentation techniques. All methods showed improvements in the performance of the NER taggers on CS data. Finally, we propose several intra-word language identification approaches to determine the language type of a mixed text and identify whether it is a named entity or not. PhD thesis

Visit

arxiv.org

Tasks

code switchinginformation extractionlanguage identificationnamed entity recognition

Tags

Computation and LanguageArtificial Intelligence

Similar

Embedded English verbs in Arabic-English code-switching in EgyptCode‐Switching: Amharic‐EnglishDialectal Arabic Code-Switching Dataset.A structural analysis of Moroccan Arabic and English intra-sentential code switching.SYNTACTIC ANALYSIS OF INTRA-SENTENTIAL ARABIC–ENGLISH CODE-SWITCHING WITHIN VERB PHRASECode-switching in Edward Said’s Out of Place Code-switching in Edward Said’s Out of Place: Questioning the status of Arabic in relation to English

Embedded English verbs in Arabic-English code-switching in Egypt

Aims: This study provides new insights into Arabic-English code-switching with

Code‐Switching: Amharic‐English

Dialectal Arabic Code-Switching Dataset.

Dialectal Arabic Code-Switching Dataset: includes the annotated two-hours Egyptian dataset from the ADI-5 development split in the MGB-3 challenge. The first Dialectal Arabic Code Switching - DACS corpus from broadcast speech. Annotated at the token-level, conside

A structural analysis of Moroccan Arabic and English intra-sentential code switching.

A phenomenon of language contact between different speech communities is that of code switching whic

SYNTACTIC ANALYSIS OF INTRA-SENTENTIAL ARABIC–ENGLISH CODE-SWITCHING WITHIN VERB PHRASE

The current study aims at analyzing the syntax of Arabic-English intra-sentential code-switching wit

Code-switching in Edward Said’s Out of Place Code-switching in Edward Said’s Out of Place: Questioning the status of Arabic in relation to English

International audience