Logo Lanfrica
  • Accueil
  • Atlas
  • Analyses
  • Documentation
  • Sign in

© 2026 Lanfrica. Tous droits réservés. Tous les droits d'auteur des ressources affichées sur le site Web Lanfrica appartiennent aux détenteurs de droits d'auteur d'origine, sauf indication contraire explicite.

From dialect features to structured data Modelling spoken Arabic varieties in the WIBARAB project

Domaine:

natural language processing

Type de record:

dataset
Créateur:
EngIriMörPro
Éditeur:
Zenodo
Hôte:avatar
The Vienna Corpus of Arabic Varieties (VICAV) is a language documentation platform that provides access to a growing collection of digital language resources. Integrating approaches from language technology and the wider field of text-oriented digital humanities, the project aims to address issues of representing heterogeneous data by providing a flexible, yet sustainable technical environment based on a modular data architecture. In addition to a bibliography of research literature, typologically similar to what can be found in the World Atlas of Language Structures (WALS) or the Database of Arabic Dialects (DAD), VICAV offers several types of data: the so far largest part of the collection comprises linguistic profiles (i.e. standardised concise descriptions of linguistic varieties), structured lists of linguistic features, sample texts, corpora of unmonitored speech, and dictionaries. The VICAV infrastructure is meant to ensure consistent encoding across projects as well as sustainable creation and publishing workflows across projects and builds largely on TEI (P5) as its underlying data model (Procházka et al. 2015).    The newest and, in terms of resources, duration, and scope, most extensive addition to the VICAV projects is the ERC Advanced Grant WIBARAB (What is Bedouin-type Arabic? 101020127-WIRARAB). It investigates the linguistic and socio-historical realities behind the millennia-old dichotomous distinction between Bedouin or sedentary dialects (Procházka 2024). The central component of WIBARAB is a linguistic feature database covering over 300 varieties of spoken Arabic. It incorporates data from hitherto little researched areas collected in campaigns in Saudi Arabia, Kuwait, Jordan, Lebanon, Sudan, Tunisia and Morocco, along with previously published material. The project has a strong focus on open access, data structures, standards and best practices, as well as data modelling.     By contrast to comparable other projects collecting linguistic data, WIBARAB has adopted a strictly text-oriented approach that has been grounded in the application of the Guidelines of the Text Encoding Initiative (TEI). While TEI (P5) provides a well-tried inventory of elements to represent morphological and lexical concepts, the issue of representing other grammatical phenomena has received less attention and modelling extra-linguistic and socio-linguistic phenomena constitutes the interesting part of this encoding challenge. Our paper will focus on the WIBARAB team’s customisation, trying to re-use as many concepts existing in the TEI guidelines as possible, document workflow steps such as the writing of a meaningful, re-usable ODD, and examine some encoding decisions by furnishing examples of phonological, morphological, syntactical, phraseological and lexical data.     Finally, we will touch on some methodological challenges in developing an efficient research-tool on the basis of this TEI-encoded, intricately structured linguistic database. It has already been used for a wide range of varied research questions such as the in-detail description of particular linguistic varieties (Torzullo 2024) as well as socio-linguisthic approaches dealing with questions of intergenerational variation, intra-speaker variation and identity (Iriarte Díez 2025a, 2025b). We will also provide a first glimpse of the evolving map-based front-end which will be integrated into the overall VICAV infrastructure. 

Visit

doi.orgzenodo.org

Licenses

Creative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similaires

QCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual FeaturesSpoken Arabic Algerian dialect identificationBUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT بناء قدرات لفهم اللهجة التونسية المنطوقة BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECTIn search of a standard: dialect variation and New Arabic features in the oldest Arabic written documentsBuilding Ontologies to Understand Spoken Tunisian DialectInterrogatives in Tunisia. With a focus on the Arabic varieties spoken in northwestern, central and southern Tunisia

QCRI @ DSL 2016: Spoken Arabic Dialect Identification Using Textual Features

The paper describes the QCRI submissions to the task of automatic Arabic dialect classification into 5 Arabic variants, namely Egyptian, Gulf, Levantine, North-African, and Modern Standard Arabic (MSA). The training data is relatively small and is automatically gen

Spoken Arabic Algerian dialect identification

BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT بناء قدرات لفهم اللهجة التونسية المنطوقة BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT BUILDING ONTOLOGIES TO UNDERSTAND SPOKEN TUNISIAN DIALECT

This paper presents a method to understand spoken Tunisian dialect based on lexical semantic. This m

In search of a standard: dialect variation and New Arabic features in the oldest Arabic written documents

International audience The few surviving pre-Islamic inscriptions in both the Arabic

Building Ontologies to Understand Spoken Tunisian Dialect

This paper presents a method to understand spoken Tunisian dialect based on lexical semantic. This m

Interrogatives in Tunisia. With a focus on the Arabic varieties spoken in northwestern, central and southern Tunisia