The Vienna Corpus of Arabic Varieties (VICAV) is a language documentation platform that provides access to a growing collection of digital language resources. Integrating approaches from language technology and the wider field of text-oriented digital humanities, the project aims to address issues of representing heterogeneous data by providing a flexible, yet sustainable technical environment based on a modular data architecture. In addition to a bibliography of research literature, typologically similar to what can be found in the World Atlas of Language Structures (WALS) or the Database of Arabic Dialects (DAD), VICAV offers several types of data: the so far largest part of the collection comprises linguistic profiles (i.e. standardised concise descriptions of linguistic varieties), structured lists of linguistic features, sample texts, corpora of unmonitored speech, and dictionaries. The VICAV infrastructure is meant to ensure consistent encoding across projects as well as sustainable creation and publishing workflows across projects and builds largely on TEI (P5) as its underlying data model (Procházka et al. 2015).
The newest and, in terms of resources, duration, and scope, most extensive addition to the VICAV projects is the ERC Advanced Grant WIBARAB (What is Bedouin-type Arabic? 101020127-WIRARAB). It investigates the linguistic and socio-historical realities behind the millennia-old dichotomous distinction between Bedouin or sedentary dialects (Procházka 2024). The central component of WIBARAB is a linguistic feature database covering over 300 varieties of spoken Arabic. It incorporates data from hitherto little researched areas collected in campaigns in Saudi Arabia, Kuwait, Jordan, Lebanon, Sudan, Tunisia and Morocco, along with previously published material. The project has a strong focus on open access, data structures, standards and best practices, as well as data modelling.
By contrast to comparable other projects collecting linguistic data, WIBARAB has adopted a strictly text-oriented approach that has been grounded in the application of the Guidelines of the Text Encoding Initiative (TEI). While TEI (P5) provides a well-tried inventory of elements to represent morphological and lexical concepts, the issue of representing other grammatical phenomena has received less attention and modelling extra-linguistic and socio-linguistic phenomena constitutes the interesting part of this encoding challenge. Our paper will focus on the WIBARAB team’s customisation, trying to re-use as many concepts existing in the TEI guidelines as possible, document workflow steps such as the writing of a meaningful, re-usable ODD, and examine some encoding decisions by furnishing examples of phonological, morphological, syntactical, phraseological and lexical data.
Finally, we will touch on some methodological challenges in developing an efficient research-tool on the basis of this TEI-encoded, intricately structured linguistic database. It has already been used for a wide range of varied research questions such as the in-detail description of particular linguistic varieties (Torzullo 2024) as well as socio-linguisthic approaches dealing with questions of intergenerational variation, intra-speaker variation and identity (Iriarte Díez 2025a, 2025b). We will also provide a first glimpse of the evolving map-based front-end which will be integrated into the overall VICAV infrastructure.