Logo Lanfrica

Managing lexical complexity ODD chaining for Arabic dialect dictionaries

Domaine:

natural language processing

Type de record:

dataset
Créateur:
MoeSchEngPro
Éditeur:
Zenodo
Hôte:avatar
The Vienna Corpus of Arabic Varieties (VICAV) is a language documentation platform hosting a varied collection of digital language resources. From its beginnings, it has adopted a text-oriented approach that has been grounded in the application of the Guidelines of the Text Encoding Initiative (TEI). The VICAV infrastructure is meant to ensure consistent encoding across projects as well as sustainable creation and publishing workflows. In addition to a bibliography of research literature, typologically similar to what can be found in the World Atlas of Language Structures (WALS) or the Database of Arabic Dialects, VICAV offers several types of data: concise descriptions of linguistic varieties, structured lists of linguistic features, sample texts, corpora of unmonitored speech, and dictionaries (Procházka et al. 2015). The published VICAV dictionaries so far cover five linguistic varieties: Baghdad, Cairo, Damascus, Tunis and Modern Standard Arabic. These dictionaries are all comparatively small, none of them containing more than 8000 entries, and constitute lexical databases with structured lexicographic information (Moerth and Schopper 2021). They are all provided with English translation equivalents (some also include German, French or Spanish translations) and share common encoding conventions. A recent addition in this series is the Shawi Dictionary, which will go online in late 2025. This is the first VICAV dictionary natively encoded in TEI Lex-0, a relatively new initiative within the TEI community, intended as a baseline encoding for lexicographic data (Tasovac et al. 2018) which meanwhile has been adopted by more dictionary projects (Salgado et al. 2019). Stemming from different projects, all VICAV dictionaries have subtle differences in their encoding requirements. Moreover, there are several other comparable lexicographic resources being developed at the ACDH-CH. To be able to document the differences between those dictionaries and yet keep them compatible, we use the method of ODD chaining, which is the process of deriving ODDs from one another(Pernes et al. 2017). In our case, we proceeded from the TEI Lex-0 ODD as the primary source and modified it to create the ACDH-CH generic-dict-schema, which serves as the baseline ODD for various VICAV and other dictionaries worked on at the ACDH-CH. Since, for example, TEI Lex-0 does not make prescribe the macrostructure of a dictionary, it is in this schema that we define that in our dictionaries examples are not embedded within entries but kept in a separate
in order to allow them to be re-used in different contexts. In a third step, the SHAWI dictionary ODD is derived from the generic-dict-schema and provides specifications adapted to the needs of the SHAWI project. Unlike in other dictionaries, the lexical profiles of tribes and their geographic context play an important role. To incorporate this information, we extend generic-dict-schema to allow typed elements within - a construct which is not needed in the other dictionaries. In our paper, we will present the dictionaries involved and our general mechanism next to discussing pros (modularization; specificity) and cons (complexity both in terms of processing overhead and modelling) of having selected this relatively complex route.