
Serbian loanverbs in Gurbet Romani (Kruševac)
Marko Simonović, Mirjana Mirić, Svetlana Ćirković, Stefan Milosavljević, Jelena Stojković
This dataset documents Serbian loanverbs attested in Romani translations of texts published on the portal krusevacgrad.rs. It focuses on verbs borrowed from Serbian into Gurbet Romani, especially on how Serbian base verbs are morphologically integrated into the Romani verbal system.
The sample consists of 13 texts, with a total of 13,479 words of Romani text. These texts were produced within two projects: “Romski grad, pravo na obrazovanje i rad – Romano foro, hakaj po sićope thaj bućaripe” and “Romski grad: Od kartonskog naselja do krova nad glavom – Romano foro: Thare artijaći mahala dži ka ćeramida umpral o šoro”. In the original publications, each Romani translation appears below the corresponding Serbian text. The texts were written by Jelena Božović and translated into Romani by Tamara Asković.
The sample contains 895 verb tokens belonging to 253 different Romani lemmas.
The dataset was developed in parallel with the previously published dataset Serbian loanverbs in Gurbet Romani (Knjaževac) (Mirić, Ćirković and Simonović 2025). While the Knjaževac dataset is based on transcripts of spontaneous speech, the Kruševac dataset is based on written Romani translations of published Serbian texts. Where possible, the annotation follows the same logic in order to allow comparison between the two Gurbet Romani samples.
Each annotated row in the dataset represents a single token of a Serbian loanverb in Gurbet Romani. The token is given together with its Romani sentence context, the corresponding BCMS sentence, source URL, sentence ID, and text ID. When a source sentence contains more than one annotated loanverb token, it appears more than once in the table, once for each token.
If there was no borrowed verb in a text segment or sentence, no token-level annotation was performed. Such rows preserve the sentence context and source metadata but leave the token-level annotation columns empty.
The form mora ‘must’, although clearly borrowed, was not annotated because it showed no agreement or tense marking in the sample. For example, in the sentence Naj amen ni kanalizacija, thatos pe kašta, snalazisamen nesar, mora the tataras e čharen ‘We do not even have sewage, we heat with wood, we manage somehow, we must heat the children’, mora was not treated as an annotated loanverb token.
The annotation identifies the Romani token, its citation form, simplified lemma, English translation, inclusion status, Romani characteristic vowel, adaptation marker, verbal form, BCMS base verb, certainty of the base verb, atypical adaptation patterns, presence or absence of a derivational suffix in the BCMS base, BCMS theme-vowel class, and the vowel preceding the Romani characteristic vowel.
Before turning to each column in the order presented, we first discuss the column Included, which marks which items the authors consider suitable for inclusion in the analysis.
The column Included indicates whether a given token is considered suitable for the main analysis of the morphological integration of Serbian loanverbs into Gurbet Romani. The value 1 marks tokens included in the analysis, while the value 0 marks tokens that are documented in the dataset but excluded from the main analysis.
A token was included if it represents a Serbian loanverb in Gurbet Romani whose form can be related to a BCMS base verb and whose Romani adaptation can be analysed in terms of the main contrast investigated in the dataset, namely the use of the Romani characteristic vowel i or o.
Tokens were excluded when they were not suitable for this analysis, even if they are relevant for documenting the broader presence of Serbian material in the Romani text. Several types of exclusion should be distinguished.
First, we excluded cases in which spelling errors made the form uncertain, that is, cases where one or more segments of the token appear to be misspelled. Examples of such excluded items are given in Table 1. By contrast, tokens were included when the issue involved only spacing or predictable assimilation processes and the intended form was recoverable with certainty. Such cases are illustrated in Table 2.
Table 1: Examples of items excluded due to misspelling
|
Token |
Intended |
Gloss |
|
pislo |
pisol |
‘write.PRS.3SG’ |
|
radisaka |
radisada |
‘work.PRF.3SG’ |
|
prenesile |
prenesime |
‘transfer.PASS.PTCPL’ |
|
Planire mis i |
Planirime si |
‘plan.PASS.PTCPL is’ |
|
ukažilkpe |
ukažil pe |
‘show up.PRS.3SG REFL’ |
|
Pnairinpe |
Planirin pe |
‘plan.PRS.3PL REFL’ |
|
održona |
podržona |
‘support.IMPF.3PL’ |
Table 2: Examples of included items with recoverable spacing or assimilation
|
Token |
Normalisation |
Gloss |
|
suočimpe |
suočin pe |
‘face.PRS.3PL REFL’ |
|
poboljšolpeo |
poboljšol pe o |
‘improve.PRS.3PL REFL DET’ |
Second, we excluded items that show atypical integration patterns of several types.
One subtype consists of forms in which the encountered Romani adaptation pattern was not analysed as containing the characteristic vowel i or o. For example, sagradame ‘build.PASS.PTCPL’ was excluded, since its characteristic vowel a is not part of the i/o contrast. The same applies to forms such as pomenusadam ‘mention.PRF.1SG’, where the relevant vowel is u. Such forms are rare and may reflect spelling errors.
Another subtype consists of passive participle forms based directly on a BCMS passive participle rather than on the infinitival or present stem of the BCMS verb. For example, postignutime ‘achieve.PASS.PTCPL’ and pomenutime ‘mention.PASS.PTCPL’ are based on the BCMS passive participles postignut and pomenut. These forms were excluded because they do not instantiate the same type of verbal-stem adaptation as forms such as postignol ‘achieve’ and pomenil ‘mention’, which are based on BCMS verbal stems and were treated as regular loanverb forms where attested.
A further subtype consists of tokens in which a derivational suffix of the BCMS base is not preserved in the Romani form. In such cases, the relevant BCMS theme-vowel material cannot be directly compared with the Romani characteristic vowel. Examples include žrtvosilje, with the lemma žrtvol, based on BCMS žrtvovati ‘sacrifice’, and diskriminisadama, with the lemma diskriminil, based on BCMS diskriminisati ‘discriminate’. With preservation of the derivational suffix, the expected forms would be of the type žrtvujil and diskriminišil.
Finally, forms were excluded when the BCMS base could not be identified with sufficient certainty for the purposes of the analysis. This is especially relevant where more than one BCMS verb could plausibly underlie the Romani token and where the choice would affect the annotation of the BCMS base, aspect, or theme-vowel class. For instance, dodaisada ‘add.PRF.3SG’, with the lemma dodail, may be connected either to BCMS dodati, dodamo ‘add.PFV’ or to BCMS dodavati, dodajemo ‘add.IPFV’.
The value 0 in the column Included therefore does not mean that a form is irrelevant, impossible, or wrongly identified as Serbian-derived. It only means that the authors do not consider it suitable for the main quantitative analysis of the relation between the BCMS base verb and the Romani i/o adaptation pattern.
The dataset contains 24 columns (A–X).
Contains the numerical identifier assigned to each row. This column preserves the original ordering of the material and can be used for sorting the table so that examples appear in their order within sentences and sentences appear in their order within texts.
Contains the loanverb token as it appears in the Romani sentence in column T. This column is empty when no borrowed verb was annotated for the relevant row.
Contains the Romani citation form of the token. Since Gurbet Romani does not have the infinitive, the citation form is given as the PRS.3SG form. Reflexive verbs are cited with the reflexive particle pe where relevant.
Contains a simplified Romani lemma. This column unifies certain formally related citation forms, including verbs with and without the reflexive particle pe, where appropriate. For example, citation forms such as desil pe ‘happen’ and održol pe ‘take place’ are entered under the simplified lemmas desil and održol, respectively.
Contains the English translation of the Romani lemma.
Indicates whether the token is included in the main analysis: 1 = yes, 0 = no. The criteria for inclusion and exclusion are discussed above.
Contains the Romani characteristic vowel in the citation form. The main values are i and o. Rare values outside this contrast are also documented, for example a in sagradal and u in pomenul. Examples of regular values include i in desil pe ‘happen’ and o in održol pe ‘take place’.
Contains the numerically binarised version of column G, just focusing on the i vs. o contrast. Verbs with the characteristic vowel i are assigned the value 1, while verbs with the characteristic vowel o are assigned the value 0. Rare cases in which the characteristic vowel is neither i nor o were not annotated in this column and were excluded from the main binary analysis.
Contains the exponent of the adaptation marker if the token contains one. Values attested in the dataset include sar, salj, sa, and sajl. For example, dobisada ‘get.PRF.3SG’ contains the adaptation marker sa, trudisalje ‘strive.PRF.3PL’ contains salj, and desisajlo ‘happen.PRF.3SG.M’ contains sajl. The value sar is attested in the verb pomosarel ‘help’, which has this element throughout the paradigm. If there is no overt adaptation marker, the value 0 is assigned, as in dobil ‘get.PRS.3SG’ or održol pe ‘take place.PRS.3SG REFL’.
Contains the binarised version of column I. Tokens with an overt adaptation marker are assigned the value 1, for example dobisada, trudisalje, and desisajlo. Tokens without an overt adaptation marker are assigned the value 0, for example dobil and održol pe.
Contains the verbal form of the token. The abbreviations follow Leipzig Glossing Rules. The values used in this dataset include:
PRS = present tense. Examples include dobin ‘get.PRS.3PL’, dobiv ‘get.PRS.1SG’, održol pe ‘take place.PRS.3SG’, and sumnjin ‘doubt.PRS.3PL’.
PRF = perfect. Examples include dobisada ‘get.PRF.3SG’, dobisadem ‘get.PRF.1SG’, dobisade ‘get.PRF.3PL’, and desisajlo ‘happen.PRF.3SG.M’.
IMPF = imperfect. Examples include zaradina ‘earn.IMPF.3PL’ and pomosarela ‘help.IMPF.3SG’.
PTCP = participle. Examples include navedime ‘state.PTCP’, podnesime ‘file.PTCP’, razočarime ‘disappoint.PTCP’, and sufinancirime ‘co-finance.PTCP’.
Contains a binarised version of column K. Perfect forms are assigned the value 1, for example dobisada ‘get.PRF.3SG’, dobisadem ‘get.PRF.1SG’, dobisade ‘get.PRF.3PL’, and desisajlo ‘happen.PRF.3SG.M’. All other forms are assigned the value 0, for example dobin ‘get.PRS.3PL’, dobiv ‘get.PRS.1SG’, zaradina ‘earn.IMPF.3PL’, and navedime ‘state.PTCP’.
Contains the infinitive form of the standard Bosnian/Croatian/Montenegrin/Serbian verb on which the Romani verb is arguably based. For example, the Romani forms dobisada ‘get.PRF.3SG’ and dobil ‘get.PRS.3SG’ are linked to BCMS dobiti ‘get’, while navedisada ‘state.PRF.3SG’ and navedil ‘state.PRS.3SG’ are linked to BCMS navesti ‘state’.
Indicates whether the BCMS base verb could be identified with certainty: 1 = yes, 0 = no. The value 1 is used when a single BCMS verb can be identified as the base, as in dobisada, linked to BCMS dobiti. The value 0 is used when more than one BCMS verb could plausibly serve as the base, or when the base cannot be identified securely. For example, dodaisada ‘add.PRF.3SG’, with the Romani lemma dodail, may be connected either to BCMS dodati, dodamo ‘add.PFV’ or to dodavati, dodajemo ‘add.IPFV’.
Identifies tokens with an atypical adaptation pattern. The typical pattern is the one in which the Romani characteristic vowel i or o appears in the position corresponding to the theme vowel of the BCMS base verb. Tokens following this regular pattern are assigned the value 0. Tokens with atypical integration patterns are assigned the value 1.
As noted above, atypical patterns include forms not analysed as containing the regular characteristic vowel i or o, forms based directly on BCMS passive participles, and forms in which a derivational suffix of the BCMS base is not preserved in the Romani form.
Indicates whether the BCMS base verb contains a derivational suffix. Verbs without a derivational suffix are assigned the value 1, as in dobiti ‘get’, navesti ‘state’, or pisati ‘write’. Verbs with a derivational suffix are assigned the value 0, as in finans-ir-ati ‘finance’, plan-ir-ati ‘plan’, or škol-ov-ati ‘school’.
Contains the theme-vowel class of the BCMS base verb. The classification follows the Database of the Western South Slavic Verb — WeSoSlav (Arsenijević et al. 2024). The values attested in this dataset include a/a, a/i, a/je, e/e, e/i, i/i, u/e, 0/e, and 0/ne.
The names of the classes correspond to the theme-vowel material in the relevant BCMS verbal forms. For example, pisati, pišemo ‘write’ belongs to the a/je class, raditi, radimo ‘work’ belongs to the i/i class, and dobiti, dobijemo ‘get’ belongs to the 0/e class.
Contains the nucleus of the syllable preceding the Romani characteristic vowel in the Romani token. Possible values include a, e, i, o, u, and r, where r functions as a syllabic nucleus. For example, the preceding vowel is o in dobil ‘get’, e in navedil ‘state’, i in finansiril ‘finance’, and r in održol pe ‘take place’.
Contains the binarised version of column R. If the preceding vowel is i, the value 1 is assigned, as in finansiril. All other values are assigned 0, as in dobil, navedil, and održol.
Contains the Romani sentence from which the loanverb token was extracted.
Contains the corresponding BCMS sentence from the source text.
Contains the URL of the source text on krusevacgrad.rs.
Contains the identifier of the sentence from which the token was extracted.
Contains the identifier of the source text.
Arsenijević, Boban, Gomboc Čeh, Katarina, Marušič, Franc Lanko, Milosavljević, Stefan, Mišmaš, Petra, Simić, Jelena, Simonović, Marko and Žaucer, Rok. 2024. Database of the Western South Slavic Verb HyperVerb 2.0 — WeSoSlav. Slovenian language resource repository CLARIN.SI, ISSN 2820-4042.http://hdl.handle.net/11356….
Mirić, Mirjana, Svetlana Ćirković and Marko Simonović. 2025. Serbian loanverbs in Gurbet Romani (Knjaževac). Zenodo dataset, version v1. DOI: 10.5281/zenodo.15642055.
This database is a result of the project “What’s in a verb? Mapping Serbian verbs borrowed into Romani” (Institute for Balkan Studies, Serbian Academy of Sciences and Arts; Department of Slavic Studies, University of Graz).
The project is financed by the Ministry of Science, Technological Development and Innovations of the Republic of Serbia in cooperation with Austria’s Agency for Education and Internationalisation (OeAD), within the programme of scientific and technological cooperation between the Republic of Serbia and the Republic of Austria, for the period 2024–2026.