Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Serbian loanverbs in Gurbet Romani (Kruševac)

Domain:

natural language processing

Record type:

dataset
Creator:
SimMirĆirMil
Publisher:
Zenodo
Host:avatar

Serbian loanverbs in Gurbet Romani (Kruševac)

 

Marko Simonović, Mirjana Mirić, Svetlana Ćirković, Stefan Milosavljević, Jelena Stojković

 

This dataset documents Serbian loanverbs attested in Romani translations of texts published on the portal krusevacgrad.rs. It focuses on verbs borrowed from Serbian into Gurbet Romani, especially on how Serbian base verbs are morphologically integrated into the Romani verbal system.

The sample consists of 13 texts, with a total of 13,479 words of Romani text. These texts were produced within two projects: “Romski grad, pravo na obrazovanje i rad – Romano foro, hakaj po sićope thaj bućaripe” and “Romski grad: Od kartonskog naselja do krova nad glavom – Romano foro: Thare artijaći mahala dži ka ćeramida umpral o šoro”. In the original publications, each Romani translation appears below the corresponding Serbian text. The texts were written by Jelena Božović and translated into Romani by Tamara Asković.

The sample contains 895 verb tokens belonging to 253 different Romani lemmas. 

The dataset was developed in parallel with the previously published dataset Serbian loanverbs in Gurbet Romani (Knjaževac) (Mirić, Ćirković and Simonović 2025). While the Knjaževac dataset is based on transcripts of spontaneous speech, the Kruševac dataset is based on written Romani translations of published Serbian texts. Where possible, the annotation follows the same logic in order to allow comparison between the two Gurbet Romani samples.

Data extraction and annotation procedure

Each annotated row in the dataset represents a single token of a Serbian loanverb in Gurbet Romani. The token is given together with its Romani sentence context, the corresponding BCMS sentence, source URL, sentence ID, and text ID. When a source sentence contains more than one annotated loanverb token, it appears more than once in the table, once for each token.

 

If there was no borrowed verb in a text segment or sentence, no token-level annotation was performed. Such rows preserve the sentence context and source metadata but leave the token-level annotation columns empty.

 

The form mora ‘must’, although clearly borrowed, was not annotated because it showed no agreement or tense marking in the sample. For example, in the sentence Naj amen ni kanalizacija, thatos pe kašta, snalazisamen nesar, mora the tataras e čharen ‘We do not even have sewage, we heat with wood, we manage somehow, we must heat the children’, mora was not treated as an annotated loanverb token.

 

The annotation identifies the Romani token, its citation form, simplified lemma, English translation, inclusion status, Romani characteristic vowel, adaptation marker, verbal form, BCMS base verb, certainty of the base verb, atypical adaptation patterns, presence or absence of a derivational suffix in the BCMS base, BCMS theme-vowel class, and the vowel preceding the Romani characteristic vowel.

 

Before turning to each column in the order presented, we first discuss the column Included, which marks which items the authors consider suitable for inclusion in the analysis.

Excluded tokens and exclusion criteria

The column Included indicates whether a given token is considered suitable for the main analysis of the morphological integration of Serbian loanverbs into Gurbet Romani. The value 1 marks tokens included in the analysis, while the value 0 marks tokens that are documented in the dataset but excluded from the main analysis.

A token was included if it represents a Serbian loanverb in Gurbet Romani whose form can be related to a BCMS base verb and whose Romani adaptation can be analysed in terms of the main contrast investigated in the dataset, namely the use of the Romani characteristic vowel i or o.

Tokens were excluded when they were not suitable for this analysis, even if they are relevant for documenting the broader presence of Serbian material in the Romani text. Several types of exclusion should be distinguished.

First, we excluded cases in which spelling errors made the form uncertain, that is, cases where one or more segments of the token appear to be misspelled. Examples of such excluded items are given in Table 1. By contrast, tokens were included when the issue involved only spacing or predictable assimilation processes and the intended form was recoverable with certainty. Such cases are illustrated in Table 2.

Table 1: Examples of items excluded due to misspelling

Token

Intended

Gloss

pislo

pisol

‘write.PRS.3SG’

radisaka

radisada

‘work.PRF.3SG’

prenesile

prenesime

‘transfer.PASS.PTCPL’

Planire mis i

Planirime si 

‘plan.PASS.PTCPL is’

ukažilkpe

ukažil pe

‘show up.PRS.3SG REFL’

Pnairinpe

Planirin pe

‘plan.PRS.3PL REFL’

održona

podržona

‘support.IMPF.3PL’



Table 2: Examples of included items with recoverable spacing or assimilation

Token

Normalisation

Gloss

suočimpe

suočin pe

‘face.PRS.3PL REFL’

poboljšolpeo

poboljšol pe o

‘improve.PRS.3PL REFL DET’

Second, we excluded items that show atypical integration patterns of several types.

One subtype consists of forms in which the encountered Romani adaptation pattern was not analysed as containing the characteristic vowel i or o. For example, sagradame ‘build.PASS.PTCPL’ was excluded, since its characteristic vowel a is not part of the i/o contrast. The same applies to forms such as pomenusadam ‘mention.PRF.1SG’, where the relevant vowel is u. Such forms are rare and may reflect spelling errors.

Another subtype consists of passive participle forms based directly on a BCMS passive participle rather than on the infinitival or present stem of the BCMS verb. For example, postignutime ‘achieve.PASS.PTCPL’ and pomenutime ‘mention.PASS.PTCPL’ are based on the BCMS passive participles postignut and pomenut. These forms were excluded because they do not instantiate the same type of verbal-stem adaptation as forms such as postignol ‘achieve’ and pomenil ‘mention’, which are based on BCMS verbal stems and were treated as regular loanverb forms where attested.

A further subtype consists of tokens in which a derivational suffix of the BCMS base is not preserved in the Romani form. In such cases, the relevant BCMS theme-vowel material cannot be directly compared with the Romani characteristic vowel. Examples include žrtvosilje, with the lemma žrtvol, based on BCMS žrtvovati ‘sacrifice’, and diskriminisadama, with the lemma diskriminil, based on BCMS diskriminisati ‘discriminate’. With preservation of the derivational suffix, the expected forms would be of the type žrtvujil and diskriminišil.

Finally, forms were excluded when the BCMS base could not be identified with sufficient certainty for the purposes of the analysis. This is especially relevant where more than one BCMS verb could plausibly underlie the Romani token and where the choice would affect the annotation of the BCMS base, aspect, or theme-vowel class. For instance, dodaisada ‘add.PRF.3SG’, with the lemma dodail, may be connected either to BCMS dodati, dodamo ‘add.PFV’ or to BCMS dodavati, dodajemo ‘add.IPFV’.

The value 0 in the column Included therefore does not mean that a form is irrelevant, impossible, or wrongly identified as Serbian-derived. It only means that the authors do not consider it suitable for the main quantitative analysis of the relation between the BCMS base verb and the Romani i/o adaptation pattern.

Summary of the columns

The dataset contains 24 columns (A–X).

A: N of example

Contains the numerical identifier assigned to each row. This column preserves the original ordering of the material and can be used for sorting the table so that examples appear in their order within sentences and sentences appear in their order within texts.

B: Token

Contains the loanverb token as it appears in the Romani sentence in column T. This column is empty when no borrowed verb was annotated for the relevant row.

C: Lemma

Contains the Romani citation form of the token. Since Gurbet Romani does not have the infinitive, the citation form is given as the PRS.3SG form. Reflexive verbs are cited with the reflexive particle pe where relevant.

D: Lemma (simplified)

Contains a simplified Romani lemma. This column unifies certain formally related citation forms, including verbs with and without the reflexive particle pe, where appropriate. For example, citation forms such as desil pe ‘happen’ and održol pe ‘take place’ are entered under the simplified lemmas desil and održol, respectively.

E: English translation

Contains the English translation of the Romani lemma.

F: Included

Indicates whether the token is included in the main analysis: 1 = yes, 0 = no. The criteria for inclusion and exclusion are discussed above.

G: Romani characteristic vowel

Contains the Romani characteristic vowel in the citation form. The main values are i and o. Rare values outside this contrast are also documented, for example a in sagradal and u in pomenul. Examples of regular values include i in desil pe ‘happen’ and o in održol pe ‘take place’.

H: Romani characteristic vowel (binarised)

Contains the numerically binarised version of column G, just focusing on the i vs. o contrast. Verbs with the characteristic vowel i are assigned the value 1, while verbs with the characteristic vowel o are assigned the value 0. Rare cases in which the characteristic vowel is neither i nor o were not annotated in this column and were excluded from the main binary analysis.

I: Adaptation marker

Contains the exponent of the adaptation marker if the token contains one. Values attested in the dataset include sar, salj, sa, and sajl. For example, dobisada ‘get.PRF.3SG’ contains the adaptation marker sa, trudisalje ‘strive.PRF.3PL’ contains salj, and desisajlo ‘happen.PRF.3SG.M’ contains sajl. The value sar is attested in the verb pomosarel ‘help’, which has this element throughout the paradigm. If there is no overt adaptation marker, the value 0 is assigned, as in dobil ‘get.PRS.3SG’ or održol pe ‘take place.PRS.3SG REFL’.

J: Adaptation marker (binarised)

Contains the binarised version of column I. Tokens with an overt adaptation marker are assigned the value 1, for example dobisada, trudisalje, and desisajlo. Tokens without an overt adaptation marker are assigned the value 0, for example dobil and održol pe.

K: Verbal form

Contains the verbal form of the token. The abbreviations follow Leipzig Glossing Rules. The values used in this dataset include:

PRS = present tense. Examples include dobin ‘get.PRS.3PL’, dobiv ‘get.PRS.1SG’, održol pe ‘take place.PRS.3SG’, and sumnjin ‘doubt.PRS.3PL’.

PRF = perfect. Examples include dobisada ‘get.PRF.3SG’, dobisadem ‘get.PRF.1SG’, dobisade ‘get.PRF.3PL’, and desisajlo ‘happen.PRF.3SG.M’.

IMPF = imperfect. Examples include zaradina ‘earn.IMPF.3PL’ and pomosarela ‘help.IMPF.3SG’.

PTCP = participle. Examples include navedime ‘state.PTCP’, podnesime ‘file.PTCP’, razočarime ‘disappoint.PTCP’, and sufinancirime ‘co-finance.PTCP’.

L: Verbal form (binarised)

Contains a binarised version of column K. Perfect forms are assigned the value 1, for example dobisada ‘get.PRF.3SG’, dobisadem ‘get.PRF.1SG’, dobisade ‘get.PRF.3PL’, and desisajlo ‘happen.PRF.3SG.M’. All other forms are assigned the value 0, for example dobin ‘get.PRS.3PL’, dobiv ‘get.PRS.1SG’, zaradina ‘earn.IMPF.3PL’, and navedime ‘state.PTCP’.

M: Base BCMS verb

Contains the infinitive form of the standard Bosnian/Croatian/Montenegrin/Serbian verb on which the Romani verb is arguably based. For example, the Romani forms dobisada ‘get.PRF.3SG’ and dobil ‘get.PRS.3SG’ are linked to BCMS dobiti ‘get’, while navedisada ‘state.PRF.3SG’ and navedil ‘state.PRS.3SG’ are linked to BCMS navesti ‘state’.

N: Base BCMS verb (certain)

Indicates whether the BCMS base verb could be identified with certainty: 1 = yes, 0 = no. The value 1 is used when a single BCMS verb can be identified as the base, as in dobisada, linked to BCMS dobiti. The value 0 is used when more than one BCMS verb could plausibly serve as the base, or when the base cannot be identified securely. For example, dodaisada ‘add.PRF.3SG’, with the Romani lemma dodail, may be connected either to BCMS dodati, dodamo ‘add.PFV’ or to dodavati, dodajemo ‘add.IPFV’.

O: Atypical adaptation pattern

Identifies tokens with an atypical adaptation pattern. The typical pattern is the one in which the Romani characteristic vowel i or o appears in the position corresponding to the theme vowel of the BCMS base verb. Tokens following this regular pattern are assigned the value 0. Tokens with atypical integration patterns are assigned the value 1.

As noted above, atypical patterns include forms not analysed as containing the regular characteristic vowel i or o, forms based directly on BCMS passive participles, and forms in which a derivational suffix of the BCMS base is not preserved in the Romani form.

P: Suffixless base

Indicates whether the BCMS base verb contains a derivational suffix. Verbs without a derivational suffix are assigned the value 1, as in dobiti ‘get’, navesti ‘state’, or pisati ‘write’. Verbs with a derivational suffix are assigned the value 0, as in finans-ir-ati ‘finance’, plan-ir-ati ‘plan’, or škol-ov-ati ‘school’.

Q: Base verb TV class

Contains the theme-vowel class of the BCMS base verb. The classification follows the Database of the Western South Slavic Verb — WeSoSlav (Arsenijević et al. 2024). The values attested in this dataset include a/a, a/i, a/je, e/e, e/i, i/i, u/e, 0/e, and 0/ne.

The names of the classes correspond to the theme-vowel material in the relevant BCMS verbal forms. For example, pisati, pišemo ‘write’ belongs to the a/je class, raditi, radimo ‘work’ belongs to the i/i class, and dobiti, dobijemo ‘get’ belongs to the 0/e class. 

R: Preceding vowel

Contains the nucleus of the syllable preceding the Romani characteristic vowel in the Romani token. Possible values include a, e, i, o, u, and r, where r functions as a syllabic nucleus. For example, the preceding vowel is o in dobil ‘get’, e in navedil ‘state’, i in finansiril ‘finance’, and r in održol pe ‘take place’.

S: Preceding vowel (binarised)

Contains the binarised version of column R. If the preceding vowel is i, the value 1 is assigned, as in finansiril. All other values are assigned 0, as in dobil, navedil, and održol.

T: Romani sentence

Contains the Romani sentence from which the loanverb token was extracted.

U: BCMS sentence

Contains the corresponding BCMS sentence from the source text.

V: Source

Contains the URL of the source text on krusevacgrad.rs.

W: Sentence ID

Contains the identifier of the sentence from which the token was extracted.

X: Text-ID

Contains the identifier of the source text.

References

Arsenijević, Boban, Gomboc Čeh, Katarina, Marušič, Franc Lanko, Milosavljević, Stefan, Mišmaš, Petra, Simić, Jelena, Simonović, Marko and Žaucer, Rok. 2024. Database of the Western South Slavic Verb HyperVerb 2.0 — WeSoSlav. Slovenian language resource repository CLARIN.SI, ISSN 2820-4042.http://hdl.handle.net/11356….

Mirić, Mirjana, Svetlana Ćirković and Marko Simonović. 2025. Serbian loanverbs in Gurbet Romani (Knjaževac). Zenodo dataset, version v1. DOI: 10.5281/zenodo.15642055.

Acknowledgements

This database is a result of the project “What’s in a verb? Mapping Serbian verbs borrowed into Romani” (Institute for Balkan Studies, Serbian Academy of Sciences and Arts; Department of Slavic Studies, University of Graz).

The project is financed by the Ministry of Science, Technological Development and Innovations of the Republic of Serbia in cooperation with Austria’s Agency for Education and Internationalisation (OeAD), within the programme of scientific and technological cooperation between the Republic of Serbia and the Republic of Austria, for the period 2024–2026.




Visit

doi.org

Languages

KlaoKrumen, TepoNdasaSar

Licenses

info:eu-repo/semantics/openAccessCreative Commons Attribution 4.0 Internationalhttps://creativecommons.org/licenses/by/4.0/legalcode

Similar

Morpho-Syntactic Descriptions in MULTEXT-East. The Case of SerbianSyntactic doubling and variation: The case of RomaniA Comparison of Two Prosody Modelling Approaches for Sesotho and SerbianDas Romani von Aja Varvara. Deskriptive und historisch-vergleichende Darstellung eines ZigeunerdialektesCOMPARISON OF ZERO AND FEW-SHOT LEARNING APPROACH USING THE LLMS FOR SENTIMENT ANALYSIS IN SERBIAN LITERATURELa diffusion de la langue-culture romani standardisée dans les écoles roumaines La diffusion de la langue-culture romani standardisée dans les écoles roumaines: Un défi à l'hétérogénéité des pratiques langagières et des positionnements identitaires des Roms

Morpho-Syntactic Descriptions in MULTEXT-East. The Case of Serbian

International audience

Syntactic doubling and variation: The case of Romani

Titre de la communication : A variationist approach to syntactic doubling: the case of Romani Intern

A Comparison of Two Prosody Modelling Approaches for Sesotho and Serbian

Das Romani von Aja Varvara. Deskriptive und historisch-vergleichende Darstellung eines Zigeunerdialektes

COMPARISON OF ZERO AND FEW-SHOT LEARNING APPROACH USING THE LLMS FOR SENTIMENT ANALYSIS IN SERBIAN LITERATURE

Goal: Assess if newer versions of LLMs offer more consistent, efficient, and potentially less biased

La diffusion de la langue-culture romani standardisée dans les écoles roumaines La diffusion de la langue-culture romani standardisée dans les écoles roumaines: Un défi à l'hétérogénéité des pratiques langagières et des positionnements identitaires des Roms

International audience D’un point de vue législatif, la Roumanie a organisé l’enseign