
This repository houses the BaseCLEA18 and DiaCLEA18 cascades, or sequences of sound changes, used for computerized forward reconstruction work upon Albanian, with associated lexicon files for various chronological strata of Latinate, Greek, and Slavic etyma. BaseCLEA18 is the baseline cascade derived from De Vaan (2018), and DiaCLEA18 is the product of cascade calibration as part of my dissertation (Marr 2026). While calibration of Albanian diachrony may continue on the GitHub repository Albanian-CFR (github.com), what is preserved here is the form at the time of final dissertation submission, at which point the accuracy of forward reconstructed predictions (in terms of matching observed outcomes) was just over 80% for these strata.
DiaSim, the software used to perform CFR with the -CLEA18 cascades, is not included, but may be found, with usage instructions, on GitHub (github.com); feel free to reach out if you need assistance. Also not included here, but included on the GitHub repository, are lexica of Indo-European inherited items, for which accuracy at the time of submission in August 2026 stood at just over 50%; these were not the main scope of the dissertation and will be subject to much further revision. Further information about the lexicon files in this package, as well as Python scripts included for maintenance and other operations upon them, may be found in the file "lex_files_README.md". The periodization policy is found in "Periodization policy.txt". Lexicon files are effectively .csv files in all but name, with each column representing a different chronological period, and the symbol '$' flagging the beginning of a comment containing citations and other notes about the form; these comments are not involved in DiaSim's forward reconstruction implementation. For more on the structure of lexicon files used by DiaSim, see Lexicon · clmarr/DiaSim Wik…. The cascade files (BaseCLEA18, DiaCLEA18) also have comments flagged by '$', and include sequences of rules in a largely SPE-equivalent format (X > Y / A __ B); for more on cascade files used by DiaSim, see Cascade · clmarr/DiaSim Wik….
The results, including statistics and diachronic derivation files for each etymon, of running the baseline cascade BaseCLEA18 (prefix base-) and the debugged cascade DiaCLEA18 (prefix dia-) upon all strata of borrowing from the point of pre-Roman contact with Greek dialects (Archaic Proto Albanian, or APA) until the latest stage of Proto-Albanian development (suffix -APArun), and from chronological strata borrowed from, respectively, Greek (-Grkrun), Latin/Romance (-Latrun), and Slavic (-Slavrun) are all included in zipped files.
Works Cited
De Vaan, Michiel. 2018. 95. The phonology of Albanian. In Handbook of Comparative and Historical Indo-European Linguistics, pages 1732–1749. De Gruyter Mouton.
Marr, Clayton G. S. 2026. Disentangling Albanian and Romance Diachrony: Electronic Neogrammarian Exploits. The Ohio State University. Ph.D. dissertation.