Logo Lanfrica

A Lemma-Level NLP Pipeline for Quantitative Chronological Analysis of Biological Semantics in the Qur'anic Corpus

Domaine:

natural language processing

Type de record:

datasetpaper
Créateur:
RajKas
Éditeur:
Zenodo
Hôte:avatar
This study presents a morphology-aware, lemma-level natural language processing (NLP) framework for quantitative analysis of biological discourse in the Qur'anic corpus and its variation between Meccan and Medinan revelation periods. A curated lexicon of biologically relevant Arabic lemmas, encompassing anatomy, cognition, creation, development, life processes, speech, and social embodiment, was integrated with a morphologically normalized Qur'anic corpus. The computational workflow combines lemma-level semantic extraction, frequency analysis, conceptual domain modeling, statistical testing, logistic regression, co-occurrence network analysis, and surah-level visualization. Biological terminology was identified in 11.9% of Meccan verses and 19.2% of Medinan verses, with a statistically significant association between revelation period and biological discourse density (χ² = 52.98, p < 0.001; Cramér's V = 0.092). The results reveal systematic chronological differentiation in biological semantics. Creation-related concepts show greater relative prominence in Meccan revelation, whereas life- and development-related concepts are more prominent in Medinan passages. The study provides a reproducible computational framework for investigating domain-specific semantic patterns in Classical Arabic texts. Associated materials: The deposited files include the research manuscript and supplementary materials containing the curated biological lexicon, complete lemma-frequency distributions, Meccan–Medinan frequency comparisons, and logistic regression coefficient estimates.