Abstract
Mining ChEMBL for natural-product (NP) leads is now routine, yet the non-random sampling of compound space — investigation bias — is rarely modelled. We present a pre-registered cheminformatics pipeline integrating LOTUS (150 753 plant compounds) with ChEMBL
Plasmodium falciparum
whole-cell records via the International Chemical Identifier Key (InChIKey), under canonical half-maximal inhibitory/effective concentration (IC50/EC50) reporting. A logistic generalised linear mixed model (GLMM) with cluster-robust standard errors by botanical family adjusts for the log-number of assays per compound. As a case study, ethnobotanically-positive African antimalarial plants (cited in ≥ 2 of five sources) are not enriched in active compounds at pChEMBL ≥ 5 once investigation bias is modelled (adjusted odds ratio [OR] = 0.42, 95% confidence interval [CI] 0.20–0.89,
p
= 0.023;
n
= 416). The effect is non-monotone across pChEMBL thresholds and modulated by the chemical pathway assigned by NPclassifier. The pipeline is reusable for other phenotypic ChEMBL endpoints and provides a template for investigation-bias-adjusted enrichment statistics on NP–bioactivity intersections.