
Electroencephalography (EEG)-based machine learning approaches for schizophrenia classification frequently report high predictive performance, yet concerns regarding dataset leakage, site confounding, and cross-domain instability remain insufficiently investigated. In this study, we performed a forensic signal-level audit of the publicly available African Schizophrenia EEG Dataset (ASZED) to evaluate the robustness and interpretability of EEG classification results under strict integrity and transferability analyses.
Using direct payload extraction from EDF recordings via raw.get_data() (MNE-Python), we computed signal-level hashes independent of EDF headers and metadata. Across 1,932 EDF records, we identified 44 exact clone families involving 490 records and 32 cross-subject payload reuse groups. The largest clone family contained 82 records spanning 26 nominally distinct subjects with sample-by-sample identical EEG payloads. Cross-phase signal reuse was also observed. Exact equality was confirmed using direct matrix comparisons (np.array_equal == True, maximum absolute difference = 0, correlation = 1.0).
Following removal of duplicated payload families, multiple machine-learning validation regimes were evaluated, including within-device classification, cross-device transfer, leave-one-subject-out (LOSO), device-stratified LOSO, permutation testing, bootstrap confidence intervals, and minimal invariant feature audits. Initial high classification performance was substantially reduced under strict domain-controlled validation. Cross-device Monte Carlo performance approached chance levels (mean AUC = 0.491 for 20→24 transfer and 0.443 for 24→20 transfer), while device-stratified LOSO produced strong inversion behavior (AUC = 0.218; inverted AUC = 0.781), indicating systematic domain-dependent sign reversal rather than stable disease generalization.
Feature-level analyses revealed substantial directional instability across acquisition domains for higher-order geometric descriptors (d_eff, radius, path length) and entropy-based metrics, whereas spectral slope preserved directional consistency between devices but did not exhibit statistically significant invariant predictive performance. Furthermore, minimal invariant models based on slope and temporal features failed to demonstrate statistically significant transferability under permutation testing.
These findings demonstrate that conventional EEG classification performance may be strongly influenced by hidden signal reuse, domain-dependent feature inversion, and acquisition-specific transformations. More broadly, the study illustrates how signal-level auditing, domain-generalization testing, and invariance analyses can identify hidden methodological structures capable of producing misleading biomarker claims.