Logo Lanfrica
  • Home
  • Atlas
  • Insights
  • Docs
  • Sign in

© 2026 Lanfrica. All rights reserved. All copyrights of the resources shown on the Lanfrica website belong to the original copyright holders, unless explicitly stated otherwise.

Davide011/ML_project_South_African_Heart_Disease

Domain:

healthcare
Creator:
Dav
Host:
Public Repository: Machine Learning & Data Mining project using the South African Heart Disease dataset. Applied PCA, Regularized Linear Regression, ANN, Logistic Regression, and Decision Trees with cross-validation for regression and classification. Includes feature scaling, EDA, and statistical tests. # Machine learning and data mining --- ## Project 🧠 South African Heart Disease Data Analysis This repository presents an in-depth analysis of the South African Heart Disease dataset, exploring how various risk factors contribute to coronary heart disease (CHD). The project combines machine learning, statistical analysis, and data visualization to investigate data structure, correlations, and predictive potential. 📊 Project Overview Dataset: Subset of the CORIS study (Rousseauw et al., 1983), including 462 individuals and 10 health-related attributes. Objective: Identify significant CHD predictors and evaluate whether linear dimensionality reduction (PCA) can capture discriminative patterns between healthy and CHD-affected individuals. 🔬 Methods & Workflow Data Preprocessing – Standardization, outlier removal, and normalization to address variable scale imbalance. Exploratory Data Analysis (EDA) – Statistical summaries, boxplots, and correlation heatmaps to visualize relationships. Dimensionality Reduction (PCA) – Performed using Singular Value Decomposition (SVD) to retain 90% of data variance. Model Evaluation – Comparison between raw and standardized PCA projections; assessment of variance explained and component significance. 📈 Findings Standardization was critical due to large variance differences across features. PCA alone was insufficient for accurate CHD classification, as clusters between classes overlapped heavily. Insights point toward Logistic Regression or nonlinear ML models (e.g., Random Forests, SVM) for improved performance. ⚙️ Skills Demonstrated Data preprocessing and feature scaling Exploratory data visualization Principal Component Analysis (PCA) Model evaluation and variance analysis Scientific communication and reporting 🧩 Technologies Python, NumPy, Pandas, Matplotlib, scikit-learn ### Group 5 Authors: * Aleksander Nagaj * Davide Ventuo * Filippo Bosi ### Dataset South african heart disease

Visit

github.com

Tags

cross-validationdataedaexploratory-data-analysisfeature-engineeringfeature-selectionlinear-regressionlogistic-regressionmatplotlibnon-linear-model+10