AI/MLEarlier work · 2023 – 2026

Heart Disease Risk Prediction

Sixth-semester coursework from 2023 that compared KNN, SVM, Random Forest, a voting ensemble and XGBoost on the widely used 1,025-row heart-disease dataset, later packaged as a Streamlit app that turns the model’s probability into a low / moderate / high risk band.

Technology stack

Pythonscikit-learnXGBoostpandasStreamlit
01

What it is

  • A training notebook comparing five classifiers on structured clinical features such as age, chest-pain type, resting blood pressure, cholesterol, maximum heart rate and ST depression.
  • A Streamlit app that loads a single joblib artifact, builds the feature vector in a fixed order, and shows the predicted probability as a risk band and gauge.
02

Why no accuracy is shown

The dataset has 1,025 rows but only 302 unique ones, and the notebook splits without de-duplicating, so most test rows have an identical copy in the training split. The accuracies it reports (up to 89.42%) are therefore not a reliable estimate of performance on new patients, and the probabilities are not calibrated.

Sources

Checked against the repositories on 22 September 2026.