Business Scenario

Diabetes is a common chronic disease caused by increasing blood sugar level in the body above a certain threshold. Diabetes is classified into type-1 and type-2, in which type-2 diabetes patients account for the majority number of total diabetes patients. Early detection is very important because it helps to reduce complications and minimize death rate.  Therefore, this app is aim to build a machine learning classification model to predict type-2 diabetes.

Techniques

  • Download data from URL and extract data: requests, zipfile
  • Data manipulation: pandas, numpy
  • Data Visualization: matplotlib, seaborn
  • Handling imbalanced datasets: imbalanced-learn
  • Encoding, scaling features, model training: scikit-learn
  • Model deploying via web app: streamlit

Workflow

1. Data Collection

  • Download and Extract Data: Dataset is downloaded from Behavioral Risk Factor Surveillance System (BRFSS) 2022, which was released on July 07, 2023. BRFSS is one of the biggest on-going health-related telephone surveys system in the world, which collects health-related risk behavior, chronic health issues and other data from people more than 18 years old living in the US and participating US territories annually.
  • Selects specific columns relevant to the analysis: We select 15 features related to diabetes in the original dataset to analyze, in which DIABETE4 is output and 14 remaining attributes are input.
  • Maps categorical values to more readable labels: because original data is represented in number value, we map them to their actual values by referencing the code book.
  • Renames columns to more descriptive names: Since original column name in 2022 BRFSS is complex, we will change names of columns into simpler strings.

2. Data Preprocessing

  • Handle duplicates, missing values
  • Handle records that have the values “Refused” or “Don’t know/Not sure”
  • Handle outliers which is our of range Q1 - 1.5 IQR and Q3 + 1.5 IQR

3. Model Building

  • Applies encoding and scaling to the data
  • Uses sampling methods to handle imbalanced datasets
  • Model Training: Defines a set of basic models (e.g., Decision Tree, Random Forest, Gradient Boosting, XGBoost, LightGBM). Fine-tuning on the best performing models and then compares them based on performance metrics and select the best one

4. Model Deploying

  • Load the best model from the pickle file
  • Uses the loaded model to make predictions on new data
  • Building a streamlit app to deploy the model online

Functionalities

  • predict type-2 diabetes probabilities of an individual from user input

  • determine which indicators are highly associated with type-2 diabetes

Demo