Employee Attrition Prediction System
completed · AI and Machine Learning
A completed independent machine learning project from January 2026 that estimates individual employee attrition risk and presents SHAP explanations alongside each score.
Problem
A classification score alone does not explain which input factors influenced an attrition-risk estimate, limiting its value for careful review.
Target users
People exploring interpretable workforce attrition modeling and review-oriented model explanations.
Why it matters
Pairing each estimate with feature attribution makes model behavior easier to inspect and question than an unexplained score.
Main features
- Individual attrition-risk scoring
- SHAP explanation attached to each score
- Flask inference interface
- Interactive dashboard
- Docker-based application packaging
System architecture
The project keeps XGBoost scoring and SHAP explanation behind a Flask interface and packages the application with Docker so the model-facing path is separated from presentation.
Data flow
Tabular employee features are prepared with Pandas, passed to an XGBoost model, explained with SHAP, and returned through Flask for dashboard presentation.
Backend architecture
Flask coordinates model inference and SHAP explanations through a focused application interface.
Frontend architecture
An interactive dashboard presents attrition-risk scores and their explanation output.
AI and ML techniques
- Supervised classification
- Gradient boosting
- Feature attribution with SHAP
Deployment approach
Docker packages the application and its runtime dependencies for reproducible setup.
Security and privacy
Employee data is sensitive, so any real-world extension would require strong access controls, privacy safeguards, data governance, and review of employment-decision risks.
Evaluation strategy
A rigorous assessment would use reproducible train and test splits and evaluate discrimination, calibration, subgroup behavior, explanation stability, and employment-decision risks.
Engineering challenges
The central engineering challenge is keeping preprocessing, prediction, and SHAP attribution consistent while presenting explanations in a usable form.
Trade-offs
SHAP improves inspectability but adds computation and does not establish that a prediction is fair, causal, or suitable for an employment decision.
Impact
The project demonstrates an end-to-end interpretable machine learning workflow connecting preprocessing, prediction, explanation, and dashboard presentation.
Future improvements
Document dataset provenance, add reproducible model evaluation, test calibration and subgroup behavior, and add automated checks for preprocessing and explanation consistency.