Responsible ML
Bias, fairness, explainability, privacy, and safety — building machine learning systems that are trustworthy, equitable, and accountable.
ML Failures at Scale
A biased human decision affects hundreds per year. A biased ML model affects millions — silently, consistently, at machine speed.
Where Bias Enters
- Historical bias — data reflects past discrimination; accurate model, unjust outcomes
- Representation bias — minority groups underrepresented → poor performance for them
- Measurement bias — proxy features (zip code, name) silently encode sensitive attributes
- Aggregation bias — one model for all subgroups may be suboptimal for each
- Deployment bias — distribution shift introduces new biases post-launch
Competing Definitions
Fairness has many definitions — and they are often mutually incompatible. Choosing a metric is choosing a value judgment.
You Cannot Have It All
Chouldechova (2017): when base rates differ across groups, it is mathematically impossible to simultaneously satisfy calibration, equal FPR, and equal FNR.
SHAP Values
Each feature gets a contribution equal to its average marginal effect across all possible feature coalitions — grounded in cooperative game theory (Shapley values).
LIME & Beyond
- LIME — local surrogate model explains a single prediction; fast but unstable
- Grad-CAM — gradient-based saliency maps for CNN image explanations
- Integrated Gradients — axiomatic attribution satisfying completeness
- Counterfactuals — minimal input change to flip prediction; actionable recourse
- Attention Visualization — token attention in transformers (but causation disputed)
Differential Privacy
A formal guarantee: the model's output changes negligibly whether or not any one person's data was included in training.
Train Without Centralizing
Combines with DP: add noise to local gradients before uploading → stronger privacy guarantee, some accuracy cost.
Adversarial Robustness
Imperceptible perturbations cause confident misclassification. FGSM: perturb input in gradient direction to maximize loss.
The Regulatory Landscape
- EU AI Act (2024) — risk-tiered: prohibited, high-risk, GPAI; conformity assessments required
- GDPR Art. 22 — right to explanation; right to contest automated decisions
- US AI EO (2023) — mandatory safety evaluations for frontier models; watermarking
- NIST AI RMF — voluntary framework: Govern, Map, Measure, Manage
- China — generative AI regs: content review, real-name, security assessments
Key Takeaways
- Bias enters at data collection, labeling, features, evaluation, and deployment
- Fairness definitions are incompatible — choosing one is a value judgment
- SHAP provides Shapley-value-based global + local feature attribution
- Differential privacy gives formal guarantees; DP-SGD applies it to neural nets
- Federated learning decentralizes training; combine with DP for stronger privacy
- EU AI Act is the most comprehensive regulation — risk-tiered compliance obligations