ML 101
M12 · L01
Capstone & What's Next

Responsible ML

Bias, fairness, explainability, privacy, and safety — building machine learning systems that are trustworthy, equitable, and accountable.

01 / 13
ML 101
M12 · L01
The Stakes

ML Failures at Scale

A biased human decision affects hundreds per year. A biased ML model affects millions — silently, consistently, at machine speed.

Why This Matters Now
ML systems approve loans, screen job applicants, diagnose diseases, and assess criminal risk. Errors are not random — they can systematically disadvantage entire groups.
02 / 13
ML 101
M12 · L01
Sources of Bias

Where Bias Enters

  • Historical bias — data reflects past discrimination; accurate model, unjust outcomes
  • Representation bias — minority groups underrepresented → poor performance for them
  • Measurement bias — proxy features (zip code, name) silently encode sensitive attributes
  • Aggregation bias — one model for all subgroups may be suboptimal for each
  • Deployment bias — distribution shift introduces new biases post-launch
03 / 13
ML 101
M12 · L01
Fairness

Competing Definitions

Fairness has many definitions — and they are often mutually incompatible. Choosing a metric is choosing a value judgment.

Demo. Parity
Equal positive prediction rates across groups
Eq. Odds
Equal TPR and FPR across groups
04 / 13
ML 101
M12 · L01
Impossibility Result

You Cannot Have It All

Chouldechova (2017): when base rates differ across groups, it is mathematically impossible to simultaneously satisfy calibration, equal FPR, and equal FNR.

Implication
Every fairness intervention involves trade-offs. "Fair" is not a binary property — it is a context-dependent choice about whose errors to minimize.
05 / 13
ML 101
M12 · L01
Explainability

SHAP Values

Each feature gets a contribution equal to its average marginal effect across all possible feature coalitions — grounded in cooperative game theory (Shapley values).

Shapley Value
\phi_i=\sum_{S\subseteq F\setminus\{i\}}\frac{|S|!(|F|-|S|-1)!}{|F|!}\bigl[f(S\cup\{i\})-f(S)\bigr]
06 / 13
ML 101
M12 · L01
XAI Methods

LIME & Beyond

  • LIME — local surrogate model explains a single prediction; fast but unstable
  • Grad-CAM — gradient-based saliency maps for CNN image explanations
  • Integrated Gradients — axiomatic attribution satisfying completeness
  • Counterfactuals — minimal input change to flip prediction; actionable recourse
  • Attention Visualization — token attention in transformers (but causation disputed)
07 / 13
ML 101
M12 · L01
Privacy

Differential Privacy

A formal guarantee: the model's output changes negligibly whether or not any one person's data was included in training.

(ε, δ)-DP Guarantee
P[\mathcal{M}(D)\in S]\leq e^\varepsilon\cdot P[\mathcal{M}(D')\in S]+\delta
08 / 13
ML 101
M12 · L01
Federated Learning

Train Without Centralizing

Local
Raw data never leaves client device or hospital
Gradients
Only model updates aggregated (FedAvg)

Combines with DP: add noise to local gradients before uploading → stronger privacy guarantee, some accuracy cost.

09 / 13
ML 101
M12 · L01
Safety & Alignment

Adversarial Robustness

Imperceptible perturbations cause confident misclassification. FGSM: perturb input in gradient direction to maximize loss.

Alignment Techniques
RLHF: train on human preferences. Constitutional AI: self-critique with principles. Red-teaming: adversarial failure finding before deployment.
10 / 13
ML 101
Knowledge Check

Check whatstuck

Four questions on responsible ML — the impossibility result, proxy bias, differential privacy, and federated learning.

Question 1 of 0
Score 0/0

11 / 13
ML 101
M12 · L01
Regulation

The Regulatory Landscape

  • EU AI Act (2024) — risk-tiered: prohibited, high-risk, GPAI; conformity assessments required
  • GDPR Art. 22 — right to explanation; right to contest automated decisions
  • US AI EO (2023) — mandatory safety evaluations for frontier models; watermarking
  • NIST AI RMF — voluntary framework: Govern, Map, Measure, Manage
  • China — generative AI regs: content review, real-name, security assessments
12 / 13
ML 101
Key Takeaways
Summary

Key Takeaways

  • Bias enters at data collection, labeling, features, evaluation, and deployment
  • Fairness definitions are incompatible — choosing one is a value judgment
  • SHAP provides Shapley-value-based global + local feature attribution
  • Differential privacy gives formal guarantees; DP-SGD applies it to neural nets
  • Federated learning decentralizes training; combine with DP for stronger privacy
  • EU AI Act is the most comprehensive regulation — risk-tiered compliance obligations
13 / 13