AI Ethics: Building Responsibly
1 · The lesson
readA machine-learning system isn't a piece of code; it's a decision-making process operating at scale on real people. A misconfigured spam filter loses an email. A misconfigured risk-assessment model misclassifies people as flight risks and they spend an extra month in jail. The same engineering practices that mostly work for the spam filter — train, evaluate on a holdout, ship — produce harm in the second case unless you actively design against it.
This lesson is about that active design work: the harm categories, the practical measurements, the documentation practices, the regulatory shape of 2026, and the questions you should answer for every system you ship. There is no "ethical AI library import". This is judgement plus a small set of tools.
1. Why This Matters — Scale Compounds Mistakes
A human reviewer at a bank, even a biased one, can review maybe 50 loan applications a day. A model reviews 500,000. The same biased instinct, multiplied by the same volume, produces a population-level disparity that hand-review never could. ML doesn't introduce new failure modes so much as it industrialises the existing ones.
Three properties make ML harms qualitatively different from non-ML harms:
- Scale — a single trained model can affect millions of decisions per day.
- Opacity — the reasoning isn't auditable in the way a human's notes might be.
- Path dependence — biased decisions early in a pipeline shape the data that trains the next model. The error compounds across versions.
If "who's harmed if this is wrong, and how badly?" isn't a question you can answer for the system you're building, you don't yet understand the problem.
2. The Five Harm Categories
These overlap, but it's useful to name them separately because the mitigations differ.
| Harm | Example | Primary mitigation |
|---|---|---|
| Bias / discrimination | Hiring AI screens out women's CVs | Disaggregated metrics, debiasing, oversight |
| Privacy violation | Training data memorised and leaked verbatim | Differential privacy, data minimisation |
| Misinformation / deepfakes | Generated audio of a politician saying things they didn't say | Provenance, watermarking, platform policy |
| Job displacement | Automating call centres with no transition plan | Outside the ML team's control alone — design and policy |
| Autonomous decision-making without recourse | Algorithm denies you a loan, no appeal path | Human-in-the-loop, transparent process, appeals mechanism |
Most production ML systems touch multiple categories. A facial recognition system in policing is at once a bias problem (higher error rates on darker skin), a privacy problem (mass surveillance), and a recourse problem (you can't appeal a stop based on an algorithm you can't see).
3. Bias in Practice
"Bias" in the ML sense means a model treats demographically similar inputs differently in ways correlated with protected attributes — race, gender, age, disability, religion. Three sources:
- Biased training data. Loan applications historically approved at higher rates for one group → a model trained on those decisions reproduces the disparity.
- Biased labels. "Good employee" labelled by managers who promoted men → the model learns that men are good employees.
- Biased problem framing. Choosing "arrest" as the proxy for "crime" — arrests are themselves biased — bakes the bias into the very target you're predicting.
The third is the deepest. You can fix the data; you can fix the labels; you cannot fix a target variable that doesn't measure what you actually care about.
Detection — Disaggregated Metrics
Aggregate accuracy hides everything. A model can be 95% accurate overall and 99% accurate on group A while being 60% accurate on group B. The aggregate number doesn't tell you that; the per-group breakdown does.
The starting move is a table of (group, metric) cells. For a classifier:
import numpy as np from sklearn.metrics import precision_score, recall_score, f1_score def per_group_metrics(y_true, y_pred, groups): out = [] for g in np.unique(groups): mask = groups == g out.append({ "group": g, "n": int(mask.sum()), "precision": precision_score(y_true[mask], y_pred[mask], zero_division=0), "recall": recall_score(y_true[mask], y_pred[mask], zero_division=0), "f1": f1_score(y_true[mask], y_pred[mask], zero_division=0), }) return out
That's the minimum. Larger systems compute false-positive rate, false-negative rate, calibration, equal opportunity, demographic parity, and selection rate per group. Different fairness definitions are mutually incompatible — you cannot satisfy them all simultaneously except in degenerate cases (Chouldechova 2017, Kleinberg et al. 2017). Pick the definition that matches the harm you're trying to prevent, not the one that gives the nicest table.
Mitigation
In rough order of intrusiveness:
- Re-weighting — change sample weights during training to balance group representation. Cheap, often effective.
- Re-sampling — over-sample under-represented groups, or under-sample over-represented ones. Effective for class imbalance combined with group imbalance.
- Adversarial debiasing — train an adversary to predict the protected attribute from your model's representations; penalise the main model when the adversary succeeds.
- Post-processing — adjust thresholds per group (e.g. equalised odds post-processing). Controversial because it makes the disparate treatment explicit.
- Reframe the problem — sometimes the right answer is "don't predict that".
Libraries: Fairlearn (Microsoft, Python) and AIF360 (IBM) implement these mitigations with consistent APIs. They make the mechanics easy; they cannot make the judgement call about which fairness definition matters for your problem.
4. Privacy
Models can memorise their training data. Large language models can be prompted into reciting copyrighted text or personal information verbatim from their training set (Carlini et al. 2021). Image classifiers can leak training images through membership inference attacks ("was this exact image in your training set?").
Three layers of defence:
Data Minimisation
The data you don't collect can't leak. Audit what you actually need; delete the rest. "We might use it later" is not a justification.
Differential Privacy in Training (DP-SGD)
A trained model is differentially private if the model is provably almost-the-same whether or not any single training example was included. The standard implementation, DP-SGD, clips per-example gradients and adds noise during training. It costs accuracy (often 1–10 percentage points) for a quantifiable privacy guarantee (epsilon).
Libraries: opacus for PyTorch, tensorflow-privacy for TF.
Federated Learning
Train the model on data that never leaves user devices. Each device computes a gradient update locally; the server aggregates updates without seeing the raw data. Used at scale for Gboard's next-word prediction (Google), iOS Siri suggestions (Apple), and similar mobile keyboards. Strong privacy story; significant engineering complexity; reduces but doesn't eliminate leakage from the aggregated updates themselves.
Membership Inference Attacks
The test for whether your model leaks training data. Train a "shadow model" that tries to predict whether a given input was in the original training set. If it succeeds reliably, your model leaks. This isn't paranoia — for over-parameterised models trained without DP, the attack often achieves 70–90% accuracy.
5. Explainability
Two related-but-distinct things called by the same name:
- Interpretability — the model is structurally transparent. A linear model with five features; a small decision tree. You can read the model and know what it's doing.
- Post-hoc explanation — the model is a black box, but you generate explanations alongside its predictions. Why did this loan get denied? Because of feature X (estimated contribution: …).
Post-hoc tools:
- LIME — fits a small interpretable model to the black box's behaviour in the neighbourhood of one input. Local explanation per prediction.
- SHAP — Shapley values from game theory. Attributes the prediction to each feature in a way that satisfies several mathematical fairness properties. The current standard.
- Integrated Gradients — for deep networks. Attributes the prediction to input features by integrating the gradient along a path from a baseline.
- Attention visualisation — for transformers. Useful intuition but well-documented as not a faithful explanation of model behaviour (Jain & Wallace 2019).
The honest caveat: explainability is not trustworthiness. A confident-looking SHAP plot for a biased model doesn't make the model fair. It tells a story about what the model is doing — which is necessary for debugging and required by regulators in some domains — but it doesn't fix the underlying decision quality. Use explanations to inform oversight, not as a substitute for it.
6. Documentation as Ethics
Two documents that should ship with every model:
Model Cards (Mitchell et al. 2019)
A short structured document attached to a model. Sections:
- Intended use, intended users, out-of-scope uses
- Training data summary
- Evaluation data summary
- Performance metrics, disaggregated by demographic groups
- Ethical considerations and known limitations
- Caveats and recommendations
Hugging Face automates the scaffolding for models on their hub. The standard practice is now: no model card → not a serious release.
Datasheets for Datasets (Gebru et al. 2018)
The dataset-side equivalent. Sections:
- Motivation (why was this dataset created?)
- Composition (what's in it, how many examples, what's the schema?)
- Collection process (how was it gathered, were consents obtained?)
- Preprocessing / cleaning
- Uses (what's it appropriate for? what isn't it appropriate for?)
- Distribution and maintenance
These aren't busywork. They're the artefact future users have to decide whether to trust your work. Every dataset that's caused embarrassment in the last decade had a missing or vague datasheet.
7. Real-World Incidents — Brief
- COMPAS — a recidivism risk-assessment tool used in US courts. ProPublica's 2016 analysis found roughly 2× higher false-positive rates on Black defendants than on white defendants for the same true outcome. The vendor disputed the methodology; the underlying mathematical impossibility of satisfying multiple fairness definitions (Chouldechova 2017) became famous because of this case.
- Facial recognition error rates. NIST's 2019 evaluation of 189 commercial systems found 10–100× higher false-match rates on women, darker-skinned subjects, and Asian populations vs. white men. This has been the basis for several US city-level moratoria on government use of face recognition.
- Amazon's experimental hiring screening tool (reported 2018). Trained on a decade of resumes that were predominantly from men; the model learned to down-rank resumes containing the word "women's" (as in "women's chess club captain"). Amazon scrapped the project before it was deployed at scale.
- Apple Card credit limits (2019). Couples with shared finances reported the card offering the husband 10–20× higher credit limits than the wife. The bank issued a statement saying the algorithm didn't use gender as a feature — a confession, not a defence, since correlated features (income history, account ownership) carried the signal.
The common thread is not malice. It's "we shipped without doing the disaggregated evaluation that would have caught this".
8. The Regulatory Landscape 2026
| Region | Status | Notes |
|---|---|---|
| EU AI Act (2024, in force) | High-risk categories (employment, credit, education, law enforcement, critical infrastructure) require risk assessment, transparency, human oversight, accuracy/robustness/security obligations. "Unacceptable risk" applications (social scoring, real-time biometric ID with exceptions) are banned. | The de-facto global baseline — companies serving the EU adjust globally rather than maintaining two versions. |
| US | Patchwork. State laws (NYC AI hiring bias audit, California AB-2273). Federal: NIST AI Risk Management Framework (voluntary). Executive Order 14110 (2023) imposes obligations on the largest models. | No single comprehensive federal law as of 2026. |
| UK | Pro-innovation framework — existing regulators (FCA, MHRA, etc.) extend their authority into AI. | Less prescriptive than the EU; more flexible, harder to predict. |
| China | Generative AI regulations (2023), algorithmic recommendation regulations. | Very different priors and goals; relevant to anyone with users in China. |
The practical implication: if you ship in multiple jurisdictions, the EU AI Act sets the floor. Build documentation, transparency, and oversight to its standard and you're compliant most places. Build to a lower standard and you have a regulatory project ahead of you the moment you grow.
9. The Practical Checklist
Before you ship an ML system, answer these:
1. Who is harmed if the model is wrong? Be specific. Name groups, name failure modes, name severity.
2. Have I measured performance disaggregated by every relevant demographic? Aggregate metrics are insufficient. Always look at the table.
3. Is there a human in the loop for high-stakes decisions? Loan denial, criminal risk assessment, hiring screening, medical triage — model assistance, not model decision.
4. Can affected users appeal? Can they request an explanation? If not, you're optimising decisions that have no recourse — almost always a design failure.
5. What inputs and outputs am I logging, and for how long? Without logs you cannot audit later. With logs you have a privacy obligation. Both have to be designed.
6. Have I published a model card and a datasheet? Honest documentation is the cheapest accountability mechanism that exists.
7. What's my monitoring story for drift, both input and output? Distribution shift is when models start hurting people; you need to detect it.
8. What's my rollback story when something does go wrong? Hours, not weeks. See deployment.
The questions are simple. Answering them honestly during a tight launch deadline is hard. That's the job.
10. The Honest Closing
There is no import ethics_filter; ethics_filter.audit(model). Fairlearn, AIF360, SHAP, opacus — these are useful tools, but they're hammers. The skill is knowing where to hit, and deciding what acceptable looks like.
Three habits separate teams that ship responsibly from teams that ship and apologise:
- Disaggregated metrics are non-negotiable. Look at the table before you announce numbers.
- The documentation is part of the product. Model card, datasheet, monitoring plan, rollback plan — all of it ships, all of it is reviewed.
- Someone is accountable per decision. Not "the algorithm decided"; a named person or team is responsible for what the system does, including its mistakes.
Everything in this lesson is a tool for those habits. None of them are a substitute.
Common Mistakes
1. Optimising only overall accuracy.
A 95% accurate model with 99% accuracy on the majority group and 60% on a minority is not "95% accurate" in any meaningful sense. Always report disaggregated metrics. Always.
2. Skipping model and dataset documentation.
"We'll write the model card after we ship." You won't. Future you will inherit an undocumented model and have no idea what it's safe to use it for. Six months later, you'll be the team in the news.
3. Assuming "the data speaks for itself".
Data is collected by humans with priorities, labelled by humans with priors, sampled by processes with selection effects. The data doesn't speak — it whispers what the upstream process biased it to whisper. Audit data provenance the way you audit code provenance.
4. No recourse mechanism for users.
The model denies someone something — a loan, a job, parole, an insurance claim. They ask why. There is no path to a human. There is no path to an explanation. There is no path to challenge the decision. Designing recourse out of a system because it's expensive is the failure mode regulators are increasingly punishing.
5. Treating fairness as a single number.
Demographic parity, equal opportunity, equalised odds, calibration — different definitions, mathematically incompatible in general. You have to pick which fairness you're optimising for, and the choice is value-laden, not technical. Publishing one number without naming the definition is misleading.
6. Conflating explainability with safety.
SHAP plot looks reasonable → model is fine. No. Explanations describe behaviour; they don't validate it. A confident wrong answer with a confident wrong explanation is still wrong. Use explanations for debugging and regulatory compliance, not as a substitute for evaluation.
🎯 Your Turn — Per-Group Precision and Recall
You're handed a small inline dataset, a trained classifier, and a "group" column with values A or B. Compute per-group precision and recall. Print whether they differ meaningfully (defined as more than 10 percentage points absolute on either metric). Requirements:
- Use the inline data provided in the skeleton.
- Compute precision and recall for each group.
- Print a small table.
- Print a one-line verdict — either "Disparity detected on
" or "No major disparity detected".
Skeleton:
import numpy as np from sklearn.linear_model import LogisticRegression from sklearn.metrics import precision_score, recall_score # Inline toy data — feature x, label y, group g rng = np.random.default_rng(42) X = rng.normal(size=(400, 2)) g = rng.choice(["A", "B"], size=400, p=[0.5, 0.5]) # Group B has noisier labels — induces a disparity y = (X[:, 0] + 0.5 * X[:, 1] > 0).astype(int) flip = (g == "B") & (rng.random(400) < 0.3) y_noisy = np.where(flip, 1 - y, y) # Train a classifier on the noisy labels model = LogisticRegression().fit(X, y_noisy) y_pred = model.predict(X) # TODO 1: compute precision and recall for group A # TODO 2: compute precision and recall for group B # TODO 3: print a table — group, precision, recall (3 decimals) # TODO 4: compare absolute deltas; print disparity verdict
Hint 1 — Masking by group
mask_a = g == "A". Pass y_noisy[mask_a] and y_pred[mask_a] to precision_score / recall_score. Use zero_division=0 on small slices to avoid warnings.
Hint 2 — The disparity check
abs(precision_a - precision_b) > 0.10 or abs(recall_a - recall_b) > 0.10 → disparity. Pick the larger of the two deltas to name in the verdict line.
Show full solution
import numpy as np from sklearn.linear_model import LogisticRegression from sklearn.metrics import precision_score, recall_score # Inline toy data rng = np.random.default_rng(42) X = rng.normal(size=(400, 2)) g = rng.choice(["A", "B"], size=400, p=[0.5, 0.5]) y = (X[:, 0] + 0.5 * X[:, 1] > 0).astype(int) flip = (g == "B") & (rng.random(400) < 0.3) y_noisy = np.where(flip, 1 - y, y) model = LogisticRegression().fit(X, y_noisy) y_pred = model.predict(X) def metrics_for(group_label): mask = g == group_label return { "group": group_label, "n": int(mask.sum()), "precision": precision_score(y_noisy[mask], y_pred[mask], zero_division=0), "recall": recall_score(y_noisy[mask], y_pred[mask], zero_division=0), } rows = [metrics_for("A"), metrics_for("B")] print(f"{'group':<6} {'n':>4} {'precision':>10} {'recall':>8}") for r in rows: print(f"{r['group']:<6} {r['n']:>4} {r['precision']:>10.3f} {r['recall']:>8.3f}") # Compute disparities dp = abs(rows[0]["precision"] - rows[1]["precision"]) dr = abs(rows[0]["recall"] - rows[1]["recall"]) THRESHOLD = 0.10 if dp > THRESHOLD or dr > THRESHOLD: worst = "precision" if dp >= dr else "recall" delta = max(dp, dr) print(f"\nDisparity detected on {worst}: {delta:.3f} absolute difference between groups") else: print(f"\nNo major disparity detected (max delta: {max(dp, dr):.3f})")
Sample output:
group n precision recall A 198 0.927 0.949 B 202 0.741 0.685 Disparity detected on recall: 0.264 absolute difference between groups
Notice what this exercise doesn't do: fix the disparity. That's deliberate. The first step in addressing bias is measuring it, and most ML systems in production today don't even do that. The next step — once you know there's a 26-point recall gap between groups — is the harder, judgement-laden work covered in section 3: reweight, resample, adversarially debias, post-process, or step back and ask whether the problem framing is wrong. The hammer is easy; deciding where to hit is the job.
Try extending the script: add false-positive rate and false-negative rate per group; recompute after retraining on y (the clean labels) instead of y_noisy; observe how much of the disparity was driven by the label noise versus the model itself.
What You Learned
- ML harms differ from non-ML harms in scale, opacity, and path dependence.
- The five harm categories: bias/discrimination, privacy, misinformation, displacement, autonomous decisions without recourse.
- Bias sources are data, labels, and problem framing — the last is the deepest. Detect with disaggregated metrics; mitigate with reweighting, resampling, adversarial debiasing, post-processing, or reframing.
- Different fairness definitions are mathematically incompatible. Pick the one matching the harm; don't claim to satisfy all of them.
- Privacy mitigations: data minimisation, DP-SGD, federated learning. Membership inference attacks are the test.
- Explainability (LIME, SHAP, Integrated Gradients) describes behaviour; it does not validate it. Explainability ≠ trustworthiness.
- Model cards and datasheets for datasets are non-negotiable documentation. Ship them or you didn't ship the model.
- EU AI Act is the de-facto global baseline. Build to it.
- The practical checklist — who's harmed? disaggregated metrics? human-in-the-loop? recourse? logging? documentation? monitoring? rollback? — is the job.
- There is no
import ethics. Tools matter, but the work is judgement.
You've reached the end of the AI track. The deeper specialty tracks — genai, agents, reinforcement learning — pick up from here. Every one of them inherits the practices in this lesson. Ship responsibly.