AI & MLDeep
Intermediate

AI Ethics & Bias in Machine Learning

12 min read

Learn
Deep Reading
Estimated 12 mins
Prereq
Intermediate
Basic ML concepts helpful
Interactive
Static Playbook
Static guide & reference tables

Bias isn't a bug — it's usually the data working as designed

Machine learning models learn patterns from historical data, and historical data reflects historical inequities, sampling gaps, and human decisions — including biased ones. A model trained on this data doesn't introduce bias out of nowhere; it faithfully learns and often amplifies whatever patterns, including unfair ones, were already present in what it was trained on.

This reframes the engineering problem: bias mitigation isn't primarily about fixing a broken algorithm, it's about recognizing that a model optimizing purely for predictive accuracy on biased historical data will reproduce that bias with high fidelity — and deciding, deliberately, whether and how to correct for that. Ignoring it isn't neutral; it's a choice to let historical inequities propagate into automated decisions at scale.

How bias enters the pipeline

Sampling bias: training data that doesn't represent the population the model will actually be used on. Facial recognition systems trained predominantly on lighter-skinned faces perform measurably worse on darker-skinned faces — not because of anything inherent to the task, but because the training distribution didn't include enough of the population being evaluated.

Label bias: the "correct answers" used to train a model can themselves encode human bias. A hiring model trained on "who got hired historically" learns whatever biases existed in past hiring decisions, treating them as ground truth rather than as decisions that may have been discriminatory.

Proxy variables: removing an explicitly protected attribute (race, gender) from a dataset doesn't remove bias if other correlated features act as proxies for it — zip code can proxy for race, name can proxy for gender or ethnicity, and a model can reconstruct the excluded signal from these correlates without ever seeing the original attribute directly.

Feedback loops: a biased model's outputs can become future training data. A predictive policing system that sends more patrols to a neighborhood generates more recorded incidents there (simply because more patrols are looking), which then reinforces the model's belief that the neighborhood has more crime — regardless of the true underlying rate.

Note

A widely studied 2016 ProPublica investigation found that the COMPAS algorithm, used by US courts to predict recidivism risk, was substantially more likely to falsely flag Black defendants as high-risk than white defendants, while white defendants who did reoffend were more often incorrectly scored as low-risk. The case became a landmark example of how a model can meet a formal definition of accuracy overall while still producing systematically unequal error rates across groups — with direct consequences for people's liberty.

Fairness metrics: there is no single definition of "fair"

A core, uncomfortable truth in fairness research: several intuitively reasonable definitions of fairness are mathematically incompatible with each other in most real situations, meaning a model literally cannot satisfy all of them at once for groups with different base rates. Engineers have to choose which notion of fairness matters for a given application.

Demographic parity requires that a model's positive prediction rate is equal across groups — the same proportion of each group gets approved for a loan, for instance. Equal opportunity requires that the true positive rate is equal across groups — among people who would actually repay a loan, an equal proportion get approved regardless of group. Predictive parity requires that, among people the model predicts positively for, the actual positive rate is the same across groups — a risk score of 0.8 should mean roughly the same real-world likelihood regardless of the person's group.

A landmark theoretical result showed that when base rates differ between groups (which they often do, precisely because of the historical inequities discussed above), it's mathematically impossible to satisfy equal opportunity and predictive parity simultaneously, except in trivial cases. Choosing a fairness metric is therefore a value judgment about what kind of error the application can least tolerate — it isn't a purely technical decision with one correct answer.

python

Common fairness metrics

MetricWhat it requiresGood fit when
Demographic parityEqual positive prediction rate across groupsAccess/allocation decisions where equal representation matters most
Equal opportunityEqual true positive rate across groupsErrors that deny a deserved benefit are the primary harm to avoid
Predictive parityEqual precision (positive predictive value) across groupsA prediction's meaning must be consistent regardless of group
Individual fairnessSimilar individuals receive similar predictionsCase-by-case consistency matters more than group-level statistics

Responsible AI practices for engineers

Audit training data before training a model, not after deployment — check group representation, label quality, and known historical biases in the data source. Disaggregate every evaluation metric by relevant subgroup, not just overall accuracy; a model can look excellent in aggregate while performing badly for a specific group that's underrepresented in the test set.

Treat proxy variables seriously — removing a protected attribute is not sufficient if correlated features remain; test whether the model's predictions correlate with the protected attribute even when it's excluded from the input. Document known limitations explicitly — model cards and datasheets, describing what a model was trained on, what populations it was validated against, and what it should not be used for, are now standard practice for responsible deployment, not optional extras.

Keep a human in the loop for high-stakes decisions — lending, hiring, criminal justice, medical diagnosis — where an automated model's error has serious consequences for a real person, rather than treating model output as a final, unreviewable decision. Monitor for feedback loops after deployment, since a model's own outputs can quietly bias the data it's retrained on later.

Note

There is no single technique that makes a model "unbiased" — fairness interventions (reweighting data, adjusting decision thresholds per group, adversarial debiasing) each involve real tradeoffs against overall accuracy and against other fairness definitions. The engineering responsibility isn't to eliminate every tradeoff, which is often mathematically impossible, but to make the tradeoffs explicit, defensible, and aligned with the actual stakes of the application.

What's next

Fairness auditing depends on the same evaluation discipline used throughout machine learning — clean train/validation/test splits, honest metrics, and skepticism about aggregate numbers. If you haven't already, revisiting Overfitting & Regularization is useful context: many of the same evaluation pitfalls that hide overfitting also hide unfair, subgroup-specific model failures.

I build these systems professionally.

Whether it's a RAG pipeline, analytics migration, or AI workflow — let's talk.

Need custom AI or MarTech setup? Let's build together.