Responsible AI: Fairness, Bias, and Interpretability Questions
Building ML and AI systems that are fair, explainable, and safe. Covers identifying and mitigating bias, fairness metrics and tradeoffs, model interpretability and explainability techniques, label-bias feedback loops, and responsible and safe development practices for production models. Emphasizes accountability and transparency as first-class design constraints.
What is a proxy variable? Give two production examples where a seemingly innocuous feature, such as ZIP code or browsing history, can proxy for a protected characteristic and cause indirect discrimination. Describe detection techniques and a concrete mitigation.
Sample Answer
Direct answer: a proxy variable is a feature that is not itself a protected attribute but is statistically correlated with one closely enough that using it produces the same discriminatory effect as using the protected attribute directly.
Structured elaboration. Proxies arise because protected attributes like race, gender, or age are embedded in the broader social and economic structure that generates most other data. ZIP code is the textbook example: because of historical residential segregation, ZIP code can correlate strongly with race in many US metro areas, so a model that uses ZIP code to price insurance or approve a loan can reproduce racial disparities even though race is never an explicit input. Browsing or purchase history is a second common proxy: shopping patterns can correlate with gender or age closely enough to leak the same signal a directly-collected demographic field would.
Worked example. A lending model drops "race" from its inputs but keeps ZIP code, years at current address, and college attended. If a regulator or auditor regresses the model's approval decisions against race using only these "neutral" features, they can often recover most of the disparity that direct use of race would have produced, because the combination of features jointly encodes the same information.
Detection. Compute the correlation (or mutual information, which also catches non-linear relationships) between each candidate feature and each protected attribute in your training population. Follow up with a leakage-style test: train a small classifier to predict the protected attribute FROM the remaining features; a leakage classifier with high accuracy is strong evidence the feature set as a whole functions as a proxy, even if no single feature has a high pairwise correlation.
Mitigation. Options in increasing order of aggressiveness: (1) keep the feature but monitor outcome disparities and correct downstream (a threshold or post-processing fix); (2) transform the feature to strip the correlated component (for example replacing raw ZIP code with a broader regional cost-of-living index that carries most of the legitimate signal but less of the demographic correlation); (3) remove the feature outright, accepting some accuracy loss, when its predictive value is small relative to its correlation with the protected attribute.
Trade-offs and pitfalls. Removing every feature that has ANY correlation with a protected attribute is usually not viable in practice, since features like income or education level are also correlated with protected attributes for the same structural reasons and often carry real predictive signal a business cannot simply discard; the goal is not zero correlation but understanding and justifying the residual correlation, and documenting that judgment.
Define disparate impact and disparate treatment in machine learning, with a concise example of each drawn from a hiring-recommendation model. Explain why one form is more likely to be regulated in certain jurisdictions, and what documentation you would keep to demonstrate compliance.
Sample Answer
Direct answer
Disparate treatment is intentionally using a protected characteristic (or a clear stand-in for it) as a factor in a decision; disparate impact is a facially neutral practice that produces a disproportionate adverse effect on a protected group, regardless of intent. In machine learning, disparate treatment shows up as a protected attribute (or an unmistakable proxy for it) being an explicit model input or an explicit rule; disparate impact shows up as an outcome gap produced by features that never mention the protected group at all. Disparate impact is the doctrine regulators actually apply to algorithmic hiring tools, because proving an algorithm's "intent" is close to meaningless, so I would keep audit records and validation evidence built around disparate impact from day one.
Structured elaboration
Disparate treatment, drawn from a hiring-recommendation model: the model, or the pipeline around it, uses gender directly as a feature, or a recruiter overrides the model's score downward specifically for candidates it knows are pregnant. The decision explicitly turns on the protected characteristic. Under United States employment law this requires showing intent (though intent can be inferred from clearly discriminatory rules), and it is illegal per se once shown, with no "business necessity" defense available.
Disparate impact, drawn from the same domain: the model screens resumes using "years of continuous full-time employment" as a strong positive feature. Nothing about gender appears anywhere in the model, but the feature systematically disadvantages applicants (disproportionately women) who took career breaks for caregiving, producing a lower selection rate for that group even though the practice is facially neutral and was not designed to discriminate. This theory does not require proving intent: Griggs v. Duke Power Co. (1971) established that a facially neutral employment practice with a disproportionate adverse effect on a protected group violates Title VII of the Civil Rights Act unless the employer can show the practice is job-related and consistent with business necessity.
Why disparate impact is the doctrine more likely to be regulated for algorithmic systems. An algorithm does not have intent in the legal sense; it optimizes a loss function against training data. Proving discriminatory INTENT behind a gradient-descent update is close to a category error, so litigation and regulation targeting algorithmic hiring tools has concentrated on disparate impact, which only requires showing an outcome gap and then litigating whether the practice is job-related. This is operationalized by the EEOC's Uniform Guidelines on Employee Selection Procedures (1978) via the four-fifths rule: if a group's selection rate is less than 80% of the highest-selection-rate group's rate, that is treated as evidence of adverse impact. More recent AI-specific regulation follows the same pattern: New York City's Local Law 144 (effective 2023) requires any employer using an "automated employment decision tool" to commission an independent annual bias audit that computes adverse-impact-style ratios by protected category and publish a summary, again a disparate-impact framing rather than a disparate-treatment one, because that is the theory that is actually enforceable against an algorithm.
Documentation to keep for compliance.
- Per-cycle selection-rate-by-group records and the computed four-fifths ratio, kept on a regular cadence (not just once at launch), matching the Uniform Guidelines' recordkeeping expectations.
- A feature inventory confirming no protected attribute, or an unambiguous proxy for one, is an explicit model input (this is your disparate-treatment defense).
- Job-relatedness / business-necessity validation evidence for any feature shown to correlate with a protected group and with disparate impact, so it can be defended if challenged.
- A model card or equivalent: training data provenance, feature list, version history, and the date and rationale of any change made to the model or its thresholds. This last point matters specifically because an undocumented change made "to fix the numbers" is exactly the kind of intervention that can convert a disparate-impact defense into a disparate-treatment problem, per the Ricci v. DeStefano (2009) line of reasoning.
- The independent bias audit report itself where required (as under NYC Local Law 144), including the methodology and the adverse-impact ratios by category, and evidence it was published and that any required candidate notice was given.
- Applicant-flow demographic records (EEO-1-style categories) at each pipeline stage, so a selection-rate gap can be traced to the specific stage that introduced it.
Worked example
Say 200 applicants from group A and 200 from group B apply, and the model recommends 120 of group A (60% selection rate) and 80 of group B (40% selection rate). The adverse impact ratio is:
adverse impact ratio=max(SRA,SRB)min(SRA,SRB)=0.600.40=0.667
0.667 is below the 0.8 four-fifths threshold, so this pattern would be flagged as evidence of disparate impact regardless of whether the model ever saw group membership as a feature; the next step is to check whether any input feature driving the gap can be justified as job-related, not to ask whether anyone "intended" the gap.
Trade-offs and pitfalls
- Removing the protected attribute as a feature does not remove disparate-treatment risk if a near-perfect proxy remains (zip code standing in for race, name standing in for national origin) and it does nothing at all for disparate-impact exposure, since disparate impact is about the OUTCOME, not the inputs.
- A common wrong turn is treating "we don't use the protected attribute" as sufficient compliance evidence. It rules out one theory of liability (disparate treatment) and says nothing about the other (disparate impact), which is the one regulators actually pursue for algorithmic tools.
- Fixing a disparate-impact finding by explicitly adjusting outcomes based on group membership can itself create disparate-treatment exposure; the business-necessity defense and careful, documented feature-level fixes are safer than ad hoc score adjustments made after the fact.
- Keep the bias-audit and compliance documentation process independent of the model-training team's own sign-off; a self-graded audit carries much less weight if the finding is ever challenged.
Define the disparate impact ratio and the 80 percent rule used in US employment-law contexts. Show how to compute the ratio from a model's predictions, discuss the limitations of the 80 percent rule, and explain when you would prefer a ratio-based test over a difference-based fairness test.
Sample Answer
Direct answer: the disparate impact ratio compares the selection rate of a protected group to that of the most-favored group; the 80 percent rule (from the EEOC's Uniform Guidelines on Employee Selection Procedures) treats a ratio below 0.8 as evidence of possible discrimination that shifts the burden to the employer to justify the practice.
Structured elaboration. Formally, DI ratio=selection rate of most-favored groupselection rate of protected group. It is a RATIO test, unlike demographic parity's DIFFERENCE test, which matters because a difference of a few points can look small in absolute terms but represent a large relative disadvantage when base selection rates are low.
Worked example (executed).
group_A_selected, group_A_total = 48, 200 # 24.0% selection rate
group_B_selected, group_B_total = 90, 300 # 30.0% selection rate
rate_A = group_A_selected / group_A_total
rate_B = group_B_selected / group_B_total
ratio = min(rate_A, rate_B) / max(rate_A, rate_B)
print(f"rate_A = {rate_A:.3f}, rate_B = {rate_B:.3f}, ratio = {ratio:.3f}")
Output (literal stdout of the code above, run fresh): rate_A = 0.240, rate_B = 0.300, ratio = 0.800 (exactly 80%). This example sits precisely on the regulatory threshold: a small additional shortfall for group A (say 47 selected instead of 48) would drop the ratio below 0.8 and trigger scrutiny under the guideline.
Trade-offs and pitfalls. (1) The 80% rule is a rule of thumb from 1978 guidance, not a statistical significance test; a small sample can swing the ratio a lot, so courts and regulators also look at statistical significance (commonly a z-test on the rate difference) alongside the ratio. (2) The ratio direction matters: always divide the SMALLER rate by the LARGER rate, never the reverse, or you will get a number above 1 and miss a real disparity. (3) Prefer the ratio test when absolute selection rates are low (a 2-point difference at 5% vs 7% is a much bigger relative gap than the same 2-point difference at 45% vs 47%); prefer a difference-based test like demographic parity difference when you specifically care about the absolute number of people affected, since a 0.79 ratio at very high volume can represent thousands of affected people while the same ratio at low volume affects a handful.
A product manager proposes a personalized onboarding flow that would use inferred protected attributes to improve relevance. As the engineer, how do you advise them to balance personalization benefits against fairness and privacy risks? List alternatives that avoid directly inferring protected attributes, and governance steps before deployment.
Sample Answer
Direct answer
Advise the product manager that inferring a protected attribute to personalize onboarding trades a soft, unvalidated relevance gain for a hard, well-documented fairness and privacy liability, and that better proxies for the actual underlying need (what the user is trying to accomplish, not who they are demographically) usually get most of the personalization benefit without ever inferring a sensitive category. If the team still wants to pursue attribute-based personalization after weighing that trade-off, it needs an explicit governance gate before launch, not a quiet ship.
Structured elaboration
Why inferred protected attributes are a worse bet than they look. Inference from behavioral or demographic signals (name, browsing pattern, device, location) is frequently inaccurate for the individual even when the underlying method is proven at a population level, meaning a meaningful fraction of users get onboarding tailored to the WRONG inferred category, which reads as worse than generic onboarding, not better. It also converts an implicit signal into an explicit one inside the product: once "this user's inferred gender/ethnicity/age bracket" exists as a stored field, it becomes discoverable, subpoenable, and reusable for purposes well beyond the original onboarding-personalization use case, which is exactly the kind of scope creep that regulators and privacy audits flag first.
Alternatives that avoid inferring the protected attribute directly.
- Explicit, opt-in preference capture: ask the user what they want out of onboarding directly ("what are you here to do"), rather than inferring who they are and guessing what they probably want. This is more honest, usually more accurate, and requires no inference model at all.
- Behavioral personalization on the CURRENT session: adapt onboarding based on what the user has actually clicked or searched for in this session, which personalizes on intent rather than identity and carries none of the same fairness risk.
- Cohort-level, not individual-level, differentiation: if there is a genuine, validated need to tailor onboarding by broad user segment (e.g. "new to this product category" versus "switching from a competitor"), capture that segment through an explicit onboarding question, not an inference model over protected characteristics.
- Progressive disclosure: default every user to the same general onboarding flow, and let INDIVIDUAL, EXPLICIT actions unlock more tailored paths as the user demonstrates a preference, rather than trying to guess it up front from an inferred attribute.
Governance steps before deployment, if attribute-based personalization is pursued anyway. Require an explicit fairness and privacy review before build starts, not after a beta ships, since redesigning a shipped personalization feature is far more expensive than reviewing a proposal. Require a documented lawful basis and, in most jurisdictions with data-protection regimes, a data protection impact assessment specifically because the feature involves inferring a special category of data. Require a measured accuracy check on the inference itself (what fraction of users would be inferred incorrectly) presented alongside the personalization benefit, so the trade-off is decided on evidence rather than optimism. Require an opt-out that is easy to find and that reverts to the generic flow, and log inference decisions so they can be audited later if a complaint or a legal inquiry arises.
Worked example
A concrete version of this conversation: the PM proposes inferring likely age bracket from signup behavior to simplify onboarding copy for older users. The alternative that gets most of the benefit without the inference: add one explicit onboarding question ("how familiar are you with [product category]: new to this / switching from another tool / experienced") and branch the copy on the ANSWER, not an inferred age bracket. This captures the actual thing the PM cared about (adjust onboarding complexity to the user's familiarity) more accurately than age would have proxied for it in the first place: plenty of older users are power users of the product category, and plenty of younger users are first-timers, so age was always an imperfect proxy for the real variable of interest. The governance cost of the explicit-question approach is near zero (no DPIA, no inference-accuracy validation, no special-category data at all), while the personalization benefit is likely HIGHER because it targets the actual causal driver of what makes onboarding land well.
Trade-offs and pitfalls
The most common wrong turn is a PM assuming "inferred" is safer than "asked directly" because the user never explicitly provided the attribute; regulators and courts generally do not treat inference as meaningfully different from collection when the outcome (a system holding and acting on a protected characteristic) is the same. A second pitfall is skipping the accuracy question entirely: teams evaluate personalization proposals on the assumed benefit and never quantify the cost of getting the inference wrong for a meaningful share of users, which flips the expected value of the feature once measured honestly. Third, "just for onboarding, not for anything else" commitments tend not to survive contact with a growth or ads team six months later who discover the inferred field already exists and want to reuse it; if the field should not be reused, the strongest control is not building it as a persistent field at all, doing the personalization decision at request time without persisting the inference.
Define demographic parity, equalized odds, and calibration (group-wise calibration). For each metric give a formal definition and a loan-approval example of how you would measure it, then state which metric you would prioritize if (a) a regulator requires equal treatment across groups and (b) downstream decisions require well-calibrated risk scores.
Sample Answer
A strong answer opens by naming the three definitions and stating plainly that they generally cannot all hold at once when base rates differ across groups.
Structured elaboration
| Metric | Formal condition | What it controls |
|---|---|---|
| Demographic parity | P(Y^=1∣A=a)=P(Y^=1∣A=b) | Equal selection rate across groups, regardless of outcome |
| Equalized odds | P(Y^=1∣Y=y,A=a)=P(Y^=1∣Y=y,A=b) for both y∈{0,1} | Equal true-positive and false-positive rates across groups |
| Calibration (group-wise) | P(Y=1∣score=s,A=a)=s for every group a | A predicted score of s means the same real-world probability in every group |
Loan example. Say a bank approves loans with a risk score.
- Demographic parity means the same fraction of applicants in each demographic group gets approved, even if the groups have different true default rates.
- Equalized odds means that among applicants who would actually repay, the approval rate is the same across groups (equal TPR), and among those who would default, the rejection rate is the same across groups (equal FPR).
- Calibration means that a 0.2 default-risk score means a genuine 20% default probability whether the applicant is in group A or group B.
Worked example. If group A has a true default rate of 10% and group B has a true default rate of 30%, a single calibrated score function will naturally assign more high scores to group B. Forcing demographic parity on top of that calibration would require either denying good group-A applicants or approving bad group-B applicants purely to match rates, which breaks calibration. This is not a hypothetical: it is the mathematical content of the impossibility result once you fix differing base rates.
Trade-offs and pitfalls. (a) A regulator asking for "equal treatment across groups" usually means demographic parity or equalized odds, not calibration, so lean there. (b) A downstream risk-scoring use case (setting an interest rate, sizing a reserve) needs calibration, because a wrongly-calibrated score misprices risk for an entire group even if selection rates look fair. (c) The most common mistake is treating these three as compatible variations on "fairness" rather than as genuinely conflicting design choices; picking one is a policy decision, not a purely technical one, and should be made with legal and business stakeholders, not unilaterally by the model team.
Unlock Full Question Bank
Get access to all 10 Responsible AI: Fairness, Bias, and Interpretability interview questions and detailed answers.
Sign in to ContinueJoin thousands of developers preparing for their dream job.