Skip to content
Menu

How to Test AI Tools for Bias in Hiring and HR

Help you audit AI hiring tools for bias, document findings, set safeguards, and create an appeals process.

Test AI tools for bias by examining their data, recommendations, errors, and effects on candidates across relevant groups. Involve legal, HR, security, and accessibility specialists, and keep a human accountable for every hiring decision.

Why Conventional Accuracy Metrics Fail in Hiring

Overall accuracy can hide poor outcomes for smaller groups. Review error types separately for each relevant group, and compare the results with qualified job-related outcomes.

Historical hiring data may contain past bias or patterns unrelated to job performance. A tool may also use indirect signals, such as employment gaps, to infer a protected characteristic.

Evaluate whether the tool predicts relevant job outcomes consistently across groups. Do not treat an apparently accurate tool as fair without examining how its errors affect different candidates.

Step 1: Define Protected Attributes and Proxy Variables

Identify the protected characteristics relevant to your location and hiring process. Map each one to the data the tool collects, receives, infers, or generates.

List possible proxies as well. Graduation dates, employment gaps, schools, organizations, skills, and word choices may reveal or act as signals for a protected characteristic even when the tool does not directly use that characteristic.

Ask the vendor:

  • What data does the tool receive?
  • What features does it generate or infer?
  • Does it use external data?
  • Can a recruiter enter data that changes the recommendation?
  • How are protected and proxy characteristics handled?
  • Can you restrict particular data fields?

Document your findings and remove inputs that are unnecessary, unreliable, or unrelated to the job.

Step 2: Select Fairness Measures for Your Context

No single measure captures every possible concern. Review selection rates, missed candidates, false selections, and the relationship between predictions and relevant job outcomes.

Consider the legal requirements, job context, candidate population, and consequences of each type of error. Explain why you selected the measures, and state their limitations.

Avoid interpreting one ratio as a complete finding. Review multiple measures alongside qualitative evidence, recruiter decisions, appeals, and observed candidate outcomes.

Step 3: Construct a Representative and Labeled Audit Dataset

An audit is only as useful as its data. Review the source of each example, the roles included, the range of legitimate qualifications, and the representation of relevant groups.

Historical applicant data may reproduce past exclusion. Compare it with current recruiting channels, available talent pools, and structured job criteria. Add synthetic or carefully reviewed examples when the existing data does not cover important situations.

Define the job-related outcome you are evaluating. Prefer reliable, structured measures of performance rather than unstructured interviewer impressions.

For each example, document:

  • Relevant job qualifications
  • The tool’s inputs and output
  • The final decision
  • The reason for the decision
  • The business or operational outcome
  • Relevant demographic information kept separately from the evaluation

Protect sensitive data and limit access to people who need it for the audit.

Step 4: Perform Intersectional and Subgroup Analysis

Review results for relevant groups and for combinations of characteristics when the data supports that analysis. A broad group can conceal different outcomes for candidates who share several characteristics.

Small subgroups may produce unstable results. Do not make unsupported conclusions from limited examples. Combine evidence, seek additional data, and clearly label findings that remain uncertain.

Ask specialists to review the analysis method. Record warnings about missing data, overlapping categories, inconsistent labels, and differences in job context.

Step 5: Audit the Full Decision Pipeline, Not Just the Model

Bias can enter before a tool produces an output and after a recruiter receives its recommendation. Review the entire process.

Start with the job advertisement, recruiting channels, screening questions, and eligibility rules. Check whether the language, requirements, or access to the process create unnecessary barriers.

Next, review resume parsing, candidate summaries, scoring, ranking, interview questions, and rejection reasons. Test whether equivalent qualifications receive consistent treatment.

Review the decision threshold. Changing the cutoff can change which candidates proceed even when the underlying scores have not changed.

Finally, examine human overrides. Record why recruiters accept or reject recommendations, whether reviewers can see relevant job information, and whether overrides follow documented rules.

Step 6: Implement Continuous Monitoring and an Appeals Mechanism

A single audit will not reveal every issue that may emerge. Review outcomes regularly and whenever the tool, data, recruiting process, job, or relevant law changes.

Set alerts for unexpected changes in selection, rejection, error, appeal, or override patterns. Assign someone responsibility for investigating alerts and documenting the outcome.

Establish a clear appeals process. Tell candidates how to request human review, what information they may provide, and when they can expect a response.

Reviewers should have access to the job criteria and relevant evidence, but not unnecessary sensitive data. Keep the tool’s internal workings confidential while providing a clear, plain-language explanation of the decision.

Track appeals and recurring concerns. Treat patterns in complaints as signals that require investigation rather than proof of a particular cause.

Before You Deploy the Tool

Ask the vendor:

  • What documentation supports the tool’s fairness claims?
  • Which groups and use cases were included in evaluations?
  • What known limitations and unsuitable uses are disclosed?
  • Which inputs and inferences can affect results?
  • How can restricted data fields be controlled?
  • How are errors and appeals handled?
  • What information is retained?
  • How can changes to the tool or underlying data be detected?
  • What support is available when an unexplained pattern appears?

Run your own checks on representative examples. Compare recommendations with structured job criteria, examine differences between groups, and record uncertainty.

Before and After Each Hiring Cycle

Before deployment:

  • Define the job-related purpose and decision rules.
  • Identify protected characteristics and possible proxies.
  • Review the data and vendor documentation.
  • Select audit measures and document their limitations.
  • Assign decision ownership and appeal responsibility.
  • Establish retention, access, and privacy controls.

During use:

  • Review recommendations and overrides.
  • Monitor errors, complaints, and unexpected outcomes.
  • Check whether candidates can complete the process accessibly.
  • Investigate changes rather than assuming they are caused by the tool.

After use:

  • Compare outcomes with job-related criteria.
  • Review group and intersectional patterns.
  • Examine appeals and recurring concerns.
  • Document remediation, retesting, and approval decisions.
  • Suspend automated recommendations when serious concerns remain unresolved.

FAQ

Q: How much data do I need for each demographic group?
A: There is no single suitable sample size for every audit. Consider the groups being compared, the variability in the data, the consequences of an error, and the limits of your analysis. Seek additional data when the available evidence cannot support a reliable conclusion.

Q: Should I use a disparate impact ratio?
A: You can use it as one part of an audit, but it does not establish fairness by itself. Review it alongside error patterns, job-related evidence, data quality, statistical uncertainty, and the purpose of the tool.

Q: Can I use the same framework for generative AI recruiting tools?
A: Partially. Add checks for factual reliability, unsupported inferences, unequal tone, inconsistent answers, inaccessible language, and different levels of detail across candidates.

Q: What should I do if the tool produces a biased recommendation?
A: Pause reliance on the recommendation, preserve the relevant records, notify the responsible team, and provide human review. Investigate the input, model, threshold, workflow, and human overrides before deciding whether to retest, reconfigure, or stop using the tool.