Skip to main content
Key Takeaways

Hidden Risk: AI hiring tools can appear fair while failing to distinguish qualified candidates from irrelevant or unqualified applicants.

Shifting Bias: Vendor bias audits cannot predict changing outcomes because model versions, roles, names, and demographics influence results.

Human Ownership: Every AI-assisted personnel decision needs a named human accountable for defending the outcome to regulators or plaintiffs.

Competence Testing: Organizations should test AI hiring systems with matched and keyword-stuffed résumés before trusting their recommendations.

Governance First: Treat AI procurement as a governance decision, scaling explainability and oversight to the consequences of each use.

Picture the last AI-assisted hiring decision your team made. A flagged candidate. A "fit score" near the bottom of the queue. Your recruiter moved on.

Now picture that candidate filing a discrimination complaint six months later and asking, in a deposition, exactly why they were rejected. Whose name is on the file? Yours? The recruiter's? The vendor's?

If your honest answer is "the algorithm flagged it," you have an accountability problem, and it is structural. No vendor will solve it for you.

Create a Free Account to Keep Reading—and Keep Leading Smarter

Unlock this piece and join a community of forward-thinking leaders discovering tools, playbooks, and insights for thriving in the age of AI.

Name*
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
By submitting this form, you agree to receive our newsletter, and occasional emails related to People Managing People. You can unsubscribe at any time. For more details, please review our Privacy Policy

I've spent the last decade as a business intelligence analyst at Amazon and Microsoft. I also spent the last year auditing eight commercial AI résumé-screening platforms. Roughly 4,000 trials, identical résumés varied only by candidate name, tested for both demographic bias and basic competence.

What I found is more uncomfortable than the standard "AI is biased" story. Several platforms praised as the most neutral are neutral the way a broken thermometer is neutral. They produce stable, demographically balanced outputs because they cannot meaningfully distinguish a qualified candidate from an unqualified one. That is the appearance of fairness laid over a tool that does not work.

I call this the "Illusion of Neutrality", and it is the dynamic senior HR leaders should worry about most, because it is the one your bias audit cannot catch.

AI is Doing Narrative Work Now

The labor market context matters here, mostly for what it reveals about motive. A September 2025 New York Fed analysis found only 1% of service firms reported AI as the reason for layoffs in the past six months, while 35% said AI led them to retrain workers.

Yet a Resume.org survey of 1,000 hiring managers, published in January, found that 59% admit they emphasize AI when explaining hiring freezes or layoffs because it "plays better with stakeholders" than citing financial constraints.

AI is being asked to do narrative work in addition to operational work. Keep that in mind, because the same instinct that blames AI for a layoff will blame AI for a bad hiring call.

Create a People Managing People account for access to exclusive content, practical templates, member-only events, and weekly leadership insights—it’s free to join.

Name*
This field is hidden when viewing the form
This field is hidden when viewing the form
By submitting this form, you agree to receive our newsletter, and occasional emails related to People Managing People. You can unsubscribe at any time. For more details, please review our Privacy Policy
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form

"Data-driven" Curdled into "Find Me the Data"

The promise of data-driven leadership was clean. Reduce reliance on gut. Surface evidence. Create a reviewable record. When the discipline holds, it works.

What I have watched happen over the last five years drew far less attention than it deserved. The tools never changed. The relationship between leaders and the tools did. The actual request stopped being "what does the data say?" and became "find me the data that supports this."

AI accelerates the shift because a model output is harder to push back on than a hiring manager's gut call. So when something goes wrong, the answer becomes "the system surfaced it." That sentence is the entire problem.

The Fairest Models Can't Tell Candidates Apart

Across eight platforms, including ChatGPT, Gemini, Copilot, and Claude, three patterns emerged.

The bias is still there, and it is less predictable. Several models penalized résumés simply because a candidate name was attached, regardless of which name. Others rewarded specific demographic groups in some roles and penalized them in others. One model split a single racial group in opposite directions by gender on the same résumé.

The bias shifts with the role and the model version. Your vendor's bias audit does not predict how the tool will behave on the next requisition.

The "fairest" models are often the least competent. When I tested whether models could distinguish a relevant résumé from an irrelevant one, several could not. One model rated a service-industry résumé stuffed with finance keywords at 92 out of 100 for a Director of Finance role. The same résumé without the keywords scored 15. The model was scoring vocabulary. Qualification never entered the evaluation. A tool like that has stopped evaluating anything at all.

The same vendor gives different answers depending on which version is running. A candidate evaluated by ChatGPT's "fast" version could score 15 points higher or lower than the same candidate on its "slow" version the same day. Most managers do not know which version they are running, and do not know it matters.

These are not edge cases. They are the operating conditions of AI hiring tools right now.

Accountability Laundering

Better decisions are only part of why these tools get adopted. They also make hard decisions easier to defend. Layoffs. Performance ratings. Hiring rejections. Compensation calls. When the model output is in the file, the person who signed off has plausible deniability they did not have before.

A 2025 ResumeBuilder survey makes this concrete. Among managers using AI for personnel decisions, 64% use it to help determine terminations, and more than one in five admit they often let the AI make the final call without human input.

The behavioral research is even more pointed. A 2025 Nature study by Köbis and colleagues found that AI agents complied with fully unethical instructions between 58 and 98% of the time depending on the model, compared with 25 to 40% for human agents given the same instructions.

Machines are far more willing than people to follow through on ethically compromised orders, which is exactly the property that makes them useful as accountability infrastructure.

The research literature has a name for the human side of this. Madeleine Clare Elish called it the "moral crumple zone", where people at the sharp end of an automated system absorb blame for failures that originated in diffuse design and governance decisions. Staff start to talk about "what the system decided" rather than "what we decided," even when they still have discretion.

The cleanest name for the full dynamic is accountability laundering. The model takes a decision that used to require defensible human judgment and converts it into an output nobody has to defend. When the decision is wrong, the failure is everyone's and no one's, and your organization loses the ability to learn from its own mistakes.

Governance, or It Doesn't Ship

The time has come for leaders to treat AI procurement as a governance decision rather than an IT decision.

Put a name on every AI-assisted call. As one AI attorney recently told CIO, the algorithm will not show up in court. The humans who developed, deployed, or used it will. Before any AI tool touches hiring, performance ratings, compensation, or terminations, require an answer to one question: when this output turns out to be wrong, whose name is on the file? Not the vendor. Not the model. A person, senior enough to defend the call to a regulator or a plaintiff's attorney.

Scale explainability to the stakes. A model that recommends training content can be a black box. A model that contributes to a termination cannot. Most organizations set their explainability standard once, at procurement. The disciplined ones reset it per use case and walk away from vendors who cannot meet the higher bar.

Run your own dual-validation audits. Vendor bias audits are necessary but insufficient. Also, test competence. Can the tool actually distinguish qualified from unqualified candidates on the roles you hire for? A few mismatched résumés. A few keyword-stuffed ones. Compare the scores. If the model cannot tell the difference, it cannot help you, no matter how unbiased it appears.

HR is the function AI vendors are courting hardest, and the function holding the bag when one of these tools produces an indefensible output. The temptation will be to treat procurement as a vendor question. Resist it.

As one labor attorney put it to The Washington Post when algorithms first started shaping layoff lists: "Don't try to pass the buck to the software."

If your answer to a flawed AI-assisted decision is "the system did it," your system has become your risk. Somewhere in your organization, a person needs to own every call these tools touch. Find out who that is before a deposition does.

Kevin Webster

Kevin Webster is a business intelligence analyst and independent researcher exploring the intersection of AI, labor ethics, and the future of work. With a decade of experience at organizations including Amazon and Microsoft, he writes about the practical and ethical implementation of AI in the modern workplace.