Skip to main content

A few months ago, I sat in on a hiring debrief where an AI screening score had been the deciding factor in an offer. The candidate scored well above the cutoff, the team felt good about it, and the req closed two days faster than average.

Six weeks later, that hire was gone.

Nobody in the room could tell me whether the score had ever been checked against a real outcome — good or bad. It had simply been trusted, the way you trust a thermometer without ever asking who calibrated it.

That’s the reason Screenz exists, and the reason this article does too.

If you’re evaluating AI screening tools, you’ve seen the accuracy numbers vendors lead with — 89% this, 94% that, “outperforms manual review.” Those numbers are usually true. They’re also usually answering a different question than the one you’re actually asking.

What “Accuracy” Usually Means (And What Screenz Measures Instead)

Most accuracy claims from AI screening vendors describe how well the tool matches a resume or profile to a job description — pulling out skills, titles, and years of experience, then ranking candidates against stated criteria. That’s real, measurable, and genuinely useful. It’s just not the same thing as predicting how someone will actually perform once they’re hired.

Those are two different questions:

  • Matching accuracy: Does this candidate’s profile resemble what we said we wanted?
  • Predictive validity: Does a high score actually correlate with strong performance or longer tenure once someone’s in the seat?

A tool can be excellent at the first and completely unproven on the second. Most vendor marketing doesn’t distinguish between them, and most buyers don’t know to ask — which is exactly how a well-marketed score ends up treated as gospel in a hiring debrief.

At Screenz, we treat these as two separate numbers, not one blended score — because collapsing them is exactly how a matching tool ends up marketed as a predictive one.

The Question We Built Screenz to Answer

There’s a name for this in the research: predictive validity — the correlation between an assessment score and a later, independent measure of performance, like manager ratings, output, or 12-month retention. It’s scored from 0 (no relationship) to 1 (perfect prediction), and in the real world, nothing hiring-related gets close to 1.

For context, meta-analytic research on structured interviews — the closest well-studied cousin to today’s AI interview agents — generally puts validity coefficients somewhere in the .28–.30 range, depending on the performance dimension measured. Those numbers reflect decades of refinement and are considered a strong result in personnel selection. It’s nowhere near what a headline number like “94% accurate” implies to a buyer who isn’t fluent in the difference — and it’s a meaningfully higher bar than most AI screening vendors are willing to be measured against.

This is the exact question we built Screenz’s Intelligence Layer to keep answering. Every applicant gets the same structured, role-specific interview. Candidates are then ranked against that criteria with the transcript and reasoning attached, so a reviewer can check the model’s work rather than take the score on faith. Once someone’s hired, we track check-in and retention signals against what the interview said about them — and feed that pattern back into how the next shortlist gets ranked. That loop is the whole point: validity isn’t a claim we make once at launch, it’s something the system keeps re-testing against outcomes as they come in.

So when a vendor says their tool predicts performance, there’s one question that cuts straight to it:

“Validated against what outcome, measured how, and over what timeframe?”

If the honest answer is “we haven’t measured that yet,” that’s not automatically disqualifying — it just means you’re buying a matching tool, not a predictive one, and your reliance on the score in close calls should be set accordingly.

Three Questions to Ask Before You Trust a Score (Screenz Included)

  1. Validated against which outcome — and how long after hire?

“Predicts success” can mean a 30-day performance flag, a 12-month retention number, or a manager’s gut-check rating. Ask which one, and ask what “success” was defined as before the study ran — not after, when it’s easy to pick the metric that makes the result look best.

  1. Validated on data like yours, or borrowed from somewhere else?

A validity study run across a vendor’s broader customer base doesn’t automatically transfer to your roles, industry, or labor market. Ask whether the evidence includes anything resembling your context — and whether the vendor offers a way to validate the tool against your own outcome data on an ongoing basis, not just at the sales stage.

  1. What’s the adverse impact ratio, and who checked it?

Under the widely used four-fifths rule, a selection rate for any group that falls below 80% of the rate for the highest-scoring group is worth investigating. A tool can have decent predictive validity and still carry disparate impact — those are separate questions, and a vendor who can only answer one hasn’t really answered either.

These aren’t hypothetical for us. They’re the three questions the Intelligence Layer is built to keep answering — not just once at launch, but on an ongoing basis as new outcome data comes in.

None of this is about doubting AI as a category. It’s about the difference between a vendor telling you their score works and a vendor showing you the evidence that it works — for outcomes that matter to your business, not just the ones that are easy to advertise.

Why We Built Screenz Around This Question

I’ll be upfront: this isn’t a neutral list. We built Screenz around an AI agent that conducts structured interviews — not one that just parses a resume — which means the predictive-validity question isn’t abstract for us. It’s the whole product. Efficiency is easy to prove and easy to market: faster shortlists, lower cost-per-hire. Outcome validity is harder to build and slower to prove, which is probably why fewer vendors lead with it.

We’d rather be judged on the harder question. If you ask us “validated against what,” we want that to be the easiest part of the conversation — not the part that gets redirected into a feature demo.

Bring This to Your Next Vendor Conversation, Screenz Included

You don’t need a statistics background to have this conversation. You need three questions and the willingness to sit with an honest “we don’t know yet,” if that’s the answer you get. Bring the three above to your current vendor, or to the next one pitching you a score, and watch whether the response is a specific number and methodology — or a confident restatement of the accuracy claim you already had.

The gap between those two answers is the whole point.

Curious what outcome-validated screening actually looks like in practice? See how Screenz approaches it.

Rob Griesmeyer

Rob Griesmeyer is the Technical Co-Founder of Screenz.ai, where he leads the development of an AI agent that conducts structured interviews and validates screening scores against post-hire outcomes. He has spent more than 14 years building ML and AI products, and previously COO at Interview Query.