Skip to main content
Key Takeaways

Metric Blindness: Usage metrics show AI adoption, but cannot reveal whether employees are improving, declining, or surrendering judgment.

Self-Assessment: Employees often misjudge AI’s impact on their performance, making surveys and self-reported benefits unreliable evidence.

Judgment Signals: AI judgment becomes measurable through accept, modify, reject patterns, seeded errors, and evaluations of plausible competing answers.

Junior Risk: AI can prevent junior employees from practicing foundational work, delaying capability development and hiding weaknesses until critical decisions.

Design First: Organizations should measure human capability before redesigning work, or automation may permanently embed unrecognized skill decline.

Your AI dashboard tells you how many people are using the tools. It cannot tell you what the tools are doing to them.

That was a recurring theme in many of the sessions I sat in on last week at AI4 in Las Vegas. 

Organizations have spent two years building reporting around seats activated, weekly active users, prompts submitted, hours reported saved. Those numbers answered a question leadership asked in 2023, which was whether anyone would actually use the thing. Adoption was the risk, so measuring adoption made sense.

Create a Free Account to Keep Reading—and Keep Leading Smarter

Unlock this piece and join a community of forward-thinking leaders discovering tools, playbooks, and insights for thriving in the age of AI.

Name*
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
By submitting this form, you agree to receive our newsletter, and occasional emails related to People Managing People. You can unsubscribe at any time. For more details, please review our Privacy Policy

That question no longer needs an answer. But the instrument built to answer it stayed in place, and is now pointed at a problem it’s not actually designed to see.

One of the sessions I sat in on was called "Beyond Adoption: Building AI Judgment in the Flow of Work". Sharahn McClung's framing was the cleanest version of this I've heard. Companies measure speed. They measure usage. Then they assume improvement. The first two are counted. The third is inferred.

Research published by BCG in June suggests the inference is wrong often enough to matter. In a study of 70 C-suite leaders and senior executives, half said they are already watching skills erode inside their own organizations, and more than 60% expect it to become a material threat within three to five years. Only one in ten had an organization-wide strategy or targeted initiatives underway. A third had not discussed it at all.

If you listen to our podcast, you’ve probably heard me talk about the need for leaders to be a bit more philosophical in these times. But here, the question becomes more practical rather than philosophical. If you wanted to know whether AI was making your people better or worse, where would you look?

Start By Ruling Out the Obvious Answer

The instinct is to ask them.

There's a reason to distrust that instinct, and it comes from a study of people who should have been the best positioned to answer.

In July 2025, METR ran a randomized trial with 16 experienced open-source developers working on 246 real issues in codebases they knew intimately. Before starting, the developers predicted AI tools would make them roughly 24% faster. They were 19% slower. Afterward, having lived through the entire experience, they still estimated they'd been about 20% faster.

Two things about that result are worth holding onto.

METR now labels the finding historical, and a follow-up in February 2026 found some evidence of speedup on more recent models. The specific slowdown number may not describe the tools your teams use today.

What survives is the second finding, which is the more useful one anyway. These were expert practitioners on familiar ground with the outcome directly observable to them, and the distance between their perception and their measured performance ran to nearly 40 points. Nothing about better models makes people better at reading their own performance.

Any measurement approach that runs through self-report inherits that problem. Engagement surveys asking whether AI is helping will produce answers. The answers will be sincere. They will also be roughly as reliable as those developers' estimates.

Two Curves, Same Shape

Picture two teams with identical usage trajectories. Both climbing steadily, both showing healthy adoption on every metric you currently collect.

One team has learned when the model is worth trusting and when it needs to be overruled. Their capability is compounding. The other has been accepting the outputs as fact for months. Their capability is draining out from underneath a rising line.

There is no version of a usage metric that separates those teams.

A Wharton working paper from January measured how far apart they actually are. Steven Shaw and Gideon Nave ran three preregistered experiments with roughly 1,400 participants across more than 9,000 trials, randomizing whether an AI assistant gave accurate or faulty answers through hidden seed prompts.

Against a baseline of participants with no AI access, accuracy climbed 25 percentage points when the AI was right. When it erred, accuracy fell 15 points below the no-AI baseline.

You read that right, below the people who never had the tool. On the tasks where the model was unreliable, access to it made participants worse than not having it at all.

Confidence went up in both conditions, including after errors.

Shaw and Nave call the behavior “cognitive surrender”, meaning adopting AI outputs with minimal scrutiny while overriding both intuition and deliberation. They frame artificial cognition as a third system operating outside the brain, one that can supplement human reasoning or replace it.

This was a lab study using reasoning problems, not employees doing their jobs, and that distinction is worth keeping straight. What it establishes is the mechanism and its magnitude under controlled conditions.

BCG's leaders describe the same pattern from inside real organizations. Almost 90% identified teams accepting AI output without stress testing as a symptom of capability decline. More than half described ownership eroding alongside it. One of them put it bluntly, saying that when the outcome isn't positive, the responsibility gets laid on the AI.

Each week, AI Signal takes one meaningful shift in AI and helps people leaders understand what changed, why it matters, and what to consider next.

Name*
This field is hidden when viewing the form
This field is hidden when viewing the form
By submitting this form, you agree to receive our newsletter, and occasional emails related to People Managing People. You can unsubscribe at any time. For more details, please review our Privacy Policy
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form

What Judgment Breaks Into

Judgment sounds like the kind of thing that resists measurement. McClung's session was useful because it refused that premise and broke the word into observable acts.

Trusting a good answer and moving on. Re-litigating correct output is its own form of waste. A team that verifies everything twice is not exercising judgment, it's failing to develop any. Efficient acceptance belongs on the list of things you want.

Overriding a confident wrong answer. The hard one. The output reads well, the model shows no hesitation, and catching the error requires enough independent knowledge to notice something is off. This is the act that fails first and the one nothing in a usage metric can see.

Knowing when not to reach for the tool at all. McClung's version of this was the most human moment in the session. She described sometimes writing a prompt and realizing partway through that all the thinking had already happened in the act of writing it, leaving nothing to do with the answer that came back.

I don't know about you, but I personally have a fear of a blank page. I look at it and I don’t know where to get started. And sometimes I find by the time that I write the prompt, I've done all my process and all of my thinking in the prompt, and I don't actually need the tool.

Each of these is a decision a person makes before any output ships. None of them improve because usage goes up. They are inputs an organization has to build, and treating them as byproducts of adoption is how you arrive at the position BCG's respondents describe.

It also explains why judgment and decision making carried the highest de-skilling risk score of any skill in their study, sitting alongside problem understanding and framing, which those same leaders rated the single most important skill to long-term organizational performance.

Three Ways to See Judgment

Ordered by how soon you could start.

Accept, modify, reject.

Every AI session generates a record of what was asked, what came back, and what the person did with it next. Accepting a correct answer and accepting a wrong one produce identical usage data and represent opposite acts of judgment. Modification is evidence somebody read the output and thought about it before shipping.

Aggregated across a team over months, that ratio stops being anecdote and becomes something you can trend.

Session Transcripts are Signal

Session Transcripts are Signal

“Every session is a transcript,” McClung said. “It says what the person asked, how they responded, how they find a problem, what the AI returned, and when they actually built it, and real time whether they accepted the answer as is, modified it, or rejected AI. So that accept, modify, reject signal is judgment made visible.

 

“Think back to the dashboard problem of usage. Accepting your correct answer and accepting a wrong answer look identical from usage. The transcript is what lets us separate them. So modifying is the evidence of someone who read and thought critically and protected their output caught in one place.”

The obstacle is access. Most HR functions cannot currently see this data, and getting to it is a conversation with IT and your AI platform owners rather than something you launch next week. The organizations that get there first will be the ones where the people team and the platform team were already talking.

The seeded error

Take a real task. Plant a known flaw in what the model returns, chosen in advance. A wrong figure, a skipped step, a claim that doesn't hold up. Does the person catch it before it moves?

The strength of this over any test is that it runs on actual deliverables and scores the same way every time. Because the error was selected beforehand, there's no grading ambiguity and you can compare one month against another cleanly.

The sensitivity is real, however. Run badly, this reads as entrapment and poisons trust immediately. It works when teams know the practice exists, understand why, and see the results used developmentally rather than punitively.

Evaluative judgment

Two responses, both fluent, both plausible, one subtly wrong. Can the person identify which, and explain why?

This is closest to what expert work demands and the most expensive to run, because building the answer key requires someone with genuine domain expertise to construct plausible-but-flawed alternatives. Be honest with yourself about that cost before committing to it.

The Versions Already Running

Some organizations have moved without waiting for perfect instrumentation.

At CNIL, France's data protection authority, managers are responsible for assessing whether employees can challenge AI outputs, not merely operate them. Critical engagement became a visible, evaluated dimension of performance.

That's the one I'd borrow first. It requires no new tooling and it moves the question into the performance system, which HR already owns. It also solves a problem BCG identifies elsewhere in their research, which is that certification rates climb while actual capability may not.

Shell rebuilt its junior learning paths so new people independently frame the problem, test assumptions, and produce a baseline analysis before AI touches the work. Reported early results included better questions and clearer reasoning behind decisions.

Some companies run AI failure drills, deliberately introducing hallucinations and curveballs into training so that questioning outputs becomes a reflex rather than an instruction. One Indian bank holds a structured AI-free working session on the first Friday of every month across all functions.

None of these are measurement exactly. They're the practice that measurement would tell you whether you need.

The Problem Arrives Late

Fifty-three percent of BCG's surveyed leaders pointed to junior talent developing more slowly.

The analytical grunt work is where, traditionally, judgment gets built. The research, the first drafts, the debugging, the slow business of decomposing a problem and discovering which parts you got wrong. Nobody enjoys those repetitions, but everybody needs them.

That work is the first thing organizations hand to AI. And the people who used to learn from it are now expected to perform without it, in some cases at a level one of BCG's respondents described as what a five to ten year employee would do.

There’s a measurement wrinkle here though. Everything above detects decline against a baseline. In experienced people, that works, because the baseline exists and you can watch it move.

In someone hired last year, there is no baseline. The capability never formed, and its absence looks identical to normal early-career performance until the moment you need them to make a call without support. You are not measuring erosion. You are measuring something that failed to appear, and none of the current instruments are built for that.

Which means the junior cohort is where this costs the most and shows up the latest.

Why the Sequence Matters

Work redesign is the initiative of the next two quarters in many organizations I talk to. Task inventories, agent pilots, org charts drawn around what machines hold and what people keep.

Those decisions calcify. They determine which capabilities are valued and which ones your people are likely to be practicing for years afterward, and which ones fall out of use permanently.

Making them on top of an unmeasured decline builds the decline into the structure. You would be deciding what humans should keep doing while holding no honest read on what your humans can still do and where they’ve atrophied.

I don't think the instrument has to be perfect before you start. The seeded error test could run inside a month in most functions. CNIL's approach needs a conversation with managers and a line in a review template.

What I'd want, before signing off on a redesign, is an answer to one question and the evidence behind it.

If you switched AI off across your team tomorrow, would your people be better than they were before you brought it in, or worse?

David Rice

David Rice is a long time journalist and editor who specializes in covering human resources and leadership topics. His career has seen him focus on a variety of industries for both print and digital publications in the United States and UK.