Skip to main content
Key Takeaways

Evidence Shift: Two trials show AI coaching succeeds with structured goals but fails to match human relationship-based development.

Task Fit: Use AI coaching for rehearsal, reflection, commitments, and preparation—not open-ended leadership development requiring human judgment.

Measure Change: Engagement statistics reveal usage, not improvement; buyers should demand validated evidence of behavior change and development outcomes.

Hybrid Value: The strongest deployment model supplements human coaches, extending structured support while preserving personal attention for complex development needs.

Who Needs Humans: Employees with low confidence and hope may benefit least from unsupervised bots, making human allocation a leadership responsibility.

In 2022, a team of researchers led by Nicky Terblanche published a randomized controlled trial that gave the AI coaching industry its founding piece of evidence. An AI chatbot coach, tested over ten months against human coaches and two control groups, helped clients reach their goals about as well as the humans did.

The finding traveled fast. It underwrote pitch decks, product launches, and a category of enterprise software that barely existed when the study began.

Four years later, Terblanche co-authored a second trial. This one compared accredited human coaches against automated AI coaches head to head, 114 coachees inside a global organization, outcomes measured with validated psychometrics across goals, motivation, resilience, and wellbeing.

Create a Free Account to Keep Reading—and Keep Leading Smarter

Unlock this piece and join a community of forward-thinking leaders discovering tools, playbooks, and insights for thriving in the age of AI.

Name*
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
This field is hidden when viewing the form
By submitting this form, you agree to receive our newsletter, and occasional emails related to People Managing People. You can unsubscribe at any time. For more details, please review our Privacy Policy

Human coaching moved the needle on every outcome they measured, and moved it substantially. The AI arm didn't beat the control group, the people who got no coaching at all, on a single primary measure."

Same lead researcher. Opposite headline. The gap between those two studies is the most useful thing a leader buying one of these tools can understand right now, because it isn't a reversal. It's a map of where AI coaching works and where it comes apart.

The 2022 study handed the AI a narrow, structured job. Move someone through a defined goal-attainment protocol. The 2026 study asked it to be the coach. One of those tasks suits a machine. The other one, on the current evidence, does not.

The Task Determines the Result

The distinction that matters is between automating a coaching function and augmenting a coaching relationship, and the two trials sit on opposite sides of it.

The 2022 protocol worked because goal attainment responds to structure. Regular check-ins, a clear framework, consistent prompting toward a stated objective. A chatbot executes that reliably and never gets tired of asking.

The 2026 trial removed the guardrails and put the AI in the seat a human coach normally occupies, working with senior leaders on the open-ended development questions that don't reduce to a protocol.

The researchers explain the difference through a co-regulation model. Coaching outcomes, on this account, come from the mutual, dynamic influence between two people in the room, the small adjustments each makes in response to the other. That exchange carried the effect sizes in the human arm. The AI arm never generated it, and the numbers reflected the absence.

One caveat belongs here. The trial tested a chatbot coach, and the newest enterprise products are built on agentic systems with memory and context engines that didn't exist when the study was designed.

Vendors will say the research tested last-generation technology, and they have a point worth stating plainly rather than leaving them to land it in a rebuttal. What the study establishes is not that no AI can coach. It establishes that structure alone doesn't reproduce what a coaching relationship does, and that any tool claiming otherwise owes you evidence.

Deployment Didn't Wait for the Evidence

While the research got more complicated, adoption accelerated.

Valence, whose AI coach Nadia reached the market in early 2023, reports deployments across Fortune 500 companies including Experian, Delta Air Lines, Kraft Heinz, and General Mills, with more than a million coaching conversations logged.

According to the company, Experian has embedded Nadia in its performance management process, with about a third of its global workforce now using it, and Delta has rolled it out to frontline leaders. BetterUp, which runs a hybrid model pairing AI with human coaches, reports that a majority of the workers it surveyed (52%) said they wanted both, the empathy of a human coach alongside the availability of AI.

Those numbers come from the vendors, and they measure engagement and satisfaction rather than development. That distinction is the whole game.

The coaching profession itself has moved more slowly. The International Coaching Federation's 2025 Global Coaching Study valued the industry at $5.34 billion, up 17% from 2023, but found only 19% of coaches had invested in new tools in the prior year and 53% reported no digital platform in their practice at all.

The people who deliver coaching for a living are hesitant about the technology their enterprise clients are buying by the seat.

Engagement Is Not the Same as Change

Here is the measurement problem in one sentence, borrowed from a 2026 buyer's guide to the category: if a platform can't show behavior change, you are paying for engagement, not outcomes.

Most AI coaching platforms report the easy metrics. Session counts. Daily active users. Net promoter scores. Nudges sent and opened. Every one of those measures whether people are using the tool, and none of them measures whether the tool is doing anything.

Kuber Sharma, senior director of product marketing for AI and automation at UiPath, who has deployed AI tools across his own team and advises enterprises doing the same, points to a leading indicator that cuts through the vanity metrics.

The signal that predicts whether AI coaching is landing is whether people initiate interactions on their own. A tool that only gets used when it’s assigned isn’t coaching anyone. Self-initiated use means the person found it worth returning to, which is the closest thing to an early sign that a behavior change goal has taken hold.

Kuber Sharma-40692
Kuber SharmaOpens new window

Senior Director, Product Marketing at UiPath

Sharma has watched the failure mode repeat. Organizations that blur the line between reinforcement and replacement, usually under cost pressure, end up with lower engagement and worse outcomes than if they had simply done less coaching more deliberately.

The tool gets credited for outcomes that were always going to happen and blamed for outcomes it never touched, and the actual question, whether anyone changed how they work, goes unmeasured.

Where the Line Sits

The research and the deployment data point in the same direction. AI performs in the bounded work, such as preparing for a hard conversation, rehearsing a negotiation, structured reflection, or reinforcing a commitment between sessions. Those are the tasks the 2022 study validated, and they are real value, available now, at a scale human coaching has never reached.

Coaching used to be a benefit reserved for executives. A tool that extends the useful parts of it to a frontline manager is worth wanting. The organizations getting this right are also disciplined about where they don't point the tool.

Worxogo, a behavioral nudge vendor for frontline teams, found across its deployments that in judgment-heavy work like fraud investigation, most managers showed up but didn't coach, an archetype the company calls the "Courtside Coach", and it stopped positioning the tool as a coaching layer there at all.

Our team learned that lesson against its own early assumption. They had expected the most judgment-heavy work to benefit most from coaching nudges. The data pointed the other way.

Anant Sood-68337
Anant SoodOpens new window

Co-founder of Worxogo

Instead, the tool treats the signal as information for the leader above instead. The decision that mattered was where to stop rather than where to deploy.

The relationship layer is where the machine stops for everyone. The part that carried every effect size in the 2026 trial, the co-regulation between two people, doesn't survive being handed to software.

An organization that treats an AI coach as a supplement to human attention is using the evidence. An organization that treats it as a replacement is buying satisfaction scores and calling them development.

The 2026 study leaves one finding that ought to change how leaders think about all of this. Success was predicted most strongly by the coachee's starting self-efficacy and hope. In other words, the people who got the most from coaching walked in already believing they could grow.

The employees who most need development support, the ones without that foundation, are precisely the ones an unsupervised bot serves worst. Coaching them well takes a person paying attention. Deciding who gets that person, and who gets handed a chatbot instead, is not a procurement question. It belongs to the people running the company.

David Rice

David Rice is a long time journalist and editor who specializes in covering human resources and leadership topics. His career has seen him focus on a variety of industries for both print and digital publications in the United States and UK.