ThinkWork

We Ran One Discovery Call Through Four AI Coaching Tools. They Didn't Agree on Anything.

Same call, same rep, same 22 minutes. Four different verdicts on whether he could run discovery at all.

Twenty-two minutes of audio. One enterprise-software rep, one mid-market prospect three years into a contract with a legacy vendor, a real discovery call from last quarter — not a demo, not a role-play. We transcribed it, stripped every identifying detail, and ran the identical transcript through four AI coaching tools that all claim to assess sales competency. We wanted one answer: given the exact same 22 minutes, would four tools agree on whether this rep can run discovery at all?

They didn't agree on whether he could run one at all.

One tool rated his discovery strong, with clean qualification and a clear next step. A second flagged discovery as his single biggest gap, citing too many closed questions early on. A third didn't isolate discovery as a competency at all — it produced one blended "call quality" score of 78 and left it there. The fourth agreed with the second tool that the early questions were closed, then contradicted it by praising the exact same three questions as "efficient, respectful of the prospect's time."

Four readings of one call. Not a rounding-error disagreement between similar products — four tools that, given identical evidence, couldn't tell you with any consistency whether to promote this rep or put him on a plan.

What actually happened, for the record

Stripped of any tool's interpretation, here's what's in the transcript:

That's the whole raw material. Everything below is what four tools did with it.

Four tools, four verdicts

ToolHow it scored discoveryWhat it actually flaggedWhat it never saw
Incumbent conversation-intelligence platformStrong (4/5)Question count and talk ratioQuestion type — never distinguished open from closed
LLM-wrapper coaching bot, default promptNeeds improvementCorrectly caught the two closed questions at minute 2–4Mis-tagged the third question — the open one — as closed, undercounting the rep
Framework-agnostic scorecard toolNot scored separately (78/100 overall)Keyword match: "switching costs" mentioned twice, logged as "objection surfaced and handled"Whether the second mention was a resolution or just a repeat
Rubric-driven assessor, calibrated taxonomyDeveloping — proficient on qualifying, gap on holding silenceSame three questions as tool two, correctly tagged; the eight-minute gap between raising and resolving the objection; the four-second silence at minute 14Nothing material in this call

Where and why they actually diverge

Three moments account for almost the entire spread.

The two closed questions. Tool one doesn't penalise them because it counts total questions asked, not question type — volume, not shape. Tool two flags the pattern correctly but then mis-tags the third question, which is open, as closed. That's not a coaching insight, it's a parsing bug: an LLM told to spot closed questions without a working definition will flag anything starting with "is" or "are," regardless of what follows. Tool four gets it right because its taxonomy has a fixed, testable definition of "closed question" that a human wrote down and validated against examples, rather than one the model improvises per call.

The switching-cost objection and its recovery. Tool three's keyword approach scores the moment "switching costs" appears a second time as resolved — it has no mechanism to check whether the second mention was actually a resolution or the prospect repeating the same concern more forcefully. Tool four is the only one that names the pattern correctly: a "delayed, self-initiated recovery," eight minutes after the objection first surfaced. It has a name for that pattern because its rubric has a slot for it. The others don't score it because they were never built with the concept.

The four-second silence. Tools one and two don't measure it at all — their models work off transcript text, and dead air isn't in a transcript unless something explicitly codes for it. Tool four flags it because its ingestion pipeline timestamps gaps and its rubric contains an explicit marker for holding silence after a high-stakes question. This is the clearest tell in the whole test: three of four tools are architecturally incapable of ever surfacing this moment, no matter how capable the underlying language model is, because they never look at what happens between the words.

The actual mechanism, not a vendor problem

None of the four tools are lying, and none are obviously worse language models than the others. The difference is whether there's a fixed rubric underneath — named competencies, defined behavioural markers, calibrated by humans against real examples — or whether the model is inventing its own categories fresh for every call. A tool with a real rubric produces the same verdict on the same evidence every time, and two humans grading against that rubric should land close to the tool's score. A tool without one produces a plausible-sounding paragraph that happens to use the word "discovery," and there's no way to check it against anything, because there's nothing to check it against.

I've spent fifteen years building rubrics for exactly this problem, and the tell is always the same: ask what the tool is scoring against, not what it's telling you.

Three questions to ask before you buy one

  1. Ask for the rubric, not the report. If the vendor can't hand you a written definition of each competency and its behavioural markers — the thing a human coach could grade against independently — the "AI coaching" is a wrapper around a plausible paragraph generator.
  2. Test it against your own scorecard. Grade one call yourself using a fixed Discovery Call Scorecard, have a peer grade it blind, and then run the tool. If you and your peer land within a point of each other and the tool doesn't, the tool's rubric isn't doing the job either — however good its prose sounds.
  3. Ask what happens between the words. Silence, recovery time, sequencing — if the tool only ingests a transcript and never times the gaps, it structurally cannot see roughly a third of what actually happened on the call.

Before you shortlist anything, it's worth running your existing tools — and whatever's pitching you next — through a proper RevOps Tech Stack Audit Checklist rather than a vendor demo. Demos are built to agree with themselves.

None of the four tools we tested is malicious or incompetent. Each is answering a different, unstated question and printing the word "discovery" on the label regardless. If your coaching programme runs on a scorecard nobody actually wrote down, you don't have four opinions about that rep. You have zero, four times over.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel — just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Confirm by email first. Unsubscribe any time.