ThinkWork

Your AI Notetaker Isn't Taking Notes. It's Building a Model of How Your Reps Sell.

Every call your AI notetaker logs is a training signal. Most sales orgs have no idea what they're teaching it.

Most sales leaders who deployed an AI notetaker made a procurement decision, not a pedagogy one. They evaluated transcription accuracy, CRM sync, Zoom compatibility, price per seat. They did not ask: what happens to our definition of good selling when every call this tool observes becomes a data point? That question sounds abstract. It is not. It has a direct line to your win rate.

What the tool is actually doing

Gong, Chorus, Fireflies, Fathom, take your pick. The core mechanic is the same: the platform attends your calls, produces a transcript, extracts a summary, tags moments (objections, competitor mentions, next steps), and over time builds a pattern-matching model calibrated to your organisation's call data. That last part is the bit that matters.

When the platform surfaces "best practice moments" or scores a call, it is comparing against a baseline. That baseline is your calls. Not some idealised external standard of what great discovery looks like. Your calls. The median rep in your team, running the average quality of conversation your org produces, becomes the reference point against which outliers are judged.

If your discovery is mediocre, the system learns mediocre. If your reps routinely skip problem-impact questions and jump to solution, the system encodes that skip as normal. If they talk 70% of the time in a first meeting, the talk-to-listen ratio your platform flags as a concern will be calibrated against a population that already talks too much.

You are not capturing performance. You are normalising it.

The passive transcription illusion

There is a comforting mental model that goes like this: the AI notetaker is just a more reliable version of a shared Google Doc. It captures what happened. You, the human, interpret it. The tool is neutral.

That model is wrong for two reasons.

First, the tool is not neutral at the summarisation layer. Every auto-generated summary involves editorial choices about what counts as a key moment, a meaningful next step, a strong buying signal. Those choices reflect the model's training, which reflects your historical call data. A rep who runs a competency-blind discovery gets a summary that reads like a competency-blind discovery is fine, because it is.

Second, most managers are not interpreting the data. They are reading the summary. The gap between "what actually happened on that call" and "what the AI said happened" collapses in practice, because the manager has fifteen other things to do and the AI summary is three bullet points. The transcript becomes the ground truth by default.

The result is a system that launders mediocre selling into tidy, professional-looking records. Structured enough to feel like rigour. Empty enough to produce nothing.

The competency problem

Here is the specific failure mode I see repeatedly. A business deploys a conversation intelligence platform, runs it for six months, generates a library of hundreds of tagged calls, and then the enablement lead tries to use it to diagnose a skill gap. They go looking for evidence. What they find is transcripts organised by deal, by rep, by quarter, and by keyword. What they do not find is any systematic answer to: can our reps run a proper qualification conversation, or are they just having friendly chats that look like qualification?

The platform cannot answer that question because nobody defined what qualification competency looks like before the data collection started. There is no rubric upstream of the tool. So the tool, faithfully, recorded everything and evaluated nothing.

This is the actual cost. Not bad transcription. Not poor CRM sync. A six-month library of data that cannot tell you what you most need to know.

If you want to know whether your team is genuinely capable of running multi-threaded discovery across economic and technical buyers, you need a defined competency standard that exists before the call, not a language model that reverse-engineers a pattern from the calls you have already made. The Human vs AI Scoring Agreement Checker is worth running at this stage precisely because it forces that definition: what are humans actually grading, and does the AI agree on the same calls? Where they diverge tells you where your standard is either undefined or inconsistently applied.

What you should have built before you bought the tool

The sequence most orgs follow: buy the tool, collect data, eventually wonder why the data is not actionable. The sequence that works: define the standard, instrument the tool to that standard, collect data that is interpretable from day one.

"Define the standard" means something specific. It means answering, in writing, questions like these:

Without those answers, your notetaker is logging activity. With them, it can log performance. The Team Skill-Gap Heatmap Generator is useful at exactly this junction: once you have a competency list, you can map your population against it and find out where the data collection should be paying most attention.

The fix is not a different tool

I want to be clear about what I am not arguing. I am not arguing that AI notetakers are bad, or that conversation intelligence platforms are oversold theatre. Some of them are genuinely useful. The transcription fidelity is good. The search functionality saves time. The async coaching workflows are real.

What I am arguing is that the tool's value is entirely contingent on the quality of the standard sitting upstream of it. A well-configured platform pointing at a well-defined competency model will surface genuinely useful patterns. The same platform pointed at undefined behaviour will faithfully record that behaviour and eventually convince you it is normal.

Switching tools does not fix this. Buying a more sophisticated platform does not fix this. Hiring a prompt engineer to write better summary templates does not fix this. The fix is deciding, in advance, what good looks like, at the competency level, and then letting the tool observe against that definition.

Most enablement teams skipped that step because it is harder than procurement. It requires actual expertise about selling, not just familiarity with the platform's admin panel. It requires someone being willing to say: this rep's discovery is weak, and here is specifically why, and here is what the data should be capturing to show improvement. That is not a software problem. It is a standards problem.

One thing to check this week

Pull five calls from your platform. Not the flagged ones. Random ones. Read the summaries. Then read the transcripts. Ask yourself: does the summary capture anything about the quality of the rep's questioning, or does it only capture the content of what was discussed?

If the answer is the latter, which it almost certainly is, you have a system that is very good at telling you what topics came up and very bad at telling you whether your rep can sell. That gap is not your vendor's fault. It is a measurement design problem, and it predates the tool.

The calls are the data. The question is what you are teaching the system to see.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel, just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Alerts start right away. Unsubscribe any time.