Field Notes: What Forty Coaching Calls Taught Us About Briefing AI Notetakers
Default prompts give you action items. A few specific instructions give you the competency evidence your 1:1s are actually short of.
Forty coaching calls, six managers, one recurring problem: the AI notetaker handed everyone a tidy list of action items, and almost none of it was usable for actually coaching the rep. "Send updated pricing by Friday." "Follow up on renewal timeline." Perfectly good for a status meeting. Useless for a 1:1 built around a competency — because coaching on a skill requires the actual words the rep used, not a summary of what the call resulted in.
This is a field-notes post, not a teardown, because the fix wasn't a different tool. It was different instructions to the same tool.
The default is built for a different job
Most AI notetakers — whatever the brand — default to meeting-productivity summarisation: decisions made, action items, next steps. That's the right shape for a project check-in. It's the wrong shape for a manager trying to coach discovery questioning or objection handling, because the default summary strips out exactly the evidence a competency-based coaching conversation runs on: exact phrasing, the sequence of what came right before and after a hard moment, and whether the rep recovered from it or just moved past it.
What we actually did
Across six managers running weekly recorded-call reviews, we changed the notetaker's custom instructions every fortnight and logged what came back, against roughly the same call mix each time (renewals, new-logo discovery, one expansion call per manager per week). Forty calls in, a clear pattern emerged: the default settings recover outcomes. A short list of specific instructions recovers behaviour — which is the only thing you can actually coach on.
Same call, two summaries
Default notetaker output, unedited:
"Rep addressed pricing concern. Prospect seemed satisfied with the response. Next steps: send contract."
Reconfigured output, same call, same minute:
"At 11:42, prospect said: 'I just don't see how we justify paying twice what we're paying now.' Rep didn't respond to the figure directly — asked 'what would justifying it need to look like, on your end?' Prospect's answer named a specific budget owner not mentioned earlier in the call. Rep did not return to that name before the call ended."
The first tells you the call went fine. The second gives you an entire coaching conversation: the rep asked a genuinely good clarifying question instead of jumping straight to a discount, which is worth naming and reinforcing — and then let a live piece of information (a named budget owner) go unaddressed, which is worth flagging before the next call. None of that survives the default summary. All of it survives the reconfigured one.
The instructions that recovered it
In roughly descending order of how much evidence each one recovered:
- "Quote, don't paraphrase, any objection, question, or number." This single instruction did more work than everything else combined — most of what gets lost is lost at the paraphrase step.
- "For any objection or pushback, note what was said in the two turns immediately before and after it." This is what surfaces recovery — or the absence of it — rather than just the fact that an objection existed.
- "Flag any named person, budget figure, or timeline mentioned once and not referenced again." This catches exactly the kind of dropped thread in the example above, which a generic summary has no reason to notice.
- "Note any pause over three seconds following a direct question, and who spoke next." This is the closest a transcript-based tool gets to seeing silence, which is otherwise invisible to it.
- "Do not summarise sentiment — describe the specific words or behaviour a human would read that way." Banning phrases like "prospect seemed satisfied" forces the model to show its working, which is usually where the actual coaching material is hiding.
| Default setting | What it discards | Reconfigured instruction | What comes back |
|---|---|---|---|
| Action-item summary | Exact phrasing of objections and questions | Quote, don't paraphrase | The actual words to coach against |
| Outcome-only next steps | Whether pushback was recovered from | Note the two turns before/after pushback | Recovery pattern, or its absence |
| Single-pass summary | Threads mentioned once and dropped | Flag names/figures/timelines mentioned once | Dropped information the rep should chase |
| Text-only transcript read | Silence and pacing | Flag pauses over three seconds + who spoke next | The closest proxy for silence a transcript tool can offer |
| Sentiment labels | The evidence behind the label | Describe behaviour, don't label sentiment | Specific, checkable coaching material |
What still doesn't come back
Being straight about the limits: none of this recovers actual tone or prosody — a notetaker reading text still can't tell you a rep sounded defensive versus just quiet, only that a pause happened. Crosstalk still gets mangled in the transcript before any instruction can act on it. And there's an honest trade-off nobody likes mentioning: a coaching-grade summary runs about three times longer than the default. If your managers were barely reading the five-bullet version, the twenty-bullet version isn't automatically an improvement — it's only worth it if someone commits to actually reading it before the next 1:1.
Setting this up this week
If you manage a coaching cadence off recorded calls, change the notetaker brief before you change anything else — it's a five-minute edit, not a procurement decision. Use the evidence it recovers to structure the actual conversation with something like the GROW Model Coaching Conversation Script, and turn the specific quotes and dropped threads into feedback the rep can act on with a Call Coaching Feedback Template rather than a vague "good energy on that call."
The tool was never broken. It was configured to help you run a status meeting, and you were trying to run a coaching one. Change the brief, not the software.