ThinkWork

How to Run a Coaching Calibration Session So Two Managers Stop Grading the Same Call Differently

Blind-grade the same call, compare scores, argue about the gaps, anchor the rubric. Skip this and your 'coaching program' is just manager opinion with a template.

Take the same twelve-minute call. Hand it to two managers who've each been coaching for three years. Tell them nothing except "score this against the rubric." Nine times out of ten, they'll land a full competency level apart on at least one line item — one calls the discovery Proficient, the other calls it Developing — and neither of them is lying, cutting corners, or bad at their job. They're each applying a private rubric that happens to share a name with the official one. Nobody has ever put the two rubrics in the same room and made them fight.

That's the entire case for calibration. Not as a nice-to-have culture exercise, but as the only mechanism that turns "we coach against a competency model" from a claim into something true. Skip it, and what you're running is manager opinion with a template stapled on.

Why two good managers diverge on the same call

It isn't a training gap in the usual sense. Both managers know the rubric. The divergence happens because a rubric is a set of words, and words like "structured," "consultative," or "handled the objection well" carry different weight depending on which calls a manager has personally been exposed to. A manager who came up on enterprise deals reads "discovery depth" against a six-stakeholder, ninety-day standard. A manager who came up on transactional SMB reads the same line against a fifteen-minute call standard. Both are applying the rubric in good faith. Both are wrong, in the sense that the org now has two competing definitions of "good," and a rep transferred between their teams gets graded against a moving target.

Add in the halo effect — managers rate their own reps a shade more generously than a stranger's rep on an identical transcript — and you get divergence that compounds every quarter it goes unaddressed. Six months in, "Proficient" means something different depending on which manager wrote it, and nobody notices until a rep gets passed over for promotion by a manager using the stricter definition, having been told for a year by their own manager that they were ready.

The calibration session, step by step

This works as a standing ritual, not a one-off. Monthly for the first quarter, then quarterly once scores stop moving.

  1. Pick one call, blind. Choose a call nobody in the room has heard, ideally from a rep outside anyone's direct team, so there's no ownership bias in the room.
  2. Score independently, in silence, before any discussion. Every manager scores every competency on the rubric alone, on paper or in the tool, with no cross-talk. This is the step people try to skip because it feels slow. It's the only step that actually produces the data you need.
  3. Reveal the scores side by side. Put every manager's score for every competency on a shared screen at once. Do not let anyone explain before the numbers are all visible — explanation before comparison lets the room anchor on whoever speaks first.
  4. Argue the gaps, competency by competency, not call by call. If three managers say Proficient on discovery and one says Advanced, that manager has to point at the specific moment in the transcript that earned the extra level. This is where the actual rubric — the shared one, not anyone's private version — gets built, one disputed line at a time.
  5. Write the anchor down. Every time a gap gets resolved, capture the specific quote or behaviour that settled it, and attach it to the rubric as a worked example. After four or five sessions you have a rubric with real anchors instead of adjectives.
  6. Track divergence over time, not just scores. The metric that matters isn't "did everyone score Proficient" — it's "how far apart were the raw scores before you calibrated." That number should shrink every session. If it isn't shrinking, your calibration sessions are venting, not converging.

What the gap actually looks like on paper

SessionAvg. score divergence across competencies (levels)Competencies with >1 level gap
Month 1 (uncalibrated)0.9Discovery, Negotiation, Qualification
Month 20.6Negotiation
Month 30.3
Quarter 2 (spot check)0.2

That trajectory is the actual proof of concept for a coaching programme. Not attendance at manager training, not a rubric document sitting in a shared drive — a measured, shrinking gap between how two managers read the identical call.

Where it breaks down in practice

The most common failure isn't skipping the session outright, it's letting the discussion step turn into consensus theatre — the most senior manager in the room states an opinion, everyone else revises their private score to match it, and you've replaced four independent judgements with one, wearing a group's clothing. The fix is structural: capture the independent scores before revealing anyone's identity alongside them, and when you argue the gaps, ask for the specific transcript evidence before you ask for the opinion. A framework like the SBI Feedback Model Cheat Sheet is useful here for a different reason than it's usually deployed — it forces the manager arguing for a higher score to cite the Situation and Behaviour before the Impact, which kills a lot of "I just think it was good" arguments on contact.

If you don't yet have a shared instrument to calibrate against, don't start with the whole 54-competency model — pick one high-friction competency and build consistency there first. A Discovery Call Scorecard is a good place to start precisely because discovery is where the divergence above tends to be worst. And once a session resolves a gap, put the outcome somewhere a manager will actually see it again — a Call Coaching Feedback Template that captures the anchor example alongside the score keeps the calibration alive between sessions instead of evaporating the moment everyone leaves the room.

What you're actually buying

Calibration sessions are slow, mildly uncomfortable, and produce no visible output for the first two rounds beyond an argument about a fifteen-minute-old call. What they buy is the ability to tell a rep the truth: that "Proficient" means the same thing regardless of who's grading them, that a promotion decision isn't a function of which manager happened to inherit them, and that the competency model on the wall is the one actually being used, not a poster next to the real one everyone's carrying around in their head.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel — just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Confirm by email first. Unsubscribe any time.