How to Run a Blind Calibration Session So Two Managers Score the Same Call the Same Way
Give two managers the same call with no names attached. If their scores land more than one tier apart, you don't have a rubric problem — you have a rubric that only exists on paper.
Every sales org with more than one manager eventually discovers the same uncomfortable fact: the rubric everyone signed off on in the enablement meeting produces different scores depending on who's holding the clipboard. This isn't a people problem, and it isn't usually a training problem either. It's what happens when a written rubric hasn't been tested against real, ambiguous evidence with more than one set of eyes on it at the same time. Blind calibration is the fastest way to find out whether your rubric is actually shared or just typed up and filed. Here's how to run one properly.
Step 1: Pick a call in the messy middle, not the obvious cases
Don't use the best call of the quarter or the disaster everyone already talks about — both score the same regardless of who's scoring, which tells you nothing. You want a call where reasonable people could land on adjacent tiers: a discovery call where the rep asked good questions but missed the follow-up that would have surfaced the real budget conversation, or a negotiation where the rep held the line on price but gave away scope without noticing. Ambiguity is the entire point. If everyone would obviously agree, you haven't tested anything.
Step 2: Strip it so nobody can pattern-match on identity
Remove the rep's name, the account name, and anything that lets a manager score the person instead of the behaviour — reputation halo works in both directions, and a manager who knows "this is Priya's call" will unconsciously score generously if Priya is known to be strong, or harshly if she's had a rough quarter. Keep the call itself intact: transcript or recording, whichever your rubric was actually written against. If you're building this scorer from scratch rather than adapting one you already have, the Discovery Call Scorecard gives you a structure to strip down to blind-safe fields.
Step 3: Give managers the rubric and nothing else — no pre-discussion
This is the step people skip, and it's the one that matters most. If managers talk about the call before scoring, or align on "what we're looking for" beforehand, you've contaminated the test — you're now measuring their ability to agree in conversation, not their independent read of the same evidence. Send the call and the rubric. Set a deadline. No group chat about it in the meantime.
Step 4: Score independently and in writing, sealed until reveal
Each manager scores every competency on the rubric against the call, alone, and submits it — a form, a doc, whatever's fastest — without seeing anyone else's scores first. Sealed matters here in the literal sense: if one manager's scores are visible before another submits, you'll get anchoring, and anchoring defeats the whole exercise. Two or three managers is enough for a first pass; five is better if you have the headcount to spare, because it's harder to write off a five-way split as one manager having an off day.
Step 5: Reveal and diff — find where the gap actually lives
Line up the scores competency by competency. Most of the time, most competencies will land within one tier of each other — that's normal variance and not worth a meeting. What you're hunting for is anywhere the gap is two tiers or more, because that's not noise, that's a rubric cell that means something different to each manager reading it. In a decent-sized sample this is usually one or two competencies, not the whole scorecard, which is good news — it means the rubric mostly works and you've found the specific place it doesn't.
Step 6: Talk only about the divergent cells
Don't relitigate the whole call or the whole rubric — that turns a focused fix into a two-hour meeting that resolves nothing. Pull up the specific competency where scores diverged, have each manager explain, in one or two sentences, what evidence in the call drove their score. Almost always the disagreement traces to language in the rubric itself: "handles pushback confidently" means something different to a manager who reads "confidently" as tone versus one who reads it as outcome. That's not a training gap in the manager. That's an ambiguous word in a document you control.
Step 7: Rewrite the offending language the same day
Don't table it for the next enablement cycle — momentum and specificity both decay fast. Rewrite the ambiguous cell immediately, using language anchored to what was actually said or done on the call you just reviewed, and retest it on the next calibration round. A rubric cell that's been through one real disagreement and one rewrite is worth more than ten cells nobody has ever stress-tested.
What the variance is actually telling you
| Gap size | What it means | What to do |
|---|---|---|
| Same tier | Rubric cell is working | Nothing — move on |
| One tier apart | Normal scoring variance | Note it, don't over-correct |
| Two or more tiers apart | Ambiguous rubric language, not a people problem | Rewrite the cell, retest next cycle |
| Consistent one-directional gap (one manager always higher/lower) | Individual calibration drift | Pair that manager on live scoring for a cycle |
Run this quarterly, not once. Rubrics drift as your product, market, and buyer conversations change, and a rubric that was tight in January can quietly go soft by September without anyone noticing until a promotion decision or a PIP gets challenged. If you're building the coaching layer that sits on top of this — what a manager actually says to a rep once a real skill gap has been identified rather than guessed at — the GROW Model Coaching Conversation Script and the Constructive Feedback Cheat Sheet for Sales Managers are worth having on hand for that follow-up conversation.
A rubric nobody has stress-tested against a real, ambiguous call is a document, not a standard. Blind calibration is what turns it into the second thing.