ThinkWork

What Happened When We Made Managers Score Their Own 1:1s Before We Scored Them

The self-scores and the external scores didn't just differ — they diverged in a consistent direction, and that direction told us more than the scores themselves.

Across four engagements this year, we tried something that made a few managers visibly uncomfortable before they'd even opened their mouths: we asked them to grade their own 1:1s against the same rubric our external assessors would use, before they saw the external grade. Forty-six managers, two hundred and fourteen recorded coaching conversations, one shared rubric across all of it. The two sets of scores didn't match. Nobody on our side expected them to — self-assessment inflation is one of the oldest findings in the coaching literature. What we didn't expect was that the mismatch had a shape. It wasn't noise. It pointed at exactly one thing, consistently, across four completely different sales organisations: these managers can't diagnose. They can coach the room. They can't coach the skill.

How we ran it

The rubric came out of the same competency taxonomy we use for rep assessment — not because 1:1s and sales calls are the same conversation, but because a coaching conversation is only as good as the specificity of what it's trying to fix. We scored six dimensions per 1:1: rapport and psychological safety, active listening, structure of the feedback (did it follow something like SBI — situation, behaviour, impact), specificity of the skill diagnosis, quality of the follow-up action, and whether the manager checked understanding before moving on. Five-point scale, each dimension. Managers scored themselves within an hour of the session, before seeing anything from us. Our assessors — human graders working from the recording, blind to the self-score — scored the same six dimensions days later.

The numbers

DimensionSelf-score avgExternal avgGapWhy the gap sits here
Specificity of skill diagnosis4.32.22.1Managers can't rate accuracy of a diagnosis they don't know is wrong
Quality of follow-up action4.02.41.6Vague diagnosis produces vague homework; managers rate effort, not precision
Feedback structure (SBI)3.82.90.9Structure is visible and checkable even without expertise
Checked understanding3.93.10.8Managers remember asking "does that make sense," graders check if the rep could actually restate it
Active listening4.13.60.5Genuinely close — this is a behaviour managers can self-observe
Rapport / psychological safety4.44.20.2Managers can accurately feel whether a room was tense or warm

Read down that gap column and the story is obvious. The dimensions where managers can perceive their own performance directly — did the room feel safe, was I listening, did I ask if that made sense — are close to accurate. The dimensions that require an external reference point — is this actually the right skill, will this action plan actually close the gap — blow out to a full two points on a five-point scale. That's not managers being dishonest. It's managers being asked to grade something they have no instrument to measure.

What the diagnosis gap looks like on tape

One example, lightly altered for anonymity. A rep had been losing deals at the proposal stage. In the 1:1, the manager told him: "Good energy on the call, just tighten up your discovery questions before you pitch." Manager's self-score on diagnosis: 5 out of 5, with a note — "identified the root cause clearly."

Our assessor, watching the same underlying call the rep had brought in, scored it a 2. The rep's discovery questions were fine. What he'd actually failed to do, three separate times across the quarter, was reconfirm who controlled budget before he moved to pricing — a distinct, nameable competency, not a discovery-technique problem at all. The manager wasn't wrong that something was off. He was pattern-matching to the nearest thing in his own vocabulary — "discovery" is a word every sales manager has — instead of the actual competency gap, which he didn't have a name for and therefore couldn't see.

That's the mechanism behind every number in that table. You can't accurately self-assess your diagnostic precision, because if your diagnosis were wrong, you wouldn't know — you'd just feel confident. The confidence is real. The diagnosis underneath it is a coin flip.

Why this matters more than the coaching itself

The instinct here is to conclude these managers need coaching training. They don't, particularly — most of them coach with warmth, structure their feedback reasonably, and check for understanding at a level that would satisfy most training programmes. What they're missing is a shared taxonomy of the specific skills their reps are supposed to be building, fine-grained enough that "tighten up discovery" and "reconfirm the economic buyer" are visibly different diagnoses rather than the same vague instinct wearing two names.

Without that taxonomy, every manager is coaching from their own playbook, which means their diagnosis is really just "what would I have done" translated into feedback language. Some of the time that happens to land on the real gap. A lot of the time — 2.1 points' worth of the time, per our data — it doesn't, and the rep goes away with homework that fixes nothing because it was never aimed at the actual problem.

What we changed after seeing this

We stopped treating the self-score/external-score gap as a grading exercise and started treating it as a diagnostic on the managers themselves. Any manager whose diagnosis-dimension gap ran above 1.5 got moved onto a structured coaching cadence built around a shared competency list, rather than more generic "coaching skills" training. Six weeks in, across the two engagements far enough along to check, the gap on that dimension had closed by roughly half — not because the managers got more empathetic, but because they finally had names for the things they were supposed to be looking for.

If you run this exercise yourself, don't be surprised by the overall inflation — everyone self-inflates. Be surprised, and pay attention, if the inflation is uneven. The dimension where your managers are most confident and most wrong is not a coincidence. It's a map of exactly which skill they need before they coach another rep.

If you want a starting rubric rather than building one from scratch, the Coaching Effectiveness Scorecard for Managers covers the same six dimensions we used here. Pair it with the SBI Feedback Model Cheat Sheet for the structure dimension, and the Skill/Will Coaching Diagnostic Framework if you suspect — as we increasingly do — that half your "coaching problem" is actually a "we've never agreed what the skills are" problem.

Nobody in these engagements set out to grade themselves generously. They graded themselves accurately on everything they could see, and confidently on the one thing they couldn't.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel — just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Confirm by email first. Unsubscribe any time.