ThinkWork

How to Audit an AI Call-Scoring Model Before You Let It Grade a Single Rep

A five-step inter-rater reliability check any manager can run in an afternoon. No data science degree required.

Before an AI model touches comp, promotion, or a performance improvement plan, somebody with a pulse needs to check its work — not because AI grading is uniquely untrustworthy, but because any grading system, human or machine, drifts silently until someone runs the numbers, and nobody runs the numbers. Here's a way to do it in an afternoon, without a data science background, using calls you already have and managers you already pay.

Why "it looked accurate in the demo" isn't an audit

Vendor demos are curated. They show you calls where the model's read is obviously right, because obviously-right calls are what you'd pick to sell software. The calls that matter for your audit are the ambiguous ones — the discovery call where the rep asked a decent question but the buyer didn't fully answer it, the negotiation where the rep folded on price but held on scope. Those are where a model's real accuracy shows up, and where most rollouts never look, because looking requires blind human grading, and blind human grading is the step everyone skips to save time.

The five steps

StepWhat you doTimeWhat it catches
1. Stratified samplePull 30 calls: 10 the model scored low, 10 mid, 10 high, across at least three reps and two call types20 minWhether you're testing the full range or just the easy cases
2. Blind human gradingTwo experienced managers independently score the same 30 calls on the same rubric, with no visibility into the AI's scores or each other's2–3 hrsWhether your own humans agree with each other before you ask if they agree with the model
3. Agreement rateCalculate human-to-human agreement, then human-to-AI agreement, using the same method for both20 minWhether the model behaves like a second grader or an outlier
4. Disagreement auditPull every call where the AI and the humans differ by more than one grade band, and read the transcript for a pattern45 minThe specific bias — keyword-matching, recency, tone-over-substance, length
5. Set a re-check cadenceDecide the bar for "good enough to deploy" and schedule the next audit, especially after any model or prompt change10 minSilent drift after the vendor pushes an update you didn't ask for

Step 1: pull a sample that can actually embarrass the model

Don't grab thirty random calls. Grab ten the model scored low, ten mid, ten high, spanning at least three reps and, if you run more than one call type, both of them. A sample stacked with obviously good or obviously bad calls will always show high agreement, because those calls don't need judgment. You're testing judgment.

Step 2: grade blind, and grade twice

Two managers, working independently, score all thirty calls against your existing rubric — not a simplified version, the real one, the one that's supposedly informing the model. Neither manager sees the AI's score. Neither sees the other's score until both are submitted. This is the step people cut, usually with "we trust our managers' judgment already," which is exactly backwards — you need to measure that judgment before you can measure the model against it.

Step 3: do the maths that actually matters

The number that matters isn't "the AI agreed with a human 82% of the time." It's how that 82% compares to how often your two human graders agreed with each other. If your managers — experienced, using the same rubric, grading the same calls — only agree with each other 75% of the time, then an AI landing at 82% agreement with either of them isn't accurate, it's suspiciously smooth, and you should be asking why a harder problem produced an easier answer. Exact-match agreement is the simplest version of this; if you want more rigour, Cohen's kappa corrects for chance agreement, but exact-match plus a look at the disagreements will tell a manager almost everything they need to know in an afternoon.

Step 4: read every disagreement, not just count them

This is the step that actually produces something actionable. Pull every call where the model and your humans differed by more than one grade band and read the transcript. Patterns worth naming, because we've seen all four in live deployments:

Step 5: set the bar, then keep checking it

Pick your threshold before you see the results, not after, or you'll rationalise whatever number you get. A reasonable bar: human-AI agreement should sit within a few points of human-human agreement — not below it, and, suspiciously, not far above it either. Then put a recurring date in the diary, particularly after any vendor update, prompt change, or model version bump. Those ship quietly and can move your accuracy without anyone telling you.

What to do with the result

If the model passes, you've earned the right to trust it on the calls that matter, and you now have a documented baseline to compare against next quarter. If it doesn't, don't panic and pull the vendor relationship — use the disagreement audit to work out whether the fix is a prompt change, a rubric clarification, or a permanent human-review layer on the ambiguous middle band of scores, which is usually the 20% of calls doing 80% of the damage to trust in the tool. Either way, write the disagreement patterns up using something structured like the SBI Feedback Model Cheat Sheet — situation, behaviour, impact — so the audit produces a document your team can act on, not a spreadsheet nobody reopens. And if your managers need to build their own fluency before they can run this audit with confidence, a Sales Metrics Literacy Quiz is a faster starting point than a lecture.

We build AI grading into our own assessment work, and we still run this audit against our own model, on a schedule, because "we designed it carefully" has never once been a substitute for "we checked."

The teams that get burned by AI call scoring aren't the ones using a bad model. They're the ones who never found out whether their model was good, because finding out required an afternoon nobody scheduled.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel — just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Confirm by email first. Unsubscribe any time.