You Inherited an AI Scorecard. You Have No Idea What It Was Built to Measure.
Most teams running AI assessment didn't write the rubric. Here's how to find out whether it's grading what you think it is.
Someone handed you a scoring system. Maybe it came with the platform contract, maybe the previous enablement lead set it up before they left, maybe a vendor pre-configured it during onboarding and called it "best practice." Either way, it is now running against real people's performance records, influencing their coaching plans, possibly feeding into compensation reviews, and you would struggle, if someone pressed you hard in a meeting, to explain why a 3 means what it says rather than a 2 or a 4. That is not a minor administrative gap. That is a liability.
The problem is not that AI scoring is unreliable. Some of it is genuinely good. The problem is that the reasoning behind a rubric does not travel with the rubric. Criteria get copy-pasted, bands get adjusted for reasons nobody wrote down, and the person who knew why "meets expectations" was set at that threshold left the business eighteen months ago. You inherited the output of someone else's thinking without the thinking.
Here is how to find out what you actually have.
Step 1: Read every criterion against real outputs, not the documentation
Start with the rubric itself. Print it, or open it in a way that lets you annotate it. Now pull ten scored transcripts or call recordings from the last thirty days, spread across your performance bands. Read each criterion, then ask: can I see the evidence for this score in the output I am looking at?
Do this for every criterion, not just the ones that feel risky. You are looking for two failure modes. The first is criteria that are so vague the score could be almost anything ("demonstrates professionalism," "shows good product knowledge"). The second is criteria that appear specific but are actually proxies, things like talk-to-listen ratio where the AI is measuring a signal that correlates with skill under certain conditions and not others.
Write a short note next to each criterion: visible and specific, vague, or proxy. Any criterion in the latter two categories needs to be interrogated before you rely on the scores it produces.
Step 2: Run a blind inter-rater check
Take five of those same transcripts. Strip out the AI scores. Give them to two experienced practitioners, ideally people who have actually done the role being assessed, with nothing but the rubric and a clean scoresheet. Ask them to score independently, with no discussion.
Then compare: human one, human two, AI.
You are not expecting perfect agreement. You are looking for pattern. Where all three cluster within one band, the criterion is probably measuring something real. Where the AI consistently scores higher or lower than both humans, you have a calibration problem. Where the two humans disagree significantly with each other, the criterion is under-defined and the AI score is noise dressed as data.
The Human vs AI Scoring Agreement Checker can structure this comparison if you want a clean view of the gaps, but even a simple spreadsheet will show you where the divergence is worst. Rank your criteria by disagreement. Address the top three before your next performance cycle runs.
Step 3: Identify criteria that require inference the model cannot make
This is the one most people skip, and it is the one that causes the most downstream damage.
A transcript is a record of what was said. It is not a record of what was meant, what the prospect was thinking, whether the rep had read the account history before the call, or whether the silence at 4:12 was a confident pause or a rep who had lost their place. AI scoring systems working from transcripts can only evaluate what is linguistically present. Anything that requires contextual inference is beyond them.
Go through your criteria and flag any that implicitly require the scorer to know something the transcript does not contain. Common examples:
| Criterion | What the transcript shows | What it cannot show |
|---|---|---|
| "Tailors pitch to prospect's specific situation" | The rep mentioned the prospect's industry | Whether that tailoring was based on research or a lucky guess |
| "Handles objections with confidence" | The rep gave a response | The rep's tone, composure under pressure |
| "Demonstrates strategic account awareness" | Keywords present or absent | Whether the rep understood the account or was pattern-matching |
Any criterion in the right-hand column is being scored on a proxy at best. That is not automatically disqualifying, but you need to know it is happening, and your calibration notes should say so explicitly.
Step 4: Spot the criteria that have drifted from what the business actually needs
Rubrics are written at a point in time. Your business is not the same business it was when this one was written.
Pull your last three quarterly commercial reviews. Look at what the leadership team identified as the primary skill gaps, the deal-stage failures, the conversion problems. Now look at your rubric. Count how many criteria directly address those stated priorities.
If your pipeline problem is that reps cannot get past procurement without re-qualifying the deal, and your rubric has six criteria about discovery and nothing about commercial navigation, the scorecard is measuring last year's problem. You will generate a lot of data about something that does not matter and nothing about the thing that does.
This is also where inherited rubrics from vendor templates are most likely to fail. A vendor builds a generic rubric optimised for the modal sales motion. Your sales motion is not modal. The criteria that were meaningful for the previous owner may be irrelevant or actively misleading for you.
Make a simple list: criteria that map to current business priorities, criteria that are neutral, criteria that map to old or irrelevant priorities. Anything in the third column is a candidate for removal or replacement before you run another full assessment cycle.
Step 5: Get a human expert to re-mark a sample before the next cycle runs
None of the above replaces a qualified second opinion on the scores themselves.
Pick fifteen assessments: five from your top performers, five from your middle band, five from the people your rubric is flagging as underperforming. Have someone with genuine domain expertise, not just a manager, but someone who can be specific about why a call worked or did not, re-mark them using the rubric as written. Compare to the AI output.
If the expert agrees with the AI on twelve of fifteen, your rubric is probably calibrated and your methodology is defensible. If the expert disagrees on eight or more, you have a problem you need to fix before those scores appear in anyone's development plan or review documentation.
The point of this exercise is not to replace AI scoring. It is to give you the ability to stand behind the scores you are using. "The system said so" is not a defensible position when someone asks why their performance rating dropped after a coaching quarter where, by their account, they improved. You need to be able to say: here is what we measured, here is why we measured it, here is how we know the bands mean what we say they mean.
The rubric you inherited might be excellent. Some of them are. But right now you cannot tell the difference between excellent and plausible, and those two things produce very different outcomes when the data gets used.