ThinkWork
Team/Manager Checklist Free

AI Scorecard Inherited Validation Checklist

A structured checklist for validating an AI scoring rubric you didn't build, before you stake your team's development decisions on its output. Covers the five most common failure points in pre-built scorecards and gives you a concrete test for each one you can run within a week.

How to use it

Work through the checklist before the next cohort is scored, not after. For each failure point, run the stated test and record your finding in the Notes column. If you hit two or more red flags, treat the scorecard as provisional and run a parallel human-scoring exercise until the gaps are fixed.

What's inside

  • Five-failure-point structure covering the most common ways an inherited scorecard breaks down
  • Observable-criteria test to catch dimensions the AI cannot actually see from a transcript
  • Band-definition consistency test using a blind double-mark exercise
  • Calibration evidence audit to check whether the scoring model was ever validated against real performance data
  • Weighting audit to identify vendor defaults masquerading as your organisation's priorities
  • Pass/fail threshold audit to confirm cut scores have a documented rationale, not an arbitrary origin
  • Red flag thresholds for each test, telling you when to pause and when to proceed with caution
  • A "decision table" summarising what to do based on how many failure points you find
  • Honest "when this checklist is not enough" note for edge cases that need deeper investigation

Who this is for

Managers, L&D leads, or sales enablement owners who have been handed a live AI roleplay scorecard they did not commission, and who need to decide whether to trust it before the next cohort is assessed.

This checklist does not require access to the vendor's model internals, training data, or source code. Every test here can be run with transcripts, the scoring rubric document, and a couple of hours from two experienced assessors.


Before you start: gather these five documents

DocumentWhere to get itRed flag if missing
The full scoring rubric, including all criteria and band descriptorsVendor or internal project folderStop. You cannot validate what you cannot read.
The criteria weighting breakdownSame sourceSuggests weights were never made explicit
Any calibration or validation report produced at setupVendor deliverablesCommon to find this was never done
The pass/fail threshold and how it was setProject sign-off documentationOften undocumented
A sample of 20+ scored transcripts with the AI's dimension-level scoresPlatform exportNeeded for tests 1, 2, and 3

Failure Point 1: Criteria the AI cannot observe from a transcript

Why it matters. Some scorecard criteria describe behaviours that require non-verbal or contextual information: eye contact, physical presence, tone of voice, whether the rep paused meaningfully. If a text transcript is the sole input to the AI, these criteria will be scored on a proxy at best, or fabricated at worst.

The test.

  1. Print the full list of scored criteria.
  2. For each criterion, ask: "Could a human assessor score this accurately from a typed transcript alone, with no audio and no video?"
  3. Mark any criterion where the honest answer is "no" or "only partially".

Red flag. More than one criterion fails this test. Even one is a problem if it carries significant weighting.

What to do. For any unobservable criterion: confirm with the vendor exactly what signal the AI uses as its proxy, or remove the criterion from the rubric. Do not leave it scored silently on an undisclosed basis.


Failure Point 2: Band definitions too vague to apply consistently

Why it matters. If two experienced assessors cannot independently reach the same score from the same transcript, the scoring instrument is unreliable. AI consistency is meaningless if the underlying scale is ambiguous.

The test.

  1. Select 8 transcripts from your sample: two that the AI rated at each of four performance levels (or two clearly above and two clearly below the pass mark if you only have two bands).
  2. Ask two assessors, independently and without seeing the AI scores, to score each transcript on two or three criteria only. Keep it manageable.
  3. Compare. Calculate the percentage of exact matches and the percentage within one band.

Red flag. Exact agreement below 60%, or within-one-band agreement below 80%.

What to do. Return to the band descriptors and rewrite any that use vague language ("demonstrates good rapport", "shows awareness of customer needs") without observable anchors. Every band level needs a behavioural example, not an adjective.


Failure Point 3: Calibration evidence was never collected

Why it matters. A well-built AI scorecard was calibrated against a reference set: transcripts scored by expert human assessors, used to check that the AI's scores correlate meaningfully with ground truth. Without this, you have no evidence the AI is measuring what it claims to measure.

The test.

  1. Ask the vendor or the internal project owner: "Do you have a calibration or validation report showing AI scores were compared against human expert scores on a held-out sample?"
  2. If yes, check the sample size (fewer than 30 is weak), the correlation statistic used, and whether the sample represented the full performance range, not just average performers.
  3. If no report exists, treat the scorecard as uncalibrated.

Red flag. No report exists, sample size under 30, or the calibration sample was drawn only from a narrow band of performers.

What to do. Run your own lightweight calibration. Have two experienced coaches score 20 transcripts independently, average their scores per dimension, then compare against the AI output. A Pearson r below 0.70 on any weighted dimension is a serious concern.


Failure Point 4: Weights reflect vendor defaults, not your competency priorities

Why it matters. Most AI roleplay scoring platforms ship with default weightings built around a generic sales framework. If nobody actively reset those weights to reflect your organisation's actual competency priorities, you are measuring something plausible-sounding but not your thing.

The test.

  1. List every scored criterion and its weight.
  2. Ask the original project owner: "Were these weights set by us, or did they come from the platform default?"
  3. If set by you: find the document or meeting record where those decisions were made.
  4. If no record exists: assume they are vendor defaults.
  5. Cross-reference the criterion weights against your internal competency framework or the skills your coaching programme treats as priorities. Do the highest-weighted criteria match your highest-priority skills?

Red flag. No documentation of weight decisions, or the top three weighted criteria do not correspond to your top three coaching priorities.

What to do. Convene a 90-minute session with your most experienced coaches, a sales manager, and (if available) the L&D lead. Agree weights from scratch using a forced-rank exercise. Reset the platform, or apply manual reweighting to any historical scores before using them for decisions.


Failure Point 5: Pass/fail thresholds with no documented rationale

Why it matters. A cut score of 70% sounds authoritative. But if nobody can tell you why 70 and not 65 or 75, the threshold is arbitrary. Arbitrary thresholds mean some competent reps are wrongly flagged and some weak reps are wrongly passed.

The test.

  1. Find the pass/fail threshold in the documentation.
  2. Ask: "Where is the evidence that this threshold distinguishes competent from not-yet-competent performance?"
  3. Acceptable answers: the threshold was set against a criterion-referenced performance standard, or against the historical score distribution of known-good performers, or through an Angoff-style panel exercise. Any of these counts.
  4. Unacceptable answers: "It seemed reasonable", "The vendor suggested it", "It was round number", or silence.

Red flag. No documented rationale, or the rationale is a vendor default.

What to do. Pull the AI scores for any cohort of reps you already know the performance outcome for (passed probation, hit quota, failed out). Check whether the scorecard threshold would have correctly classified them at a rate above 70%. If it does, you have retrospective evidence. If it does not, the threshold needs resetting before it is used to make pass/fail decisions.


Decision table

Red flags foundRecommended action
0Proceed. Document your validation steps and review again after next cohort.
1Proceed with caution. Fix the gap within 30 days. Do not use scores for high-stakes decisions until resolved.
2-3Treat as provisional. Run parallel human scoring for the next cohort. Present findings to the project owner with a remediation plan.
4-5Do not use for consequential decisions. Escalate to the commissioning manager. A partial rebuild is likely needed.

When this checklist is not enough

This checklist tells you whether the scorecard is structurally sound enough to trust. It does not tell you:

  • Whether the underlying language model has systematic bias against certain accents, dialects, or communication styles. That requires a fairness audit across demographic subgroups, which needs a larger sample and a different methodology.
  • Whether the criteria are measuring the right things for your specific sales motion, as opposed to a generic one. That is a content validity question and requires a job task analysis.
  • Whether the platform's scoring is stable over time, meaning the same transcript scored six months apart returns the same result. That is a reliability-over-time check and needs a longitudinal sample.

If any of those three matter for your context, treat this checklist as the first gate, not the last.


Notes column template

Use this in a working document during your validation.

Failure PointTest Run?FindingRed Flag?Action RequiredOwnerDue Date
1. Unobservable criteria
2. Vague band definitions
3. Missing calibration
4. Vendor default weights
5. Undocumented threshold
Stay current

Get told when new AI Assessment, Scoring & Assurance resources land.

Tick the topics you care about, then subscribe. Alerts start straight away, no confirmation step. No digest spam, just a note when something genuinely useful is added.

Choose topics below (or leave blank for everything). Unsubscribe any time.