AI Scoring Calibration Session Guide for Enablement Owners
A structured, repeatable process for stress-testing your AI scorecard against human judgment before scores reach a learner. Run this before every new cohort to catch systematic drift early and document agreed fixes to thresholds or rubric language.
How to use it
Pull a stratified transcript sample using the criteria below, run blind human marking before anyone sees AI scores, then compare results using the drift-logging table. Work through the decision rules to determine whether a divergence pattern calls for a rubric rewrite or a prompt change, and record your agreed adjustments before the cohort goes live.
What's inside
- Pre-session setup checklist covering sample size, stratification criteria, and blind-marking logistics
- Session agenda with time allocations for each stage
- Blind human marking sheet with per-competency scoring columns
- Drift-logging table for capturing systematic versus random disagreement patterns
- Named failure modes (over-scoring empathy, under-scoring pacing, binary collapse) with worked examples
- Decision rules: rubric rewrite versus prompt change versus threshold adjustment
- Agreed-adjustments documentation template
- "When NOT to run this session" notes for contexts where calibration adds noise rather than signal
What this is for
This guide runs you through a calibration session you hold before a scored cohort goes live. Its job is to find systematic drift between AI scores and human judgment, and fix the root cause, before learners ever see a number.
It is not a dispute-resolution tool. If you already have scores in the wild and a learner is contesting a result, use the Human vs. AI Score Disagreement Protocol instead. This guide is upstream of that problem.
When NOT to run this session
- Cohort has fewer than 8 participants. Below that threshold, calibration overhead exceeds the benefit. Use straight human marking instead.
- You have no trained human markers available. Calibration against untrained raters produces noise, not signal.
- The AI rubric is brand new and has never been piloted on real transcripts. Run a single pilot pass first; calibration presupposes a rubric that is at least internally consistent.
- You are re-running an identical programme with an identical AI model and no rubric changes since the last calibration. Annual recalibration is sufficient in stable conditions.
Pre-Session Setup (complete 48 hours before the session)
1. Build a stratified sample
Pull transcripts that cover all three performance bands you expect in the real cohort. A random sample will over-represent the middle.
| Band | Target share of sample | Selection rule |
|---|---|---|
| Low (bottom 20% by supervisor estimate) | 30% | Pick transcripts where supervisor flagged a coaching moment |
| Mid (middle 60%) | 40% | Random selection from the pool |
| High (top 20%) | 30% | Pick transcripts supervisor rated as strong |
Minimum sample size: 12 transcripts for a cohort of 8-25 learners. Add 2 transcripts per additional 10 learners above 25.
Strip all AI scores before distributing. Human markers must not see them.
2. Assign markers
Minimum two human markers per transcript. Three is better if you have the capacity. Markers should have direct experience of the skill being assessed (a manager who coaches discovery calls marks discovery call transcripts, not an L&D coordinator who has never run one).
3. Prepare materials
Each marker needs:
- A clean copy of the rubric with behavioural anchors for each score point
- The blind marking sheet (template below)
- Clear instruction: score independently, do not discuss until the comparison stage
Session Agenda
Allow 3 hours for a 12-transcript sample with two markers. Scale time proportionally.
| Time | Stage | Owner |
|---|---|---|
| 0:00-0:15 | Briefing: purpose, ground rules, what counts as a valid disagreement | Enablement owner |
| 0:15-1:00 | Blind human marking (silent, independent) | All markers |
| 1:00-1:15 | Break, enablement owner compiles scores | Enablement owner |
| 1:15-1:45 | Reveal AI scores, populate drift-logging table live | Enablement owner + markers |
| 1:45-2:30 | Drift analysis: separate systematic from random disagreement | Full group |
| 2:30-2:50 | Apply decision rules, agree adjustments | Enablement owner leads |
| 2:50-3:00 | Document agreed changes, assign owners | Enablement owner |
Blind Marking Sheet
One row per competency scored in your rubric. Markers complete this before AI scores are revealed.
| Transcript ID | Competency | Marker 1 Score (1-4) | Marker 2 Score (1-4) | Marker 3 Score (1-4) | Human Average | AI Score | Gap (Human avg minus AI) |
|---|---|---|---|---|---|---|---|
| T-01 | Opening hook | ||||||
| T-01 | Needs questioning | ||||||
| T-01 | Empathy response | ||||||
| T-01 | Pacing | ||||||
| T-01 | Close attempt |
Repeat for each transcript. Compute human average after all markers have submitted independently.
Two-band gap threshold: A gap of 2 or more points (on a 1-4 scale) between human average and AI score on any single competency flags for the drift analysis. A one-band gap is noted but not automatically escalated.
Drift-Logging Table
After revealing AI scores, populate this table. Its purpose is to show you whether disagreements are random (scattered across competencies and transcripts) or systematic (the AI consistently drifts in one direction on one competency).
| Competency | Number of gaps flagged | Direction of AI drift (over/under vs human) | Consistent across 3+ transcripts? | Consistent across both performance bands? |
|---|---|---|---|---|
| Opening hook | ||||
| Needs questioning | ||||
| Empathy response | ||||
| Pacing | ||||
| Close attempt |
A drift is systematic if it appears in the same direction on the same competency across 3 or more transcripts. Random disagreement (different transcripts, different directions, no pattern) does not require a rubric or prompt change.
Named Failure Modes
These are the patterns most commonly found during calibration. Recognising them by name speeds up the conversation.
Over-scoring empathy. The AI reads acknowledgement phrases ("I understand", "that makes sense") and scores them as empathy regardless of whether the content that follows demonstrates genuine understanding. Human markers score the behaviour; the AI scores the vocabulary. Fix: tighten rubric language to require a behavioural indicator, not just the phrase.
Under-scoring pacing. AI models score from transcript text. Silence, deliberate pausing, and the rep allowing the prospect to finish are invisible in text. AI will systematically under-score pacing if the rubric is written to capture audio behaviour. Fix: either remove pacing from AI-scored competencies or rewrite the rubric to use text-visible proxies (word count per turn, number of clarifying questions per rep turn).
Binary collapse. The AI treats a 1-4 scale as a binary pass/fail, distributing scores at 1 and 4 with almost nothing at 2 or 3. This is almost always a prompt issue, not a rubric issue. Fix: adjust the prompt to require the model to apply the full scale and justify middle scores.
Band compression at the top. The AI rarely awards 4s on subjective competencies even when human markers consistently do. Fix: check rubric anchor language. If the 4-point descriptor contains words like "exceptional" or "exemplary", the model is likely interpreting those conservatively. Replace with observable behaviours.
Interrater drift between human markers. Before concluding the AI is wrong, check whether your human markers disagree with each other as much as they disagree with the AI. If human interrater agreement is below 70% on a competency, that competency's rubric anchor needs rewriting before you can meaningfully calibrate AI against it.
Decision Rules
Once you have completed the drift-logging table, apply these rules in order.
Rule 1: Is human interrater agreement below 70% on the flagged competency? Yes: rewrite the rubric anchor before any other action. No AI fix will compensate for an ambiguous rubric. No: proceed to Rule 2.
Rule 2: Is the drift on a competency that requires audio or non-verbal signal (tone, pacing, silence)? Yes: either remove it from AI scoring or rewrite it with text-visible proxies. Do not attempt to fix this with a prompt change. No: proceed to Rule 3.
Rule 3: Is the AI using the full scoring scale, or collapsing to binary? Collapsing: this is a prompt issue. Adjust the prompt instruction to enforce mid-scale scoring with justification. Re-run five test transcripts before committing the change. Full scale: proceed to Rule 4.
Rule 4: Is the drift consistent in direction across 3 or more transcripts and across both performance bands? Yes: this is a rubric language problem. The AI is reading the anchor differently from your markers. Rewrite the anchor with more specific behavioural language. Reconvene for a mini-calibration (six transcripts only) to validate the fix. No: this is random disagreement. Document it, monitor the next cohort, and take no structural action now.
Agreed-Adjustments Documentation Template
Complete this before closing the session. This becomes the version-controlled record of your rubric state.
| Competency | Type of change | Old threshold or rubric language | New threshold or rubric language | Who implements | Deadline | Validation method |
|---|---|---|---|---|---|---|
| Rubric rewrite / Prompt change / Threshold adjustment |
Store this alongside your rubric version history. If scores from a previous cohort were produced under the old settings, note the version boundary clearly so learner scores are not compared across calibration changes without a caveat.
Recalibration Triggers
Run this session again (outside the standard pre-cohort schedule) if any of these occur:
- The AI model or model version changes
- The rubric is edited for any reason
- A new facilitator joins as a human marker for the first time
- Post-cohort data shows a two-band gap rate above 10% on any single competency
- A significant change in the learner population (new sales motion, new product category)