ThinkWork
Team/Manager Framework Free

Human vs. AI Score Disagreement Protocol (2-Band+ Gaps)

A step-by-step triage framework for when a human marker and an AI score diverge by two bands or more on a sales roleplay or call assessment. Covers the five root causes, a decision tree with four possible outcomes, and ready-to-use language for communicating the result to the learner and to HR.

How to use it

Pull this out the moment a 2-band-plus gap is flagged, before anyone acts on either score or feeds either result into a performance record. Work through the triage sequence in order, use the decision tree to reach one of the four named outcomes, then use the communication scripts verbatim or close to it. Review the "flag for rubric" outcome log quarterly as a standing calibration input.

What's inside

  • Definition of a qualifying disagreement (2-band-plus gap) and why the threshold matters
  • The five named causes of disagreement with honest descriptions of how each one appears in practice
  • A linear triage sequence (six steps) to diagnose which cause is in play
  • A decision tree mapping each diagnosis to one of four outcomes: accept AI, accept human, escalate to blind third mark, or flag criterion for rubric review
  • Word-for-word script for explaining the outcome to the learner without undermining the programme
  • Word-for-word script for communicating the outcome to HR or a people team
  • A named failure mode for each triage step so you know when you are going wrong
  • A short "when NOT to use this protocol" note covering low-stakes developmental scores

What counts as a qualifying disagreement

A 2-band-plus gap means the human and AI scores land at least two full performance bands apart on the same criterion, using the same rubric, for the same assessed moment. Examples: AI scores "Developing," human scores "Exceeding." AI scores "Not Yet," human scores "Proficient."

A 1-band gap does not trigger this protocol. Treat it as normal variance and note it in your calibration log for pattern review. A 2-band-plus gap means one of five specific things has gone wrong, and you must find out which before either score is recorded, shared, or acted on.


The five causes of a 2-band-plus gap

1. Criterion ambiguity. The rubric descriptor is written loosely enough that two competent, good-faith markers can read it differently and both be correct relative to the text. Common in soft-skill criteria: "builds rapport," "demonstrates curiosity." The rubric is the problem, not the markers.

2. Transcript fidelity failure. The AI scored from a transcript, and the transcript contains errors: missed words, misattributed turns, dropped filler that carried meaning (pauses, affirmations), or a cropped recording. The AI's input was not the same conversation the human heard.

3. Rubric drift. The human marker has unconsciously shifted their personal standard since the rubric was last calibrated. Often directional: consistently lenient or consistently harsh compared to the anchor examples. The human is scoring a version of the rubric that no longer matches the written one.

4. Model blind spots. The AI model has a documented or discoverable limitation: it cannot score for tone or prosody from text alone, it underweights domain-specific vocabulary your team uses, or it performs poorly on calls below a certain length. The AI is technically functioning but the criterion falls outside what it can reliably assess.

5. Human marker bias. The human score is influenced by knowledge of the learner (halo or horn effect), by recency bias from a previous call, or by the emotional content of the roleplay scenario. The human is not scoring the evidence in front of them.


Triage sequence

Work through these steps in order. Stop when you have a diagnosis.

Step 1: Check the transcript fidelity. Pull the transcript the AI used. Play back the recording against it for the specific assessed moment. Look for dropped turns, misattributions, or cuts.

  • If you find material errors, the cause is transcript fidelity failure. Go to the decision tree.
  • Failure mode: skipping this step because you trust the transcription vendor. Even 94% accuracy rates produce meaningful errors on short, high-stakes exchanges.

Step 2: Check the criterion text. Read the rubric descriptor for the criterion in question out loud. Then read it again as if you were a new marker with no prior context.

  • If two reasonable readings of the text produce different band-level interpretations, the cause is criterion ambiguity. Go to the decision tree.
  • Failure mode: assuming you know what the criterion means because you wrote it. The issue is whether a new reader would agree.

Step 3: Check for model blind spot. Consult your AI vendor's documented limitations or your own validation record for this tool. Ask: does this criterion require something the model cannot assess from text (tone, pacing, physical presence)?

  • If yes, the cause is a model blind spot. Go to the decision tree.
  • Failure mode: assuming the model can assess everything because the vendor did not explicitly say it cannot. Absence of a disclaimer is not a capability guarantee.

Step 4: Run a rubric drift check on the human marker. Pull the marker's last 10 scores on this criterion. Compare the distribution against the programme baseline. A marker who is more than 0.5 bands above or below the cohort mean on a criterion, consistently, is showing drift.

  • If the marker's recent distribution is significantly off-baseline, the cause is likely rubric drift. Go to the decision tree.
  • Failure mode: treating this as an accusation rather than a calibration question. Frame it as a data check, not a performance conversation.

Step 5: Check for context contamination. Ask the marker directly: did you have prior knowledge of this learner's performance history when you scored? Did the scenario involve a topic you find personally charged? Is this call adjacent to one you scored differently and remember clearly?

  • If yes to any, the cause is likely human marker bias. Go to the decision tree.
  • Failure mode: not asking because it feels confrontational. This is a standard calibration question, not an allegation.

Step 6: No clear cause found. If steps 1-5 produce no diagnosis, escalate to a blind third mark regardless. Document that the gap was undiagnosed.


Decision tree

DiagnosisOutcome
Transcript fidelity failureRescore AI from corrected transcript. If gap closes to 1 band or fewer, accept corrected AI score. If gap persists, go to blind third mark.
Criterion ambiguityFlag criterion for rubric review. Suspend scoring on this criterion until reviewed. Use blind third mark as interim score for this learner.
Model blind spotAccept human score. Document the limitation formally. Remove criterion from AI scope until resolved with vendor.
Rubric drift (human)Recalibrate marker against anchor examples. Rescore with recalibrated marker. Accept recalibrated human score.
Human marker biasBlind third mark. Do not inform third marker of either prior score or the context contamination.
UndiagnosedBlind third mark. Flag for protocol review.

Blind third mark rules: The third marker must not see either prior score. Give them only the rubric and the assessed material. After they score, the outcome is: if third mark agrees with AI, accept AI. If third mark agrees with human, accept human. If third mark is a third different score, take the median band and escalate the criterion for rubric review.


Communication scripts

Explaining the outcome to the learner

Use this once you have reached a final score outcome. Do not discuss the disagreement before you have one.

"Your assessment on [criterion] went through an additional review step. Where the initial scores diverged by more than our threshold, our protocol requires us to resolve that before we record anything. The final score is [band], based on [accepted score source]. That is the score that stands. If you'd like to talk through what it reflects in your assessed call, I'm happy to do that now."

Do not say: "The AI got it wrong," "The human marker got it wrong," or "We're not sure which to trust." None of those are accurate and all of them erode confidence in the programme.

Communicating the outcome to HR or the people team

"A 2-band disagreement was flagged between the AI and human scores for [learner] on [criterion] on [date]. We ran the disagreement protocol. The diagnosed cause was [cause]. The resolved outcome is [band/score], reached by [method: corrected rescore / recalibrated human / blind third mark]. That is the score entered into the record. The criterion has [been flagged for rubric review / had the model limitation documented / no further action required]."

Do not volunteer uncertainty about the programme's reliability. State the process, the diagnosis, and the outcome. If HR asks whether the scores are reliable, the accurate answer is: "The protocol exists precisely because we treat any significant disagreement as a signal to investigate rather than arbitrate. The score in the record has been reached through a defined process."


When NOT to use this protocol

Do not apply this to low-stakes, developmental-only scores where no result feeds into a performance record, promotion decision, or pay review. In a pure learning context, a 2-band gap is useful diagnostic information in itself. Show the learner both scores, explain that different markers read the evidence differently, and use it as a conversation about the criterion. Reserve this protocol for assessed scores that carry formal weight.


Quarterly maintenance

Keep a log of every disagreement that passes through this protocol. Review it quarterly for:

  • Criteria that appear repeatedly (criterion ambiguity or model blind spot signal)
  • Markers who appear repeatedly (rubric drift or bias pattern signal)
  • Outcomes that consistently go to blind third mark (rubric or training issue)

The log is your early warning system. A single disagreement is an operational event. A pattern is a programme fault.

Stay current

Get told when new AI Assessment, Scoring & Assurance resources land.

Tick the topics you care about, then subscribe. Alerts start straight away, no confirmation step. No digest spam, just a note when something genuinely useful is added.

Choose topics below (or leave blank for everything). Unsubscribe any time.