ThinkWork
Interactive Leader Free · no sign-up

Human vs AI Scoring Agreement Checker

If you score conversations with AI, this is the number nobody has calculated: how often does the model agree with a trained human marking the same thing against the same rubric? Put the two sets of scores in and find out. Everything runs in your browser, nothing is uploaded, and no sign-up is required.

Your numbers

Format: label, AI score, human score. One per line. Labels are optional; commas, tabs or spaces all work.
4 for a 0-4 rubric, 5 for 1-5, and so on
Set to 2 on a 0-10 scale for a comparable read

Method & assumptions

  • Agreement is computed per paired score, not per call, so you can use it on one call across many criteria or on many calls against one criterion.
  • Rank correlation is Spearman's rho with average ranks for ties.
  • A 'band' defaults to 1 point. If your scale is 0 to 10, set it to 2 for a comparable read.
  • The marker must be blind to the AI's score or the number is meaningless.
  • One sample marked once is a prompt to look closer, not a reliability figure.

How to use it

Take a handful of conversations your AI has already scored and have someone experienced mark the same ones blind, without seeing the AI's output. That blindness matters: a marker who can see the AI's score will drift toward it. Paste both sets in, one line per criterion, and read the within-one-band figure first, then the rank correlation.

For enablement leaders

Is the AI scoring your team actually right?

If you already grade calls with AI, the odds are nobody has independently checked those scores, and they are deciding who gets coached. Send one call your AI has already graded: a trained assessor marks it blind against your own rubric, then we compare where we agree and where we do not. Free, and you keep the marking either way.

Get one call marked blind Not scoring with AI yet? Start here