If you score conversations with AI, this is the number nobody has calculated: how often does the model agree with a trained human marking the same thing against the same rubric? Put the two sets of scores in and find out. Everything runs in your browser, nothing is uploaded, and no sign-up is required.
Take a handful of conversations your AI has already scored and have someone experienced mark the same ones blind, without seeing the AI's output. That blindness matters: a marker who can see the AI's score will drift toward it. Paste both sets in, one line per criterion, and read the within-one-band figure first, then the rank correlation.
If you already grade calls with AI, the odds are nobody has independently checked those scores, and they are deciding who gets coached. Send one call your AI has already graded: a trained assessor marks it blind against your own rubric, then we compare where we agree and where we do not. Free, and you keep the marking either way.