ThinkWork

The Rep Who Talks Less Is Scoring Lower. That Is Your AI Judge's Problem, Not Your Rep's.

LLM-as-judge systems reliably score longer, more elaborate responses higher, even when a concise answer is objectively better.

Your top discovery rep asks two questions, gets the economic buyer talking for twelve minutes, and closes the call with a clear next step agreed. Your middling rep asks nine questions, interrupts twice, and fills every silence with another prompt. The AI scorer gives the second rep a higher competency rating. You coach the first rep to "be more thorough." You have just made your team worse.

This is not a hypothetical. It is a predictable output of a failure mode called verbosity bias, and most enablement and L&D teams deploying AI call-scoring have never heard of it.

What Verbosity Bias Actually Is

Verbosity bias is well-documented in the academic literature on LLM evaluation. When an LLM is used as a judge, it consistently rates longer, more elaborate responses as higher quality, independent of whether the length adds any informational value. The original Stanford and Anthropic research on this goes back to 2023; it has been replicated across model families. The effect is not subtle. In head-to-head comparisons, a verbose but incorrect answer will frequently outscore a concise correct one.

The intuitive explanation: LLMs are trained on human-generated feedback, and humans, when asked to rate responses, tend to interpret length as effort and effort as quality. That preference gets baked into the model's reward signal. The judge inherits the bias of the training raters.

In a sales call context, this translates directly. Longer transcript, more rep turns, more words per turn: the AI reads these as signals of engagement, thoroughness, and skill. The rep who asks precisely and then shuts up produces a shorter transcript. The system scores her lower. You now have a metric that actively penalises controlled, buyer-centric discovery.

It Does Not Travel Alone

Verbosity bias compounds with two other known LLM-judge failure modes that are also under-discussed in sales tech circles.

Position bias is the tendency of an LLM judge to rate content more highly when it appears early in the input, or in the first position in a comparison. If your scoring pipeline feeds the transcript sequentially and the rep says something strong in the first two minutes, that moment carries more weight than an equally strong moment at minute eighteen. A rep who front-loads talk and fades is structurally advantaged over one who builds to a precise insight.

Self-preference (sometimes called self-enhancement bias) appears when the model generating the score is the same family as the model being evaluated, or when the judge has been fine-tuned on outputs similar to the responses it is rating. In practice for teams building internal pipelines: if you used GPT-4 to generate your ideal response examples, and GPT-4 to score rep responses against them, you have a closed loop that rewards responses that sound like GPT-4. That is not the same as rewarding responses that close business.

Together, these three biases push scores toward: longer, earlier, and more AI-sounding. None of those attributes correlate with winning.

The Specific Damage This Does to Coaching

Here is what I see when teams have been running biased AI scoring for six months or more without auditing it.

Coaching direction AI scoring producesWhat actually improvesNet effect
Ask more questions per callAsk fewer, better questionsReps become chattier
Elaborate on responsesGive concise, specific answersReps over-explain
Fill silences with follow-up promptsHold silence after a strong questionReps interrupt the thinking pause
Front-load your value propositionBuild to it through discoveryReps pitch earlier

Each of these individually is correctable. All four running simultaneously for two quarters, reinforced by a score that managers can see, is a cultural problem. Reps optimise for what gets measured. If the measure is broken, the optimisation is broken.

The particularly nasty edge case: your best reps, the ones with enough experience to feel when a call is going well, will often score lowest. They are the ones comfortable with silence. They are the ones who ask one precise question instead of three vague ones. They are the ones who stop when they have what they need. The AI flags them as underperforming. A nervous manager starts coaching them. You are now degrading the people whose behaviour you should be replicating.

The Fix Is Not a Better Model

Switching from GPT-4o to Claude 3.5 or to whatever the current frontier model is will not solve this. Verbosity bias is present across model families. What changes the outcome is the structure of the scoring prompt, specifically whether the judge is asked to rate overall impression or to evaluate specific, observable, defined behaviours.

"Rate the quality of this rep's discovery questioning" is an impression prompt. It activates verbosity bias because the model has no anchor other than what a good answer sounds like, and a long answer sounds more like a good answer.

"Did the rep ask an open question that invited the buyer to describe the business impact of the problem? Yes or no, and quote the line." is a behaviour prompt. It gives the judge a specific thing to look for. Length is irrelevant. Position in the transcript is irrelevant. The model either finds the behaviour or it does not.

This is the principle behind The Mastery Standard's approach to rubric construction: every competency is defined at the observable behaviour level, not the impression level. The judge is not asked "how good was the discovery?" It is asked whether specific named behaviours were present, when, and with what effect. That structure substantially reduces the surface area for verbosity bias to operate.

If you are running an internal pipeline and you have inherited a rubric you have never stress-tested, the fastest diagnostic is to run two transcripts through your scorer: one where a rep asks many questions and generates a long, busy transcript, and one where a rep asks two precise questions and the buyer talks for ten minutes. Check whether the second rep scores lower. If she does, your rubric is measuring volume, not skill.

The Human vs AI Scoring Agreement Checker is useful here. If your human raters and your AI scorer systematically diverge on short, buyer-heavy calls, you have found your verbosity leak. That divergence pattern is more diagnostic than any overall agreement percentage.

A secondary check worth running: pull your Cold Call Talk-to-Listen Ratio Scorecard benchmarks and compare them against what your AI scorer actually rewards. If the scorer's top-rated calls have a rep talk-to-listen ratio above 50%, something is wrong. Best-in-class cold call discovery runs closer to 35-40% rep talk time. The scorer should be penalising verbosity, not rewarding it.

What Good Rubric Design Looks Like in Practice

A behaviour-anchored rubric entry for discovery questioning does not say "demonstrates strong questioning skills." It says something like:

  1. Rep asked at least one open question about business impact (not product fit).
  2. Rep waited for the full answer before speaking again.
  3. Rep did not answer their own question or add a clarifying prompt before the buyer responded.
  4. Rep used the buyer's language in the subsequent question.

Each of those is scorable from a transcript without any inference about quality. An LLM judge can find them or not find them. Verbosity is orthogonal to all four.

The same principle applies to roleplay assessment. If you are scoring simulated calls and the rubric allows the model to form an overall impression rather than check specific events, you are measuring how much the output resembles a good call rather than whether it contains the behaviours that make calls good.

The difference sounds minor. The coaching downstream is not.

Your AI scorer is a mirror. If the mirror is warped, everything you build from what you see in it is going to be slightly wrong, and the error accumulates. The reps who need the most help may look fine. The reps doing the most important things may look like they need work. Six months of that, at scale, is expensive. Not just in training budget but in quota.

New posts

Get new posts in your inbox.

A fresh post most mornings. No digest spam, no course funnel, just the post, and one click to stop. Prefer a reader? Subscribe by RSS.

Alerts start right away. Unsubscribe any time.