AI Scoring Blind Spot Register: What a Transcript Cannot Show
A structured checklist of every sales competency that transcript-based AI scoring physically cannot assess, with the specific failure mode each gap produces and the supplementary evidence needed to close it. Use this before extending trust to any AI score, to draw a permanent boundary around what the model is and is not being asked to judge.
How to use it
Run this register when onboarding any AI scoring tool or when expanding its scope to new call types. For each blind spot, decide whether your team can gather the supplementary evidence listed: if you cannot, mark the competency as unscored and note it on the rep's review record. Revisit the register whenever you change call format, add video reviews, or update your scorecard.
What's inside
- Full taxonomy of transcript-invisible signals: tone, pace, silence, physical presence, camera behaviour, pre-call research quality
- Infer-only competencies that transcripts mask: genuine curiosity vs. scripted delivery, emotional regulation, real-time listening vs. wait-to-talk
- Named failure mode for each blind spot: the specific score that looks valid but isn't
- Supplementary evidence column: exactly what would close each gap
- Hold trigger guidance: when absence of that evidence should pause the score entirely
- A ready-to-use manager verdict field for each competency
- Usage boundary statement to anchor the register in your QA process
Purpose and Scope
This register exists to draw a permanent boundary. Transcript-based AI scoring is useful. It is not complete. The failure mode to fear is not a wrong score that looks wrong: it is a wrong score that looks right, because the model had enough words to produce a confident output even though the thing being measured was never in the words.
Use this register before extending trust to any AI score. It is not a dispute tool (see the Disagreement Protocol for that). It is not a scorecard audit (see the Inherited Validation Checklist for that). It is a prior constraint: a map of what the model cannot see, so that a manager does not accidentally treat absence of evidence as evidence of quality.
Hold trigger rule (applies throughout): If a competency is marked BLIND SPOT and the supplementary evidence column cannot be populated for a given call, the competency score must be suspended: not zeroed, not averaged, suspended, and the rep's review record must note the gap explicitly.
How to Read the Table
- Competency: The skill or behaviour being assessed.
- Why transcript fails: The structural reason a word-for-word record cannot capture this.
- Failure mode: The specific false reading the AI produces.
- Supplementary evidence: What would actually close the gap.
- Hold trigger: The condition under which the score must be suspended.
The Blind Spot Register
Section 1: Paralinguistic Signals (Voice and Pacing)
| Competency | Why transcript fails | Failure mode | Supplementary evidence | Hold trigger |
|---|---|---|---|---|
| Vocal warmth and rapport tone | Transcripts record words, not delivery. A warm sentence and a cold one are identical text. | Rep scores high on "rapport language" despite a flat, transactional tone that killed trust. | Audio recording reviewed by a human scorer using a defined tone rubric. | No audio available, or audio reviewed by AI alone without human calibration. |
| Pacing and patience | Transcript timestamps exist but do not indicate whether pace felt rushed or confident. | Rep scores high on "clear communication" despite speaking at a pace that overwhelmed the buyer. | Audio review noting average sentence rate in key segments (discovery, objection, close). | No audio, or timestamps present but not reviewed against a pacing benchmark. |
| Strategic use of silence | Silence produces no transcript content. A five-second pause after a closing question does not appear. | Rep scores low on "engagement" because word count drops at a critical moment, the opposite of the truth. | Audio or video review; note pauses longer than 3 seconds at key junctures. | No audio/video, or call was phone-only with no recording. |
| Filler word frequency and self-correction | Transcription software often cleans filler words. Even when retained, pattern recognition requires a baseline. | Rep appears fluent on transcript; actual delivery was hesitant and credibility-damaging. | Human audio review against a filler-word baseline set at calibration. | Audio not reviewed by a human scorer. |
Section 2: Physical and Visual Presence (Video Calls)
| Competency | Why transcript fails | Failure mode | Supplementary evidence | Hold trigger |
|---|---|---|---|---|
| Camera engagement and eye contact | Video content is not transcribed. | Rep scores high on "professionalism" despite being visibly distracted or off-camera during key buyer statements. | Video review using a defined visual engagement rubric (eye contact %, posture, background quality). | Call was audio-only, or video was not recorded, or video was not reviewed by a human. |
| Environment and setup professionalism | Transcript captures no visual context. | Rep receives a full quality score on a call where a chaotic background or poor lighting damaged buyer confidence. | Video review noting environment quality at call start. | Video not available or not reviewed. |
| Non-verbal affirmation (nodding, facial response) | Not transcribable. | Rep scores low on "active listening" because verbal acknowledgements were sparse, even though sustained non-verbal engagement was present throughout. | Video review noting visual acknowledgement frequency in discovery phase. | Video not available. |
Section 3: Pre-Call Competencies
| Competency | Why transcript fails | Failure mode | Failure mode | Hold trigger |
|---|---|---|---|---|
| Pre-call research quality | Research happens before the call. Its quality is inferred from question relevance, but a rep can ask relevant questions from a brief template without meaningful research. | Rep scores high on "discovery quality" because questions used the buyer's industry terminology: sourced from a two-minute LinkedIn skim, not genuine preparation. | CRM check: were account notes, prior meeting summaries, and known pain points logged before call start? Manager spot-check of pre-call prep notes. | No pre-call record exists and the score is being used for a consequential decision (promotion, PIP). |
| Agenda-setting intent | A stated agenda in the transcript looks identical whether the rep set it because they planned the call or because a script prompt told them to. | Rep scores high on "call structure" with no genuine command of where the call was going. | Pre-call plan reviewed by manager: did the rep document a call objective and success criteria before dialling? | No pre-call objective on file. |
Section 4: Infer-Only Competencies (Behaviour That Transcripts Mask)
These competencies are the most dangerous category. The transcript contains enough content to generate a confident score. That score is not measuring what it claims to measure.
| Competency | Why transcript fails | Failure mode | Supplementary evidence | Hold trigger |
|---|---|---|---|---|
| Genuine curiosity vs. scripted question delivery | A curious follow-up question and a template follow-up question are textually indistinguishable. | Rep scores high on "discovery depth" despite running a question list with no real interest in the answers. | Audio or video review: does the rep's follow-up respond to what was actually said, or does it pivot to the next question regardless? Manager calibration session using 2-3 example calls. | No human review of question sequencing against buyer response content. |
| Real-time listening vs. wait-to-talk | Transcript shows the rep did not interrupt. It cannot show whether the rep was processing what the buyer said or mentally rehearsing their next line. | Rep scores high on "active listening" despite missing a materially important buyer signal that appeared in the transcript and was never addressed. | Human reviewer flags: did the rep's response directly reflect content the buyer introduced, or did it revert to a prepared track? | No human review of response-to-signal mapping. |
| Emotional regulation under pressure | A rep who is rattled by a hard objection may still produce grammatically correct sentences. | Rep scores high on "objection handling" despite an audible change in composure that the buyer noticed and that damaged the call. | Audio review at objection moments: note pace change, filler word spike, or sentence fragmentation relative to baseline. | No audio review of objection sequences. |
| Buyer-led vs. rep-led pacing | Transcript shows who spoke and for how long. It does not show whether the rep was following the buyer's energy or overriding it. | Rep scores high on "consultative selling" despite consistently redirecting before the buyer had finished forming a thought. | Audio/video review noting interruption patterns and topic control transitions. Human scorer rates which party was setting conversational direction. | No audio/video or no human review of control transitions. |
| Authenticity of empathy statements | "That makes sense" and "I hear you" appear in transcript identically whether delivered with genuine acknowledgement or as a reflexive filler before the next pitch point. | Rep scores high on "emotional intelligence" on a call where the buyer explicitly noted feeling unheard in the closing phase. | Audio/video review: does a stated empathy response change the rep's subsequent direction, or does it function as a transitional filler? | No human review of empathy response behaviour. |
Manager Verdict Field
For each competency assessed in a review cycle, record one of three verdicts:
| Verdict | Meaning |
|---|---|
| SCORED | Supplementary evidence was available and reviewed. AI score stands with human confirmation. |
| ADJUSTED | AI score overridden or modified based on supplementary evidence. Reason documented. |
| SUSPENDED | Supplementary evidence was not available. Score removed from the rep's record for this competency this period. Gap noted. |
A suspended score is not a zero. It is a gap. Treat it as a coaching prompt, not a penalty.
Permanent Boundary Statement
Pin this to your QA process documentation:
This AI scoring tool assesses linguistic content: what was said, in what sequence, against defined criteria. It does not assess how it was said, what happened before the call, what the buyer observed visually, or whether the rep's internal process matched their surface behaviour. All competencies in the Blind Spot Register require supplementary human evidence before a score is treated as valid.
That boundary does not diminish the tool. It keeps the tool honest.