AI Scoring Criteria Constructability Audit: Can the Model Actually See It?
A criterion-by-criterion rubric that pressure-tests every item on your AI roleplay scorecard against four mechanical dimensions before a single live assessment runs. The output is a written record of which criteria are safe to automate, which need a human overlay, and which should be cut from the AI layer entirely.
How to use it
Pull your current AI roleplay scorecard and run each criterion through the four dimensions in order, scoring 1-3 on each and recording the named failure mode where it applies. Total the scores, apply the pass/flag/remove decision rule, and use the completed audit as your sign-off document before the model goes live with real reps.
What's inside
- Four scored dimensions: signal availability, signal ambiguity, context dependency, and inference distance
- A 1-3 scale for each dimension with precise descriptors and named failure modes at each level
- A composite scoring table that converts dimension scores into a pass, flag, or remove recommendation
- A worked example row showing how a real criterion looks when audited end to end
- A named failure mode glossary covering the six most common constructability errors
- A pre-rollout sign-off block for the scorecard owner to formalise the audit record
- A "when NOT to run this audit" note for contexts where the tool does not apply
- Decision guidance on what to do with flagged criteria rather than simply cutting them
What This Is For
Before your AI model grades a single rep on a roleplay, someone needs to answer one question for every criterion on the scorecard: can the model actually see the evidence it needs to score this?
Most rubric owners skip this step. They build competency lists, write descriptors, and assume the transcript will contain what they need. Often it does not. The result is scores that look authoritative but are partly invented, partly noise, and occasionally inverted, where a confident, fluent rep scores higher than a careful, accurate one because the model is pattern-matching on surface features it can observe rather than substance it cannot.
This audit forces that question into the open, criterion by criterion, before rollout.
When NOT to Use This Audit
- Your AI tool scores audio or video directly, not transcripts. The dimensions below are calibrated for text-based transcript scoring.
- You are auditing a human observation rubric with no AI component.
- Your "AI scoring" is a keyword match or rule-based system. Apply a simpler signal-presence check instead.
The Four Dimensions
Each criterion on your scorecard gets scored 1-3 on each dimension. Lower is better on all four.
Dimension 1: Signal Availability
Is the evidence the model needs actually present in the transcript?
| Score | Description | Named Failure Mode |
|---|---|---|
| 1 | The criterion produces clear, direct textual evidence every time it is performed. A model can find it without inference. | None. Proceed. |
| 2 | The evidence is present in most transcripts but may be implicit, distributed across turns, or expressed in varied forms that require paraphrase recognition. | Scattered Signal. The model may miss evidence that a human reader would accumulate across the conversation. |
| 3 | The evidence is absent from the transcript by design, occurs in paralanguage (tone, pause, pace), or lives in the rep's internal reasoning rather than their words. | Invisible Criterion. The model is scoring a proxy at best, a fiction at worst. |
Dimension 2: Signal Ambiguity
Could the same words, in different contexts, indicate opposite quality levels?
| Score | Description | Named Failure Mode |
|---|---|---|
| 1 | The language markers for high and low performance are reliably distinct. A skilled and an unskilled rep would not produce the same words. | None. Proceed. |
| 2 | Some phrases appear in both strong and weak performances depending on delivery or intent. A model scoring on words alone will sometimes get this backwards. | Surface Mimic. Reps who sound like they are doing the right thing get credit. Reps who do the right thing quietly do not. |
| 3 | High-quality performance on this criterion routinely produces identical or similar text to low-quality performance. The discriminating signal is not linguistic. | Polarity Blind. Scores will be random with respect to actual competency. |
Dimension 3: Context Dependency
Does a correct response depend on information the model was never given?
| Score | Description | Named Failure Mode |
|---|---|---|
| 1 | The correct behaviour is the same regardless of prospect type, deal stage, product variant, or prior conversation history. The model needs no external context to score accurately. | None. Proceed. |
| 2 | Correct behaviour varies by context, but the relevant context is contained within the transcript itself and a well-prompted model can read it. | In-Transcript Context. Manageable, but the model prompt must explicitly instruct contextual reading. Flag for prompt review. |
| 3 | Correct behaviour depends on information outside the transcript: the rep's territory, the prospect's prior objections from a different call, internal pricing rules, or product knowledge the model was not trained on. | Context Blindness. The model will apply a generic standard to a situation that demands a specific one. Scores will be systematically wrong for non-standard cases. |
Dimension 4: Inference Distance
How many logical steps separate the raw text from the conclusion the model must reach?
| Score | Description | Named Failure Mode |
|---|---|---|
| 1 | One step. The model reads the text and the score is directly apparent. "Did the rep state a next step with a date?" requires a single look. | None. Proceed. |
| 2 | Two to three steps. The model must combine observations, apply a rule, and then score. Manageable, but compounding error risk increases with each step. | Inference Stack. Each step that is wrong compounds. A model that misreads step one will produce a confident, wrong answer at step three. |
| 3 | Four or more steps, or the inference requires theory of mind (what was the rep trying to do? what did the prospect actually mean?). | Conjecture Scoring. The model is not scoring the rep. It is generating a plausible story about the rep. |
Composite Score and Recommendation
Add the four dimension scores. The maximum is 12. The minimum is 4.
| Total | Recommendation | What To Do |
|---|---|---|
| 4-6 | PASS | Safe to include in the AI scoring layer as written. |
| 7-9 | FLAG | Include with modifications. See the Flag Actions table below. |
| 10-12 | REMOVE | Cut from the AI scoring layer. Move to human review or drop from the rubric. |
Flag Actions by Dimension
When a criterion scores 2 or 3 on a specific dimension, use the corresponding action before marking it as conditionally safe.
| High-scoring dimension | Minimum action before conditional pass |
|---|---|
| Signal Availability (score 3) | Remove from AI layer. No prompt fix resolves absent evidence. |
| Signal Ambiguity (score 3) | Remove from AI layer unless you can provide labelled examples that cover the full range of ambiguous cases. |
| Context Dependency (score 3) | Remove from AI layer unless the missing context can be injected into the scoring prompt systematically for every call. |
| Inference Distance (score 3) | Remove from AI layer. Decompose the criterion into simpler observable sub-criteria if the competency matters. |
| Any dimension at score 2 | Revise the criterion descriptor to reduce ambiguity, then re-score before rollout. |
Worked Example
Criterion: "Rep demonstrated genuine curiosity about the prospect's situation."
| Dimension | Score | Reason |
|---|---|---|
| Signal Availability | 3 | Curiosity is an internal state. Its presence is not reliably signalled by question volume alone. A rep can ask five questions mechanically. |
| Signal Ambiguity | 3 | The same follow-up question can be curious or scripted. The words do not distinguish them. |
| Context Dependency | 2 | What counts as sufficient curiosity varies by deal stage and buyer type, though stage is partially readable in transcript. |
| Inference Distance | 3 | Concluding genuine curiosity requires theory of mind. |
| Total | 11 | REMOVE |
Action: Replace with two observable sub-criteria: "Rep asked at least one follow-up question directly based on the prospect's prior answer" (scoreable) and "Rep restated the prospect's stated priority before proposing next steps" (scoreable). Curiosity as a holistic construct moves to manager observation.
Failure Mode Glossary
| Failure Mode | Plain description |
|---|---|
| Invisible Criterion | The thing being scored is not in the transcript. |
| Scattered Signal | The evidence exists but is distributed in a way models routinely miss. |
| Surface Mimic | Confident-sounding language earns credit that skilled-but-quiet language does not. |
| Polarity Blind | High and low performers produce indistinguishable text. Scores are inverted as often as correct. |
| Context Blindness | The model applies a universal standard to a situation that requires a specific one. |
| Conjecture Scoring | The model narrates a plausible interpretation and presents it as an observation. |
Audit Sign-Off Block
Complete this for every scorecard before the AI layer goes live.
`` Scorecard name: Version: Total criteria audited: PASS: FLAG (conditional pass after revision): REMOVE: Criteria removed from AI layer and reason: Criteria moved to human overlay and reason: Revised criteria re-scored before conditional pass: Y / N Audit completed by: Date: Reviewed by (second sign-off if criteria affect performance rating): Date: ``
One Rule to Keep
If you cannot write down, in one sentence, exactly what a model would need to read in a transcript to give this criterion a high score, the criterion is not ready for AI scoring. Write that sentence first. If you cannot, the answer is human review.