AI Roleplay Scoring Vendor Vetting Checklist (50 Questions)
Fifty structured questions to put to any AI scoring vendor before you trust its output with real learner data, organised across five risk areas specific to AI-scored sales assessment. Each question includes a plain-English note on what a weak or evasive answer looks like and why it matters.
How to use it
Send this checklist to the vendor before the final evaluation call and ask for written answers, not live demo responses. Score each answer using the weak/evasive signals noted below each question, then total the flags by risk area to identify where to probe hardest. Any section with three or more flags is a reason to pause the procurement or escalate to a contractual safeguard.
What's inside
- 50 numbered questions across five risk areas, each with a plain-English "weak answer looks like" note
- Section 1: Model training data provenance (10 questions)
- Section 2: Scoring rubric transparency (10 questions)
- Section 3: Calibration and drift protocols (10 questions)
- Section 4: Human-override and appeals process (10 questions)
- Section 5: Data privacy and portability (10 questions)
- A flag-counting scoring guide to interpret results by risk area
- A pre-call preparation note explaining how to run the vendor session
How to run the vendor session
Send the checklist in advance and request written answers. Verbal answers given in the heat of a demo are easy to walk back. Written answers are a matter of record and force the vendor to be precise. On the call, probe any answer that hedges, redirects to a roadmap, or answers a different question than the one you asked. Flag any response that would be a problem if it were still true twelve months after go-live.
Flagging: Mark each question Y (satisfactory) or F (flag). Count flags per section. Three or more flags in any single section is a procurement risk. Five or more flags in total across sections is a strong signal to pause until resolved in writing.
Section 1: Model Training Data Provenance
What was the model trained on, and is it fit for your context?
1. What datasets were used to train the scoring model? Weak answer: "Proprietary data" with no further detail. You need to know if the training corpus reflects your industry, language register and deal type.
2. Were real sales roleplay recordings used in training? If so, were participants informed and did they consent? Weak answer: "Standard terms of service covered it." That is not consent. It is a legal and ethical liability.
3. What is the approximate size of the labelled training dataset for the specific scoring tasks this model performs? Weak answer: Any refusal to give an order of magnitude. Small labelled sets produce brittle models.
4. How were training labels produced? Human annotators, expert raters, or synthetic? Weak answer: "A mix" with no breakdown. Synthetic labels trained on synthetic labels compound error.
5. What sales methodologies or frameworks were the human raters trained on when labelling? Weak answer: "General sales best practice." If the raters had no framework, the labels are subjective.
6. How does the model handle non-native English speakers and regional accents? Weak answer: "The model performs well across accents." Ask for accuracy disaggregated by accent group. Aggregated accuracy hides bias.
7. How does the model handle speech disfluencies such as filler words, restarts, and overlapping speech? Weak answer: "The ASR layer cleans that up." Disfluency handling is a scoring question, not just a transcription question. A rep who recovers well from a stumble should not be penalised.
8. Has the model been tested on your specific industry's vocabulary and deal context? Can you share those results? Weak answer: "It generalises well." Generalisation claims without domain-specific test results are marketing.
9. What languages is the model validated for, not just available in? Weak answer: A list of available languages. Availability and validation are different things. Ask for per-language accuracy data.
10. When was the training data last refreshed, and how often is the model retrained? Weak answer: "Continuous improvement." Ask for a concrete schedule. Models trained on 2021 data may not reflect current buying behaviour or objection patterns.
Section 2: Scoring Rubric Transparency
Do you understand what the model is actually measuring?
11. Can you provide the full scoring rubric in writing, criterion by criterion? Weak answer: Any partial view or demo-only access. If you cannot read the rubric in full, you cannot validate it.
12. For each criterion, how is the score computed? Is it a keyword match, a semantic similarity score, a classification output, or something else? Weak answer: "AI-powered assessment." That is a description of nothing.
13. What is the score range for each criterion, and what does each point on the scale represent behaviourally? Weak answer: A 1-5 scale with no behavioural anchors. Without anchors, the score is uninterpretable.
14. Can the rubric be customised to reflect our internal sales competency framework? Weak answer: "Yes, via the settings panel" without clarifying whether customisation affects the underlying model or just the label on the output.
15. What happens to historical scores if a rubric criterion is updated mid-cohort? Weak answer: "Scores update automatically." Automatic backdating destroys longitudinal comparability. Scores should be versioned.
16. Can you produce inter-rater agreement data showing how closely the AI scores align with a human expert benchmark on the same recordings? Weak answer: Any refusal or "we don't publish that." Inter-rater reliability is the single most important validity metric for any scoring system.
17. What is the model's known failure mode for each criterion? Where does it score unreliably? Weak answer: "It's highly accurate." Every model has failure modes. A vendor who cannot name them has not tested for them.
18. How does the model score silence or pauses? Is thinking time penalised? Weak answer: Vague reassurance. In complex sales, deliberate pausing is a skill. You need to know if the model treats it as one.
19. How does the model handle a rep who correctly identifies that a standard technique is inappropriate in context and adapts? Weak answer: "The model rewards correct technique." Adaptive judgement is a higher-order skill. If the rubric only rewards scripted behaviour, it will penalise your best reps.
20. Is the rubric the same across all customers, or is it independently calibrated per client? Weak answer: "It's standardised for consistency." Standardisation across industries with different sales motions is a validity problem, not a feature.
Section 3: Calibration and Drift Protocols
Will scores stay accurate and consistent over time?
21. How do you detect when the model's scoring has drifted from its baseline accuracy? Weak answer: "We monitor performance." Ask for the specific drift metric and the threshold that triggers a review.
22. How often is the model recalibrated against a held-out human-rated dataset? Weak answer: "As needed." That means never until something breaks.
23. What is your process when a recalibration changes score distributions? Are affected learner records updated, flagged, or left as-is? Weak answer: Silent updates. Any recalibration that retroactively changes scores without notification invalidates the learning record.
24. How do you handle scoring consistency across different simulation scenarios that test the same underlying competency? Weak answer: No evidence of cross-scenario consistency testing. If Scenario A and Scenario B both test discovery questioning, they should produce comparable score distributions.
25. Can you provide historical accuracy data over time, not just a point-in-time validation report? Weak answer: A single validation study. A single study tells you nothing about drift.
26. Who is responsible for monitoring model performance post-deployment? Is it an automated system, a named team, or the customer? Weak answer: "Our platform monitors everything automatically." Automation surfaces anomalies. Humans have to act on them. Ask who acts.
27. What is your process for managing a known scoring error after it has affected a cohort of learners? Weak answer: No documented process. This will happen. The absence of a process is a governance failure.
28. Do you run A/B comparisons between model versions before deploying updates to production scoring? Weak answer: "We test before release." Ask for the testing methodology and sign-off criteria.
29. Is scoring deterministic? If the same recording is scored twice, will it always produce the same result? Weak answer: Evasion or "approximately yes." Non-deterministic scoring cannot be defended in a learner appeal.
30. How do you account for the fact that ideal sales behaviour shifts as market conditions change? Weak answer: "The model is regularly updated." Ask who decides what counts as ideal behaviour now, and how that decision is documented.
Section 4: Human-Override and Appeals Process
What happens when the AI gets it wrong?
31. Can a human reviewer override an AI score on any criterion? Weak answer: Override is unavailable or restricted to certain tiers. If a human cannot override the AI, the AI is ungovernable.
32. Is there a formal appeals process for learners who dispute their score? Weak answer: "Managers can review." A review is not an appeal. An appeal has defined grounds, a timeframe, and a documented outcome.
33. What is the SLA for resolving a score dispute? Weak answer: No defined SLA. "We aim to respond promptly" is not a commitment.
34. When a human overrides an AI score, does that override feed back into model retraining? Weak answer: "Not currently." Overrides that are not used for training are wasted signal.
35. Who has the authority to trigger a full re-score of a cohort if a systematic error is identified? Weak answer: No clear answer or "contact support." You need a named escalation path.
36. How is a human override documented in the audit trail? Weak answer: No audit trail or override not separately logged. Overrides must be traceable for governance and legal reasons.
37. If a learner's score is changed on appeal, is the original score also retained in the record? Weak answer: "The record is updated." Both scores should be retained. Deleting the original score destroys the audit history.
38. Can calibration sessions be conducted where your internal L&D team scores recordings independently and compares results to the AI? Weak answer: Not supported in the product. Without calibration access, you cannot validate the AI against your own standard.
39. What training or guidance do you provide to human reviewers to ensure consistency in override decisions? Weak answer: "Managers use their judgement." Inconsistent human overrides introduced without a standard make the overall system less reliable, not more.
40. Is there a mechanism to flag a criterion as unreliable for a specific scenario type, so scores on that criterion are excluded or caveat-marked? Weak answer: No such feature. Edge cases exist in every deployment. You need a way to quarantine unreliable outputs.
Section 5: Data Privacy and Portability
Who owns the data, and can you leave?
41. Where is learner data stored, and in which jurisdiction? Weak answer: "Secure cloud infrastructure." Ask for the specific country and cloud provider.
42. Are audio recordings retained after scoring? If so, for how long and for what purpose? Weak answer: "Standard retention for quality assurance." Ask for the exact duration and whether recordings are used for model retraining.
43. Can we specify that our learner recordings are excluded from any model training or improvement activity? Weak answer: "That's our default" without a contractual commitment. Get it in writing.
44. Who has access to individual learner score data within your organisation? Weak answer: "Authorised staff." Ask for the specific roles and how access is governed and logged.
45. What is your process for handling a data subject access request from a learner? Weak answer: No documented process. Under UK GDPR, this is not optional.
46. What happens to our data if we terminate the contract? What is the deletion or export process and what is the timeline? Weak answer: "Data is available on request." Ask for the contractual commitment on timeline and format.
47. Can we export all learner scores and metadata in a portable, machine-readable format? Weak answer: PDF export only. If you cannot export structured data, you cannot migrate to a different system or run your own analysis.
48. Do you have ISO 27001 certification or equivalent? Can you share your most recent penetration test summary? Weak answer: "We take security seriously." Ask for the certification number and the test date.
49. In the event of a data breach, what is your notification commitment to us as the data controller? Weak answer: "We follow legal requirements." Ask for the specific hours-to-notification commitment.
50. If the vendor relationship ends or the company is acquired, what contractual protections exist to prevent our data being used by a third party? Weak answer: "Standard acquisition terms apply." This should be an explicit clause in your contract, not covered by a catch-all.
Flag Summary
| Section | Questions | Flags (F) | Risk Level |
|---|---|---|---|
| 1. Training Data Provenance | 1-10 | /10 | |
| 2. Rubric Transparency | 11-20 | /10 | |
| 3. Calibration and Drift | 21-30 | /10 | |
| 4. Human-Override and Appeals | 31-40 | /10 | |
| 5. Data Privacy and Portability | 41-50 | /10 | |
| Total | /50 |
3+ flags in any section: Raise as a contractual condition before sign-off. 5+ flags total: Pause procurement and request a written remediation plan. 8+ flags total: Do not proceed without legal review and independent validation.