Five Heads of Enablement on What Their AI Rollout Actually Broke
None of them blamed the model. All five blamed something they'd built years before the AI showed up.
I spoke to five heads of enablement about AI tools that stalled, got quietly scaled back, or died in a re-negotiation with finance before renewal. Names, companies, and identifying details are composited and changed — this isn't a name-and-shame piece, it's a pattern piece — but the substance of what each of them said is real, pulled from conversations I've had over the past year. None of them blamed the model. All five, without prompting, ended up describing something their organisation had left undefined years before any AI tool showed up.
1. The 600-rep SaaS org that automated a checklist nobody agreed on
"We rolled out AI call scoring across four regions in six weeks. Fast. Everyone was thrilled with the speed. Then the EMEA and AMER leads got in a room and realised their reps were being scored against different definitions of 'good discovery' — EMEA's came from a Challenger-trained manager, AMER's came from whoever wrote the onboarding deck in 2021. The AI didn't cause that split. It just scored both sides fast enough that the split became impossible to ignore."
This is the most common version of the failure: the tool didn't break anything, it exposed something that was already broken and had been getting away with it because nobody was measuring consistently enough to notice.
2. The scale-up that couldn't tell the model "no"
"Our AI flagged 'strong discovery' on calls our best manager would have torn apart in a coaching session. We went back and forth with the vendor for two months. Eventually we figured out why: we'd never actually written down what discovery means at each level for us. The vendor's default rubric was fine, generic, and completely misaligned with what we actually reward. We were arguing with a mirror."
Without an org-specific standard to hand the vendor, the AI defaults to whatever definition ships in the box. That's not a vendor failure. It's the customer skipping the one step — define your own competency, don't inherit one — that makes any of this trustworthy.
3. The enterprise team running two AI tools that disagreed with each other
"Sales had one AI tool scoring calls. CS had a different one. Both vendors used the phrase 'objection handling' in their output. Neither could tell you what the other one meant by it. We had a rep transfer from SDR to AE and his 'objection handling' score dropped fifteen points overnight — not because he got worse, because the two tools weren't measuring the same thing and calling it the same name."
This one is a naming problem wearing a data problem's clothes. Two rubrics, same label, no shared definition, and a rep's history effectively wiped every time they cross a tool boundary.
4. The org that never built a baseline before switching the AI on
"We turned the tool on and three weeks later a VP asked 'is this actually working, are we better than we were.' Nobody could answer, because nobody had scored anything before the tool existed. We had no before. So every conversation about ROI became a debate about vibes, because the only numbers we had were the ones the new tool produced, and you can't prove improvement against a baseline of zero."
This is the quietest killer because it doesn't show up as a complaint about the tool. It shows up eighteen months later as "we can't justify the renewal," because nobody can show the line moving, and a line needs two points, not one.
5. The team where managers didn't trust a score they couldn't explain
"Our managers stopped using the coaching prompts within two months. Not because they were wrong — because managers couldn't explain to a rep why the AI landed on a 3 instead of a 4, so they stopped bringing it into 1:1s. A manager who can't defend a number in a room won't use that number, full stop, no matter how good the model behind it is."
Adoption didn't fail because reps rejected it. It failed because the managers — the people actually meant to run the coaching loop — had no rubric language to translate a score into a conversation.
What's actually the same story five times
| What they said broke | What they actually meant |
|---|---|
| "The AI's discovery score is inconsistent" | No shared definition of discovery existed before the AI, region to region |
| "The vendor's rubric doesn't match how we coach" | Nobody wrote down the org's own standard before buying |
| "Two tools disagree" | Same competency name, two undocumented definitions |
| "We can't prove ROI" | No baseline was scored before rollout — nothing to compare against |
| "Managers stopped trusting the scores" | No rubric language existed for managers to translate a number into a coaching conversation |
Every row starts as an AI complaint and ends as a measurement-infrastructure gap that predates the AI by years. The tool didn't create the gap. It just runs at a speed and scale that makes an undefined competency impossible to paper over the way a once-a-quarter manual QBR review always managed to.
The fix none of the five had done first
All five, independently, described the same retrofit: stop the rollout, get sales leadership in a room, write down the actual competencies the org coaches to, agree what each level looks like, score a sample of historical calls by hand to set the baseline, then turn the AI back on against that. Every one of them said some version of "we should have done that before we bought anything." A Competency-Based Onboarding Curriculum Framework is a reasonable place to start if you're building that definition from scratch rather than retrofitting it after a failed rollout — it forces the "what does level 2 look like" conversation before you've spent a renewal cycle finding out the hard way.
If the immediate problem is that managers have stopped using the coaching data, the fix usually isn't more data, it's a fixed rhythm for using what already exists — a Manager Coaching Cadence Checklist gets that loop running again without waiting for the rubric project to finish.
The pattern underneath the pattern
An AI tool is a measurement instrument. Point a very precise instrument at an undefined thing and you get a very precise-looking number that means nothing, delivered with total confidence. That's worse than no measurement at all, because it looks finished. None of the five people I spoke to had a model problem. All five had spent years running enablement on shared vocabulary that felt agreed but had never actually been written down, and the AI was just the first tool literal enough to notice.