The head of sales enablement had a dashboard that showed her exactly how every rep performed.
Eighty-eights. Seventy-twos. Ninety-fours. The numbers were clean, comparable, and easy to report upward. She could see trends, track improvement, and identify who was struggling.
But when she sat down with a manager to coach a specific rep who had scored an 88, she ran into a problem the dashboard couldn’t solve. The manager didn’t know why the rep had lost 12 points. Was it hesitation? Pacing? A product fact stated incorrectly? A missing compliance disclosure?
The number alone couldn’t tell them. And without the “why,” they couldn’t coach the rep effectively.
This is the story of how one legal services firm discovered that a single score wasn’t enough — and what they built instead.
The setup: two reps, one score
The company provides legal support and document services to law firms and corporate legal departments. Their sales team sells complex service packages — litigation support, e-discovery, managed review, and compliance consulting — to sophisticated buyers who expect precision.
Their training platform scored every practice session on a 0–100 scale. The score was a composite: tone, pacing, objection handling, product knowledge, and compliance language all collapsed into one number.
The problem became visible when two reps finished practice sessions on the same day. Both scored an 88.
Rep A had lost points on pacing and hesitation. She knew the products cold, but she sounded uncertain. Her coaching need was clear: confidence, delivery, and speed.
Rep B had delivered a smooth, confident performance. His pacing was excellent. His objection handling was textbook. But he had stated a service capability that didn’t hold up — and the platform hadn’t flagged it because it wasn’t checking content.
Two reps. One score. Completely different coaching needs.
“The number told me they were the same. They weren’t. One needed help with delivery. The other needed to learn what our services actually include. Those are not the same conversation, and the score gave me no way to tell them apart.”
The cost of collapsed scores
The firm had been using the score-based platform for 18 months. In that time, the head of sales enablement had noticed patterns she couldn’t explain:
- Reps with strong scores were still making factual errors on live calls
- Managers struggled to give specific, actionable coaching — they knew a rep had scored low, but not what to fix
- The same mistakes were showing up across multiple reps, but the score data couldn’t reveal the pattern
The underlying problem was structural. The score collapsed everything — tone, pacing, product accuracy, compliance — into one number, and once it was collapsed, you couldn’t get the detail back out.
When a rep lost points, the manager didn’t know whether they lost them on tone or on a product fact. And without that distinction, coaching was generic: “improve product knowledge” instead of “here’s the specific claim you got wrong, and here’s the citation.”
The intervention: verdicts instead of just a score
The firm replaced their score-only platform with EOS (akaeos.com), which adds a second layer underneath the score: claim-level verdicts.
Instead of asking only “how well did this session go,” EOS asks, for each specific claim a rep made: was that claim actually accurate? Every claim receives one of three verdicts:
- Supported — the claim matches current product and policy documentation
- Contradicted — the claim conflicts directly with that documentation
- Insufficient — the claim is too vague to verify as stated, or it’s missing a qualifier, exception, or required disclosure that changes whether it’s fully accurate
The Insufficient verdict turned out to be where most of the real coaching value lived — and it’s also the category a plain score can’t represent at all.
Why “Insufficient” changed everything
Consider a rep who says a service offers “full document review coverage.” The claim is directionally true — the service does include document review. But if it’s only true for certain case types, and the rep didn’t mention the limitation, the claim isn’t fully supported either. It’s something in between: right on the main point, incomplete on the detail that actually matters to the customer.
A binary right/wrong system has no good place to put that claim. It either gets marked Supported — which lets a real gap through uncorrected — or Contradicted — which overstates the problem and makes the rep defensive about something they mostly got right. Flagging it as Insufficient says exactly what happened: the core claim held up, and here’s the specific condition that got dropped.
That distinction mattered even more for compliance language specifically. On several calls, EOS flagged claims as Insufficient not because the rep said anything false, but because a required disclosure — a standard scope-of-engagement caveat their compliance team requires on certain service lines — was never mentioned at all. That’s a different kind of gap than a wrong fact, and in a legal services firm, it’s often the more expensive one: a scope dispute six months into an engagement traces back to a caveat that was never said out loud.
That’s a coaching note a manager can actually act on. “Always mention the case-type limitation” is a concrete instruction. “Improve product knowledge” is not.
What they found in the first month
The firm ran a pilot with 30 reps.

Three patterns emerged:
- Scores didn’t tell the full story. Several reps with high overall scores had a meaningful number of Insufficient claims — directionally right, but missing critical qualifiers. The old score-based system had been giving them passing grades without revealing the gaps.
- The most common errors were in the middle category. Reps rarely made outright false (Contradicted) claims. They frequently made claims that were mostly right but Insufficient — missing a key exception, limitation, or disclosure. These were the errors most likely to surface later as client dissatisfaction or scope disputes.
- Managers could finally coach specifically. Instead of saying “work on your product knowledge,” managers could say “when you discuss e-discovery coverage, always mention the data-volume limit.” Using EOS’s manager dashboard, the verdicts gave them the specificity they needed.
“Before, I could see that a rep had room to improve. I couldn’t tell anyone what to improve. Now I can. ‘This claim came back Insufficient — here’s the condition you left out.’ That’s a coaching conversation that actually changes behavior.”
- Is your training platform giving you a score — or a diagnosis? Our 12-page report, Beyond Roleplay: The Rise of the Sales Knowledge Engine, covers claim extraction, the fact-grounded verification pipeline, and the verdict system that replaces vague scores with actionable coaching. [Download the whitepaper →]
Why verdicts beat a single score
A score tells you how a session went. It doesn’t tell you why — and “why” is the only part that’s actually coachable.
Two reps can land on the same score for completely different reasons. One lost points on delivery. The other lost points on product accuracy. Those are not the same coaching conversation, and a score alone gives a manager no way to tell them apart.
Verdicts add the layer a score is missing. They don’t replace the score — they sit underneath it, providing the detail that makes coaching possible. For a legal services firm where precision matters and partial truths can create scope disputes, that distinction was everything.
The takeaway for sales leaders
If your training platform gives you a single number and nothing else, you have a problem you might not see until a coaching conversation goes nowhere.
A score tells you that something needs improvement. It doesn’t tell you what. Verdicts — Supported, Contradicted, Insufficient — tell you exactly what a rep got right, what they got mostly right but incomplete, and what they got wrong outright. They turn “improve product knowledge” into “here’s the specific claim you missed, and here’s the citation.”
That’s the difference between generic feedback and coaching that actually changes behavior.
- Stop guessing what your reps got wrong. EOS replaces single scores with specific, actionable verdicts — Supported, Contradicted, or Insufficient — for every claim a rep makes. Start free with up to 5 seats at akaeos.com, or [download the full whitepaper].
Leave a Reply