A.I🇺🇸, All Articles

How One Software Company Stopped LLM Hallucinations in Sales Training

The VP of Sales had a problem that shouldn’t have existed.

His reps were acing AI roleplay training. Empathy scores were high. Objection handling was textbook. The platform’s feedback was consistently positive — strong rapport, excellent pacing, clear communication.

And yet, lost-deal reviews kept surfacing the same failure mode. A rep would confidently state a product capability that didn’t exist. An “unlimited” tier that was actually capped. An integration that was still in beta. A pricing condition that had changed three months ago.

The reps sounded excellent. The product information was wrong.

The training platform had no way of knowing — because it was checking for plausibility, not truth.

This is the story of how one B2B software company — we’ll call them Vertex Cloud — fixed that gap.

The problem: plausible-sounding wrong answers

Vertex Cloud sells a B2B SaaS platform to mid-market and enterprise customers. Their product is complex: multiple tiers, dozens of features, pricing that varies by seat count and contract length, and integrations with third-party tools that change quarterly.

Their AI training platform could generate practice conversations from uploaded documents in seconds. Reps practiced with AI customers, got scored on empathy and objection handling, and could repeat scenarios as many times as they wanted.

But the platform had a fundamental blind spot: it only checked whether a statement sounded reasonable.

When a rep practiced a call and confidently stated that the Enterprise plan included “unlimited API calls” — when the pricing guide clearly said 10 million per month — the platform gave full marks. The content of the claim was never evaluated.

The system was grading confidence, not correctness. And because the platform rewarded fluency, reps were being trained to sound more certain about information that wasn’t always true.

The cost of plausible wrong answers

Vertex Cloud tracked the impact for six months before the intervention:

  • 14% of lost deals involved buyers who discovered inaccurate product information during the evaluation
  • 9% of support escalations came from customers who were sold capabilities that didn’t exist
  • New reps took an average of 10 weeks to reach certification — and many still failed internal product knowledge assessments after certifying

The VP of Sales put it simply:

“Our reps were getting better at sounding confident. But they weren’t getting better at being correct. The platform couldn’t tell the difference, and neither could we.”

The root cause was structural. The training platform’s underlying LLM had never read the company’s actual pricing guide. It only knew what “unlimited API calls” sounded like based on patterns learned from other companies’ products. When a rep said it, the model registered: that sounds plausible — and moved on.

Plausible isn’t the same as true. But to a general-purpose LLM, they look identical.

The intervention: replacing plausibility with verification

Vertex Cloud replaced their plausibility-checking platform with EOS (akaeos.com) — an approach that doesn’t rely on a model’s general knowledge of “how software products usually work.”

EOS did three things differently:

  1. Built a company-specific knowledge base. Instead of relying on the model’s general training data, EOS ingested Vertex Cloud’s actual product documentation — pricing guides, feature matrices, integration lists, and compliance language. This became the only source of truth for verification.
  2. Checked every claim against that source. Every factual claim from a practice call — every feature, price, limit, and integration — was checked against the company’s own documents, not against what the model thought was true.
  3. Returned three verdicts, not one score. Supported (matched the documentation, with a citation), Contradicted (directly conflicted with it), or Insufficient (too vague, or missing a required qualifier).

Every contradicted or insufficient claim triggered a personalized, citation-backed quiz for that specific rep.

The shift was fundamental. The old platform asked: “Does this sound plausible?” EOS asked: “Is this actually true — and here’s the document that proves it?”

The 90-day results

Vertex Cloud ran the pilot with 60 reps over 90 days.

Contradicted claims per session dropped 71% (2.4 → 0.7) — the headline result. Lost deals attributed to inaccurate product information fell a comparable 70% (14% → 4.2%) over the same period.

The more important shift was in what the training program was actually measuring. Before, the dashboard showed hours of roleplay completed, empathy scores, and objection-handling ratings — it measured that training happened. After, using EOS’s manager dashboard, it showed which product claims were most frequently misunderstood, which reps had the most gaps, and a certification trail for every rep on every product tier.

One rep put it this way:

I thought I knew our product cold. The first week on this system, I got flagged on three things I’d been saying wrong for months. The old system never caught it because it never checked. It just told me I sounded good.

Why this worked

The old platform’s LLM had never read the company’s pricing guide. It only knew what “unlimited API calls” sounded like based on patterns from thousands of other companies. When a rep said it, the model thought: that’s a common thing software companies say — and gave a passing score.

That’s the core problem with general-purpose LLMs in sales training: they check for plausibility, not truth. A statement can sound perfectly reasonable while being completely wrong. “Our warranty lasts 48 months.” “We integrate directly with Platform X.” “This feature is included in the Professional plan.” Every one of those statements is believable. Every one of them could also be false — and a general-purpose model has no way of knowing which.

What Vertex Cloud built instead was a system that starts from its own documents, not from a model’s memory of how similar products work. As their pricing and features change, the knowledge base gets re-synced — so a verification made in July still reflects what’s actually true in October.

The takeaway for sales leaders

If your AI training platform is using a general-purpose LLM to evaluate your reps’ product knowledge, you have a problem you might not see until a deal dies.

That model has never read your pricing guide. It doesn’t know your latest feature matrix. It has no idea what your legal team approved last week. It only knows what sounds plausible — and in sales, plausible wrong answers can sound exactly like correct ones.

The fix isn’t a smarter language model. It’s better evidence: a fact-grounded system that starts from what your company actually publishes and checks every claim against that source, not against a model’s general assumptions.

Leave a Reply

Your email address will not be published. Required fields are marked *