Home / Customers / Vulnerability ML
RiskQualityLeader AIAnalyst AI
FCA-regulated Insurer · 1m+ calls modelled · Multi-shot LLM classification

30% increase in vulnerability detection accuracy.

Manual QA samples 1% of calls. Traditional statistical models hit 80% accuracy on the 100% they score. MOJO-CX's multi-shot classification lifts that accuracy by 30 percentage points — on every call, every day.

The vulnerability detection arms race in regulated industries has three tiers: manual sampling (1% coverage, high human accuracy but unauditable at scale), statistical models (100% coverage, 80% accuracy), and now multi-shot LLM classification (100% coverage, 80% + 30pp = effectively 100% effective accuracy). The third tier is what closes the Consumer-Duty maths gap.

+30%
Accuracy lift vs statistical-only baseline
100%
Of calls scored — vs 1% manual sample
800k+
Calls modelled and audit-evidenced

FCA Consumer Duty applies to 100% of customer interactions. Traditional QA samples 1% of calls. The maths gap was the problem the entire regulated industry was carrying — and the FCA had made clear that sampling alone was no longer a defensible operating model for vulnerability identification.

Existing speech-analytics vulnerability models hit roughly 80% accuracy on the 100% of calls they scored. Good — but the 20% miss rate at scale meant hundreds of vulnerable-customer interactions per week going un-flagged across a typical book.

  • Multi-shot LLM classification built into MOJO-CX's Auto-QA pipeline. The same conversation is scored against multiple model prompts, and the consensus is used as the signal — dramatically reducing false negatives on the edge cases that matter most.
  • 100% call coverage retained. No regression to sampling, no trade-off between accuracy and scale.
  • Customer-bespoke vulnerability taxonomies. The Insurer's risk team defined what "vulnerable" meant in their specific book and product mix; the model trained against it.
  • Audit-evidenced for FCA Consumer-Duty inspection. Every flag carries the conversation timestamp, the language patterns that triggered it, the agent action, and the disclosure made (if any).
  • +30 percentage points in vulnerability detection accuracy vs statistical-only baseline
  • 100% call coverage maintained — 800k+ calls modelled in the deployment period
  • Audit-ready Consumer-Duty evidence trail running by default on every interaction
  • Hundreds of vulnerable interactions per week moved from "missed" to "flagged and actioned"
  • Risk team confidence in the operating model materially improved at the next FCA touchpoint
Manual sampling at 1% and statistical models at 80% accuracy were both leaving vulnerable customers unprotected. Multi-shot classification on 100% closed the gap. It's the difference between "we have a vulnerability policy" and "we can show you we acted on it, on this call from last Tuesday".
How Leader AI + Analyst AI drove the result

Leader AI scores the calls. Analyst AI clusters the misses and tunes the model. Continuous improvement, automated.

Leader AI runs the multi-shot vulnerability classification on every conversation. Analyst AI reads the model's lower-confidence cases, clusters them by topic and pattern, and surfaces the edge cases for risk-team review — generating the next iteration of the classifier without a data-scientist in the loop. The accuracy compounds with every model refresh.

Meet Leader AI and Analyst AI

See your own vulnerability coverage gap through MOJO-CX.

Two-week Analyst AI Auto Discovery on one queue. We benchmark your current detection coverage and show you the gap multi-shot classification closes.