Back to Blog

QA Operations

Why Sampling 2 Percent of Calls Isn't QA

Tobias Lindqvist 7 min read
Why Sampling 2 Percent of Calls Isn't QA

Ask most QA managers how many calls their team reviews each month, and you'll get an honest answer: somewhere between 2 and 5 percent. At a contact center processing 20,000 calls a month, that means at most 1,000 calls get any human attention. The other 19,000 go unreviewed. Management calls it "statistically representative." Operations leaders who've watched agent performance drift know it isn't.

What "2%" Actually Means in Practice

A QA lead who reviews 10 calls per agent per month is considered thorough. At that rate, an agent handling 400 calls a month has 390 calls that nobody examined. If that agent started adding unnecessary hold time three weeks ago, or stopped confirming customer contact details before closing, the first QA review that catches it might come 40 to 60 days after the behavior started. By then, the habit is set. Coaching feels like punishment.

The sample selection problem compounds this. QA teams don't pull calls at random -- they pull flagged calls (already escalated), short calls (easier to review), and calls from the same small pool of agents they've always reviewed. The distribution of which calls get human attention is biased from the start. High-performing agents get fewer reviews because nothing flags them. Mid-range agents who need the most development fall below the noise floor.

The Three Things That Hide in the Other 98%

1. Agent Behavior Drift

Individual agent performance doesn't fall off a cliff -- it drifts. An agent who scored consistently well in their first six months starts shortening call closings, stops using the confirmation script, begins transferring calls that should have been resolved. Each deviation is small. Sampled at 2%, none of it trips a threshold. By the time it shows up in customer satisfaction data or escalation rates, the pattern has been running for months.

Full-coverage scoring catches drift within days. Not because AI is a better judge than a trained QA analyst -- but because it scores every call every day, which a QA analyst cannot.

2. Compliance Gaps

For contact centers in regulated industries -- financial services, healthcare, insurance -- compliance exposure lives in the unreviewed calls. An agent who occasionally skips a required disclosure isn't going to surface in a 2% sample. At 100% coverage, that pattern shows up immediately: a specific agent, a specific disclosure, a score of 0 on that criterion every third call. That's a coaching flag, not a regulatory finding.

The math is straightforward. If an agent fails a required disclosure on 8% of calls, and you're sampling 2% of that agent's calls, you have roughly a 14% chance of catching even one violation in a given month. Your QA program is giving regulatory risk a place to hide.

3. Churn Signals

The language patterns that predict customer churn appear, on average, around three minutes into a support call. "I've called about this three times." "I'm actually considering switching." "This is the last time I'm going to deal with this." These phrases are not in the 2% your team reviewed. They're in the transcript pile nobody has time for.

Churn signals don't require nuanced judgment to detect -- they require coverage. A model scoring every transcript for competitor mentions, cancellation language, and escalation demand patterns can surface these signals within minutes of call completion. The information was always there. It just needed to be read.

What Full Coverage Changes -- and What It Doesn't

Automated 100% scoring is not a replacement for human QA judgment. A rubric is only as good as the criteria it encodes, and those criteria need to be written by people who understand the business and the customer relationship. The AI applies the rubric consistently and at scale -- it does not write the rubric.

What changes with full coverage is the role of human reviewers. Instead of spending 80% of QA capacity transcribing and rating calls, QA leads spend that time on the hard calls: the borderline scores that need a second look, the coaching conversations with agents in the bottom quartile, the rubric refinements that keep the scoring model calibrated. The AI processes 20,000 calls. The QA team acts on what the AI found.

Starting the Transition

The most common objection to full-coverage scoring is that ops teams don't have time to act on 20,000 data points. This misunderstands the workflow. Full coverage doesn't mean reviewing 20,000 calls -- it means having reliable signal on 20,000 calls so you know which 40 to look at.

A coaching queue built from 100% scoring surfaces the five agents in the bottom quartile this week, ranked by which criterion they failed most consistently. That's an actionable list. A QA lead working from a 2% sample produces a list shaped by sampling bias and scheduling constraints, not by actual agent need.

The 2% sample was a practical constraint, not a quality standard. The constraint no longer exists. What remains is a choice about whether to update the program.

See It In Action

Voicemarrow grades every call, not the 2% you had time to pull

Connect your call platform, configure your rubric, and let the platform surface what your current program misses. Takes about 30 minutes to set up a pilot.

Request a Demo

Related reading: