When someone asks for "quality control" on an outsourced process, they almost always picture the same thing: review everything. Every case, every record, every call. It sounds rigorous. In practice it's impossible to sustain and, worse, it improves nothing. A serious quality process doesn't review everything: it reviews a well-built sample and does something with what it finds.

The gap between auditing a sample and "reviewing when there's time" is what separates quality control from a feeling of quality.

Reviewing everything is not quality control

100% inspection has three problems. The first is cost: reviewing every transaction costs almost as much as producing it, so it doubles the price of the process. The second is fatigue: whoever reviews thousands of identical cases stops seeing errors after two hours. The third is the worst: reviewing everything gives you an error count, but it doesn't tell you why errors happen or how to avoid them. It becomes a tally, not a tool.

Sampling done well flips that logic. It reviews fewer cases, but with real attention, and uses each error to understand the process. It isn't out to catch culprits: it's out to find patterns.

What an audited sample is

An audited sample is a subset of transactions chosen to represent the whole, reviewed against a fixed rubric by someone who didn't do the work. Three words do the heavy lifting there: represent, rubric and independence.

Represent means the sample looks like the universe. If most of the volume is simple cases and a fraction is complex, the sample has to keep that proportion, or at least review each group separately. A sample drawn only from the easy cases tells you nothing useful.

Independence means the person auditing isn't the one who did the work or their direct manager. Nobody audits their own team's output harshly, and that leniency isn't bad faith: it's human. That's why quality is measured better from a separate role.

How the sample is built

Sample size is not a fixed percentage. It's not "review 5%" out of habit. It depends on total volume, on how variable the process is, and on how much confidence you need in the result. Higher volume lowers the percentage you need; higher variability raises it. Setting a percentage without looking at those variables is as arbitrary as not sampling at all.

On that base there are two practical decisions:

  • Stratify by transaction type. A back office is rarely homogeneous: there are new entries, corrections, exception cases. It's better to sample each stratum separately, because errors behave differently in each one.
  • Choose at random within each stratum. If the auditor picks what to review, they end up reviewing the usual. Random selection is what keeps the sample from biasing itself.

The rubric: what gets scored and how heavily

Without a written rubric, two auditors score the same case differently and the number loses meaning. The rubric defines what counts as an error and how much each one weighs. The most important distinction is between critical and non-critical errors.

A critical error breaks the result for the end customer or violates a rule: a mis-keyed field that triggers a wrong charge, a case closed without the mandatory validation. A non-critical error affects form, not substance: a differently formatted field, a poorly worded note. Counting them together, at the same weight, hides what matters. One case can have three formatting errors and still be well resolved; another, a single critical error and be wrong.

A quality score that treats every error the same rewards neatness over correctness.

Do something with the error, don't just count it

This is where most quality programs fall short. They measure, report a percentage and take it to the monthly meeting. The number goes up or down and nobody knows why. A control that works does three more things.

First, it calibrates: every so often several auditors review the same case and compare scores, so that "done well" means the same thing to everyone. Second, it looks for root cause: if the same error repeats, it's almost never the agent being careless; it's usually an ambiguous instruction, a system that allows the error, or a badly documented step. Third, it closes the loop: the finding comes back as a change to the procedure or as concrete feedback, not as a generic telling-off.

That loop is also what connects quality to the SLA you signed: the quality score only means something if it was defined in the agreement with the same clarity as response time.

The sampling mistakes that invalidate the control

Three show up again and again. Sampling only recent work, because it's what's at hand, and losing the pattern that was visible over the full month. Announcing in advance which cases will be audited, so the team prepares exactly those. And changing the rubric mid-period, so this month's number can't be compared with the last. All three have the same effect: they produce a figure that looks like control but isn't.

Where smartBPO fits

In the back-office processes we run, quality isn't a review we do when there's time to spare: it's a role separate from the one doing the work, with a rubric agreed with the client and a sample built by transaction type, not by whatever is in plain sight. We report errors separating critical from formatting, and we take each pattern to its cause —instruction, system or documentation— instead of leaving it as a percentage on a dashboard. We don't promise zero errors; no process has that. What we aim for is that errors get seen in time, understood, and don't repeat for the same reason two months running.