A customer support operation almost measures itself. There's a visible queue, a wait time, a contact that comes in and goes out. In back office none of that exists. Work arrives by email, by file, by a system inbox, and it can sit there for three days without anyone noticing. That's why so many outsourced back office operations get reported with indicators borrowed from the contact center that explain nothing.
The problem isn't a shortage of data. It's that the first thing is almost always missing: a clear definition of what counts as a unit of work.
First: define the unit of work
Before talking about productivity you need to be able to finish this sentence: "one unit is ___". It sounds obvious and it's where most agreements break. A processed invoice is not the same as an invoice with ten lines and three attachments. A reconciliation of twenty movements is not one of two hundred.
A usable definition has three parts:
- What a unit is. A document, a record, a case, a complete request. Named as it appears in the system, not as people call it in conversation.
- When it starts and when it ends. It starts when the item lands in the team's inbox, not when someone opens it. It ends when it reaches a verifiable final state, not when the analyst marks it done.
- What complexities exist. Two or three types, not fifteen. Simple, with exception, with external validation. Each type is measured separately because they don't compare.
If that definition isn't written before go-live, the first monthly report will be an argument about the denominator. It's the same discipline that applies to any process before handing it over: it's in what to document before outsourcing a process.
The four indicators that hold the operation up
With the unit defined, four figures are enough to run a back office operation. Everything else is supporting detail.
- Output per effective hour. Units completed divided by the hours actually spent on the process, not by payroll hours. Reported separately by complexity type.
- Cycle time. How long a unit takes from arrival to closure. Read in percentiles, not averages: the average hides precisely the cases that generate complaints.
- First-time accuracy. Share of units completed correctly with no rework or return. Not measured on what the team reported; measured on an audited sample.
- Age of the open inventory. How many units are open and for how long. This is the indicator that flags trouble before it shows up in the other three.
If the monthly report shows output and not the age of the backlog, it's showing half the operation.
The backlog: the indicator ignored until it hurts
A team can close more units every month and still be falling behind, if intake grows faster. Output rises, the client is pleased, and three months later a batch of old cases surfaces that has already produced complaints, interest charges or an audit finding.
That's why the backlog is read by age and not as a single number. A hundred cases opened yesterday are not a hundred cases opened three weeks ago. The practical approach is to group by day ranges and watch how the oldest band moves month to month. If that band grows while output also grows, the process has a leak: cases nobody can close because a decision, a data point or an access is missing, and that rotate between inboxes without resolving.
Those cases deserve their own category and their own owner on the client side. They aren't provider productivity: they're dependencies, and confusing the two contaminates any evaluation.
Accuracy is not perceived quality
In back office, error has a direct and delayed cost. A badly captured field doesn't produce an immediate complaint; it produces a discrepancy at close, a bank rejection or a reprocess two weeks later, when nobody remembers who did it.
Measuring accuracy by asking the team to report its own errors doesn't work. You audit a sample, with written criteria and a reviewer who isn't the person who did the work. It's worth separating critical errors — those affecting money, a third party or a legal obligation — from cosmetic ones, because a single average mixes them and buries the one that matters. The full mechanics are in quality control by audited sampling.
What isn't worth measuring
Three indicators show up often in back office reports and add little:
- Logged hours. They say people were there, not that work came out. In a process with a defined unit, output already covers it.
- Average time per unit as an individual target. Pushing on it makes hard cases get left for later. The old backlog grows and nobody understands why.
- Total volume without a complexity split. A month with more simple cases looks like a more productive month, and it isn't.
The unit shapes the price
When the unit of work is well defined and has complexity types, the service can be priced per transaction instead of per FTE, and the incentive changes: the provider gains by improving the process, not by adding people. When the unit is fuzzy or volume swings hard, an FTE scheme stays more honest for both sides. The comparison between schemes is in BPO pricing models.
That same map of units and complexities is what lets you decide later which steps to automate and which to leave alone, on judgement rather than fashion. We cover it in AI automation in BPO.
Questions for the monthly review
- What is the written definition of a unit, and does it match what the system uses?
- Is output split by complexity type?
- How is the oldest band of the backlog moving compared with last month?
- What share of open cases is blocked by a dependency on our side?
- Does accuracy come from an audited sample or from the team's self-report?
- Are critical errors counted separately from cosmetic ones?
How smartBPO works it
On back office processes we agree the unit definition and the complexity types during transition design, and we validate them against the real states in the client's system before committing to any figure. We report output, cycle time by percentile, accuracy on an audited sample and the age of the open inventory, with cases blocked by client dependencies counted separately so the conversation is about the process and not about whose fault it is. When the oldest band of the backlog grows, we bring it to the table with the cause classified before it turns into an audit finding. And once the unit is stable, we propose revisiting the pricing scheme, because measuring properly is what allows charging for output instead of for presence. We don't promise zero errors: we propose that errors are caught by sampling, classified by impact and corrected with a date.