Almost every outsourcing contract starts the same way: an annex with twenty metrics, a dashboard nobody opens after month three, and a monthly meeting where somebody reviews the average of something. A year in, when someone asks whether the operation is working, nobody can answer with data.
The problem is not measuring too little. It is measuring things that change no decision.
The average trap
The average is the most comfortable metric and the one that hides the most. An average first response time of four minutes can mean every user waited four minutes, or that 80% waited thirty seconds and 20% waited twenty. The second situation is the one that generates complaints, and the average makes it invisible.
That is why percentiles are better. The p90 —the time that 90% of cases did not exceed— says far more about the real experience than the mean. If the average is four minutes and the p90 is eighteen, you have a load distribution problem, not a capacity problem.
The six metrics that actually decide
1. First response time (p90)
How long a user waits before someone answers. It is the first thing they perceive and the first thing that breaks when volume rises. Measured at p90, not as an average.
2. First contact resolution
What percentage of cases close without the user having to write again. It is the metric that correlates best with satisfaction and cost at the same time: every reopening is an interaction paid for twice. If volume rises and first contact resolution rises with it, the team is learning. If volume rises and it falls, the team is saturated.
3. SLA compliance
Not resolution time itself, but the percentage of cases within the agreed commitment. The difference matters: a resolution time that drops while SLA compliance also drops means the easy cases are being closed fast and the hard ones abandoned.
4. CSAT, always next to its response rate
A CSAT of 98% on a 4% response rate says nothing: the people who answered were the delighted and the furious, and the delighted were more numerous. CSAT without its response rate beside it is a decorative number. Always publish them together.
5. Level 2 escalation rate
What percentage of cases level 1 cannot resolve. It is a thermometer for documentation and training, not for agent quality. A rate that has not dropped by month three almost always means the knowledge base is incomplete, not that the team is weak.
6. Cost per resolved case
Not cost per hour or cost per agent. Cost per case actually closed. It is the only one that lets you compare like with like between a low-volume month and a peak season, and the only one that reveals whether automation is genuinely saving anything or just moving work around.
In back office operations, the sixth comes with a seventh: error rate per batch, measured on an audited sample rather than on self-reporting.
What has to be defined before month one
Metrics are useless if they are agreed after go-live. Three things get defined during diagnosis, not on the fly:
- The baseline. What you compare against. Without a starting number, any month-three result is an anecdote.
- Who measures. If the provider reports its own metrics with no sample audit, the client is measuring trust, not performance.
- The cadence, and what happens on a miss. A metric with no agreed consequence is a data point, not a commitment.
Metrics are part of the contract
The difference between an operation that improves and one that merely reports is whether the numbers carry consequences defined in advance. A dashboard is a management tool when someone is obliged to act on it, and an ornament when nobody is.
At smartBPO we define these metrics during diagnosis, before the first agent touches the operation, and we report them from month one. Not because it is textbook good practice, but because it is the only way the month-six conversation is about how to improve rather than about what went wrong.