Almost every outsourced operation reports a satisfaction number, and almost none can explain what moves it. It gets asked at the end of the contact, drifts a few points between one month and the next, and in the governance meeting someone asks to "work on satisfaction" without anyone being able to say what there is to work on. The problem is rarely the metric chosen. It's that it's collected badly, read in aggregate, and used to decide things it can't decide.
The three don't measure the same thing
CSAT measures satisfaction with one specific interaction. It's asked right after the contact and answers "how was this". It's useful for finding where a flow breaks, because it ties to a case, a reason and an agent. Its limit is that it's heavily influenced by the outcome: a flawless agent who has to deny a request gets a low score, and that isn't a service problem.
NPS measures the relationship with the brand, not the interaction. The recommendation question summarises the customer's accumulated experience with the product, the price, the billing, the logistics and, in some proportion, support. It's a legitimate metric and a poor outsourcing-contract metric: a provider doesn't control the product or the commercial policy. It works as business context, not as a grade for the service delivered.
CES measures the effort the customer spent getting their issue resolved. How many times they had to repeat the case, how many channels they touched, how many days they waited, how many steps they were asked for. Of the three it's the most actionable for a service operation, because effort depends mostly on how the flow is designed, and flow design can be changed.
If you can only sustain one well-collected metric, in support effort usually pays better than recommendation.
The problem isn't which one, it's how it's collected
A badly collected metric is worse than none, because it generates meetings about noise. The frequent defects:
- Low, skewed response rate. The very angry and the very pleased answer. The average of two extremes describes nobody.
- Wrong moment. Asking at the end of the call measures the conversation; asking days later measures whether the problem actually stayed solved. They're different questions and it helps to know which one is being asked.
- The agent requesting the survey. If the person who handled the case invites the rating, the number stops being a measurement and becomes a negotiation.
- Inconsistent scales. Five points in one channel and ten in another, or a scale change mid-year, makes the series incomparable.
- Bilingual operations without adjustment. The same scale is read differently by market. Comparing markets without accounting for that leads to false conclusions about the team.
- An undefined universe. If it isn't written which contacts are eligible for a survey, the universe ends up being the convenient one.
Satisfaction and quality are not the same thing
Satisfaction is the customer's voice. Audited quality is process compliance. You need both, and they often contradict each other: flawless cases with low scores, poorly handled cases with high scores because the customer got what they wanted. That contradiction isn't a measurement error, it's the most useful information the dashboard produces — as long as someone looks at it case by case instead of averaging it away. The audit method that makes the cross-reference possible is in quality control by audited sampling.
Which is why the number only becomes actionable when it arrives with a reason attached. Without a reason field, or open comments classified with the same taxonomy the operation uses, a drop in satisfaction doesn't say whether what failed was the service, the product, the wait, or a policy the customer doesn't accept.
Before putting it in the SLA
Putting satisfaction in the SLA with a financial penalty is tempting and usually goes badly. An indicator that depends on the customer's voluntary response and on the outcome of the case, with money attached, pushes predictable behaviour: not sending the survey on hard cases, asking for good scores, closing cases as resolved to trigger the send. If it's going in, four things should be written first: the universe of eligible contacts, the minimum sample for the month to count, exclusions agreed in writing, and which part of the result depends on the provider and which doesn't. It's the same criterion we apply to the rest of the agreement in how to write an SLA you can meet.
How to read it so it's useful
The aggregate average is the worst possible format. Four cuts that almost always reveal something:
- By channel. The aggregate hides precisely the channel that's failing; expectations of time and tone aren't the same on voice, chat and email, as we covered in support channels.
- By contact reason. Ranking dissatisfaction by reason separates what the operation can fix from what only the client can fix upstream.
- By recontact and escalation. Cases that came back or moved up a tier concentrate the low scores and explain a good part of the drop.
- By distribution, not by mean. It matters whether the month fell because many bottom scores appeared or because the middle moved. Those are different problems.
And a loop back is needed: someone accountable for contacting low-score cases within a defined window, with an action rather than an apology. Without that loop, measuring satisfaction is documenting dissatisfaction more precisely. The indicator set that supports this reading from day one is in what to measure from the first month.
What to ask a provider
- Which metric do they report, on what scale, and over which universe of contacts?
- What's the response rate and how has it moved over recent months?
- Who sends the survey and at what point in the case lifecycle?
- How do they classify open comments, and who reads them?
- How do they separate dissatisfaction with the outcome from dissatisfaction with the service?
- What happens to a low-score case, within what time, and who answers for it?
How smartBPO works it
We define a single primary metric per operation with the client and write down its universe, scale and exclusions before go-live, so the series is comparable from the first month. Where the process allows it we prefer to measure customer effort and complement it with transactional satisfaction, because effort points at where to intervene. The survey is triggered by the system, never requested by the agent. We classify open comments with the same reason taxonomy we use in the volume report, cross every low score against the quality audit of the same case, and bring cuts by channel, reason and recontact to the governance meeting instead of an average. Low-score cases have an owner and a contact deadline. We don't promise a satisfaction level: we propose how to measure it without bias, which part of the result depends on the operation and which on the client's own process, and to review it with the evidence of the case in front of us.