The monthly report lands and the SLA was missed. Service level came in below what was agreed, or resolution time slipped, or audited quality dropped two months running. The typical reaction is one of two things: a firm email asking for an action plan, or opening the contract to work out how much can be deducted. Neither one, on its own, fixes the operation.
A breach is a data point, not a verdict. What decides the outcome is what happens over the following four weeks, and the order things are done in matters more than it looks.
Step one: verify the number is measured properly
Before discussing responsibility, you need to be sure the figure says what both parties think it says. A sizeable share of the breaches that reach the table are measurement disagreements, not operational failures.
- Universe. Does it include every contact, or are there agreed exclusions? A channel added mid-quarter can be contaminating the series.
- Clock. When does it start and when does it stop? Cases waiting on the end customer, or on an internal approval, are usually the blind spot.
- Hours. An indicator calculated over 24 hours when the operation runs twelve gives a different and false result.
- Source. If the provider reports from its tool and the client measures from its own, there will almost always be a gap. Someone has to decide which one governs.
If that conversation wasn't settled up front, the problem starts in the design of the agreement rather than its execution. It's exactly what we cover in how to write an SLA you can meet.
Step two: separate the cause before demanding a plan
Asking for an action plan without knowing the cause produces a generic document: more training, more follow-up, more commitment. Useless. Real causes almost always fall into one of five categories, and each demands a different response.
- Demand. Volume came in above what was sized, or its distribution across time slots changed. That's a shared forecasting problem, not a matter of team effort.
- Capacity. The headcount exists but isn't available: absence, attrition without replacement, off-line time miscalculated. That's the point we go through in shrinkage in BPO.
- Competence. People are there but aren't resolving. Incomplete ramp, an outdated knowledge base, new case types with no procedure.
- Client dependency. Slow approvals, a system outage, access that arrived late, a product change with no notice. It happens more than anyone admits.
- Design. The SLA was never achievable with the structure and budget agreed. Hard for both sides to admit, and common.
If the cause sits in the design or in a client dependency, applying a penalty improves nothing: it just moves money while the indicator stays where it was.
Penalties: what they're for and what they aren't
A well-set penalty is a signal, not a revenue line. It does two legitimate things: it forces the issue up the provider's internal ladder and it offsets part of the client's operational damage. What it doesn't do is correct the cause.
There's a side effect worth anticipating. If the deduction is large and recurring, the provider starts optimising the indicator instead of the service: closing cases fast to pull resolution time down, reclassifying reasons, pushing contacts into a channel that isn't measured. The indicator improves and the end customer is worse off. That's why healthier schemes combine a moderate deduction with a recovery mechanism on deadlines, not the other way round.
The shape of the penalty also depends on how the service is charged. In an hourly or FTE model the deduction hits the provider's margin; in a per-transaction model it can end up punishing precisely the volume that needs processing. Worth re-reading the pricing model before setting the percentage.
The recovery plan that actually works
A useful plan is short, has dates, and can be verified without waiting for month-end. Six elements:
- Scope. Which indicator, in which queue or channel, since when. Not the whole service is broken; naming the exact piece cuts the noise.
- Cause with evidence. Not a hypothesis: the data behind it. Demand curve against staff available, audit of failed cases, system downtime log.
- Actions with owners. Each action with a name and a date, including the ones that belong to the client.
- A weekly interim metric. An indicator that moves before the monthly SLA does, so you know within seven days whether the plan is working.
- Return date. When the agreed level is expected back, with an explicit test for compliance.
- What happens if not. The next step agreed in advance: structural review, change of account lead, renegotiation of the indicator, or start of exit.
That plan needs a forum where it gets reviewed, with people who can decide. Without a governance body with clear cadence and authority, the plan turns into an email chain. We set it out in BPO contract governance.
The client's share
Uncomfortable but necessary: in a good share of breaches there's a contribution from the client side. Approvals that take days, product information that changes with no notice, a promotional peak nobody flagged, two tools that don't talk to each other. If the recovery plan only lists provider tasks, it's probably incomplete and will fail in month two.
Another piece that's often weak is quality evidence. Discussing performance without a sampled case audit is discussing perceptions; with one, the conversation moves from "the service feels bad" to "in these fourteen cases this step failed". That shift is what unblocks things. It's in quality control by audited sampling.
When it really is time to leave
Changing provider costs time, accumulated knowledge and operational risk. It's worth it when one of these conditions holds, not before: the recovery plan failed with a clear cause and actions that weren't executed; the provider argues about measurement instead of working the problem; the breach repeats across different indicators, which points to something structural rather than one queue; or there was a risk event rather than a performance one, such as a security or compliance failure. If you get there, the exit is planned with the same discipline as the entry: switching providers without breaking operations.
Questions for the recovery meeting
- Are we measuring the same thing? Show me the universe, the exclusions and the clock stops.
- What's the cause, and what data supports it?
- Which part of the breach depends on us as the client?
- Which specific actions, with owner and date, change that data?
- Which interim indicator are we looking at next week?
- On what date do we return to the agreed level, and what do we do if we don't?
How smartBPO works it
We agree the universe, the exclusions and the clock stops for every indicator with the client before go-live, so that a bad month is argued over the figure and not over how it was calculated. When an indicator drops, we bring the classified cause and the evidence behind it to the meeting, including the causes that are ours and the ones that sit with the client, with the same candour. The recovery plan comes with an owner, a date and a weekly metric that allows correction before month-end. We prefer a moderate penalty scheme paired with recovery deadlines, because a heavy deduction pushes people to dress up the indicator and that degrades the service. And we write into the contract what happens if the plan doesn't work, including an orderly transfer of the operation. We don't promise that no month will ever miss: we propose that when it does, the cause is visible and the way back has a date.