A backlog is a debt. It is not a long list or a bad month: it is work that was already promised, that has already spent someone's patience, and that keeps accruing interest in the form of complaints, chase-up calls and errors caused by age. Treating it as a temporary delay is why so many operations live with one for years.

The usual reaction is to ask for effort: overtime, weekends, a recovery sprint. Sometimes that works once. It almost never works twice, because effort does not fix whatever created the pile, and the operation is back in the same place six weeks later with a more tired team.

A backlog is not measured in cases

The figure quoted in meetings is usually the number of pending cases. On its own that number says nothing useful, because it mixes three-minute cases with two-hour ones. The unit that works is pending work hours: volume by case type, multiplied by the standard handling time for each type. With that conversion, twenty thousand cases might be two weeks of operation or four months, and the conversation changes completely.

That requires standard times by case type, which is exactly what is discussed in measuring back office productivity. Without that base, any recovery plan is a promise with no arithmetic behind it.

There is a second figure that matters just as much: age. A backlog of a thousand cases averaging three days old is a scheduling problem. The same volume averaging seventy days is a credibility problem, and it has probably generated its own traffic by now: every old case produces emails, calls and escalations that consume capacity from the very operation that should be closing it.

First work out which kind of backlog it is

There are only two, and they are handled differently.

An event backlog comes from something specific: a peak, a system outage, a migration, a group of people leaving at once. Daily intake is below daily capacity, but a pile was left behind. It clears with a temporary top-up and it goes away.

A structural backlog is a different animal: more work arrives than the operation can process on a normal day. The pile grows on its own, every day, even when nobody does anything wrong. Here a temporary top-up is money lost: the pile is cleared and immediately rebuilds.

The test that separates them is simple and uses the last few weeks of data: daily intake against daily closures, excluding overtime. If average intake exceeds average closures, the backlog is structural, and what needs fixing is the sizing rather than the team's morale. That calculation is in how to size a support operation.

Clearing a structural backlog with overtime means paying every month for a sum nobody did once.

Close the tap before emptying the bucket

No recovery holds if intake stays the same. Before launching the plan it is worth looking at what is arriving and why, with the disposition catalogue in hand — the subject of contact disposition codes. It is common to find that part of the volume should not exist at all: duplicate requests because the requester cannot see the status of their case, forms that arrive incomplete and get sent back, the same file passing three times through the same queue.

That avoidable work is usually the most profitable part of the plan, because a case you stop receiving is worth more than a case closed quickly. The levers for cutting it are set out in how to reduce contact volume.

Work order is a decision, not a default

Oldest first looks fair and is often right, because it stops old cases ageing indefinitely. But it is not always the best answer. A backlog gets ordered by explicit criteria, agreed with the client and written down:

  • Risk. Cases with regulatory, financial or security impact go first, whatever their date.
  • Age. Within each risk group, oldest first. This is the rule that stops a case being forgotten forever.
  • Cases that generate work. Files that are producing chase-ups and calls cost twice over while they stay open.
  • Homogeneous blocks. Grouping cases of the same type cuts context switching and lifts output without demanding more effort.

What does not work is leaving the order to each person's judgement. When that happens everyone takes the easy cases, and the backlog turns into a store of difficult work whose average age rises while the count falls.

The levers and what each one costs

There are four ways to clear a pile, and none of them is free:

  1. More hours from the current team. Fast to switch on and the highest yield per person, because they already know the process. It also has the shortest limit: sustained over time it produces absence and resignations, which is precisely what caused many backlogs in the first place.
  2. Temporary reinforcement. Works if the case type is quick to learn. Training time and the ramp-up curve have to be deducted, because in the first weeks that reinforcement produces less and consumes time from the people who already know the work. That deduction is covered in agent training and ramp-up.
  3. Simplifying the criteria, by exception and in writing. Sometimes part of the backlog can be resolved with a shorter validation without taking real risk. That is legitimate, but it has to be a documented client decision with defined scope and expiry, not a silent loosening of the standard.
  4. Automating a segment. Useful where there is a large, repetitive subset: data extraction, validation against a source, pre-classification. It is rarely ready in time for the current crisis, but it prevents the next one. What to automate and what not to is in automation with AI agents.

What breaks when you rush

A recovery plan measured only by cases closed pushes people to close them badly. The case leaves the queue, comes back two weeks later as a complaint and re-enters as a new case: the backlog was not reduced, it was reclassified.

Which is why quality auditing should go up during a recovery, not down. And it should look specifically at cases closed under the recovery plan, not at a general sample. The method is in quality control by audited sampling. A plan that closes fast and generates rework is not recovering anything: it is moving the problem to another month.

How to tell whether the plan is working

The daily pending count is the worst available indicator: it moves with intake and says nothing about whether recovery is progressing. Three readings are more useful.

The first is the daily balance: closures minus intake. While it stays positive the pile shrinks, and its size tells you how many days are left. The second is the age of the oldest case, which is what the client perceives and what reveals whether the agreed work order is being respected. The third is the curve: a real plan shows a sustained slope, not a jump at the start followed by a plateau. When the curve flattens early, it is almost always because the difficult work is what remains and the plan did not account for it.

If the backlog has already caused service level breaches, the conversation with the client follows its own protocol, separate from the operational plan: it is in what to do about an SLA breach.

The backlog nobody sees

Some piles appear on no dashboard because they sit outside the official queue: untagged emails in a shared inbox, cases sent back to the client's team waiting for information, requests a supervisor is holding for review. When measuring the real backlog it is worth counting them, however uncomfortable, because those are the ones that resurface just as the plan is declared finished.

It is also worth checking whether part of the delay sits outside the operation. A case that has waited forty days for a client approval will not be fixed by adding provider headcount, and reporting it inside the same metric hides where the real bottleneck is.

How smartBPO works it

When we take on an operation with a backlog, the first thing we do is convert it into work hours by case type and compare intake against closures without overtime, so we know whether we are facing an event pile or a capacity shortfall. We agree the work order in writing before starting, including which risk cases skip the queue. We raise quality auditing on what gets closed during the plan, not afterwards. We report daily balance, age of the oldest case and the closure curve instead of the count of the day. And we separate what is stuck in our queue from what is waiting on the client's side, because those are two different problems and only one of them is solved with capacity.