A storm knocks out power in the area. The internet provider has a four-hour outage. A road gets blocked and half the morning shift doesn't make it in on time. None of this is exceptional, in Colombia or anywhere else a service centre operates. What is exceptional is finding it written into the contract before it happens.
Continuity is the part of the agreement people read after it has already failed. Nobody asks about it in the first meeting: it doesn't sell, it lengthens the negotiation, and talking about outages while signing feels like bad manners. Later it becomes the tensest conversation of the year, with the operation stopped and two versions of who was supposed to do what.
Continuity isn't redundancy
The two get mixed up often. Redundancy means duplicating a resource: a second link, a generator, a mirrored server. Continuity is a decision made in advance about what gets sustained, in what order and within what timeframe, knowing that not everything can be sustained at once.
A plan promising that everything carries on as normal isn't a plan, it's a statement of intent. A useful plan says the opposite: during a contingency this gets handled first, that gets degraded, and this other thing stops. A user-safety case doesn't wait; a satisfaction survey does. Writing that hierarchy down while nothing is on fire, with the client in the room, is half the work.
The other half is timing. How long the operation can take to return to a minimum service, and how long to return to normal. Those are two different numbers and both should be agreed, not estimated by the provider alone.
A continuity plan that doesn't prioritise isn't a plan: it's a wish list on letterhead.
The scenarios that actually happen
Plans tend to imagine catastrophes and forget the everyday. Most real interruptions in an outsourced operation fit into a short, undramatic list:
- Power. Grid cuts, scheduled utility maintenance, internal building faults.
- Connectivity. The main link drops, or quality degrades to the point where voice is unusable even though chat keeps working.
- Third-party platforms. The client's CRM, telephony or ticketing tool goes down. The provider controls none of it and still has a team sitting there.
- Mass absence. Weather, public order, transport, health outbreaks. It hits the shift, not the infrastructure.
- Security or access incidents. Blocked credentials, a VPN down, a change on the client's side nobody announced.
- Building inaccessible. Rare, and the fastest way to find out whether the plan existed.
Worth noting that three of those six aren't the provider's fault. That doesn't make them irrelevant: the operation stops either way, and somebody has to decide what happens with the team in the meantime.
Site, home, or both
Work from home tends to get presented as the default contingency plan. It works, but it comes with conditions worth checking before taking it for granted.
The first is that homes fail too: residential power and internet are, on average, less stable than a site with backup. The second is equipment: sending someone home without a headset, without an assigned machine and without a support channel isn't continuity, it's dispersal. The third is information security: some processes can't leave a controlled environment, either by agreement with the client or because of the nature of the data. If that's the case, the plan has to solve it with something else, not with an improvised exception on the day of the incident.
The reverse logic applies as well. A fully remote operation needs an answer for the scenario where several people's homes fail at once, which is exactly what happens when the event is regional. There the backup is a physical site, however small, or a second city.
Coordination breaks before technology does
In nearly every badly handled incident, the technology recovers before the communication does. The client's team hears about it from an end user, not from their provider. Or they hear three hours later, once someone posted in a channel nobody is reading, because that channel is precisely the one that went down.
A continuity plan that works answers four questions with actual names attached: who declares the contingency, who they notify, within what timeframe, and through which alternate channel. That alternate channel matters: if notification depends on corporate email and the incident is connectivity, the notification doesn't exist. A mobile number written into the contract is inelegant and highly effective.
Then comes cadence: how often the status gets updated while the event is open, even when the update is "still no fix". Silence during an outage does more damage than the outage.
What happens to the SLA meanwhile
This is where most contracts turn vague. The force majeure clause gets copied from a template, ends up broadly worded, and covers anything the provider would rather not answer for. The opposite is preferable: a narrow definition, with examples, and an explicit rule about the metrics.
That rule can be to exclude the hours of the event from the period's calculation, or to measure them separately and report them. What shouldn't happen is the event disappearing from the monthly report: erase it and the trend stops being readable, and nobody learns anything. This connects directly to how you write an SLA you can meet: a metric with no exception rule gets missed for reasons outside anyone's control and loses its authority for the rest of the year. None of this is legal advice; the actual wording should be reviewed by a lawyer in each case.
A plan that isn't rehearsed doesn't exist
The difference between a provider that has continuity and one that has a continuity document shows up in a single question: when was the last drill, and what went wrong.
There are three levels of rehearsal and all three are useful. The tabletop one gathers the people responsible and walks through the scenario out loud, moving nothing; it's good for discovering that two people each thought the decision belonged to the other. The partial one moves a small group into the alternate mode for a few hours. The full one relocates the whole operation, and it's the only one that surfaces the real bottlenecks — like the generator starting up but not powering the floor's air conditioning.
What matters isn't the frequency agreed, but that a written report with findings and owners follows, and that the report reaches the committee with the client. It's natural material for contract governance: with no forum to review it, the drill becomes an annual formality.
What to get in writing
- The list of scenarios covered and, explicitly, the ones that aren't.
- The service priority during a contingency: what gets handled, what gets degraded, what gets suspended.
- The agreed timeframes to reach minimum service and to return to normal.
- The notification protocol: who declares, to whom, within what time, through which alternate channel, and at what update cadence.
- How metrics are treated during the event and how they're recorded in the report.
- The drill calendar and the commitment to share the findings.
- The conditions for work from home as a backup: equipment, minimum connectivity, security rules, and excluded processes.
How smartBPO works it
We build the scenario matrix with the client before launch, not after the first incident, and we prioritise services while nothing is burning: what holds, what degrades, what stops. We put in writing who declares a contingency on each side, with a name and an alternate channel, plus an update cadence for as long as the event stays open. We combine site and work from home according to what each process allows, rather than applying one rule to the whole operation. We rehearse, and we share the rehearsal report, including what didn't work. And we don't offer perfect availability: we offer an agreed, measured and reviewable response, which is the only thing that holds up when the event arrives without warning.