ACloud.Solutions

ISO 27001 and compliance

A business continuity tabletop exercise that does not need a conference room

The business continuity plan is fourteen pages. It has a call tree with two people who have left, an RTO of four hours that nobody has tested, and a section on relocating to an alternative site, which is notable for a company that does not have a site.

A business continuity tabletop exercise is how you find that out in two hours rather than during an incident. It needs no venue, no consultant and no technology. Five people, a scenario, and somebody writing down what nobody can answer.

What a business continuity tabletop exercise is

A facilitated discussion, not a technical test. You describe a situation, then ask the people in the room what they would do, in what order, and who they would tell. Nothing is switched off and nothing is restored. The output is a list of things that turned out not to work.

Two hours is right. Under ninety minutes and you get through one phase. Over two and a half and attention goes.

Five people is right: whoever holds technical recovery, whoever talks to customers, whoever can authorise spending, whoever handles contracts, and a facilitator who is not answering questions. In a company of forty that is typically the technical lead, a support lead, a director and whoever holds contracts, with you facilitating.

If you are running the ISMS alone you should not be both facilitator and principal respondent, so ask somebody else to facilitate from a script you wrote. It works better than it sounds, because a facilitator who does not know the answers asks better follow-up questions.

Choosing a scenario

One scenario. Not three. The temptation is to cover everything and the result is that nothing gets tested properly.

Ransomware on the production platform. The strongest default. It touches recovery, communications, legal notification, customer contracts and the decision about payment, which is a decision nobody wants to make for the first time under pressure.

Loss of the identity provider. Underrated and increasingly the right choice. If your directory is unavailable, nobody can sign into anything, including the tools you would use to coordinate the response. This surfaces the dependency nobody maps, and it pairs with break-glass accounts, which is exactly the control the scenario tests.

Key person unavailable. Uncomfortable and the most valuable at this size. If the one person who knows how deployment works is unreachable for a fortnight, what stops. In a one-person IT function this is a genuine risk that belongs in your register honestly rather than a hypothetical.

Supplier failure. Your hosting provider has a regional outage, or a critical SaaS supplier is breached. Tests whether you know what depends on what, and whether the contract says anything useful.

Pick the one that makes you most uncomfortable. That is where the findings are.

Running it

Three phases, roughly forty minutes each.

Phase one, detection and triage. How do we find out. Who is told first. What do we know at this point, and what do we assume. Push on the difference between those two, because that is where incidents go wrong.

Phase two, response. What do we do in the first hour. Who decides. What do we tell customers, when, and who writes it. Do we have an obligation to notify anyone, within what period, and who makes that call.

Phase three, recovery and after. How do we get back. How do we know we are back. What do we tell people afterwards, and what do we do differently.

Inject one complication partway through. The person who would normally handle this is on a flight. The status page is hosted on the thing that is down. The backup is fine but nobody has ever restored one. Complications are where the useful findings come from, because the plan covers the straightforward path.

Two rules that keep it honest. Nobody says "we would just" without saying who and how. And "I would check the runbook" is followed immediately by "show me", which is where the exercise usually stops being comfortable.

The findings that come out every time

There is a pattern, and it is remarkably consistent.

Nobody knows where the runbook is, or the version they find is out of date. This finding appears in essentially every first exercise, and it is the argument for documentation people actually use.

The contact list is stale. Personal phone numbers nobody has, a supplier contact who left, an escalation path into a company that has been acquired.

Nobody has restored a backup. The jobs are green. Green means the backup job succeeded, not that a restore works, and those are different claims.

The recovery time objective was invented. Somebody wrote four hours because it sounded reasonable. Ask what the actual recovery steps are and time them honestly, and it is frequently longer. Worth checking whether any customer contract commits you to a figure, because a contractual number and an aspirational number in your own plan are very different things and the contract wins.

The notification clock is not understood. Who decides whether an incident is notifiable, on what basis, and within what period. The ICO's personal data breach guidance sets out the obligation, and the exercise is where you find out whether anyone in the room knows it.

Everything routes through one person. Which is the honest finding in a small company, and the useful output is not to pretend otherwise but to identify the two or three things that must be documented well enough for somebody else to do them.

Recording it

The record is the evidence, and it is short.

Date, scenario, participants and roles, the timeline as discussed, findings, and actions with owners and dates. Two pages. The findings and actions are the part that matters; the narrative is context.

Then the actions have to go somewhere that gets reviewed, which is your corrective action log rather than a document in a folder. A finding raised in an exercise and never closed is worse than not having exercised, because you documented knowing about it.

What it satisfies

Annex A 5.29 covers information security during disruption and 5.30 covers ICT readiness for business continuity, and both expect the arrangements to be tested. A two-hour tabletop with a written record satisfies "tested" for a company of this size. It is not a full failover test and it does not claim to be.

Annually is the usual commitment, and it is the one that lapses first because nothing prompts it, per the surveillance audit pattern. Book next year's the day you finish this one.

The book covers business continuity at SME scale, including how much plan is enough, and the security and compliance work includes facilitating one of these, which solves the problem of not being able to facilitate and answer at the same time.