Summary. Most business continuity plans are written to satisfy a customer questionnaire and are never tested, which means they fail in exactly the circumstances they were written for. A plan that works is built from a real dependency map, states who decides what without waiting for a meeting, and has been exercised by the people who will use it. This guide covers the build: the business impact analysis that determines what must be recovered and how fast, the dependency mapping that identifies single points of failure, the crisis structure and decision authority, communications with employees and customers and regulators, recovery procedures for technology and facilities and people, the contractual and regulatory obligations a disruption triggers, the insurance that responds, and the testing that makes the difference between a document and a capability.
A distribution company's warehouse management system fails on a Tuesday morning. The vendor's data center has a hardware failure; restoration takes fourteen hours.
The company has a business continuity plan. It is 68 pages, it was written by a consultant three years ago to satisfy a customer's vendor questionnaire, and it is in a shared drive folder nobody can find because the person who organized the drive left.
What happens over those fourteen hours:
No one knows who is in charge. The COO is on a plane. The IT director is managing the vendor. Operations managers at four facilities make independent decisions, three of them different.
The company cannot ship, because the system holds the pick lists and nobody has a paper fallback. Two thousand orders queue.
Customers are not told for six hours, and when they are told, three different people give three different estimates, one of which is wrong by a factor of three.
Two customer contracts contain notice provisions requiring written notification of any service disruption within four hours, with liquidated damages for failure. Nobody reads them until the following week.
The company's cyber policy has a twelve-hour waiting period and would not have responded anyway, because this was a hardware failure rather than a security event; the contingent business interruption coverage that would have applied was declined at renewal to save premium.
And the backup, which existed, had never been tested for restoration, and the restoration took nine of the fourteen hours because nobody had done it before.
The plan was not wrong. It was irrelevant, because it described a structure nobody had rehearsed, for scenarios nobody had mapped, using contacts nobody had verified.
The difference between a plan and a capability is testing.
Step one: the business impact analysis
Start with what the business actually does, not with a list of disasters.
Identify business processes. Order intake, fulfillment, manufacturing, billing, payroll, customer support, regulatory reporting, and so on. Be specific — "operations" is not a process.
For each process, determine:
- The financial impact of disruption, per hour and per day, and how it escalates. Revenue loss, penalties, expedited costs, and recovery expense.
- The non-financial impact — regulatory consequences, customer attrition, reputational damage, safety, and contractual breach.
- The maximum tolerable downtime — the point beyond which the disruption causes damage the business cannot absorb.
- The recovery time objective (RTO) — the target for restoring the process, set inside the maximum tolerable downtime.
- The recovery point objective (RPO) — how much data loss is acceptable, which drives backup frequency. An RPO of four hours means backups every four hours; an RPO of zero means real-time replication and a substantially larger budget.
- Seasonality — a disruption in the peak week may be ten times as costly as the same disruption in a slow month, and the plan should say so.
Then rank. Most organizations discover that three to six processes are genuinely critical and the rest can wait days. That ranking is the plan's foundation, and without it every recovery effort competes with every other one at the worst possible moment.
A caution. Business owners asked to set an RTO will say "immediately" for everything. Force the tradeoff by attaching the cost: an RTO of four hours for a process requires redundant infrastructure that an RTO of two days does not. Make them choose with the price attached.
Step two: risk assessment and dependency mapping
The dependency map is the most valuable artifact in the entire exercise, and most organizations have never built one.
For each critical process, identify every dependency:
- People — who performs it, who else can, and what happens if the one person who knows it is unavailable. Single-person dependencies are the most common and most correctable vulnerability in any organization.
- Technology — applications, infrastructure, data, network, and the interfaces between them.
- Facilities — locations, equipment, utilities, and physical access.
- Suppliers and vendors — including the ones behind the ones you know about. A company that depends on a SaaS provider that depends on a single cloud region has a dependency it did not choose.
- Data — where it lives, who holds it, and how it is recovered.
- Third-party services — payment processing, telecommunications, logistics carriers, and outsourced functions.
- Records and documents required to operate.
- Regulatory approvals whose lapse would stop operations.
Then identify the single points of failure — dependencies with no alternative — and decide, explicitly, whether to accept, mitigate, transfer, or eliminate each. Documenting a decision to accept a risk is a legitimate outcome and is far better than not having identified it.
Threat assessment. For each, likelihood and impact:
- Technology — system failure, data loss, ransomware and other cyber events, and vendor outage.
- Facilities — fire, flood, storm, earthquake, utility failure, and loss of access.
- People — key person loss, mass absence, labor action, and workplace violence.
- Supply chain — supplier failure, logistics disruption, and raw material shortage.
- Financial — customer default, credit facility loss, and fraud.
- External — regulatory change, litigation, product recall, and reputational events.
Weight by likelihood, but plan by effect. A plan organized around causes needs a chapter for every scenario. A plan organized around effects — loss of a facility, loss of a system, loss of people, loss of a supplier — covers nearly everything with four playbooks, and it works for the scenario nobody anticipated.
Step three: the crisis management structure
The single most important element is knowing who decides, before there is anything to decide.
Define the team, by role rather than by name, with named incumbents and at least two alternates each:
- Crisis leader — with authority to commit resources and make operational decisions without further approval up to a stated threshold. Frequently the COO or a senior operations executive rather than the CEO, who has an external role to play.
- Operations lead.
- Technology lead.
- Communications lead.
- Legal.
- Human resources / people safety.
- Finance, for emergency expenditure authority.
- Facilities and security.
- A scribe, whose only job is to maintain a timeline of decisions and actions. This role is always omitted and always needed — the timeline is the record for the insurance claim, the regulator, and the after-action review.
Define authority explicitly. What can the crisis leader decide alone? What requires the CEO? What requires the board? What spending authority exists without normal approval? An organization that has to convene a committee to authorize a $50,000 emergency expenditure at 2 a.m. does not have a plan.
Define activation criteria and levels. A tiered structure works well: Level 1, a localized issue handled by the function; Level 2, a significant disruption requiring cross-functional coordination; Level 3, an enterprise crisis with executive and board involvement. State who can declare each level and what happens when they do — because the most common failure is not the wrong response, it is nobody declaring anything for six hours.
Establish the communication channel — a bridge line, a chat channel, and a location, physical or virtual — with the details on a card people carry, because the intranet may be the thing that is down.
Battle rhythm. In a sustained event: scheduled status calls at defined intervals, a standard reporting format, a shared situation log, and defined handoffs for events lasting more than a shift. Improvised coordination degrades fast after twelve hours.
Step four: communications
Prepare before the event. Communications drafted during a crisis are late, inconsistent, and frequently create legal exposure.
Employees first. They will hear about it regardless, they are the most credible source for everyone else, and they need to know what to do. Maintain an out-of-band contact list — personal phone and email, kept current — because a company email outage makes the internal directory useless. Use a mass notification tool if the workforce is large or distributed. Tell them what happened, what the company is doing, what they should do, and when they will hear more.
Customers. Identify who communicates with each account, prepare holding statements, and — critically — check the contractual notice obligations in advance. Many service agreements require notification of a disruption within a defined period, and several require specific content. Build a schedule of those obligations by customer, and keep it with the plan.
Suppliers, where the disruption affects them or where their help is needed.
Regulators, where a reporting obligation exists — and there may be several, with different clocks: a cyber incident report to a sector regulator, a securities disclosure, a health or safety report, an environmental release notification, or a data breach notification with a state-specific deadline.
Insurers, immediately, under every potentially applicable policy.
The public and the media. One spokesperson. A prepared holding statement that says what is known, what the company is doing, and when there will be more. Do not speculate about cause, do not estimate a restoration time you cannot support, and do not minimize. A dedicated status page — hosted independently of the company's own infrastructure — is far better than nothing and is what customers will look for.
Legal review of external communications, balanced against speed. The compromise is pre-approved templates for the predictable scenarios, so review is a check rather than a drafting exercise.
Two rules that prevent most communications damage: say only what you know, and update on a schedule even when there is nothing new. Silence is interpreted as concealment, and an estimate that slips three times destroys credibility that took years to build.
Step five: recovery procedures
Technology.
- Backups, on a schedule matched to the RPO, with offline or immutable copies — because ransomware encrypts the connected backups first. The widely used discipline is three copies, on two different media, with one off-site and one offline or immutable.
- Restoration testing, at least annually, timed. A backup that has never been restored is a hypothesis. The most common finding in a real event is that restoration takes far longer than anyone assumed.
- Failover capability for critical systems, with the failover procedure documented and exercised.
- Documented recovery runbooks — the actual steps, in order, with credentials accessible to more than one person and stored where they can be reached if the primary systems are down.
- Alternate access — how people work if the network is unavailable.
- Vendor escalation contacts and contract support terms, with the service level commitments and the escalation path documented and current.
- Manual workarounds for the critical processes. What does fulfillment do without the warehouse system? What does billing do without the ERP? A one-page paper procedure per critical process is cheap and is what actually keeps the business moving.
Facilities.
- Alternate work locations, which for many businesses now means remote work — with the equipment and network capacity to support it, tested.
- Alternate production or fulfillment capability, including reciprocal arrangements with other facilities or contract capacity.
- Equipment replacement sources and lead times.
- Physical access to records and to critical materials.
- Utilities — generators, fuel contracts, and how long they last.
People.
- Safety first, always, including evacuation and shelter procedures, accountability for headcount, and support for affected employees.
- Cross-training for critical roles, which is the mitigation for the single-person dependency identified in the mapping.
- Succession for key positions, documented, including delegation of signing authority.
- Payroll continuity — how people get paid if the payroll system or the office is unavailable. This is frequently overlooked and is the thing employees care about most in a prolonged event.
- Employee assistance for events with a human toll.
Supply chain.
- Alternate suppliers identified and, ideally, qualified in advance — because qualifying a new supplier during a disruption takes weeks the business does not have.
- Safety stock for critical inputs, sized against the realistic replacement lead time.
- Visibility into subtier suppliers, at least for the components with no alternative.
- Contract terms — force majeure, allocation provisions, and whether the supplier may prioritize other customers.
Step six: the legal layer
A disruption triggers obligations that operations teams do not know exist. Map them in advance.
Contractual obligations:
- Notice provisions — the deadlines and content required by customer, supplier, lender, and lease agreements. Build a schedule.
- Service level agreements — the credits or penalties that accrue, and whether there is a cap or a termination right at a threshold.
- Force majeure clauses, in both directions. Read them: they are not standard, the enumerated events differ, and many require prompt written notice as a condition of relief, with a mitigation obligation and, in some, a termination right after a defined period. A party that fails to give timely notice may forfeit the excuse entirely.
- Lender covenants — reporting obligations, material adverse change provisions, and financial covenants that a disruption may breach.
- Insurance policy conditions — notice, proof of loss, and mitigation obligations.
Regulatory obligations vary by industry and by event: cyber incident reporting with clocks measured in hours in several regimes; data breach notification with state-specific deadlines; securities disclosure of material events; health and safety reporting; environmental release notification; and sector-specific requirements for financial institutions, health care providers, utilities, and government contractors.
Litigation and evidence. Where the event may lead to claims — a customer loss, an injury, a data breach — issue a litigation hold immediately and preserve the systems, logs, and communications that would otherwise be overwritten or discarded during recovery. Recovery activities routinely destroy the evidence needed to establish what happened, and forensic imaging before restoration is frequently the right sequence even though it delays recovery. Decide that tradeoff deliberately, with counsel, rather than by default.
Privilege. Where the event will produce claims or regulatory scrutiny, structure the investigation under counsel from the beginning. An after-action review written for operational improvement is discoverable and will be read as an admission; the same analysis conducted under counsel for the purpose of legal advice may not be. Both are worth having, and they should be separate documents.
Contract review before the event, not after. Pull every material agreement and extract the notice deadlines, the force majeure terms, the SLA credits, and the termination triggers into a one-page schedule kept with the plan. Doing this during a crisis is how the four-hour notice obligation in the opening example was missed.
Step seven: insurance
Business interruption coverage — part of a property policy — responds to loss of income from a direct physical loss to covered property. Two consequences follow, and both surprise people: a cyber event or a vendor outage with no physical damage generally does not trigger it, and a communicable disease event generally does not either, following extensive litigation on the point.
Key terms to understand:
- The period of restoration and any extended period of indemnity after operations resume.
- The waiting period or deductible, expressed in hours or days.
- Contingent business interruption, covering loss from a supplier's or customer's disruption. Confirm whether it requires physical damage at the third-party location and whether unnamed suppliers are covered, because many policies cover only scheduled ones.
- Civil authority and ingress/egress coverage.
- Service interruption coverage for utility failure, which frequently requires damage to the utility's property and excludes transmission lines unless specifically endorsed.
- Extra expense coverage for the costs of continuing operations.
- The limit, which should be based on a worksheet reflecting the actual recovery period rather than an annual figure.
Cyber insurance covers what business interruption does not: network interruption from a security event, dependent network interruption from a vendor's security event, data restoration, forensics, notification, extortion, and liability. Read the waiting period, the indemnity period, the sublimits, and the panel requirements, and understand that a hardware failure at a vendor is usually neither a cyber event nor a physical loss — which is exactly the gap in the opening example.
Other coverages to review against the dependency map: property, including the treatment of off-premises equipment and property in transit; supply chain and trade disruption; product recall; crime; and D&O for the claims that follow a badly managed event.
Claim preparation. Business interruption claims are documentation-intensive and are frequently underpaid because the insured cannot substantiate the loss. Assign someone to track incremental costs and lost revenue from hour one, in a format the policy's proof of loss requirements will accept, and consider a forensic accountant early for a substantial claim.
Step eight: testing
This is what separates a plan from a document, and it is the step that gets cut.
A tiered program:
Plan review — annually, and after any material change to the business, the systems, or the vendor set. Verify contacts, verify vendor terms, verify that the critical process list still reflects reality.
Tabletop exercises — two to four hours, quarterly or semiannually. A scenario is presented, the team works through the response, and a facilitator injects complications. This is the highest-return activity in the entire program relative to its cost. Vary the scenario — a facility loss, a system outage, a ransomware event, a key supplier failure, a senior leader unavailable — and vary who is in the room, because the plan must work when the primary people are the ones affected.
Functional tests — actually perform a component: restore a backup and time it, fail over a system, activate the notification tool, run a shift on the manual workaround.
Full simulations — annually for mature programs, involving multiple functions and running for hours.
After-action review after every exercise and every real event: what worked, what did not, what was missing, and specific assigned actions with owners and dates. An after-action review that produces no assigned actions is a meeting.
Update the plan from what testing reveals. Nearly every exercise finds an outdated contact, an assumption that does not hold, or a dependency nobody mapped.
The most common findings, in rough order: contact lists are stale; nobody knows who declares an incident; restoration takes far longer than assumed; the manual workaround does not exist; the notice obligations in customer contracts were never extracted; and the plan is stored somewhere that is unavailable during the event.
Step nine: keeping it usable
Format matters. A 68-page document is not usable at 2 a.m. Build:
- A one-page activation card — who to call, how to declare, who is on the team, and the bridge line. Physical copies, and on people's phones.
- Role-specific checklists — one page per role, with the first ten actions.
- Scenario playbooks — four to six pages each, by effect rather than by cause.
- The reference document — dependency maps, contact lists, contract schedules, vendor terms, insurance summary — maintained but not read during an event.
Store it where it will be available — offline copies, an out-of-band location, and printed copies for the crisis team. A plan that lives only on the network that is down is not a plan.
Assign ownership. One person accountable for maintenance, with a review calendar. Programs without an owner degrade within eighteen months.
Integrate it with vendor management, so that a new critical vendor triggers a dependency assessment, and with onboarding, so that new team members learn their role.
And keep it proportionate. A thirty-person company does not need what a multinational needs. It needs: a dependency map, a one-page activation card, tested backups, an out-of-band contact list, a schedule of contractual notice obligations, and a tabletop exercise once a year. That is a day of work per year and it addresses the great majority of the risk.
A short case study
A specialty pharmaceutical distributor with 180 employees rebuilds its program after a near miss.
The impact analysis identifies four critical processes: order intake, cold chain storage and fulfillment, controlled substance recordkeeping, and customer support. Cold chain has a maximum tolerable downtime measured in hours, not days, because product spoils and because a lapse creates a regulatory problem as well as a commercial one.
The dependency map surfaces three single points of failure: one person who understands the temperature monitoring integration; a single refrigeration service vendor with no alternate under contract; and a warehouse management system hosted by a vendor in one region.
Decisions. Cross-train two additional people on the monitoring system and document the runbook. Qualify a second refrigeration vendor and put a standby agreement in place. Negotiate multi-region failover into the WMS contract at renewal, and build a paper picking fallback in the interim.
The crisis structure is defined by role with two alternates each, an activation card is distributed, and a scribe role is created.
The legal schedule extracts notice obligations from 34 customer agreements. Nine require notification within specified periods, four as short as two hours for temperature excursions. This schedule alone justifies the project.
Insurance is reviewed: contingent business interruption is added for the two scheduled critical suppliers, the cyber policy's waiting period is reduced, and the business interruption worksheet is rebuilt against a realistic 30-day recovery rather than the prior estimate.
Testing. A tabletop on refrigeration failure at the primary facility, which reveals that nobody knows who is authorized to approve emergency transport at 3 a.m. A functional test of backup restoration, timed at eleven hours against an assumed four — which drives an infrastructure change. A notification tool activation test with a 71 percent contact rate, which drives a contact data cleanup.
Eight months later, a regional power event takes the primary facility offline for nineteen hours. Generators run, the crisis team activates within twenty minutes, customers are notified within the contractual windows, product is moved to the secondary facility under the standby transport arrangement, and the incremental cost is tracked from the first hour. The insurance claim is paid substantially in full.
The total investment in the program was roughly one full-time-equivalent quarter of effort plus modest incremental insurance premium. The event would have cost multiples of that without it.
Conclusion
Three points carry the weight.
Map dependencies, and plan by effect. The dependency map is the artifact that makes everything else possible, and it finds the single-person and single-vendor failures that cause most disruptions. Playbooks organized around effects — lost facility, lost system, lost people, lost supplier — cover the scenarios nobody predicted.
Define decision authority before the event. The most common failure in a real disruption is not a wrong decision; it is that nobody was authorized to make one for six hours. Name roles, name alternates, state spending authority, and define who declares.
Test, or you have a document. Restore a backup and time it. Activate the notification tool. Run a tabletop where the primary decision-maker is unavailable. Every one of those exercises finds something, and the things they find are the things that would otherwise be discovered at the worst possible moment.
Frequently asked questions
How long should a plan be? The usable part should be a one-page activation card, a one-page checklist per role, and four to six scenario playbooks of a few pages each. Everything else — dependency maps, contact lists, contract schedules, insurance summaries — is reference material that is maintained but never read during an event.
Who should own it? One named person, with a review calendar and executive sponsorship. Programs owned by a committee decay; programs owned by a consultant who has left decay faster.
How often should we test? A tabletop every six months, a functional test of at least one component annually (restoration is the obvious candidate), and a plan review annually and after any material change. That cadence is achievable for a company of almost any size.
Does our customer's vendor questionnaire count as a plan? No. It documents that a plan exists. The gap between a questionnaire answer and a capability is exactly the gap described at the top of this guide, and customers increasingly ask for evidence of testing rather than for the document.
What is the minimum viable program for a small company? A dependency map, a one-page activation card with an out-of-band contact list, tested backups, a schedule of contractual notice obligations, and one tabletop a year. That is roughly a day of effort annually and it addresses most of the realistic risk.
Does business interruption insurance cover a cyber outage? Generally not, because it requires direct physical loss to covered property. Network interruption from a security event is covered by a cyber policy, and an outage caused by a vendor's hardware failure with no security event and no physical damage may be covered by neither. That gap should be identified deliberately and either accepted or insured.
Should we pay a ransom? That is a decision for the crisis team with counsel and law enforcement involved, and it turns on whether backups are viable, whether data was exfiltrated, whether the payment is lawful under sanctions rules, and what the insurer's policy requires. The decision framework should be agreed before an event, because it cannot be reasoned through calmly at hour four.
When should we invoke force majeure? Only after reading the specific clause, because they are not standard. Most require prompt written notice as a condition of relief, most impose a mitigation obligation, and many give the other party a termination right after a defined period — so invoking it may cost the contract.
The first four hours
For the team that has to run an event without having rehearsed one, here is the order of operations.
Minutes 0–15. Someone declares. That is the whole first step, and it is the one most often delayed. Convene the team on the bridge line. Confirm safety of people first; everything else waits behind that.
Minutes 15–60. Establish what is known and — explicitly — what is not. Start the scribe's timeline. Assign a single owner for each workstream: technical recovery, customer communication, employee communication, vendor escalation, and legal and insurance. Set the next status call time and hold it.
Hour 1–2. Notify insurers under every potentially applicable policy. Pull the contractual notice schedule and identify which customer notifications are due and when. Issue the employee communication. Post a holding statement if customers are affected. Begin tracking incremental costs.
Hour 2–4. Assess whether a regulatory reporting clock has started, and if so, which. Decide whether a litigation hold is warranted and issue it if so — before recovery activities overwrite the evidence. Confirm the recovery estimate with the technical team and communicate it externally only when it can be supported. Escalate to the board if the event meets the criteria.
Throughout. Update on the schedule even when there is nothing new. Keep the timeline. Do not let the crisis leader also be the person performing the technical recovery. And plan the handoff before anyone has been awake for eighteen hours, because sustained events are lost in hour twenty, not hour two.
One more thing about the after-action review. Hold it within two weeks, while memories are fresh and before the organization has decided what the story was. Run it as a factual reconstruction from the scribe's timeline rather than as a discussion, invite the people who did the work rather than only the people who supervised it, and separate the operational review from any privileged legal analysis. Then assign the actions with owners and dates, and check them at the next review. An organization that runs one honest after-action review learns more than one that writes three plans.
Related articles
- Business Continuity and Crisis Response Checklist — the plan in checklist form.
- Crisis Management and Business Continuity Toolkit — the full roadmap.
- Data Breach and Incident Response Toolkit: From Detection to Notification — the cyber-specific playbook.
- Cybersecurity Program Toolkit — the controls that prevent one category of event.
- Negotiating a Commercial Insurance Program: A Practical Guide — business interruption, contingent BI, and the terms that decide the claim.
- Handling an Insurance Claim After a Property Loss — documenting and presenting the claim.
- Handling a Product Recall: A Practical Guide — the crisis structure applied to a specific event.
- Vendor Cybersecurity Diligence Checklist — assessing the dependencies the map identifies.
- Litigation Hold and Evidence Preservation Checklist — preserving evidence while recovering.
- Contract Lifecycle Toolkit: From Term Sheet to Termination — extracting the notice and force majeure obligations in advance.
This guide is provided for general informational purposes and does not constitute legal or insurance advice. Regulatory reporting obligations, contractual notice requirements, and insurance coverage vary substantially by industry, jurisdiction, and policy language. Consult qualified counsel and a licensed broker when building or activating a plan.