Summary. Most business continuity programs are documents rather than capabilities, and the difference surfaces at the worst possible moment — the plan describes a structure nobody has rehearsed, the contact list is two years old, the backup has never been restored, and the whole thing lives on the network that is down. This toolkit builds the capability: the impact analysis that determines what must be recovered and how fast, the dependency map that finds single points of failure, the crisis structure that names who decides and with what spending authority, the communications and contractual notice obligations extracted in advance, the recovery capability, the insurance and the gaps between policies, and the exercise regime.
What this toolkit is for, and who should use it
Three observations organize the field. The most common failure in a real event is not a wrong decision; it is that nobody was authorized to make one for six hours. The second is a dependency nobody had mapped — a single person, a single vendor, a single cloud region. And the third is a contractual notice obligation nobody had extracted, discovered a week after the deadline.
This toolkit is for operations leaders, general counsel, and risk managers at a company of any size — with a note at the end on what a proportionate program looks like for a small one.
Roadmap at a glance
- Business impact analysis.
- Dependency mapping and single points of failure.
- The crisis structure — roles, authority, and activation.
- Communications, prepared in advance.
- The legal layer — contractual notice, force majeure, and regulatory clocks.
- Technology recovery.
- Facilities and people.
- Supply chain.
- Insurance — what responds and what does not.
- Testing.
- Keeping it usable — format, storage, and ownership.
- The first four hours, and the questions companies ask.
Stage 1 — Business impact analysis
Start with what the business does, not with a list of disasters.
List the business processes specifically. For each, estimate the financial impact per hour and per day and how it escalates; identify the non-financial impact — regulatory consequences, customer attrition, safety, and contractual breach; set the maximum tolerable downtime; set the recovery time objective inside it; set the recovery point objective, which drives backup frequency; and note seasonality, because the same disruption in the peak week may cost ten times what it costs in a slow month.
Then rank. Most organizations find three to six processes are genuinely critical and the rest can wait days. Without that ranking, every recovery effort competes with every other at the worst moment.
Force the tradeoff by attaching cost to the RTO, because business owners asked to set one will say "immediately" for everything.
Stage 2 — Dependency mapping
The most valuable artifact in the exercise, and most organizations have never built one.
For each critical process, identify: people — who performs it, who else can, and what happens if the one person who knows it is unavailable; technology — applications, infrastructure, data, network, and interfaces; facilities — locations, equipment, utilities, and access; suppliers and vendors, including the ones behind the ones you know about; data — where it lives and how it is recovered; third-party services — payments, telecommunications, logistics, and outsourced functions; and records and regulatory approvals required to operate.
Mark every single point of failure and decide explicitly whether to accept, mitigate, transfer, or eliminate each. Documenting a decision to accept is a legitimate outcome; not having identified it is not.
Build playbooks by effect — lost facility, lost system, lost people, lost supplier — rather than by cause. Four playbooks cover nearly everything, including the scenario nobody predicted.
Stage 3 — The crisis structure
Name the team by role with at least two alternates each: crisis leader, operations, technology, communications, legal, human resources and safety, finance, facilities and security, and a scribe whose only job is to maintain a decision and action timeline. The scribe role is always omitted and always needed — the timeline is the record for the insurance claim, the regulator, and the after-action review.
State the decision authority explicitly. What can the crisis leader decide alone? What requires the CEO or the board? What emergency spending authority exists without normal approval? An organization that must convene a committee to authorize a $50,000 expenditure at 2 a.m. does not have a plan.
Define activation levels and who can declare each — localized, significant, and enterprise.
Establish the communication channel — a bridge line, a chat channel, and a location, with the details on a card people carry, because the intranet may be the thing that is down.
Define the battle rhythm for sustained events: status call intervals, a standard reporting format, a shared situation log, and shift handoffs. Plan the handoff before anyone has been awake for eighteen hours, because sustained events are lost in hour twenty rather than hour two.
Stage 4 — Communications
Employees first. They will hear about it regardless, they are the most credible source for everyone else, and they need to know what to do. Maintain an out-of-band contact list — personal phone and email, kept current — and deploy and test a mass notification tool.
Customers — identify who communicates with each account, prepare holding statements, and know the contractual obligations below.
Suppliers, insurers, lenders, and regulators, each with its own trigger.
The public and the media — one spokesperson, a prepared holding statement, and a status page hosted independently of the company's own infrastructure. Say only what you know, do not speculate about cause, do not give a restoration estimate you cannot support, and update on a schedule even when there is nothing new. Silence reads as concealment, and an estimate that slips three times destroys credibility.
Pre-approve templates with legal, so review during an event is a check rather than a drafting exercise.
Stage 5 — The legal layer
Extract the contractual notice obligations in advance from every customer, supplier, lender, and lease agreement into a one-page schedule kept with the plan — deadlines, required content, and method. Several agreements require notice of a disruption within hours, and companies routinely discover this the following week.
Read the force majeure clauses in both directions. They are not standard: the enumerated events differ, most condition relief on prompt written notice, most impose a mitigation obligation, and many give the other party a termination right after a defined period. Invoking it may cost the contract.
Map the regulatory reporting clocks that a disruption or an incident triggers — cyber incident reporting measured in hours in several regimes, data breach notification with state-specific deadlines, securities disclosure of material events, health and safety reporting, environmental release notification, and sector-specific requirements for financial institutions, health care providers, utilities, and government contractors.
Litigation and evidence. Where claims are foreseeable, issue a litigation hold immediately and preserve the systems, logs, and communications that recovery would otherwise overwrite. Recovery routinely destroys the evidence needed to establish what happened, and forensic imaging before restoration is frequently the right sequence even though it delays recovery. Decide that tradeoff deliberately with counsel.
Privilege. Structure any investigation likely to produce claims under counsel from the beginning, and keep the operational after-action review separate from the privileged legal analysis. Both are worth having.
Stage 6 — Technology recovery
Backups on a schedule matched to the RPO, with offline or immutable copies, because ransomware encrypts connected backups first. The working discipline is three copies, two media, one off-site and one offline or immutable.
Restoration tested and timed at least annually. A backup that has never been restored is a hypothesis, and the most common finding in a real event is that restoration takes far longer than anyone assumed.
Failover capability for critical systems, exercised. Documented recovery runbooks with credentials accessible to more than one person and stored where they can be reached when the primary systems are down. Vendor escalation contacts and contract support terms, current.
A one-page manual workaround per critical process. What does fulfillment do without the warehouse system? What does billing do without the ERP? Cheap, and it is what keeps the business moving.
Stage 7 — Facilities and people
Alternate work locations, which for many businesses now means remote work — with equipment and network capacity tested rather than assumed.
Alternate production or fulfillment capacity, including reciprocal arrangements and contract capacity, and equipment replacement sources with lead times.
Utilities — generators, fuel contracts, and how long they last.
Safety first, always: evacuation and shelter procedures, headcount accountability, and support for affected employees.
Cross-training for the single-person dependencies the map identified, and succession for key roles with delegated signing authority documented.
Payroll continuity — how people get paid if the payroll system or the office is unavailable. Frequently overlooked, and the thing employees care about most in a prolonged event.
Stage 8 — Supply chain
Alternate suppliers identified and, ideally, qualified in advance, because qualifying one during a disruption takes weeks the business does not have.
Safety stock for critical inputs sized against realistic replacement lead times, not against normal ones.
Visibility into subtier suppliers, at least for components with no alternative.
Contract terms — force majeure, allocation provisions, and whether the supplier may prioritize other customers, which is what actually happens in a broad disruption.
Stage 9 — Insurance
Business interruption responds to loss of income from a direct physical loss to covered property. Two consequences surprise people: a cyber event or a vendor outage with no physical damage generally does not trigger it, and a communicable disease event generally does not either.
Check: the period of restoration and any extended period of indemnity; the waiting period; contingent business interruption for supplier and customer disruption, and whether unnamed suppliers are covered; civil authority and ingress/egress; service interruption, which frequently requires damage to the utility's property; extra expense; and whether the limit came from a worksheet rather than from a percentage of revenue.
Cyber insurance covers what business interruption does not — network interruption from a security event, dependent network interruption, data restoration, forensics, notification, and extortion. Read the waiting period, the indemnity period, the sublimits, and any mandatory vendor panel, which is worth knowing before an incident rather than during one.
Assign someone to track incremental costs and lost revenue from hour one, in the format the proof of loss requires, because BI claims are frequently underpaid for lack of substantiation.
Agree the ransom decision framework in advance — backups, exfiltration, sanctions legality, and insurer requirements — because it cannot be reasoned through calmly at hour four.
Stage 10 — Testing
Plan review annually and after any material change to the business, systems, or vendors.
Tabletop exercises every six months — two to four hours, a scenario, a facilitator injecting complications. This is the highest-return activity in the program relative to its cost. Vary the scenario and vary who is in the room, including a scenario in which the primary decision-makers are unavailable.
Functional tests annually: restore a backup and time it, fail over a system, activate the notification tool, or run a shift on the manual workaround.
After-action review after every exercise and every real event, with specific assigned actions, owners, and dates, checked at the next review. A review that produces no assigned actions is a meeting.
The most common findings, in order: stale contact lists; no one knows who declares an incident; restoration takes far longer than assumed; the manual workaround does not exist; the contractual notice obligations were never extracted; and the plan is stored somewhere unavailable during the event.
Stage 11 — Keeping it usable
Format matters. A sixty-page document is not usable at 2 a.m. Build a one-page activation card, role-specific checklists with the first ten actions, effect-based playbooks of a few pages, and a separate reference document holding the dependency maps, contact lists, contract schedules, vendor terms, and insurance summary.
Store offline copies and printed copies for the crisis team.
Assign an owner with a review calendar; programs without an owner degrade within eighteen months.
Integrate with vendor management, so a new critical vendor triggers a dependency assessment, and with onboarding, so new team members learn their role.
Keep it proportionate. A thirty-person company needs a dependency map, a one-page activation card, tested backups, an out-of-band contact list, a schedule of contractual notice obligations, and a tabletop once a year. That is a day of work annually and it addresses the great majority of the risk.
Stage 12 — The first four hours, and the questions companies ask
Minutes 0–15. Someone declares — the step most often delayed. Convene on the bridge line. Confirm safety of people first.
Minutes 15–60. Establish what is known and explicitly what is not. Start the timeline. Assign a single owner for technical recovery, customer communication, employee communication, vendor escalation, and legal and insurance. Set the next status call and hold it.
Hour 1–2. Notify insurers under every potentially applicable policy. Pull the contractual notice schedule and identify what is due when. Issue the employee communication. Post a holding statement. Begin tracking incremental costs.
Hour 2–4. Assess regulatory clocks. Decide whether a litigation hold is warranted and issue it before recovery overwrites evidence. Communicate a restoration estimate only when it can be supported. Escalate to the board if the criteria are met.
"How long should the plan be?" The usable part is a page per role plus four to six playbooks. Everything else is reference material.
"Does our customer's vendor questionnaire count as a plan?" No. It documents that a plan exists, and customers increasingly ask for evidence of testing rather than for the document.
"Does business interruption insurance cover a cyber outage?" Generally not; it requires direct physical loss. That gap should be identified deliberately and either accepted or insured.
"What is the minimum viable program?" A dependency map, an activation card with out-of-band contacts, tested backups, a contractual notice schedule, and one tabletop a year.
Master resource index
Articles
- Business Insurance and Coverage Disputes: CGL, E&O, Cyber, and D&O
- Premises Liability for Property Owners and Businesses
- Transportation and Logistics Law: The Carmack Amendment, Broker Liability, and FMCSA Compliance
Guides
- Preparing a Business Continuity and Crisis Management Plan
- Handling a Product Recall: A Practical Guide
- Handling an Insurance Claim After a Property Loss
- Negotiating a Commercial Insurance Program: A Practical Guide
Checklists
- Business Continuity and Crisis Response Checklist
- Insurance Program Review Checklist
- Vendor Cybersecurity Diligence Checklist
- Litigation Hold and Evidence Preservation Checklist
Related toolkits
- Cybersecurity Program Toolkit
- Data Breach and Incident Response Toolkit
- Product Safety and Recall Toolkit
- Insurance Coverage Toolkit
External and primary sources
- ISO 22301 (business continuity management systems); NIST SP 800-34 (contingency planning); NIST SP 800-61 (incident handling)
- HIPAA Security Rule contingency plan standard, 45 C.F.R. § 164.308(a)(7); GLBA Safeguards Rule, 16 C.F.R. Part 314
- SEC cybersecurity incident disclosure requirements for public companies; sector-specific incident reporting rules; state data breach notification statutes
- NAIC Insurance Data Security Model Law as enacted, including its 72-hour notification requirement
This toolkit is educational and not legal advice. Regulatory reporting obligations, contractual notice requirements, and insurance coverage vary substantially by industry, jurisdiction, and policy language. Consult qualified counsel and a licensed broker when building or activating a plan.