The wrong way to build a business continuity plan is to start with backup technology and ask how quickly IT can restore it. The better starting point is a business decision: how much interruption, transaction loss and service degradation can the organisation tolerate before the impact becomes unacceptable?

Once leadership answers that question, recovery architecture becomes easier to design. RTO, RPO, alternative processing, disaster recovery and resilience testing should all follow the acceptable-loss decision rather than define it.

The Two Numbers Every Plan Turns On: RTO and RPO

RTO and RPO are often treated as technical settings, but they represent different forms of business loss. They should therefore be approved by the process owner and risk leadership before IT selects the recovery method.

  • Recovery Time Objective: RTO is the target time within which a process, application or service must be restored after disruption. It answers: how long can the business operate without this capability before the impact becomes unacceptable?

  • Recovery Point Objective: RPO is the maximum acceptable amount of data loss measured backwards in time. It answers: how much recent transaction or operational data can the business afford to lose and reconstruct?

  • Maximum Acceptable Outage: MAO is the outer business limit beyond which continued disruption causes unacceptable consequences. It should sit above the recovery objective rather than being confused with the target IT restoration time.

  • Minimum Service Level: Some processes do not need full restoration immediately. The business may accept a temporary manual or degraded operating mode if it preserves critical outcomes until full service returns.

RTO and RPO Should Not Automatically Match

A service may require restoration within one hour but tolerate only a few minutes of lost transactions. Another process may need to be available within four hours but can be reconstructed from the previous night’s data.

Those differences create very different infrastructure requirements. A low RPO may require continuous replication, while a low RTO may require pre-provisioned capacity, automated failover and tested application dependencies.

Do Not Let IT Set the Numbers Alone

Recovery objectives determine cost, architecture and operational complexity. If IT chooses a five-minute RPO because replication technology can deliver it, the organisation may spend heavily to protect data that the process owner could recreate without material impact.

Conversely, accepting a long recovery period because the infrastructure is difficult to restore can hide a business risk rather than resolve it. Where security, resilience and operating risk need to be assessed together, IT security consulting should facilitate the decision rather than start by recommending a recovery product.

Business Impact Analysis Without Guesswork

A useful business impact analysis does not ask every department whether its processes are “critical”. Most departments will say yes. It identifies what happens as disruption continues and establishes when the consequences cross an agreed threshold.

Measure Impact Over Time

  • Customer impact: Determine when customers lose access, miss contractual service levels or begin abandoning the service.

  • Financial impact: Estimate transaction loss, revenue interruption, recovery cost, contractual penalties and the effect of delayed settlement or billing.

  • Regulatory impact: Identify deadlines, reporting obligations or critical services where prolonged interruption can create non-compliance.

  • Operational impact: Assess backlog growth, manual processing capacity and dependencies on staff, facilities or suppliers.

  • Data impact: Determine which lost records can be recreated and which transactions would be impossible or risky to reconstruct.

  • Safety and public-service impact: For critical infrastructure, government and operational technology environments, service interruption may create effects beyond financial loss.

Use Time Bands Instead of One Impact Score

A process that causes little damage after 30 minutes may create a severe problem after eight hours. BIA workshops should therefore assess impact over several disruption periods rather than assign one static “high, medium or low” criticality score.

This also exposes natural priorities. A process whose impact becomes unacceptable after two hours belongs ahead of one that remains tolerable for a full business day, even if both are labelled critical.

Map Dependencies Before Setting Recovery Targets

Business recovery is often slower than the application recovery target because the process depends on identity services, networks, payment gateways, external data, integration middleware, specialised staff or a third party.

The BIA should therefore identify upstream and downstream dependencies before approving RTO and RPO. Otherwise, the organisation can successfully recover one application while the business service remains unusable.

Setting Recovery Objectives Per Process

A business continuity plan should not apply one RTO or RPO to the entire company. Recovery objectives should be assigned at process or service level and then translated into application, infrastructure and data requirements.

Start With the Business Process

  1. Define the service: Identify the customer or operational outcome that must continue, rather than beginning with an application name.

  2. Identify the disruption limit: Establish when financial, customer, regulatory or operational damage becomes unacceptable.

  3. Set the RTO below that limit: Leave enough margin for escalation, validation and transition back into normal operations.

  4. Set the RPO from data tolerance: Determine how many minutes or hours of activity can realistically be reconstructed without creating control or customer problems.

  5. Define the minimum operating mode: Decide whether the process can operate manually, with reduced capacity or through an alternative channel during recovery.

Challenge Unrealistic Objectives

Every process owner may want zero downtime and zero data loss, but those targets imply high-cost architecture and extensive operational controls. A useful workshop asks what would actually happen if recovery took 15 minutes, two hours or one day, and what additional protection is justified by the difference.

The objective should be the minimum recovery level that keeps business impact within tolerance, not the fastest number technology can theoretically deliver.

Architecture Choices That Follow From the Objectives

Once RTO and RPO are approved, the architecture can be evaluated against them. The recovery technology should demonstrate that it can meet the business objective under realistic failure conditions.

Backup Frequency Follows the RPO

  • Longer RPO: Periodic backups may be sufficient where several hours of recent data can be recreated safely.

  • Short RPO: Frequent snapshots, log shipping or continuous replication may be required where transaction loss must be tightly limited.

  • Near-zero RPO: Synchronous or highly frequent replication may be necessary, but the design must consider corruption and cyberattack replication as well as hardware failure.

Recovery Topology Follows the RTO

  • Cold recovery: Lower standing cost but longer restoration time because infrastructure and services need substantial preparation after disruption.

  • Warm standby: Some infrastructure and data remain ready, reducing restoration time while avoiding the full operating cost of a continuously active environment.

  • Hot or active recovery: Infrastructure runs continuously and can support rapid failover, but it introduces greater cost, complexity and synchronisation requirements.

Backups Are Not a Disaster Recovery Plan

A backup proves that another copy of data exists. A disaster recovery capability proves that the organisation can restore the application, infrastructure, configuration, identity dependencies and data into a usable service within the approved time.

NCA’s Essential Cybersecurity Controls require backup and recovery requirements to include periodic restoration testing, not simply scheduled backup jobs.

Integration Dependencies Need Their Own Recovery Design

Modern enterprise services usually depend on several connected systems. Recovering ERP without identity, API management, middleware or payment connectivity may satisfy a server-level objective while failing the business objective.

Where continuity depends on several applications and integration paths, enterprise systems integration should map recovery sequencing and dependencies as part of the architecture rather than after the DR environment has already been selected.

The same logic applies when systems move to hosted or hybrid environments. A cloud migration strategy should define resilience assumptions, region or zone dependencies, backup ownership and recovery responsibilities instead of assuming that cloud hosting automatically provides business continuity.

Testing: Tabletop, Partial and Full

A continuity plan that has never been exercised is an assumption. Testing should progress from decision-making exercises to technical restoration and finally to realistic service recovery where the risk profile justifies it.

Tabletop Exercises Test Decisions

A tabletop exercise walks business, technology, cyber, legal, communications and management teams through a disruption scenario without actually taking production systems offline.

  • Test escalation: Confirm who can declare a disaster and activate the plan.

  • Test decision rights: Check who chooses between failover, manual operations and continued outage.

  • Test communication: Validate employee, customer, regulator and third-party communication paths.

  • Test assumptions: Challenge contact lists, supplier availability, alternate locations and dependence on unavailable staff.

Partial Tests Prove Components

Partial testing can restore selected databases, applications, infrastructure or processes without performing a complete enterprise failover. It provides useful technical evidence while limiting operational risk.

These exercises should still measure elapsed recovery time and restored data point against the approved RTO and RPO. “The restore worked” is not enough if it took twice the permitted time.

Full Tests Prove the Operating Model

A full resilience test validates the complete chain from incident declaration through failover, service validation, business operation and eventual return to normal service.

Not every process needs a full production-scale test at the same frequency. Test depth should follow criticality, regulatory expectations, architecture changes and lessons from previous exercises.

Regulatory Expectations in Saudi Arabia

Saudi entities should not treat business continuity as an isolated technology standard. Cybersecurity, sector regulation, cloud requirements and critical-system obligations can all affect the resilience design.

NCA Essential Cybersecurity Controls

The current ECC 2-2024 contains a dedicated Cybersecurity Resilience domain. It requires entities within scope to identify, document, approve and implement cybersecurity requirements for business continuity management.

At a minimum, the controls require continuity of cybersecurity systems and procedures, response plans for cyber incidents that may affect business continuity and disaster recovery plans. The requirements must also be reviewed periodically.

For teams mapping a continuity programme into the wider control environment, the related guide to NCA essential cybersecurity controls should be used to connect BCP and disaster recovery evidence to the broader NCA compliance model.

Critical Systems Require Deeper Resilience

NCA guidance for critical systems goes further. It includes development of disaster recovery capability for critical systems and requires the organisation to define outage and data-loss limits as part of the recovery design.

Operational technology controls also require BCP, BIA, RTO and RPO considerations and periodic testing or simulation exercises such as tabletop exercises.

SAMA-Regulated Organisations Have Additional BCM Requirements

For organisations regulated by the Saudi Central Bank, SAMA’s Business Continuity Management Framework explicitly links continuity planning to predetermined RTO, RPO and MAO values. It also expects critical technology infrastructure to have an approved IT disaster recovery plan aligned with the BIA.

SAMA also expects relevant key service providers supporting critical activities to maintain continuity plans and have those plans tested at least annually. This means supplier resilience cannot be excluded simply because the service is contracted out.

The same concern applies outside financial services whenever a supplier supports a critical process. A structured third party risk management process should capture supplier recovery objectives, testing evidence and escalation contacts rather than stopping at contract and cybersecurity questionnaires.

Plan Maintenance and Ownership

A continuity plan becomes stale when it is treated as a document owned only by risk or IT. Recovery assumptions change whenever the organisation adds applications, vendors, locations, interfaces, employees or new regulatory obligations.

Assign Business Ownership

  • Process owner: Approves the impact tolerance, minimum operating level and business recovery priority.

  • Technology owner: Demonstrates that applications, infrastructure and data can meet the approved objectives.

  • Risk or BCM owner: Maintains the methodology, coordinates testing and challenges inconsistent recovery targets.

  • Cybersecurity owner: Ensures cyber incidents, compromised credentials, ransomware and security-control dependencies are included in recovery scenarios.

  • Third-party owner: Maintains evidence that critical suppliers can support the required recovery model.

Trigger Reviews When the Environment Changes

Do not wait for an annual review if a major ERP replacement, cloud migration, acquisition, outsourcing arrangement or critical integration changes the process architecture.

The decision rights around those changes should sit within an it governance framework that requires business continuity impact to be considered before major technology changes are approved.

Where ownership is fragmented across business, technology, cybersecurity and risk, IT governance advisory can help define who sets recovery tolerances, who accepts residual resilience risk and who is accountable for closing failed-test findings.

Business Continuity Plan Readiness Checklist

Use this checklist to test whether the organisation has a recovery capability or simply a collection of backup and policy documents.

  1. Define critical services: Identify business outcomes that must continue and map the processes, applications, people, facilities and third parties supporting them.

  2. Measure disruption impact: Assess financial, customer, regulatory, operational and data consequences over several outage durations rather than assigning a single criticality score.

  3. Approve RTO and RPO: Make the business owner approve the acceptable interruption and data-loss targets before technology selects the recovery architecture.

  4. Set minimum operations: Define what degraded or manual service can operate before full technology restoration is complete.

  5. Map dependencies: Identify identity, network, integration, data, supplier and infrastructure dependencies needed for the end-to-end service.

  6. Match architecture to objectives: Confirm that replication, backups, standby environments and failover designs can meet the approved targets.

  7. Test restoration: Restore real systems and data periodically and measure actual recovery performance rather than relying on backup-job success.

  8. Exercise decision-making: Run tabletop scenarios covering cyberattack, infrastructure failure, supplier outage and facility loss.

  9. Test third parties: Obtain evidence that critical providers understand and can support the organisation’s recovery requirements.

  10. Record recovery gaps: Treat missed RTOs, incomplete restores and failed communications as remediation items with named owners and deadlines.

  11. Review after change: Reassess the BIA and recovery assumptions after material system, supplier, process or organisational changes.

  12. Retain test evidence: Keep approvals, exercise results, recovery timings, lessons learned and remediation records for management and regulatory review.

If this checklist exposes uncertainty about acceptable loss, architecture or testing evidence, use our five-stage methodology to separate diagnosis, risk decisions, solution design and implementation rather than beginning with a disaster recovery product.

For early planning, published consulting cost ranges can also help establish whether the requirement is a focused continuity assessment, a wider governance programme or a technical recovery implementation before detailed procurement begins.

Make Acceptable Loss the First Business Continuity Decision

A useful business continuity plan does not promise that nothing will fail. It states how much outage and data loss the organisation accepts, identifies the services that must recover first and proves that the architecture, people and suppliers can meet those limits.

Start with the BIA and force each critical process to defend its RTO, RPO and minimum service level. Then build and test the recovery design against those decisions. When the acceptable-loss boundary is explicit, resilience investment becomes a business choice that leadership can challenge and approve rather than a collection of backup technologies IT hopes will be sufficient during a crisis.

FAQ about business continuity plan 

What is RTO in a business continuity plan?

Recovery Time Objective, or RTO, is the target time within which a business process, application or service should be restored after a disruption. The value should come from the business impact analysis and sit below the point where continued interruption becomes unacceptable. It should not simply reflect how quickly the current infrastructure happens to recover.

What is RPO in disaster recovery?

Recovery Point Objective, or RPO, defines the maximum acceptable amount of data loss measured backwards from the disruption. An RPO of several hours may support periodic backups, while a much shorter RPO can require frequent or continuous replication. The correct target depends on whether lost transactions can be reconstructed safely and economically.

What is the difference between a BCP and a disaster recovery plan?

A business continuity plan explains how critical business services continue during and after disruption, including people, facilities, suppliers, manual processes and technology. A disaster recovery plan focuses more specifically on restoring technology services, infrastructure, applications and data. The DRP should therefore support the wider recovery objectives defined by the BCP and business impact analysis.

How often should a business continuity plan be tested?

Testing frequency should follow service criticality, regulatory requirements, architecture changes and previous test results. Organisations should combine tabletop exercises with technical recovery and restoration tests. SAMA-regulated entities also need to consider specific sector requirements, including expectations around critical service providers. A plan should also be retested after material technology or operating-model changes.

What does NCA require for business continuity and disaster recovery?

NCA’s ECC 2-2024 requires entities within scope to define, document, approve and implement cybersecurity requirements for business continuity management. At minimum, those requirements include continuity of cybersecurity systems and procedures, response plans for cybersecurity incidents that could affect business continuity and disaster recovery plans. The implementation must also be reviewed periodically.