Operational Resilience: Mapping Services and Impact Tolerances
Deciding How Much Disruption Your Customers Can Actually Absorb
Operational resilience programmes usually begin with a list of systems and a set of recovery time objectives that were agreed years ago by people who no longer work there. The numbers look precise, they were never derived from customer harm, and nobody has tested whether the platform can meet them.
The regulatory shift over the last few years asks a different question. Not how quickly can you restore a database, but how long can a customer be unable to make a payment before real harm occurs, and can you prove you would stay inside that limit under a severe but plausible disruption. Setting an operational resilience impact tolerance you can defend means starting from the outside and working in, which is the opposite of how most technology organisations think.
What is a critical business service?
Something the customer or the market receives, defined by outcome rather than by system or department.
The distinction sounds semantic and it changes everything. Payments processing is a system. Being able to make a payment is a service. Once the service is defined that way, its dependencies cross every internal boundary, which is exactly the point: disruption reaches customers through whichever dependency fails, not only through the one your team owns.
Why does the service definition decide everything downstream?
Because mapping, tolerance, testing, and investment all inherit its scope.
Define the service too narrowly and you exclude the dependency that will actually cause the outage. Define it too broadly and the tolerance becomes meaningless because it aggregates unrelated failure modes. Get it wrong and every artefact built on it inherits the error, which is why the definition deserves more senior attention than it usually receives. The Basel Committee's Principles for operational resilience, published in March 2021, take a principles-based approach precisely so institutions can apply this thinking to their own structures rather than following a template.
How many services should a bank have?
Usually between ten and thirty, depending on size and complexity.
Fewer than ten and the tolerances become so broad that nothing fails them. More than fifty and the mapping and testing burden makes the programme unaffordable, while genuinely critical services get lost among marginal ones. Draw the line at services whose disruption would cause intolerable harm to customers, threaten market integrity, or breach a regulatory obligation, and be prepared to defend the exclusions as well as the inclusions.
Is your resilience programme organised around systems or around what customers receive?
What exactly is an impact tolerance?
The maximum disruption a service can sustain before causing intolerable harm, expressed from the outside in.
It is not an internal target and not an aspiration. It is a limit, set by reference to customer and market harm, that the institution commits to staying within under severe but plausible conditions. That makes it a board-level statement about risk appetite with direct engineering consequences.
How is it different from a recovery time objective?
A tolerance describes tolerable harm, while an RTO describes an internal restoration target.
| Aspect | Impact tolerance | Recovery time objective |
|---|---|---|
| Perspective | Customer and market | Internal technical |
| Scope | End-to-end business service | Component or system |
| Set by | Business owner, board approved | Technology and operations |
| Meaning | Limit that must not be breached | Target to aim for |
| Consequence of breach | Governance and supervisory issue | Missed internal objective |
The relationship runs one way. Tolerances constrain the RTOs of every component in the chain, and if the arithmetic of component RTOs cannot deliver the service tolerance, you have a gap to remediate rather than a tolerance to relax quietly.
What should the tolerance be measured in?
Whatever expresses harm for that service, which is often time and not always.
Elapsed time is the default and it fits most access and payment services. Some services need volume, such as the number of payments not processed, or value, or the number of customers affected. Others need more than one dimension, since a short outage affecting everyone and a long outage affecting a few are different harms. Choose deliberately and state the metric, because a tolerance expressed in the wrong unit produces testing that measures the wrong thing.
How do you set a tolerance you can defend?
From evidence about harm, owned by the business, approved by the board, and revisited after testing.
Start with what actually happens to customers as disruption extends: missed obligations, unpaid bills, inability to access funds, market positions unhedged. Use complaint data, past incident experience, and payment timing obligations as evidence rather than judgment alone. Then have the accountable business owner set the tolerance and the board approve it, since it is a statement about acceptable harm rather than a technical parameter. Technology's job is to say what is currently achievable and what closing the gap would cost.
Why are first-pass tolerances usually wrong?
Because they are set optimistically and then contradicted by the first honest test.
Most institutions set generous initial tolerances, discover under testing that the platform cannot hold them, and then face a choice between remediation and revision. That is a healthy process if it is transparent and a governance problem if the tolerance quietly moves to match the tested result. Record the original tolerance, the tested outcome, and the decision, so the trail shows judgment rather than accommodation. Metrics that boards can actually use are discussed well in operational resilience metrics for boards.
How should service mapping be done?
Across every layer the service depends on, derived from live data rather than drawn by hand.
| Layer | What to capture | Common gap |
|---|---|---|
| People | Roles and skills required, including out of hours | Single individuals with unique knowledge |
| Process | Manual steps, approvals, workarounds | Undocumented workarounds that carry the service |
| Technology | Applications, platforms, data stores, networks | Shared components under multiple services |
| Data | Datasets required, their currency and integrity | Data quality dependencies nobody mapped |
| Facilities | Locations, connectivity, physical access | Assumption that remote work always substitutes |
| Third parties | Providers and their own critical dependencies | Nth-party layers with no visibility |
How deep should mapping go?
Deep enough to find the single points of failure, which usually means further than the first pass.
Stop when you have identified every dependency whose failure could breach the tolerance, not when the diagram looks complete. That typically means going below the obvious application layer into shared platform services, identity, network paths, and provider dependencies. Shared components are where the surprises are, because a queue service or an identity provider under several critical services concentrates risk invisibly, which is the same discovery problem addressed in third-party and concentration risk.
How do you keep the map current?
By deriving it from telemetry and deployment metadata rather than maintaining documents.
A hand-drawn map is stale within a quarter and misleading within two. Require deployed components to declare the service they support, derive dependency edges from observed calls and infrastructure definitions, and reconcile the derived map against the documented one so drift becomes a work item. That is the shift from static maps to live dependency data described in moving from critical-process maps to live dependencies, and it depends on the instrumentation covered in this guide to observability for core systems.
Could you produce today's dependency map for a critical service without a workshop?
Talk to Digiqt about deriving service maps from live telemetry
How do you test against tolerance?
With severe but plausible scenarios that name a real failure and measure the service outcome.
| Scenario | Tests | Frequency |
|---|---|---|
| Regional infrastructure failure | Failover capability against tolerance | Annually per critical service |
| Critical third-party outage | Substitutability and degraded mode | Annually |
| Data corruption requiring restore | Recovery integrity and elapsed time | Annually |
| Cyber incident with untrusted environment | Rebuild rather than restore | Annually |
| Key person unavailability | People dependency | Periodically |
| Simultaneous partial degradation | Combined effects and detection | Periodically |
What counts as severe but plausible?
Extreme enough to challenge the service, grounded in something that has happened somewhere.
The test of plausibility is whether you can point to a comparable event in the industry, and the test of severity is whether the scenario could actually breach the tolerance. Scenarios that neither stretch the platform nor could ever occur are wasted cycles. CPMI-IOSCO's guidance on cyber resilience for financial market infrastructures, published in June 2016, frames the expectation as anticipating threats, responding rapidly, and achieving faster and safer target recovery objectives, which is a useful reminder that recovery capability is the thing being tested rather than the incident narrative.
What happens when you cannot stay inside tolerance?
You record the gap, remediate on a dated plan, or revise the tolerance with stated justification.
All three are legitimate. What is not legitimate is discovering the gap and neither fixing nor disclosing it. Log it with an owner, an investment estimate, and a target date, report it through governance, and track closure. Where the required investment is disproportionate, the honest answer may be a revised tolerance with compensating measures, and a supervisor will generally accept a reasoned position far more readily than an unevidenced one.
How does this connect to investment decisions?
It converts resilience spending from a technical argument into a business one.
This is the underappreciated benefit of the whole exercise. Once a tolerance is board approved and testing shows a gap, the investment case writes itself in terms the executive committee already accepted, which is a far stronger position than requesting redundancy funding on architectural grounds. It also works in reverse: services with generous tolerances and adequate current capability should not receive further hardening, which lets you redirect spending honestly. Where the gap is about failover capability specifically, the engineering choices and their costs are set out in multi-region failover design and zero data loss.
What governance and reporting is expected?
A documented self-assessment, board approval of services and tolerances, and evidence of testing and remediation.
Maintain a self-assessment that identifies the critical services, states the tolerances and their basis, describes the mapping, records scenario testing and results, lists gaps with remediation plans, and shows board approval. Keep it current rather than annual, since the questions arrive when an incident happens rather than on your reporting cycle. Third-party dependencies belong inside it, using the tools set out in the FSB's December 2023 toolkit for identifying critical third-party services and monitoring systemic dependencies, and resolution-related capabilities should be consistent with what you tell resolution authorities, which is covered in recovery and resolution technology capabilities.
How should delivery be phased?
Services first, then tolerances, then mapping, then testing, then remediation.
| Phase | Duration | Deliverable |
|---|---|---|
| Service identification | 1 to 2 months | Ten to thirty services defined by outcome, with owners |
| Tolerance setting | 1 to 2 months | Metric, value, evidence, and board approval per service |
| Mapping | 3 to 5 months | Dependencies across all layers, derived from live data |
| Scenario design and testing | 2 to 4 months | Severe but plausible scenarios executed and measured |
| Gap remediation | Ongoing | Dated plans, owners, investment cases |
| Self-assessment and reporting | Continuous | Living document, board review cycle |
Do not start with mapping, tempting though it is for a technology team. Mapping without agreed services produces an enormous diagram nobody can act on, whereas mapping scoped by ten well-defined services is tractable and immediately useful.
Which metrics prove the programme is real?
Tolerance coverage, tested coverage, gaps and their ageing, map freshness, and breaches in live incidents.
Report the share of critical services with board-approved tolerances and a stated metric. Report tested coverage separately, since an untested tolerance is an assertion. Track open gaps, their remediation dates, and their ageing, because ageing gaps are the clearest sign a programme has stalled. Measure map freshness as the share of dependencies derived automatically within a recent window. And record actual tolerance breaches during real incidents, which is the only unambiguous evidence of whether the tolerances and the platform agree.
Operational resilience work earns its keep when it changes what the institution builds and funds. If the tolerances are approved, the testing is honest, and the gaps are visible with owners and dates, the programme is doing its job. If it produces a self-assessment nobody has read since sign-off, it is documentation with a resilience label on it.
Frequently Asked Questions
What is a critical business service?
Something a customer or the market receives, such as making a payment or accessing funds, rather than a system or a department. The outcome defines it, not the technology delivering it.
How is an impact tolerance different from a recovery time objective?
A tolerance is the maximum disruption the outside world can absorb before intolerable harm. An RTO is an internal target for restoring a component. Tolerances constrain RTOs, not the reverse.
How many critical business services should a bank have?
Usually between ten and thirty. Too few makes tolerances meaningless, and too many makes mapping and testing unaffordable while obscuring what genuinely matters.
What should a tolerance be measured in?
Whatever expresses harm for that service, commonly elapsed time, but sometimes transaction volume, value, or number of customers affected. Some services need more than one dimension.
Who should set the tolerance?
The business owner accountable for the service, with evidence about customer harm, approved at board level. Technology sets what is achievable, not what is tolerable.
What does severe but plausible mean?
A scenario extreme enough to challenge the service but grounded in real possibility, such as a region failure, a critical provider outage, or data corruption requiring restore.
What happens when testing shows you cannot stay inside tolerance?
You record the gap, remediate with a plan and a date, or change the tolerance with justification. The gap is the finding, and hiding it is what turns a technical issue into a governance failure.
How do you keep service maps current?
Derive them from live telemetry and deployment metadata rather than maintaining diagrams, because a documented map is out of date within a quarter of being signed off.



