Technology

Operational Resilience: Mapping Services and Impact Tolerances

|Posted by Hitul Mistry / 31 Aug 26

Deciding How Much Disruption Your Customers Can Actually Absorb

Operational resilience programmes usually begin with a list of systems and a set of recovery time objectives that were agreed years ago by people who no longer work there. The numbers look precise, they were never derived from customer harm, and nobody has tested whether the platform can meet them.

The regulatory shift over the last few years asks a different question. Not how quickly can you restore a database, but how long can a customer be unable to make a payment before real harm occurs, and can you prove you would stay inside that limit under a severe but plausible disruption. Setting an operational resilience impact tolerance you can defend means starting from the outside and working in, which is the opposite of how most technology organisations think.

What is a critical business service?

Something the customer or the market receives, defined by outcome rather than by system or department.

The distinction sounds semantic and it changes everything. Payments processing is a system. Being able to make a payment is a service. Once the service is defined that way, its dependencies cross every internal boundary, which is exactly the point: disruption reaches customers through whichever dependency fails, not only through the one your team owns.

Why does the service definition decide everything downstream?

Because mapping, tolerance, testing, and investment all inherit its scope.

Define the service too narrowly and you exclude the dependency that will actually cause the outage. Define it too broadly and the tolerance becomes meaningless because it aggregates unrelated failure modes. Get it wrong and every artefact built on it inherits the error, which is why the definition deserves more senior attention than it usually receives. The Basel Committee's Principles for operational resilience, published in March 2021, take a principles-based approach precisely so institutions can apply this thinking to their own structures rather than following a template.

How many services should a bank have?

Usually between ten and thirty, depending on size and complexity.

Fewer than ten and the tolerances become so broad that nothing fails them. More than fifty and the mapping and testing burden makes the programme unaffordable, while genuinely critical services get lost among marginal ones. Draw the line at services whose disruption would cause intolerable harm to customers, threaten market integrity, or breach a regulatory obligation, and be prepared to defend the exclusions as well as the inclusions.

Is your resilience programme organised around systems or around what customers receive?

Talk to Digiqt about defining critical business services

What exactly is an impact tolerance?

The maximum disruption a service can sustain before causing intolerable harm, expressed from the outside in.

It is not an internal target and not an aspiration. It is a limit, set by reference to customer and market harm, that the institution commits to staying within under severe but plausible conditions. That makes it a board-level statement about risk appetite with direct engineering consequences.

How is it different from a recovery time objective?

A tolerance describes tolerable harm, while an RTO describes an internal restoration target.

AspectImpact toleranceRecovery time objective
PerspectiveCustomer and marketInternal technical
ScopeEnd-to-end business serviceComponent or system
Set byBusiness owner, board approvedTechnology and operations
MeaningLimit that must not be breachedTarget to aim for
Consequence of breachGovernance and supervisory issueMissed internal objective

The relationship runs one way. Tolerances constrain the RTOs of every component in the chain, and if the arithmetic of component RTOs cannot deliver the service tolerance, you have a gap to remediate rather than a tolerance to relax quietly.

What should the tolerance be measured in?

Whatever expresses harm for that service, which is often time and not always.

Elapsed time is the default and it fits most access and payment services. Some services need volume, such as the number of payments not processed, or value, or the number of customers affected. Others need more than one dimension, since a short outage affecting everyone and a long outage affecting a few are different harms. Choose deliberately and state the metric, because a tolerance expressed in the wrong unit produces testing that measures the wrong thing.

How do you set a tolerance you can defend?

From evidence about harm, owned by the business, approved by the board, and revisited after testing.

Start with what actually happens to customers as disruption extends: missed obligations, unpaid bills, inability to access funds, market positions unhedged. Use complaint data, past incident experience, and payment timing obligations as evidence rather than judgment alone. Then have the accountable business owner set the tolerance and the board approve it, since it is a statement about acceptable harm rather than a technical parameter. Technology's job is to say what is currently achievable and what closing the gap would cost.

Why are first-pass tolerances usually wrong?

Because they are set optimistically and then contradicted by the first honest test.

Most institutions set generous initial tolerances, discover under testing that the platform cannot hold them, and then face a choice between remediation and revision. That is a healthy process if it is transparent and a governance problem if the tolerance quietly moves to match the tested result. Record the original tolerance, the tested outcome, and the decision, so the trail shows judgment rather than accommodation. Metrics that boards can actually use are discussed well in operational resilience metrics for boards.

How should service mapping be done?

Across every layer the service depends on, derived from live data rather than drawn by hand.

LayerWhat to captureCommon gap
PeopleRoles and skills required, including out of hoursSingle individuals with unique knowledge
ProcessManual steps, approvals, workaroundsUndocumented workarounds that carry the service
TechnologyApplications, platforms, data stores, networksShared components under multiple services
DataDatasets required, their currency and integrityData quality dependencies nobody mapped
FacilitiesLocations, connectivity, physical accessAssumption that remote work always substitutes
Third partiesProviders and their own critical dependenciesNth-party layers with no visibility

How deep should mapping go?

Deep enough to find the single points of failure, which usually means further than the first pass.

Stop when you have identified every dependency whose failure could breach the tolerance, not when the diagram looks complete. That typically means going below the obvious application layer into shared platform services, identity, network paths, and provider dependencies. Shared components are where the surprises are, because a queue service or an identity provider under several critical services concentrates risk invisibly, which is the same discovery problem addressed in third-party and concentration risk.

How do you keep the map current?

By deriving it from telemetry and deployment metadata rather than maintaining documents.

A hand-drawn map is stale within a quarter and misleading within two. Require deployed components to declare the service they support, derive dependency edges from observed calls and infrastructure definitions, and reconcile the derived map against the documented one so drift becomes a work item. That is the shift from static maps to live dependency data described in moving from critical-process maps to live dependencies, and it depends on the instrumentation covered in this guide to observability for core systems.

Could you produce today's dependency map for a critical service without a workshop?

Talk to Digiqt about deriving service maps from live telemetry

How do you test against tolerance?

With severe but plausible scenarios that name a real failure and measure the service outcome.

ScenarioTestsFrequency
Regional infrastructure failureFailover capability against toleranceAnnually per critical service
Critical third-party outageSubstitutability and degraded modeAnnually
Data corruption requiring restoreRecovery integrity and elapsed timeAnnually
Cyber incident with untrusted environmentRebuild rather than restoreAnnually
Key person unavailabilityPeople dependencyPeriodically
Simultaneous partial degradationCombined effects and detectionPeriodically

What counts as severe but plausible?

Extreme enough to challenge the service, grounded in something that has happened somewhere.

The test of plausibility is whether you can point to a comparable event in the industry, and the test of severity is whether the scenario could actually breach the tolerance. Scenarios that neither stretch the platform nor could ever occur are wasted cycles. CPMI-IOSCO's guidance on cyber resilience for financial market infrastructures, published in June 2016, frames the expectation as anticipating threats, responding rapidly, and achieving faster and safer target recovery objectives, which is a useful reminder that recovery capability is the thing being tested rather than the incident narrative.

What happens when you cannot stay inside tolerance?

You record the gap, remediate on a dated plan, or revise the tolerance with stated justification.

All three are legitimate. What is not legitimate is discovering the gap and neither fixing nor disclosing it. Log it with an owner, an investment estimate, and a target date, report it through governance, and track closure. Where the required investment is disproportionate, the honest answer may be a revised tolerance with compensating measures, and a supervisor will generally accept a reasoned position far more readily than an unevidenced one.

How does this connect to investment decisions?

It converts resilience spending from a technical argument into a business one.

This is the underappreciated benefit of the whole exercise. Once a tolerance is board approved and testing shows a gap, the investment case writes itself in terms the executive committee already accepted, which is a far stronger position than requesting redundancy funding on architectural grounds. It also works in reverse: services with generous tolerances and adequate current capability should not receive further hardening, which lets you redirect spending honestly. Where the gap is about failover capability specifically, the engineering choices and their costs are set out in multi-region failover design and zero data loss.

What governance and reporting is expected?

A documented self-assessment, board approval of services and tolerances, and evidence of testing and remediation.

Maintain a self-assessment that identifies the critical services, states the tolerances and their basis, describes the mapping, records scenario testing and results, lists gaps with remediation plans, and shows board approval. Keep it current rather than annual, since the questions arrive when an incident happens rather than on your reporting cycle. Third-party dependencies belong inside it, using the tools set out in the FSB's December 2023 toolkit for identifying critical third-party services and monitoring systemic dependencies, and resolution-related capabilities should be consistent with what you tell resolution authorities, which is covered in recovery and resolution technology capabilities.

How should delivery be phased?

Services first, then tolerances, then mapping, then testing, then remediation.

PhaseDurationDeliverable
Service identification1 to 2 monthsTen to thirty services defined by outcome, with owners
Tolerance setting1 to 2 monthsMetric, value, evidence, and board approval per service
Mapping3 to 5 monthsDependencies across all layers, derived from live data
Scenario design and testing2 to 4 monthsSevere but plausible scenarios executed and measured
Gap remediationOngoingDated plans, owners, investment cases
Self-assessment and reportingContinuousLiving document, board review cycle

Do not start with mapping, tempting though it is for a technology team. Mapping without agreed services produces an enormous diagram nobody can act on, whereas mapping scoped by ten well-defined services is tractable and immediately useful.

Which metrics prove the programme is real?

Tolerance coverage, tested coverage, gaps and their ageing, map freshness, and breaches in live incidents.

Report the share of critical services with board-approved tolerances and a stated metric. Report tested coverage separately, since an untested tolerance is an assertion. Track open gaps, their remediation dates, and their ageing, because ageing gaps are the clearest sign a programme has stalled. Measure map freshness as the share of dependencies derived automatically within a recent window. And record actual tolerance breaches during real incidents, which is the only unambiguous evidence of whether the tolerances and the platform agree.

Operational resilience work earns its keep when it changes what the institution builds and funds. If the tolerances are approved, the testing is honest, and the gaps are visible with owners and dates, the programme is doing its job. If it produces a self-assessment nobody has read since sign-off, it is documentation with a resilience label on it.

Frequently Asked Questions

What is a critical business service?

Something a customer or the market receives, such as making a payment or accessing funds, rather than a system or a department. The outcome defines it, not the technology delivering it.

How is an impact tolerance different from a recovery time objective?

A tolerance is the maximum disruption the outside world can absorb before intolerable harm. An RTO is an internal target for restoring a component. Tolerances constrain RTOs, not the reverse.

How many critical business services should a bank have?

Usually between ten and thirty. Too few makes tolerances meaningless, and too many makes mapping and testing unaffordable while obscuring what genuinely matters.

What should a tolerance be measured in?

Whatever expresses harm for that service, commonly elapsed time, but sometimes transaction volume, value, or number of customers affected. Some services need more than one dimension.

Who should set the tolerance?

The business owner accountable for the service, with evidence about customer harm, approved at board level. Technology sets what is achievable, not what is tolerable.

What does severe but plausible mean?

A scenario extreme enough to challenge the service but grounded in real possibility, such as a region failure, a critical provider outage, or data corruption requiring restore.

What happens when testing shows you cannot stay inside tolerance?

You record the gap, remediate with a plan and a date, or change the tolerance with justification. The gap is the finding, and hiding it is what turns a technical issue into a governance failure.

How do you keep service maps current?

Derive them from live telemetry and deployment metadata rather than maintaining diagrams, because a documented map is out of date within a quarter of being signed off.

Sources

Read our latest blogs and research

Featured Resources

Technology

Automated Regression Testing for Core Banking and Payment Releases

A practical guide to core banking test automation: what actually needs regression coverage, how to test payment messages and date-dependent logic, automating the legacy layer, and keeping a suite people trust.

Read more
Technology

Third-Party and Concentration Risk in Financial Cloud Estates

How to manage third party concentration risk cloud dependencies create, covering dependency discovery, nth-party visibility, substitutability, ongoing monitoring, scenario testing, and board reporting.

Read more
Technology

Banking Multi-Region Failover With Zero Data Loss Requirements

How to design banking multi region failover RPO targets that survive physics, covering replication models, quorum and split-brain prevention, in-flight payment handling, failover testing, and real cost trade-offs.

Read more

About Us

We are a technology services company focused on enabling businesses to scale through AI-driven transformation. At the intersection of innovation, automation, and design, we help our clients rethink how technology can create real business value.

From AI-powered product development to intelligent automation and custom GenAI solutions, we bring deep technical expertise and a problem-solving mindset to every project. Whether you're a startup or an enterprise, we act as your technology partner, building scalable, future-ready solutions tailored to your industry.

Driven by curiosity and built on trust, we believe in turning complexity into clarity and ideas into impact.

Our key clients

Companies we are associated with

Life99
Edelweiss
Aura
Kotak Securities
Coverfox
Phyllo
Quantify Capital
ArtistOnGo
Unimon Energy

Our Offices

Ahmedabad

B-714, K P Epitome, near Dav International School, Makarba, Ahmedabad, Gujarat 380051

+91 99747 29554

Mumbai

C-20, G Block, WeWork, Enam Sambhav, Bandra-Kurla Complex, Mumbai, Maharashtra 400051

+91 99747 29554

Stockholm

Bäverbäcksgränd 10 12462 Bandhagen, Stockholm, Sweden.

+46 72789 9039

Malaysia

Level 23-1, Premier Suite One Mont Kiara, No 1, Jalan Kiara, Mont Kiara, 50480 Kuala Lumpur

Lewes

16192 Coastal Highway, Lewes, Delaware 19958, USA

software developers ahmedabad
ISO 9001:2015 Certified

Call us

Career: +91 90165 81674

Sales: +91 99747 29554

Email us

Career: hr@digiqt.com

Sales: hitul@digiqt.com

© Digiqt 2026, All Rights Reserved