Threat-Led Penetration Testing for Financial Institutions
Running a Red Team Exercise That Tells You Something You Did Not Know
Most banks buy a lot of security testing and learn very little from it. Scanners find missing patches, annual penetration tests find the same categories of weakness in a different application each year, and none of it answers the question a board actually asks: if a capable adversary targeted our payment operation next month, would we notice, and would we stop them.
Threat-led penetration testing exists to answer exactly that. It is structured, intelligence-driven, run against production, and deliberately measured on what your defenders detect rather than on a list of vulnerabilities. Running it well is mostly a matter of governance and follow-through, because the technical execution is the part your testing provider already knows how to do.
How is threat-led testing different from a penetration test?
It simulates a specific adversary against critical business functions in production, rather than probing a defined system for weaknesses.
| Dimension | Penetration test | Threat-led test |
|---|---|---|
| Scope | A named application or environment | Critical business functions end to end |
| Environment | Frequently test or staging | Production |
| Driver | A checklist and a scope document | Tailored threat intelligence on your actual adversaries |
| Target set | Technology | People, process, and technology together |
| Defender awareness | Usually informed | Deliberately unaware |
| Primary output | Vulnerability list | Attack paths, detection gaps, and response performance |
| Success measure | Findings closed | Learning and maturity improvement |
What does the TIBER-EU framework define?
A phased process with five defined teams, developed by the ECB and EU national central banks and published in May 2018.
TIBER-EU stands for Threat Intelligence-Based Ethical Red-Teaming in Europe. It was approved by the ECB's Governing Council, published in May 2018, updated in 2024 to align with digital operational resilience requirements, and has been adopted across more than twenty European jurisdictions. Tests mimic the tactics, techniques, and procedures of real-life attackers against an entity's critical functions across people, processes, and technologies, and proceed through initiation and scoping, threat intelligence gathering, red-team testing, and reporting and remediation. The United Kingdom runs a comparable threat intelligence-led scheme, CBEST, overseen by its authorities, and the structural pattern is deliberately similar across regimes, so a programme built to one framework transfers with modest adaptation.
Why is there no pass or fail?
Because a pass would be meaningless and a fail would make institutions hide findings.
TIBER-EU explicitly emphasises learning outcomes and advancing cyber maturity rather than pass or fail scoring. That design choice is what makes honest testing possible: if a test could fail, every incentive would push toward narrow scope, forewarned defenders, and quiet findings. Preserve that principle internally too. If your executive committee treats a successful simulated compromise as a performance failure rather than as the expected result of a well-scoped test, your programme will quietly degrade into theatre within two cycles.
Does your leadership treat a successful red team as a finding or as a failure?
Talk to Digiqt about setting up a threat-led testing programme
Who does what during a test?
Five roles, with the control team carrying most of the organisational risk.
| Team | Role | Knows the test is happening |
|---|---|---|
| Blue team | The institution's defenders and responders being measured | No |
| Threat intelligence provider | Produces the targeted threat landscape and scenarios | Yes |
| Red team | Executes the simulated attack | Yes |
| Control team | Small internal group coordinating and containing the test | Yes |
| Authority cyber team | Oversees framework compliance where a regulator is involved | Yes |
Why does the control team determine whether the exercise works?
Because it manages scope, guardrails, escalation, and the moment when a real incident and a test collide.
The control team decides when to intervene, holds the kill switch, and handles the situation where the blue team escalates a genuine incident during the test window or where the red team's activity risks customer impact. Staff it with people senior enough to stop the exercise and technical enough to understand what is happening, keep it deliberately small, and rehearse the escalation path before day one. Also decide in advance how you will distinguish a real attack from the test, because that ambiguity is genuinely dangerous and it has caught out well-run programmes.
What should the blue team learn, and when?
Everything, immediately after the test, in a session designed for learning rather than accountability.
The value of the unaware response evaporates if it turns into a blame exercise afterwards. Run a joint replay where the red team walks through the attack path and the blue team walks through what they saw, when, and why they interpreted it as they did. The most useful findings usually come out of that comparison: an alert that fired and was closed, a signal that existed in a tool nobody watches overnight, a runbook step that assumed information the responder did not have.
How should scope be set?
Around critical business functions in production, with defined guardrails rather than defined targets.
Scope by outcome, not by system. The scenario should be an objective an adversary would pursue, such as effecting an unauthorised payment, exfiltrating customer data at scale, or degrading a critical service, and the red team should be free to find the path. Naming specific systems as targets invites the team down a route you already worry about, which is precisely where you least need assurance.
Why must the test run against production?
Because your test environments do not have your people, your monitoring, or your accumulated configuration drift.
Real adversaries face your actual estate, including the legacy service nobody documented and the monitoring gap that exists because a tool was migrated last quarter. Testing a staging environment measures a system you built for testing. Production testing needs guardrails: prohibited actions defined in writing, a kill switch reachable at any hour, data handling rules for anything the red team encounters, and pre-agreed legs-up so that a blocked path does not consume the engagement. Those controls make production testing routine rather than reckless.
How do you keep the exercise from wasting itself on one obstacle?
With planned legs-up and time-boxed objectives agreed with the control team.
If the red team spends three weeks on initial access and never reaches the interesting part of the estate, you have paid to learn one thing you probably already suspected. Agree in advance that after a defined period the control team may grant assumed access at a specified level so the exercise can proceed to lateral movement, privilege escalation, and objective execution. Record where the leg-up was given, because that itself is a finding about where your controls held.
How does threat intelligence drive the scenarios?
By identifying the adversaries plausibly interested in your institution and the techniques they actually use.
A useful targeted threat intelligence report names relevant threat actors, describes their observed tradecraft, maps it to your attack surface, and proposes scenarios grounded in that pairing. A poor one recycles generic industry reporting into a document nobody reads. Insist on specificity about your institution: your suppliers, your technology choices, your public footprint, your staff exposure. Then use the intelligence to prioritise which paths get exercised, since no engagement covers everything and the choice of what to simulate is the most consequential decision in the whole programme. Patch and exposure timing feeds this directly, as discussed in this analysis of vulnerability exploitation windows.
Is your threat intelligence specific to your institution or recycled industry reporting?
How should findings be handled?
As engineering work with owners, dates, and architectural consequences, not as a report to be accepted.
Separate findings into three classes and treat them differently. Detection gaps go to the security operations backlog with a specific detection to build and a test to prove it. Architectural weaknesses, such as a flat network or an over-privileged service account, go into engineering roadmaps with the recognition that they may take quarters. Process and people findings go to the owners of those processes. Then track all of them to closure with the same rigour you apply to production defects. Programmes fail here more than anywhere else: findings are accepted, partially remediated, and reappear in the next cycle, which is how institutions end up paying for the same lesson three times. Remediation of segmentation and privilege findings usually points back to the model described in this guide to zero-trust security architecture.
What makes a finding worth the cost of the exercise?
One that changes an architecture decision or creates a detection you did not have.
A credential in a file share is a fix. An attack path that shows your payment operation is reachable in three steps from a marketing laptop is a strategic finding, and it is what you paid for. Judge each engagement on how many findings of the second kind it produced, and if the answer is zero, question the scope rather than concluding you are secure. Testing whether the adversary could reach your recovery estate is a particularly valuable objective, since immutability that a red team can defeat is not immutability, which connects directly to ransomware resilience and immutable backup design and to the cost dynamics in what inflates ransomware recovery costs.
How does this fit alongside other assurance?
As the apex of a pyramid, not a replacement for continuous testing.
Continuous vulnerability management and configuration scanning handle volume. Regular penetration testing covers application and infrastructure depth. Purple team work builds detections collaboratively and is often the cheapest way to close gaps found by red teaming. Threat-led testing validates the whole system, including your people, once or twice a year. Continuous control validation sits underneath, checking that the detections you built still fire. Each layer feeds the next, and running only the top layer produces expensive theatre while running only the bottom produces confidence without evidence. Framing the programme against a recognised structure such as the NIST Cybersecurity Framework, now at version 2.0 released in February 2024, helps the layers stay coherent and gives you common vocabulary with auditors.
What about third parties and shared infrastructure?
Scope them deliberately, with contractual permission, because your critical functions run partly outside your perimeter.
Critical functions frequently depend on providers you cannot test unilaterally, and testing a third party without authorisation is not an option. Address it in three ways: include the third-party interface and your own controls around it in scope, obtain contractual rights to assurance evidence and where possible joint testing, and treat concentration explicitly. The FSB's December 2023 toolkit on third-party risk management provides tools for identifying critical third-party services and for monitoring systemic dependencies and concentration risk, which is the right frame for deciding which providers justify the negotiation effort.
How do you budget and schedule a programme?
One major exercise per year for most institutions, with intelligence refresh, remediation cycles, and purple team work between.
| Activity | Typical cadence | Main cost driver |
|---|---|---|
| Targeted threat intelligence | Annually, refreshed before each test | Specificity and analyst time |
| Red-team engagement | Annually, six to twelve weeks elapsed | Team size and scenario breadth |
| Replay and findings workshop | Immediately after each test | Internal time, not vendor cost |
| Remediation delivery | Continuous | Architectural findings dominate |
| Purple team detection building | Quarterly | Internal capability |
| Regulatory reporting where applicable | Per framework cycle | Documentation quality |
Budget for remediation, not just testing. A programme that funds the exercise and not the fixes produces a well-documented list of unaddressed weaknesses, which is worse than not testing because it establishes that you knew.
Which metrics show the programme is maturing?
Detection rate along the attack path, time to detect and contain, repeat findings, and remediation closure.
Measure how far the red team progressed before detection, and track that across cycles, since improvement means detection moving earlier in the chain. Report time from first red team action to detection and to containment. Count repeat findings from previous cycles as a direct measure of remediation effectiveness, and treat any repeat as an escalation rather than a note. Track closure rate and ageing of findings by class. And record the number of new detections built as a result of each exercise, because that is the durable capability the programme leaves behind after the report is filed.
Threat-led testing is the only assurance activity that tells you how your institution behaves under a realistic attack rather than how your systems score against a checklist. It earns its cost when the findings change architecture and detection, and it wastes it entirely when the report is accepted, circulated, and quietly filed.
Frequently Asked Questions
How is threat-led testing different from a penetration test?
A penetration test looks for weaknesses in a defined system. Threat-led testing simulates a specific adversary against critical business functions in production, across people, process, and technology.
What does the TIBER-EU framework define?
A structured process with initiation and scoping, threat intelligence, red-team testing, and reporting and remediation phases, run by five defined teams, with learning rather than pass or fail as the outcome.
Why is there no pass or fail?
Because a pass would be meaningless and a fail would push institutions toward narrow scope and hidden findings, which destroys the value of testing honestly.
Why does the test have to run against production?
Because attackers do. A test environment tells you about a configuration you built for testing, not about the systems, people, and monitoring that actually protect your critical functions.
Who should know a test is happening?
A small control team, deliberately excluding the defenders being tested. The blue team's unaware response is the primary thing being measured.
How do you manage the risk of testing live systems?
Named guardrails, a real-time kill switch, agreed prohibited actions, a control team with authority to stop, and pre-planned legs-up so a blocked path does not waste the engagement.
What makes a finding worth the exercise?
One that changes an architectural decision or a detection capability, not another credential in a file share. Value comes from the attack path, not the individual weakness.
Why do these programmes stall?
Remediation. Findings are accepted, partially fixed, and reappear in the next cycle, so track them to closure with owners and dates like any other engineering work.



