Evaluating Quantum Computing for Portfolio and Risk Simulation
The Honest Version of the Quantum Finance Business Case
Quantum computing arrives on a CTO's desk in a predictable form. A vendor deck showing a portfolio optimisation solved on quantum hardware, a competitor's press release about a research partnership, a board member asking what the institution's position is, and an internal enthusiast with a proposal. The pressure to have an answer is real and the quality of available information is poor, because the interesting results are published by parties with an interest in the outcome.
Evaluating quantum computing portfolio optimization proposals well is mostly a matter of asking a small number of unglamorous questions, of which one matters more than all the others combined.
What is actually being claimed?
Usually that a quantum method solved a problem, which is different from solving it better.
Almost every demonstration in this space is genuine as stated and much narrower than it reads. A portfolio problem was solved on quantum hardware. That claim is compatible with the problem being small enough to solve on a laptop instantly, the quantum run taking longer, the comparison being against an untuned classical method, or the timing excluding the substantial cost of encoding the problem and reading out the result. None of that makes the work dishonest, and all of it makes the headline useless for a procurement decision.
Why is the vendor claim rarely the published result?
Because the qualifications live in the paper and the headline lives in the deck.
Research papers in this field are typically careful, stating problem size, hardware conditions, comparison method, and limitations. By the time the result reaches a commercial audience the qualifications have been compressed out. The practical remedy is to read the underlying work rather than the summary, and to ask for it if it does not exist. A proposal with no reproducible underlying result is a research partnership rather than a technology purchase, and it should be evaluated on that basis.
Has anyone read the paper behind the deck you were shown?
Which finance problems are candidates?
Four families, with different levels of theoretical support.
| Problem family | Finance use | Quantum formulation | Honest status |
|---|---|---|---|
| Constrained combinatorial optimisation | Portfolio selection with cardinality and lot constraints | Formulated as a quadratic unconstrained binary optimisation problem | Formulation mature, hardware advantage unproven at real scale |
| Stochastic simulation | Market and credit risk simulation, scenario generation | Amplitude estimation style approaches | Theoretical speedup argued, hardware requirements substantial |
| Derivative pricing | Path-dependent instrument valuation | Related simulation formulations | Same constraint as above |
| Machine learning and sampling | Feature selection, certain sampling problems | Various quantum machine learning proposals | Least settled, claims most contested |
The first row is where most vendor activity concentrates, because portfolio selection with real-world constraints is a genuinely hard combinatorial problem that institutions currently solve with heuristics and accept approximate answers for. That makes it a fair target. It also makes the classical comparison harder, because the incumbent is a well-tuned heuristic rather than an exact method, and beating it is a higher bar than the marketing implies. The wider issue of models lagging portfolio reality is examined in capital models that lag portfolio change.
What do the two hardware approaches suit?
Annealing targets optimisation directly; gate-based machines are general purpose and harder to build.
Quantum annealing is purpose-built for optimisation problems expressed in a particular binary quadratic form, which maps naturally onto portfolio selection. It is available at larger scale today and its advantage over strong classical methods remains contested. Gate-based quantum computing is the general-purpose model that the simulation and pricing speedups depend on, and it is further from the scale those algorithms require. A bank evaluating proposals should know which model a given claim rests on, because the two have different maturity, different vendors, and different problem fit, and conflating them is common in summarised material.
What does current hardware actually support?
Less than the coverage suggests, and NIST's own framing is instructive.
NIST's quantum information science material describes quantum computers as a new kind of machine that can, in theory, simulate the fundamentally quantum nature of matter, alongside quantum sensors for ultra-high-precision measurement and quantum networks spanning cities and nations. The phrase worth noticing is "in theory". A national measurement science agency describing the field in those terms is a reasonable calibration point against commercial material describing production readiness.
Why is qubit count the wrong headline number?
Because it describes the hardware's size rather than its usable capability.
What determines whether a machine can solve a problem is the combination of error rates per operation, how long the qubits maintain coherence, which qubits can interact with which others, and how much overhead error correction imposes to produce reliable logical operations from unreliable physical ones. A larger machine with worse error characteristics can be less useful than a smaller, cleaner one. When a proposal leads with qubit count and does not address error rates or correction overhead, that omission is the most informative thing in the document.
How should a bank evaluate a specific claim?
With six questions, asked in order.
| Question | Why it matters | What a good answer looks like |
|---|---|---|
| What classical method was it compared against? | Determines whether the comparison is meaningful | A named, current, seriously tuned solver |
| Who tuned the classical baseline? | Untuned baselines produce flattering results | Someone with an interest in the classical side performing well |
| What size and structure was the problem? | Small instances prove nothing about scale | Realistic dimensions, stated explicitly |
| Was timing end to end? | Encoding and readout often dominate | Total wall-clock including all data movement |
| Is the result reproducible? | Single runs of stochastic methods mean little | Repeated runs with distribution reported |
| What happens as the problem grows? | Scaling behaviour is the whole question | Measured scaling across sizes, not asserted |
Why is the classical baseline the whole evaluation?
Because the only question that matters is whether it beats what you would otherwise do.
A quantum result is interesting to a physicist regardless of what classical computers can achieve. It is interesting to a bank only if it produces a better answer, or an equally good answer faster or more cheaply, than the tuned classical approach the institution would use instead. That comparison is frequently weak in published work, sometimes because constructing a genuinely strong classical baseline is expensive and unrewarding for the party running the experiment. When evaluating a proposal, fund the classical side of the comparison yourself and treat it as the control. It is also worth noting how often the outcome of this exercise is a materially improved classical solution, which is a good result even though nobody writes a press release about it. Compute economics belong in the same analysis, as covered in cloud cost optimisation for workloads at scale.
What is worth doing now?
Formulation work, which pays off regardless of what the hardware does.
The most valuable output of a quantum exploration in a bank is usually not quantum. Expressing a portfolio problem precisely, with its real constraints written down rather than embedded in the heuristic that currently approximates it, produces immediate benefit: classical solvers perform better on a clean formulation, the constraints become reviewable by risk and business stakeholders, and previously undocumented assumptions surface. If quantum hardware becomes useful, that formulation transfers directly. If it does not, the institution has a better-specified problem and a better classical answer.
The same applies to scenario and simulation work, where the discipline of stating the model precisely has value independent of what executes it, as discussed in operationalising an economic scenario generator and in the interest rate simulation machinery described in ALM and IRRBB systems.
Why does formulation work pay off either way?
Because the constraint set is the asset, and it currently exists only inside a heuristic.
In most institutions the real portfolio constraints are partly documented, partly encoded in a solver's configuration, and partly maintained as folklore by two people. Writing them down completely is a prerequisite for any new approach, quantum or otherwise, and it exposes inconsistencies that have been quietly shaping decisions. That is a substantial deliverable from a programme whose headline objective may never be met, and it is the reason a modest exploration can be justified even by a sceptic. The same knowledge concentration problem is examined in the legacy skills shortage.
Are your real portfolio constraints written down, or encoded in a solver nobody has revisited?
Talk to Digiqt about problem formulation that outlasts the hardware question
How does model risk management apply?
Fully, and the current guidance is not the one most institutions still cite.
Any output that informs a decision is a model output, regardless of the hardware that produced it. The Federal Reserve's SR 26-2, issued 17 April 2026, is the current model risk management guidance and supersedes SR 11-7 from 2011 and SR 21-8 from 2021. It sets out a risk-based approach tailored to an organisation's model risk profile and applies primarily to organisations above thirty billion dollars in assets. A quantum-derived portfolio allocation or risk number therefore needs the same development standards, validation, and challenge as any other model, which is the machinery described in model risk management platform design.
Why is non-determinism a validation problem?
Because a solver that gives different answers each time has to have that variability characterised.
Heuristic and stochastic optimisers, quantum and classical alike, return different solutions on repeated runs. That is acceptable in a decision process only if the distribution of outcomes is understood, bounded, and disclosed. Validation therefore has to cover repeated-run behaviour, sensitivity to input perturbation, and what happens when the solver returns a solution that violates a constraint, which is a real failure mode in approximate methods. Explaining a heuristic result to a challenge function is its own difficulty and it is not unique to quantum, as the discussion in signal, noise, and model risk sets out, and the governance parallels with newer techniques generally are covered in generative AI and banking model risk.
Is this the same as the cryptography question?
No, and conflating them distorts both.
Post-quantum cryptography has a deadline driven by data interception today, standardised algorithms already published, and work that must happen regardless of whether quantum computing ever becomes commercially useful for optimisation. That is a mandatory programme with a defined scope, covered in crypto-agility in banking systems. Quantum computing for portfolio and risk work is a discretionary research allocation with an uncertain payoff. They share a technology and nothing else, and institutions that fund them as one initiative tend to under-resource the mandatory half.
What would a defensible programme look like?
Small, time-boxed, benchmarked, and honest about its own conclusions.
| Phase | Duration | Deliverable |
|---|---|---|
| Problem selection | 1 month | One or two problems where current answers are known to be approximate |
| Formulation | 2 to 3 months | Complete, documented constraint set, reviewed by risk and the business |
| Classical baseline | 2 months | A seriously tuned classical solution, treated as the control |
| Literature and vendor assessment | 1 month | Underlying papers reviewed, claims mapped to hardware model |
| Bounded experiment | 3 to 4 months | Like-for-like comparison including end-to-end timing |
| Decision point | Fixed date | Continue, pause, or stop, with the classical improvement banked either way |
| Watching brief | Ongoing, low cost | Defined checkpoints on hardware and algorithmic progress |
The decision point is the part that distinguishes a research programme from a permanent enthusiasm. Set it before starting, define what evidence would justify continuing, and be willing to record a negative result, since a well-run experiment that concludes the technology is not yet useful for the institution's problems is a successful experiment.
Which metrics matter?
Improvement over the tuned classical baseline, end-to-end time, solution quality distribution, constraint violation rate, and cost per solved instance.
Report improvement against the baseline as the only headline that counts. Report end-to-end wall-clock including encoding and readout rather than kernel time. Report the distribution of solution quality across repeated runs rather than the best run observed. Report constraint violation rate, since an infeasible answer produced quickly is not an answer. And report cost per solved instance in currency, because at some point the comparison against classical compute becomes an economic one rather than a scientific one.
Quantum computing may eventually change how portfolios are optimised and how risk is simulated. It has not yet, the honest evidence says so, and the useful position for a CTO is neither dismissal nor a programme built on a vendor deck. Formulate the problems properly, tune the classical baseline, watch the field at low cost, and keep the cryptography work, which has a real deadline, entirely separate.
Frequently Asked Questions
Which finance problems are genuine quantum candidates?
Constrained portfolio selection, scenario and Monte Carlo style risk simulation, derivative pricing, and certain sampling and machine learning problems, all of which map onto known quantum formulations.
Why is qubit count the wrong headline number?
Because usable capability depends on error rates, connectivity, coherence time, and error correction overhead, so raw counts say almost nothing about what a machine can actually solve.
Why is the classical baseline the whole evaluation?
Because a quantum result only matters if it beats a well-tuned classical method on the same problem, and many published comparisons use classical baselines that were not seriously optimised.
What questions should be asked about a vendor benchmark?
What classical solver it was compared against, how that solver was tuned, the problem size and structure, whether timing was end to end including encoding and readout, and whether it is reproducible.
What is worth doing now regardless of hardware progress?
Formulating the problems properly, since a clean constrained formulation improves classical solver performance immediately and transfers directly if quantum hardware becomes useful.
How does model risk management apply to quantum-derived results?
Fully. Outputs used in decisions remain subject to model risk governance, which currently means the Federal Reserve's revised guidance in SR 26-2 for institutions in its scope.
Why is non-determinism a model validation problem?
Because a solver that returns different answers on repeated runs needs its variability characterised and bounded before its output can support a decision anyone must defend.
How should a CTO position quantum investment?
As a small, time-boxed research allocation with defined checkpoints, kept separate from the post-quantum cryptography work, which has a genuine deadline and a different urgency.



