Same Average Wealth, Missing Risk: The Scenario Truncation Trap

You speed up a portfolio calculation by keeping a few likely market continuations instead of hundreds. Average terminal wealth barely changes. Trading costs are accounted for, weights satisfy the constraints, and the result looks stable. Yet the acceleration may have removed precisely the events that made risk measurement necessary.
A comparison of average returns cannot detect this reliably. Even matching one adverse quantile does not guarantee that losses beyond it, or drawdowns along the way, have survived. The counterexample below preserves average wealth and fifth-percentile wealth exactly while reducing measured Expected Shortfall almost ninefold.
This is a separate problem from choosing an allocation algorithm or adding a CVaR constraint to a portfolio. Here we examine the input distribution: which future states the optimizer is allowed to see, and which have quietly been excluded.
How the probability distribution changes

The immediate motivation is a September 30, 2026 preprint by Ramachandran, Iyer, and Jain. Their bond model prunes successor states of a Markov chain. Section 3 reports that a 20-basis-point grid with ten successors retains roughly one-third of conditional probability mass and approximately halves conditional yield standard deviations. A 40-basis-point grid with 25 successors retains about 92%. These are diagnostics of that approximation, not recommended settings for arbitrary portfolios. Preprint, version 1, Section 3.
Consider the mechanism independently of the bond model. Let the original successor probabilities be and the retained set be . Its probability mass is:
After removing other states, probabilities are usually normalized to for the retained outcomes. They sum to one again. However, this distribution now describes the market conditional on entering the retained set. The condition disappears from the interface even though it changes the problem.
Write the removed mass as . For terminal wealth :
If the retained and removed portions have equal means, truncation leaves average wealth completely unchanged. The removed portion can nevertheless contain both a large gain and a large loss. They offset in the mean, but not in a measure of adverse outcomes.
Consequently, a report saying "94% of probability was retained" does not answer a question about the worst 5% of outcomes. We need to know which 6% disappeared. A finer grid is not sufficient either: under a fixed computational limit on successors, finer partitioning can leave more probability mass outside the retained set. The parameters need to be tested jointly.
An exact counterexample

Take five predefined wealth paths. Initial capital is 100 USD, followed by two observation dates. Their probabilities are deliberately constructed for illustration. This is neither a market estimate nor a backtest, and it does not reproduce the preprint's experiments.
| Probability | Initial, USD | Date 1, USD | Terminal, USD | Maximum drawdown |
|---|---|---|---|---|
| 47% | 100 | 99 | 99 | 1% |
| 47% | 100 | 101 | 101 | 0% |
| 4% | 100 | 70 | 100 | 30% |
| 1% | 100 | 60 | 60 | 40% |
| 1% | 100 | 140 | 140 | 0% |
The accelerated model retains the two most probable paths. Their combined mass is 94%; normalization gives each a probability of 50%. Full-model expected terminal wealth is 100 USD. The truncated model also yields 100 USD: the symmetric extreme outcomes and the recovery path did not change the mean.
Define terminal relative loss as . A positive value means capital was lost. VaR at 95% is the smallest loss at which cumulative probability reaches 95%. ES at that confidence level is the average loss within exactly the worst 5% of probability mass.
That qualification matters for discrete scenarios. Selecting every row whose loss is at least VaR and averaging them can include far more than 5% of the probability. The boundary scenario must be included only partially. Discrete distributions and probability mass at the boundary are addressed in the classic treatment by Rockafellar and Uryasev, Section 2.
In the full model, the worst 1% has a 40% loss. The next 4% of probability comes from the outcome with a 1% loss. Therefore:
Mean maximum drawdown requires a different calculation: first find each path's largest decline from its running wealth peak, then weight those drawdowns by their probabilities. The path 100 → 70 → 100 has no terminal loss, but contributes to drawdown. This is another reason to retain more than the terminal-wealth distribution.
Download the Python example. It uses only the standard library, calculates with Fraction, and checks results through exact equalities. There is no random sample or Monte Carlo error.
python3 portfolio-scenario-truncation-tail-risk.py
Synthetic exact-probability example; initial wealth = 100 USD
Retained probability mass: 94.00%
law mean_USD q05_USD VaR95_loss ES95_loss mean_MDD P(MDD>=20%)
full 100.00 99.00 1.00% 8.80% 2.07% 5.00%
truncated 100.00 99.00 1.00% 1.00% 0.50% 0.00%
Exact arithmetic assertions: PASS
The mean, fifth-percentile wealth, and VaR match. However, ES falls from 8.80% to 1.00%, and mean maximum drawdown from 2.07% to 0.50%. The probability of a drawdown of at least 20% becomes zero solely because every relevant path was removed. The script neither optimizes weights nor models fees: it isolates the distribution error before those complications are added.
Trading costs do not restore a missing tail

A rebalancer that accounts for expenses can solve the wrong problem correctly. It sees only the retained continuations, estimates a trade's future benefit, and compares that benefit with its cost. If a severe scenario is absent from the input, flawless fee accounting does not put it back.
Start the expense check with a simple transaction. Selling 10,000 USD of one asset and buying 10,000 USD of another creates 20,000 USD of total traded notional. A charge of ten basis points on each side costs 20 USD. Assume a separate cash reserve pays the fee, so that the tariff check is not mixed with the available-cash calculation.
If the report defines turnover using only the sold amount, it will show 10,000 USD. Either convention can work, but the cost rate must match the denominator. Combining one-sided turnover with a rate intended for two-sided notional introduces an error before market modeling starts. Likewise, weight changes must be measured from actual pre-trade holdings: price movements mean those weights already differ from the previous targets.
Next, incorporate expenses into the wealth path itself. Deducting them only from terminal returns is insufficient for checking intermediate limits, because payment reduces available wealth when it occurs. The order executor also needs cash-reserve, quantity-rounding, and operation-sequencing checks. This extends the discussion of automated ETF rebalancing, while leaving scenario validation as a separate requirement.
Finally, compare several cost levels in two ways: first under a frozen policy, then after reoptimization. The first comparison measures an existing decision's sensitivity. The second reveals how the decision changes. Otherwise, lower turnover after raising the tariff can be mistaken for evidence that the original algorithm was robust.
Evaluate the policy beyond the reduced tree

A practical protocol should separate two errors. First, how does truncation alter the measured risk of a fixed policy? Second, how does it change the selected policy and that policy's results? Changing both at once makes the cause of any discrepancy difficult to identify.
Freeze the weights or rebalancing rule and evaluate it under the reduced and richer distributions. Compare the mean, adverse quantiles, ES, the distribution of maximum drawdown, and breaches of specific limits. For a conditional model, repeat the check by state: a reassuring average across dates can conceal an error immediately before a crisis transition.
Then train or optimize two policies using different approximations, but evaluate both on the same independent set of richer paths. When changing several parameters, preserve the same market realizations for the competing policies. This supports comparisons of their difference on each path instead of two noisy averages. The evaluation stage must not adapt the policy using future observations.
For a continuous state, specify in advance what happens when an evaluation path leaves the grid. Nearest-node lookup, interpolation, and a fallback rule define different policies. Record how often the fallback is used and how much it contributes to losses. Silently excluding inconvenient paths repeats the original truncation error.
A related principle of evaluating decisions on original scenarios appears in Zhuang and colleagues' October 17, 2025 paper on scenario reduction with CVaR. Their application is a virtual power plant, not portfolio returns. The useful methodological point is to assess reduction through the consequences of its selected decision, not solely through similarity between input datasets. Sections II-B and II-G.
A richer simulator still remains a model. A fresh random sample from the same misspecified generator does not establish market validity. Bootstrap and outcome intervals help characterize uncertainty in available data, but cannot restore events excluded from the model itself.
What makes an approximation acceptable

There is no universally sufficient scenario count. Acceptance should be defined through the decision and the limits being protected. The same distribution error has different consequences for an unleveraged portfolio and for a strategy facing a hard liquidation threshold.
| Check | What the report should retain |
|---|---|
| Retained mass | Values before normalization by state and date, including adverse states |
| Truncation sensitivity | Joint changes in grid width, successor count, and deletion threshold |
| Terminal losses | Quantiles and ES with explicit probabilities, confidence level, and boundary handling |
| Risk within the period | Drawdown distribution, limit-breach frequency, and wealth-observation interval |
| Execution | Consistent units for expenses and turnover, insufficient cash, and fallback usage |
Set tolerances before examining the results. If an ES constraint determines the allocation, a small error in mean wealth is insufficient: check whether the ES error changes the portfolio's eligibility. Uncertainty in the tail estimate should also be visible. With few independent extreme observations, extra decimal places do not add confidence.
Keep rare stress paths separate from probabilistic forecasts when their probabilities are unknown. They answer "would the policy survive this path?", but do not by themselves provide an ES estimate. Assigning convenient weights and calling the result calibrated risk merely replaces one hidden assumption with another.
The bond preprint tests simulated zero-coupon bonds, proportional expenses, and an expected-terminal-wealth objective; the authors acknowledge the absence of evaluation on realized market history. That limits transferability. Limitations, Section 5. Applying this protocol to cryptoassets is our engineering inference, not an effect measured in that study.
The synthetic example establishes a narrow result: a reduced distribution can pass mean and VaR checks perfectly while giving a misleading picture of tail losses and drawdowns. Before using scenarios for rebalancing, validate the paths and metrics capable of changing the decision about acceptable risk.
ผู้เขียน
Trading-systems engineer
Trading-systems engineer building bots since 2017: cross-exchange arbitrage (connected up to 30 venues), cointegration-based pairs arbitrage across spot and futures, scalping, news and sentiment-driven strategies, trend algorithms, and portfolio management and balancing algorithms. Also builds sub-millisecond order execution, big-data warehouses, backtesting engines, AI agents, and trading interfaces (incl. open-source profitmaker.cc). Stack: JS/TS, Python, Rust/Zig/Go, DevOps, backend, frontend, architecture.