📝

Draft article

This article is visible to admins and superusers only. Sign in with an authorized account.

← Voltar aos artigos
October 8, 2026
5 min read

Generated volatility surfaces: where realism stops being enough

Generated volatility surfaces: where realism stops being enough
#options
#implied-volatility
#static-arbitrage
#generative-models
#hedging

A smooth implied-volatility surface can correspond to inconsistent option prices. A surface that passes price checks can produce poor scenarios for tomorrow's market. Good scenarios still need to become a hedge that helps after costs. These are three separate reasons to reject a model.

With the GARCH family, we forecast conditional return variance. In the discussion of PINNs and Heston, we solve a valuation problem under specified dynamics. The question here is different: which checks must a generator of an entire option surface pass before it is used for portfolio valuation and hedging?

The discussion separates findings from three August and September 2026 preprints, an original reproducible counterexample, and a proposed acceptance protocol. None of the papers' models were retrained here, and no exchange experiment was performed.

An arbitrage penalty is not a guarantee

A smooth surface undergoes a local geometric inspection that reveals a concealed fold

Latent Flow Matching for Arbitrage-Aware Implied Volatility Surface Generation was first submitted on August 1; the reviewed v2 is dated August 4, 2026. It generates individual surfaces without conditioning on current history. An autoencoder receives penalties for financial constraint violations, while flow matching models the latent distribution. The main experiment gives the described call-spread penalty zero weight: adding it did not improve the result on the authors' data. Version history, Sections 2.2 and 3.6.1.

The training set contains 1,000 preprocessed SPX surfaces from 2020–2023 on a grid of 32 moneyness values and 16 maturities. This defines the result's scope; it is not a finding for arbitrary option markets. Section 3.1.

The authors report that 90.8% of generated surfaces pass all three tested grid constraints, across five runs with 5,000 surfaces each. Strengthening penalties to achieve a 99.9% passing rate worsens similarity to the data distribution, particularly in the upper tail. A high acceptance rate alone therefore does not select the best risk generator. Sections 3.5.4 and 3.6.2, Tables 5 and 7.

A newer study, Universal Diffusion Models for Implied Volatility Surfaces, v2 of September 25, adds a useful check: among the compared generators, the ordinary MSE variant achieved the lowest test violations, ahead of the explicit arbitrage-penalty variant. The multi-stock experiment uses one random seed; the authors call for multiple runs to investigate penalty–training interactions. This supports ablation, not abandoning constraints altogether. Sections 3.4, 4.1, and 5.

My practical inference is to report rejection frequency, violation magnitude, and distributional changes after repair together. Removing difficult surfaces can produce a tidy sample that no longer represents stress. A projection onto an admissible set should be evaluated by comparing original and repaired scenarios, including their effect on a specific portfolio's value. Rejection sampling should disclose which regimes it removes.

A penalty optimizes the average behavior of a learned function. It does not prove a restriction for every latent vector, strike, and maturity. Checking one collection of nodes also says nothing automatically about interpolation between them. The authors explicitly limit their conclusions to a fixed grid and do not guarantee every sample. Section 4.

Check prices, boundaries, and total variance

Geometric ribbons on uneven supports illustrate slope and convexity restrictions

For clarity, start with European calls, zero rates and dividends, one underlying, and a common settlement convention. Let spot be SS and let CiC_i price a call with strike KiK_i and a fixed expiration. Necessary checks include:

max⁡(S−Ki,0)≤Ci≤S,si=Ci+1−CiKi+1−Ki,−1≤si≤0,si+1≥si.\max(S-K_i,0)\le C_i\le S, \qquad s_i=\frac{C_{i+1}-C_i}{K_{i+1}-K_i}, \qquad -1\le s_i\le0, \qquad s_{i+1}\ge s_i.

The first check bounds an individual contract. The second bounds a vertical spread's price. The last checks convexity: the slope must not become more negative as strike increases. Dividing by the actual spacing matters on an irregular grid. Finite-price restrictions and their relationship to executable spreads are developed by Gerhold and Gülüm, Theorem 3.1.

These are necessary local checks, not a complete surface certificate. Consistent boundary values, calendar restrictions, and the extension beyond the grid also matter. Under our assumptions, for example, C(0,T)=SC(0,T)=S, while C(K,T)C(K,T) approaches zero as strike grows without bound. A smooth picture establishes neither property.

With nonzero rates and dividends, first align discount factors, forwards, and price units. In the standard proportional-dividend setting, the calendar check uses fixed forward log-moneyness k=log⁡(K/FT)k=\log(K/F_T) and total variance:

w(k,T)=Tσimp2(k,T),∂Tw(k,T)≥0.w(k,T)=T\sigma_{\mathrm{imp}}^2(k,T), \qquad \partial_T w(k,T)\ge0.

The restriction concerns total variance, not annualized volatility itself. The assumptions and proof appear in Gatheral and Jacquier, Lemma 2.1. A raw call-price maturity-monotonicity check cannot be transferred unchanged to every dividend or interest-rate regime.

An original counterexample uses two smiles that are flat across strikes: volatility is 40% at a quarter-year maturity and 25% at one year. Total variance rises from 0.0400 to 0.0625. Falling annualized volatility is compatible with the calendar restriction here. Increasing total variance can be interpolated between these maturities.

This matters for the August paper: its calendar penalty penalizes declining volatility itself, imposing a stronger condition than necessary. The authors acknowledge this. If a model smooths away an admissible declining volatility term structure, an improved penalty may represent a lost useful scenario. Section 2.2.2.

A working validator should rerun checks after reversing normalization, interpolation, and extrapolation. Convexity is especially sensitive: a small level error can become a large second-difference error. Tolerances should be expressed in price and spread units, with numerical error accounted for separately.

A midpoint violation can disappear at executable quotes

Three option contracts form a butterfly inside separate bid and ask corridors

Consider a wholly synthetic slice with S=100S=100, half a year to expiration, and a contract multiplier of one. All prices below use the same monetary units.

Strike Midpoint Bid Ask
90 12.00 11.20 12.80
100 9.00 8.20 9.80
110 5.00 4.20 5.80

Each midpoint satisfies its individual bounds and has a finite positive implied volatility: approximately 21.11%, 31.97%, and 30.95%. Prices decrease with strike. But adjacent slopes are −0.3 and −0.4, violating convexity.

One long 90 call, two short 100 calls, and one long 110 call produce a nonnegative payoff at every terminal underlying price. Buying this butterfly at the midpoints has negative cost:

Bmid=12−2⋅9+5=−1.B_{\mathrm{mid}}=12-2\cdot9+5=-1.

If those prices were simultaneously available for buying and selling, a trader would receive one unit initially and a nonnegative terminal payoff. Actual purchases use asks and sales use bids:

Bexec=12.8−2⋅8.2+5.8=2.2.B_{\mathrm{exec}}=12.8-2\cdot8.2+5.8=2.2.

This arbitrage disappears at the quoted prices. A positive cost for one butterfly does not establish consistency of the entire chain. The example therefore checks more: prices of 12.50, 8.50, and 5.00 lie inside the spreads and equal expected payoffs under an explicit nonnegative underlying distribution with mean 100. The script verifies that distribution using exact fractions. For these contracts, this certifies the existence of a consistent model rather than merely failing to find a trade.

Now retain the midpoints but narrow every half-spread to 0.10. The executable butterfly costs −0.60. Charging 0.05 for each of its four contracts leaves a negative cost of −0.40. Under this idealized setup, the trade receives 0.40 initially and has a nonnegative payoff. It requires available size on every leg, simultaneous execution, and identical expiration and settlement references; collateral financing and other charges are absent here.

Download the reproducible Python example. Only the standard library is needed:

python3 generated-volatility-surfaces-static-arbitrage.py

It also checks a separate monotonicity violation, the calendar counterexample, and Black–Scholes inversion. The inversion only changes coordinates from price to volatility: a positive IV for each option does not certify the chain's joint consistency. These synthetic numbers do not measure the frequency of market arbitrage.

A valid snapshot does not establish valid dynamics

Valid individual surfaces are connected by price and volatility paths with a missing transition

A surface has two distinct time axes: time to expiration within today's snapshot, and the observation date of the snapshot itself. A calendar restriction along the first axis does not test the market's joint evolution along the second.

In Diffusion models for dynamic volatility surface generation and data-driven hedging, AD-Seq-Vol jointly generates the underlying return and the next surface, updating history along a trajectory. Its FT variant is fine-tuned with static-violation penalties. The first version was submitted on September 11; the reviewed v3 is dated September 17, 2026. The authors report almost zero static violations after fine-tuning, while explicitly leaving dynamic no-arbitrage restrictions to future work. Version history, Sections 2.2 and 5.

Adaptedness means the next step uses history already available. It does not establish a consistent pricing measure for the chosen joint dynamics. Connecting admissible snapshots does not create that guarantee.

A historical scenario generator also need not set the underlying's mean return equal to the risk-free rate. It may model the physical distribution of future markets, including risk premia. Risk-neutral price consistency and predictive accuracy under the physical distribution are separate requirements. Testing the martingale property directly under a historical distribution without specifying the measure would be a mistake.

For scenario use, I propose checking joint changes in the underlying, IV level, and smile slope; conditional tails following stressed days; and error accumulation over multiple steps. Track the same contract as well: its strike stays fixed while remaining maturity declines. A constant moneyness cell on a moving surface generally represents a different strike. Substituting contracts can break a hedge even when every snapshot has perfect geometry.

Hedging needs a separate acceptance test

Scenarios and independent observations enter a separate hedging test apparatus with trading frictions

The September paper evaluates hedging separately: training uses SPX data through June 2018, with evaluation from July 2018 to February 2023. Conditional scenarios support one-period hedge selection; the objective penalizes position changes using half the bid–ask spread as a cost coefficient. This is substantially more informative than matching surface images. Sections 3 and 4.

However, the reported error in Section 3 is the target's change minus a fitted intercept and the hedge's value change. Execution payments are not separately subtracted in that formula. A turnover penalty in optimization does not turn this metric into the net result of a self-financing trading account. A replication should explicitly track the cash account, financing, actual position changes, and payments. Error definition following Equation 12.

Why can equally realistic generators produce different hedges? In the simplest case, with no costs and one hedging instrument, the minimum-variance coefficient is:

ϕ∗=Cov⁡(ΔV,ΔH∣Ft)Var⁡(ΔH∣Ft).\phi^*=\frac{\operatorname{Cov}(\Delta V,\Delta H\mid\mathcal F_t)} {\operatorname{Var}(\Delta H\mid\mathcal F_t)}.

Correct marginal IV distributions do not determine this conditional covariance. A generator may reproduce volatility levels while misrepresenting precisely the joint moves needed for hedging.

The proposed acceptance process produces four separate results:

  1. Price admissibility: static-violation frequency and magnitude at nodes, between them, and in the tails actually used; results before and after repair.
  2. Scenario quality: temporal evaluation of conditional distributions and joint moves, including crisis periods; comparisons against surface persistence and a simple factor model.
  3. Decision quality: identical available instruments and constraints for all hedges, delta and delta–vega comparisons, and error and tail loss on subsequent real observations.
  4. Execution feasibility: turnover, spreads, fees, financing, available liquidity, and net cash results under the same computational budget.

Thresholds for each result should be fixed before selecting the winning generator. Passing static checks permits the next evaluation stage; a hedge's usefulness requires its own decision test. Neither an architecture's name nor a smooth surface establishes that transition.

blog.disclaimer

Authors

Eugen Soloviov
Eugen Soloviov

Trading-systems engineer

Trading-systems engineer building bots since 2017: cross-exchange arbitrage (connected up to 30 venues), cointegration-based pairs arbitrage across spot and futures, scalping, news and sentiment-driven strategies, trend algorithms, and portfolio management and balancing algorithms. Also builds sub-millisecond order execution, big-data warehouses, backtesting engines, AI agents, and trading interfaces (incl. open-source profitmaker.cc). Stack: JS/TS, Python, Rust/Zig/Go, DevOps, backend, frontend, architecture.

Newsletter

Fique à frente do mercado

Assine nossa newsletter para insights exclusivos sobre trading com IA, análises de mercado e atualizações da plataforma.

Respeitamos sua privacidade. Cancele a inscrição a qualquer momento.