Synthetic order books: test the evaluator before trusting the data

Suppose a generator reproduces the spread distribution and the volume distribution at every price level. Its best quote also follows a convincing path. Is the synthetic book ready for execution research?
Before answering, try a different question: can you deliberately damage the data without worsening those scores? A validation system that accepts a known defect has revealed a gap in its coverage, even if the generator under review is sophisticated.
The useful distinction is between three objects: the distribution of each variable, the way variables move together, and the behavior of a trading policy using those variables. The following example separates them with eight invented snapshots. It requires no model training and makes no claims about observed trading performance.
Matching distributions does not recover joint behavior

Write the displayed ask volumes as . A histogram of answers how frequently each best-ask size occurs. A histogram of answers the corresponding question one level deeper. Neither tells us whether both levels become thin together. Sorting each column independently discards that pairing.
There is a second loss of information. Even the complete distribution of snapshots cannot distinguish a long thin-liquidity episode from alternating thin and thick states if both sequences contain the same snapshots. An algorithm deciding whether to wait needs some description of transitions, not just the set of available states.
In LOB-ID, published on 13 August 2026, Bacalum and coauthors extract features from L2 windows with a trained DeepLOB network and compare their distributions. Their deeper-book rearrangement evades the tested statistical and impact-response checks while MIND detects it. The experiment excludes LOB-Bench's discriminator. The study covers five HKEX equities overall; the joint-structure attack uses three, over six test days. It evaluates L2 structure, not full message-level execution. LOB-ID, sections 3, 4.3.2 and 5.
The architecture itself is covered in our DeepLOB article. Here the engineering question is which changes an evaluator can see.
LOB-Bench already includes conditional statistics, response functions and a trained discriminator; describing it as a collection of histograms would misrepresent the benchmark. Its downstream prediction test also compares training with and without generated data on held-out real observations. These are distinct evaluation components. LOB-Bench, sections 4–4.3, version 2.
The practical inference is narrower than "statistics do not work": a specified collection of statistics establishes similarity only in the aspects it measures.
An exact check: the same volume, a different sweep cost

Construct eight ask-side snapshots with prices fixed at 100.00, 100.01 and 100.02 USD. One tick is 0.01 USD. The third level always contains four units. Keep the bid side fixed below the best ask. There are two arrangements of the first two levels:
| Snapshots | Reference: | Depth permutation: |
|---|---|---|
| First four | ||
| Last four |
Every individual volume column has exactly the same distribution in both datasets. The best-ask volume sequence is also unchanged. What changes is the pairing: small best-ask volume coincides with small deeper volume in the reference and large deeper volume in the permutation.
Now buy four units from each frozen snapshot, starting from an untouched book every time. This is a static depth probe, not eight sequential trades. Measure average purchase price above the best ask, in ticks per unit.
At a reference snapshot containing , one unit costs zero extra ticks, one costs one extra tick, and two cost two extra ticks. The premium is ticks. At it is ticks. Equal weighting gives 0.75 ticks.
After the depth permutation, costs ticks and costs . The average falls to 0.50 ticks. Reaching the third level changes from four snapshots out of eight to none. No marginal volume histogram can reveal this difference because those histograms are identical by construction.
A separate permutation alternates the original complete snapshots: thin, thick, thin, thick, and so on. This preserves both the snapshot distribution and the static sweep-cost distribution. Yet best-ask volume now changes at all seven adjacent transitions, compared with one transition in the reference.
| Diagnostic | Reference | Depth permutation | Time permutation |
|---|---|---|---|
| Per-level volume distributions | Same | Same | Same |
| Mean premium, ticks per unit | 0.75 | 0.50 | 0.75 |
| Fraction reaching level 3 | 1/2 | 0 | 1/2 |
| Fraction of adjacent transitions changing | 1/7 | 1/7 | 1 |
The downloadable standard-library Python example verifies these equalities with exact fractions. It contains no historical observations, queue reconstruction, fees, latency or simulated market response. The counterexample proves insufficiency of these marginal checks; it does not measure the error of a particular commercial generator.
Why means and covariances can also be fooled

A learned representation does not settle how its distribution should be compared. FID compares Gaussian approximations using means and covariances. MIND instead averages squared Wasserstein distances along random one-dimensional projections of the embeddings. This distinction is defined in the original MIND paper; its results concern the chosen representation, not a guarantee that every property of the original data is retained. MIND, sections 2–3.
An independent scalar example makes the limitation of moment matching explicit. Give equal probability to each entry in:
Both means are zero and both population variances are one. Their fitted Gaussian distributions are identical, so the Gaussian moment distance is zero. Nevertheless, half the mass of sits at zero, while has no mass there.
For equally weighted scalar samples, sort both arrays and average the squared pairwise differences:
The script checks this calculation too. It is a scalar teaching example, not an implementation of the multidimensional LOB-ID pipeline. The exact zero belongs to the constructed distributions, not to the paper's empirical attack.
There is another logically separate limitation. Imagine a feature extractor that always returns only the spread. Our two depth arrangements have the same spread, so their embeddings are identical. No distance applied afterward can recover the omitted volume relationship. A better distributional comparison cannot compensate for a representation that removes the relevant information.
This suggests two independent questions for any learned evaluator: does its representation retain the defect of interest, and does its distance detect that defect once retained? A failure at either stage is enough to produce a misleadingly good score.
Test the evaluator before scoring the generator

The following acceptance protocol is an engineering proposal based on the counterexamples above, not a reproduction of the authors' experiments.
First, write down the object being approved. A collection of independent snapshots, a sequence in event time, a sequence in clock time and an interactive simulator are different deliverables. Define the represented depth, sampling rule, instrument and intended decision. Otherwise the phrase "realistic synthetic data" has no testable scope.
Next, build deliberate corruptions with explicit invariants:
| Alteration | Preserve | Ask the evaluator to detect |
|---|---|---|
| Reorder complete snapshots | All snapshot statistics | Changed transitions and episode lengths |
| Rearrange deeper volume observations | Per-level marginals, fixed best quotes | Changed dependence across levels |
| Change time gaps only | Ordered book states | Changed behavior in clock time |
| Repeat a small set of complete windows | Plausible individual windows | Lost diversity and repeated paths |
These are diagnostic constructions, not claims that every rearranged sequence is an exchange-valid event stream. Check structural constraints separately. An invalid price ordering is useful for testing an invariant checker, but it should not become the easy shortcut through which a temporal evaluator "detects" a sophisticated dependency failure.
Then establish a real-versus-real reference. Compare disjoint historical blocks selected for the same intended use. Report both ordinary sampling variation and differences across regimes. Choose block separation with the window length and decision horizon in mind: two overlapping windows are poor candidates for independent evidence. Use repeated blocks to describe uncertainty rather than interpreting a single score as a universal pass threshold.
Freeze the feature extractor, preprocessing, sample count, window construction and projection settings before comparing generator versions. Otherwise a changed score has two possible causes: the data changed, or the measuring procedure changed. Keep a record of both versions so the comparison is reproducible.
Finally, reserve some corruptions for a later check. If you repeatedly design the evaluator around the same damaged samples, passing those samples becomes another development objective. A useful report states which defects were detected, which escaped, and which were never tested. "The evaluator rejected the depth permutation but missed the clock-time distortion" is actionable; "quality score 0.9" without a defined scale is not.
Execution readiness needs a separate check

Our exact example already provides a small bridge to execution: a volume pairing changes the cost of a fixed-size static sweep. It does not establish how a market would respond after that sweep. A sequence that looks plausible without intervention may still be unsuitable for evaluating an order that changes subsequent states.
A newer preprint, dated 15 September 2026, makes a useful distinction here. Linna and coauthors compare forecasts before and after mechanically valid message injections, then use controlled simulation and historical event alignment for validation. They explicitly interpret the result as model-implied scenario-conditioned impact, not an unrestricted causal market effect. Scenario-conditioned market impact modeling, sections 3.1, 4.5 and 6.
For a proposed execution application, freeze a small set of policies and the assumptions they require. A static sweep probes available depth. Waiting probes transitions and replenishment assumptions. A passive order also requires a specified matching and queue model. Record those dependencies before comparing outputs; do not silently change the execution rules to make one generated dataset easier to use.
Evaluate the quantities attached to that use: completion, time to completion, cost distributions and the response to different order sizes. Compare like-for-like starting conditions and separate ordinary states from the thin-liquidity states in which the policy actually makes consequential decisions. Our implementation shortfall article explains the cost measurement side; here those measurements serve as additional acceptance criteria for the synthetic environment.
An actionable release decision might read: "Accepted for static depth stress tests within the evaluated instruments and size range; temporal waiting behavior remains unvalidated." That statement defines a usable scope. Extending it to passive execution or adaptive execution requires evidence for the additional mechanisms.
The immediate experiment is inexpensive: run one deliberate depth permutation and one time permutation through the existing evaluator. Their purpose is to locate an observable blind spot before a favorable aggregate score becomes permission to trust a trading result.
Auteurs
Trading-systems engineer
Trading-systems engineer building bots since 2017: cross-exchange arbitrage (connected up to 30 venues), cointegration-based pairs arbitrage across spot and futures, scalping, news and sentiment-driven strategies, trend algorithms, and portfolio management and balancing algorithms. Also builds sub-millisecond order execution, big-data warehouses, backtesting engines, AI agents, and trading interfaces (incl. open-source profitmaker.cc). Stack: JS/TS, Python, Rust/Zig/Go, DevOps, backend, frontend, architecture.