Five Agents, One Opinion: How to Test the Value of Trading Debates

Five trading agents recommend buying an asset. One discusses fundamentals, another news, a third the chart. Their agreement looks persuasive until we discover that everyone read the same earnings summary, and the last two changed their decisions after seeing the first response. There are five votes; the number of independent reasons for the trade remains unknown.
Our reviews of AI Hedge Fund, TradingAgents, and AutoHedge examined roles and the path from analysis to a decision. The next step is to measure what discussion itself contributes. That requires saved responses from before the debate, a control without communication, and accounting for extra computation. This article separates research evidence, a mathematical illustration, and a proposed protocol. We have not run a trading experiment with LLMs for this article.
A convincing explanation and a useful decision need separate measurements

In Multi-Agent Debate for Explainable Trading, reasoning scores correlated with Sharpe at just 0.07, p = 0.29, across 210 runs. A separate intervention limiting convergence improved Sharpe by 0.14 over 35 scenarios, p = 0.028. The authors omitted trading costs and acknowledge possible historical leakage through model training. These are simulation results, not demonstrated earnings. Full text, Sections 3.2, 3.4.2, and 5.1.
This does not imply that agents should argue indefinitely. Random prompt changes can also produce opposing recommendations. Distance between portfolios does not establish which argument is correct. A narrower question is useful: did a verifiable objection improve a decision that would otherwise have remained uncorrected?
When Debate Helps, published on October 3, 2026, separates the availability of a correct proposal from its selection. The authors change the evidence presented while holding proposals fixed and test its effect on the answer. Their experiments concern reasoning tasks; applying the result to trading requires a separate experiment. Full text, Sections 3 and 5 and Appendix A.9.
For a trading system, this distinction is a useful research tool. First test the factual layer: was a number extracted correctly, does the news concern the right company, and was the publication available at decision time? Then test the forecast: does that fact help predict a predefined outcome? Finally test the portfolio: does the benefit survive constraints and execution? Success at the first level does not measure the third.
This extends the question in our Jev bot analysis: which component actually produced the result? Here, exchanging arguments is the component to switch on and off.
Independence starts before agents see each other's messages

I propose distinguishing three properties. Input diversity means different observations. Output diversity means different forecasts or weights. Weak error dependence means agents do not fail together too often. None establishes the others.
Two news stories can paraphrase one press release. Two agents reading the same document can find different calculation errors. Different portfolio weights can reflect different risk limits alone. Counting websites, roles, or dissenting answers therefore does not individually measure independent information.
Freeze every agent's initial response before discussion. A minimal record should contain:
| Field | What it lets you check |
|---|---|
| Document availability time and original-source identifier | Whether information arrived after the decision or sources repeat one another |
| Verifiable claim and location in the document | Whether the reference supports that particular fact |
| Initial forecast and confidence | What the agent decided before seeing its peers |
| Revised forecast and reason for the change | Which objection is associated with the revision |
| Verification outside the debate transcript | Whether a factual or calculation error was corrected |
The last field does not require another LLM. A table entry can be checked against the document; arithmetic can be recalculated in code. Our article on extracting signals from earnings calls covers obtaining features from text. Here, the record traces what discussion does to those features.
Another citation of the same fact adds no observation. However, a new calculation from existing data can be useful: one agent might notice that another confused a percentage change with percentage points. The record would then contain a corrected operation, rather than “three colleagues agreed.”
A reduction in disagreement after that correction is reasonable. A reduction without a newly verified reason warrants investigation. It does not yet prove imitation: several agents might independently fix the same mistake. This is why the record complements outcome measurement instead of replacing it.
Why five votes can be equivalent to one and a half

Consider our own mathematical illustration. Suppose each of n forecasts has a zero-mean error, equal variance σ², and equal pairwise correlation ρ. For the average forecast:
Var(average error) = σ² × [1 + (n − 1)ρ] / n
n_eff = n / [1 + (n − 1)ρ]
The first expression sums n variances and n(n − 1) covariances, then divides by n². The second defines the number of independent forecasts giving the same variance of the mean. It is a variance equivalence, not a count of facts or an estimate of voting accuracy.
With n = 5 and ρ = 0.6, n_eff ≈ 1.47. Increasing the number of agents to ten at the same correlation gives only 1.5625. A common systematic bias creates a further problem: averaging does not remove that bias.
To examine majority voting separately, define an artificial model. Each participant is correct with probability 0.6. With probability ρ, everyone receives one shared random answer with that accuracy; otherwise, their answers are independent. In this particular construction, the correlation between correctness indicators is exactly ρ.
| Specified correlation | Independent-vote equivalent by variance | Probability of a correct majority |
|---|---|---|
| 0.0 | 5.000000 | 0.682560 |
| 0.3 | 2.272727 | 0.657792 |
| 0.6 | 1.470588 | 0.633024 |
| 1.0 | 1.000000 | 0.600000 |
This is an exact enumeration of 32 vote combinations, without random simulation. The Python standard-library script checks total probability, each participant's accuracy, every pairwise correlation, and the variance of the mean:
python3 multi-agent-debate-independent-evidence.py
The table shows the effect of this chosen dependence structure. It does not describe actual LLM behavior or imply profitability. A different joint distribution with the same pairwise correlations can produce a different majority accuracy. In an actual experiment, error correlation must be estimated using held-out outcomes; agreement during one trading session is insufficient.
A controlled experiment: preserve the proposals, change the discussion

The following is a proposed protocol that we have not run. Its observation unit is a predefined event or decision time. Save the initial proposals for each event, then send the same set through several processing variants.
| Variant | What happens after the initial response | Purpose |
|---|---|---|
| Single agent | Use the response of a participant chosen in advance | Establish a starting point |
| Independent ensemble | Average initial forecasts without communication | Isolate the benefit of multiple responses |
| Independent review | Each agent checks its own response without seeing peers; average the revisions | Control for extra computation |
| Debate | Agents see peer arguments and revise their responses; average the revisions | Measure the contribution of interaction |
For the last two variants, fix the models, data, rounds, risk constraints, and the same token ceiling. Record actual consumption too: identical ceilings do not imply identical costs. Keep the averaging rule unchanged, or an aggregator change becomes mixed with the debate effect.
If an LLM judge is needed, add another pair: the same judge applied to initial responses and to responses after discussion. Improvement over a simple average cannot by itself reveal whether discussion or the judge helped. Access to new documents can be a separate intervention; do not bundle it with a change in communication.
For forecasts, predefine a verifiable outcome and horizon. For portfolios, also specify the available instruments, limits, rule converting forecasts to positions, and execution. An equally weighted portfolio under the same constraints, with cash accounted for, remains an essential reference: a complex system can lose to a simple allocation.
Choose prompts and stopping rules on a development period. Freeze them before testing the next time period. Historical LLM tests still face possible knowledge of later events from training; prospectively collecting answers before outcomes arrive gives stronger evidence.
Repeated generations measure model variability but do not create new market events. Five responses to one news story cannot count as five independent confirmations. Estimate uncertainty in differences between variants while accounting for shared dates and overlapping horizons, for example through time blocks; define grouping before seeing results. Invalid and late responses follow the same predefined handling rule, including their costs.
This experiment estimates the protocol's effect inside a specified simulator. It does not establish why the market moved or prove statistical independence of sources.
An extra round needs to justify its bill

Evaluate the economics of discussion alongside quality. The Agentic Trading survey, revised on October 4, 2026, examines temporal splits, costs, and execution rules in agent evaluation. It provides context for the protocol, rather than independent evidence that debate works. Full survey.
The cost of all model calls for one decision can be written without assuming current prices:
C_model = Σq (T_in,q × P_in,q + T_out,q × P_out,q) / 1,000,000
T denotes billed tokens for a call; P is the corresponding price per million tokens. Additional pricing categories, where applicable, become separate terms. Include the judge, retries after errors, and verification agents. Add tool costs separately.
For identical initial capital A and one comparison period:
Δ_net = Δ_PnL_gross − Δ_trading_costs − Δ_compute_and_tools
Δ_net_bps = 10,000 × Δ_net / A
Use the same currency throughout. If PnL already includes trading costs, do not subtract them twice. If discussion latency is reflected in a later execution price, do not charge the same effect again as an arbitrary latency penalty. Report the proportion of missed decision deadlines separately.
The final report should compare debate with the independent ensemble, then with independent review, at several predefined budgets. Include turnover, risk, net outcome, token use, and response time. The benefit might be a lower drawdown, but that comparison must account for cash holdings and market exposure.
Test the stopping rule too: for example, end discussion when the next round adds no verified fact or correction. That is an experimental candidate, not a universal recommendation. A cheap calculation correction may justify another round; a long repetition of the common thesis increases the bill without a measured gain.
The practical first step is to save initial responses and add a control without communication. If an advantage appears only against a single agent, the useful component may be the ensemble. If it survives comparison with independent review under comparable resources and after costs, there is reason to investigate debate's contribution further. Until then, five agreeing votes describe the protocol, not five independent pieces of evidence.
Research: Huang et al., arXiv:2609.29701v1 — arXiv submission history lists 2026-08-31; Zhao et al., arXiv:2610.04686v1 — 2026-10-03; Xia et al., arXiv:2605.19337v2 — revised 2026-10-04. Versions checked on 2026-10-08. Illustrations are conceptual; the voting table is our own synthetic calculation.
Penulis
Trading-systems engineer
Trading-systems engineer building bots since 2017: cross-exchange arbitrage (connected up to 30 venues), cointegration-based pairs arbitrage across spot and futures, scalping, news and sentiment-driven strategies, trend algorithms, and portfolio management and balancing algorithms. Also builds sub-millisecond order execution, big-data warehouses, backtesting engines, AI agents, and trading interfaces (incl. open-source profitmaker.cc). Stack: JS/TS, Python, Rust/Zig/Go, DevOps, backend, frontend, architecture.