One Model's Opinion Is Not Market Intelligence
A single AI has one set of blind spots and no way to detect them. How a panel of frontier models — 1,237 runs across 154 cities — turns disagreement into a usable signal.
The reasonable version of "isn't this just AI?" is not a complaint about AI. It's a complaint about singularity of view. One model, however capable, has one training distribution, one set of priors and one characteristic way of being wrong — and no mechanism for noticing any of it. Ask it twice and it will agree with itself, which feels like corroboration and isn't.
That's a genuine problem for a document heading to an investment committee, and it needs a structural answer rather than a reassuring one.
The panel
The LLM Council fans a single city out to a panel of frontier models — currently thirteen on the roster, spanning Anthropic, Google, OpenAI, xAI, DeepSeek, Qwen, MiniMax, Moonshot, Z.AI and others — each producing a full, independent analysis of the same market against the same schema.
To date: 1,237 individual member runs across 154 city consensuses.
Each member returns the complete structured analysis, not a vote. So you can open any single model's full reasoning and read what it thought and why — the panel is inspectable, not a black box that emits an average.
Why the numbers are computed, not written
This is the part that matters most and gets least attention.
A naive panel implementation asks a model to "summarise what the others said," which lets a persuasive narrative drift away from the underlying figures. Ours doesn't work that way:
- The numbers are consolidated deterministically. Median scores. Majority-vote verdicts. Frequency-ranked, deduplicated competitor and gap lists. Per-band majorities for capacity demand. This is arithmetic, and it is auditable.
- Only then does a chairman model write the prose — given the locked figures and every member's reasoning.
- The prose is spliced back in such that it cannot perturb a single number, and the merged object is re-validated against the schema.
If the chairman's output fails validation, the system falls back to the deterministic skeleton. The narrative serves the numbers; it can never quietly rewrite them.
Disagreement is the product
The instinct is to treat model disagreement as a defect to be smoothed away. It's the opposite — it's the most honest signal available.
When eleven models converge on Strong Go for a market, that convergence means something. When they split, that split is telling you the market is genuinely ambiguous — that reasonable, well-informed analysis can land in different places. A single confident answer would have destroyed that information and replaced it with false precision.
So the consensus report carries a consensusMeta block: member count, verdict agreement, and per-score bands showing the median alongside the actual min–max spread across the panel. A committee can see at a glance whether it is looking at a settled question or a contested one — before deciding how much weight the verdict deserves.
That's the thing a solo model structurally cannot provide. Not a better answer — an honest estimate of how confident anyone should be in the answer.
What it costs to be this thorough
Worth being straightforward: this is expensive, which is why it's a deliberate feature rather than the default path. Convening a full panel and synthesising a consensus runs a couple of dollars per city in model costs, against roughly thirty cents for a single grounded analysis.
Nearly every venue market in the index has been through it — 154 consensuses against 155 cities — because for a decision of this size, the marginal cost of a second, third and eleventh opinion is trivially small next to the cost of being confidently wrong about a market you then spend eight figures entering.
The honest limits
Models share biases. Frontier models are trained on overlapping data, so a panel is not a set of fully independent observers. Consensus reduces idiosyncratic error; it does not eliminate systematic error that every model shares. A market that the entire internet describes inaccurately will be described inaccurately by all thirteen.
Agreement is not truth. Eleven models agreeing tells you the claim is consistent with the general weight of available evidence. That is genuinely valuable and it is not the same as verified.
It doesn't fix recency. If an event is too recent to have been written about, no number of models will find it — thirteen models reading sources that don't exist yet miss exactly the same news as one. For a market that changed structurally last week, the panel needs the same thing a single model does: a human holding the output and the headlines side by side.
The takeaway
The right question to put to any AI-derived market view isn't "which model produced this." It's "how would I know if it were wrong?"
A single model can't answer that about itself. A panel — with deterministic numbers, an inspectable member-by-member trail, and a published spread — at least tells you where the uncertainty actually sits.
Book a demo and we'll convene the council live on a market you know well, so you can judge the panel against your own knowledge.
See this run on the markets you’re actually weighing.
A short live session on your shortlist — we run the engine on the call and you keep the output.
Book a demo