All insights
Product in Action10 August 20264 min read

How to Tell If Your Market Intelligence Is Lying to You

Four tests any AI-generated market analysis should survive before it reaches a committee — and what our own output scores against them, including where it doesn't.

A magnifying glass resting on printed data tables
Photo by Markus Winkler / Unsplash

The reasonable objection to any AI-generated market analysis is that it might be confidently wrong. It's the right instinct. A fluent paragraph asserting a day-delegate rate is indistinguishable, on the page, from a researched one — and a plausible fabrication is more dangerous than an obvious gap, because nobody checks it.

So rather than assert that ours doesn't hallucinate, here are the four tests worth applying to any such tool, and what our output actually scores against them.

1. Does it name things a local would recognise?

Generic output is the tell. "Several established four- and five-star properties compete in the corporate events segment" is a sentence that could describe ninety cities and has no informational content.

The test: does the analysis name the actual operators, districts and rates — in the vocabulary of your category? A venue read should be arguing about day-delegate rates and delegate bands; a hotel read about ADR, RevPAR and the brand-tier gap; a restaurant read about daypart demand, average check and cuisine saturation. A tool that describes all three in the same generic language isn't analysing any of them.

Our Manchester venue read names ETC.venues, Manchester Central, The Edwardian, The Lowry, Kimpton Clocktower and Victoria Warehouse, puts the gap in the 150–300 capacity tier within a five-minute walk of Spinningfields, and benchmarks Grade A rent at £43–45/sqft against London's £80+. The Singapore read names Marina Bay Sands, Suntec, Raffles and Capella, then locates the gap in daylight-flooded 50–100 capacity space that hotels relegate to windowless basement ballrooms.

You can check every one of those against your own knowledge of the city in about a minute. That's the point — a specific claim is falsifiable, and falsifiable claims are the only ones worth having.

2. Is it willing to return an answer you don't want?

The most important property of a scoring system is that it can say no — in every category, not just the ones where saying no is easy:

VerdictVenueHotel 5-starHotel boutique
Conditional Go5886123
No Go558240
Hold273532
Strong Go15513

Only 5 of 208 cities clear Strong Go for a 5-star hotel. Mean venue demand score: 59.9, range 22 to 92.

Lagos is the cleanest example of the discipline. It has one of the clearest quality chasms in the index — a genuine, specific gap in premium standalone venue supply — and it still returns No Go, because breakeven occupancy lands at 82% against a 55% target. The gap is real and the model fails anyway. A tool optimised to please would have called that a Conditional Go and licensed six figures of wasted diligence.

3. Does it distinguish what it knows from what it inferred?

A tool that presents every claim with equal confidence is hiding its weakest inputs among its strongest.

Every analysis carries a Grounded / Mixed / Estimated badge. Of the 155 venue analyses: 58 Grounded, 96 Mixed, 1 Estimated. That majority-Mixed result is not flattering, and it's the honest number — most cities carry some estimated inputs, and saying so is more useful than uniform confidence would be. Deep Dive reports go further and maintain an explicit unverifiedClaims list, surfaced in the UI and the PDF: the model flagging its own soft claims rather than smoothing them.

4. Is it one model's opinion, or a panel's?

A single model has a single set of blind spots and no way to know it. The LLM Council fans a city out to a panel of frontier models and synthesises a consensus. Currently 1,237 council member runs across 154 city consensuses — and critically, the numbers in a consensus are computed deterministically (median scores, majority-vote verdicts) before a chairman model writes the prose, so the narrative cannot move the figures.

Divergence is the useful output. When the panel splits on a city, that disagreement is a legitimate signal about genuine uncertainty in the market — information a single confident answer would have destroyed.

Where this still doesn't save you

Two honest limits.

Grounding lags breaking news. A grounded analysis is only as current as its sources, and the sources that would carry a days-old structural announcement haven't been written yet — there is no property research on last week's news. A refresh updates slow-moving market data well and is blind to very recent structural change. If a market on your shortlist changed structurally this quarter, price that event yourself; no tool's sources have caught up with it.

Single-run precision is softer than it looks. Re-run any city a few weeks apart and the verdict, the confidence and the shape of the case hold — but a headline market-size estimate can move materially between runs. Treat the verdict and the direction as robust; treat a single run's third significant figure as indicative, and never quote it to a committee as if it were surveyed.

The takeaway

The four questions above work on any vendor, including us. Ask them before the output reaches a committee — and be more suspicious of the tool that scores perfectly on all four than of the one that tells you where it's weak.

Book a demo and run these tests against a live analysis of your own market on the call.

See this run on the markets you’re actually weighing.

A short live session on your shortlist — we run the engine on the call and you keep the output.

Book a demo

More insights