Does an AI Visibility Score Actually Mean Anything?
Half the fights about AI-search metrics come from filing two different things under one word. There are real measures — server logs, referral traffic, bookings, the “how did you hear about us?” line on your booking form — and there are built statistics: the prompt-panel scores that count mentions, citations and sources. One is ground truth. The other is a constructed proxy. Both are useful; confusing them is where the trouble starts.
I run both on hotels I built myself, so this is the honest version: which numbers to trust as outcomes, which to treat as early signals, what actually moves them (controlled experiments — and yes, a few grubby tactics), and how far the measurement still has to go.
Two kinds of number, routinely confused
When someone says “AI visibility is unmeasurable” and someone else says “our score went up 30%,” they are usually both right, because they are talking about different instruments. It is worth separating them on the table before arguing about either.
Ground truth. Things that happened in the world: a log line, a session, a guest typing your name, a booked room.
- · Trust them as outcomes.
- · Lagging, and often under-counted.
- · You cannot fake them; you also cannot rush them.
Constructed proxies. You write a panel of prompts, run them, and count how often you are mentioned, cited, and beside whom.
- · Trust them as leading indicators.
- · Early, and movable.
- · Synthetic by construction — their honesty is in the method.
The useful question is never “is the score a booking?” — of course it is not. It is “does the built statistic move before the real measure, and do the two agree when I act?” When they do, you have something close to attribution. The whole rest of this guide is making the two families check each other.
The real measures: what you can take to the bank
These are the numbers a sceptic already trusts, ordered from the cash register backwards. Each tells you something true and hides something important — which is exactly why no single one of them is enough.
Tells you: The only number that pays the bills, and the final word on whether any of this worked.
Hides: It is also the quietest about cause. A booked room rarely says which channel sent it, and AI’s fingerprints are usually wiped off by the time the reservation lands.
Tells you: One line on the booking form, asked of every guest — the cheapest causal read you will ever get. On Hotel Ranque, 21 of 52 answers said AI search, against 27 for Google.
Hides: Self-reported and lossy: plenty of guests never realise an assistant planted the name, so it undercounts. Treat it as a floor.
Tells you: The hidden channel. Roughly 94% of people who ask an assistant still circle back to Google before booking, so demand the AI created shows up as someone typing your name.
Hides: Real and large, but almost always misfiled. Your analytics credits “branded search” for a sale the model actually set up.
Tells you: The clicks that arrive straight from an answer. When ChatGPT began embedding hotel links on 7 May 2026 these jumped 62% overnight across a 17,000-hotel panel and doubled week over week — net-new, not reshuffled.
Hides: Only the visits where the guest clicked through. The far larger group who read the answer and went straight to Google never appears here.
Tells you: The bedrock. The ChatGPT-User, Claude-User and PerplexityBot user-agents in your raw access logs are the only direct proof a model fetched your page — no sampling, no inference.
Hides: A fetch is not a recommendation. It tells you the model read you, not that it told a guest about you. Everything above is downstream of this.
Whatever number you land on, it is too low
Notice what every real measure in that list has in common: they all err in the same direction. Logs catch the model’s fetch but not the recommendation it made afterwards. Referral catches the guest who clicked but not the larger crowd who read the answer and opened a fresh tab. Analytics buries the rest in branded and direct, because roughly 94% of people loop back to Google before they book. There is no instrument anywhere in the stack that over-reports AI. When every gauge you own reads low, the true value is not somewhere in the middle — it is past all of them.
And the thing being undercounted is not small. Getting on for a billion people put a question to ChatGPT every week, a base that more than doubled in a year — and Gemini is right behind it, approaching a billion a month, before you even add Perplexity, Copilot and the rest. Travel discovery is happening at that scale whether or not your dashboard has caught up. The expensive mistake is not over-investing in a vanity metric; it is reading “AI is 0.9% of my traffic,” shrugging, and missing that your instrument simply cannot see most of the channel.
This is also why, in a black box, the quality of your measurement beats the quantity of it. You cannot instrument the inside of a stranger’s private ChatGPT session, ever. So the highest-fidelity probe you have is the guest’s own memory — and asking “how did you hear about us?” on the booking form is not a soft consolation prize, it is frequently the most direct read of the channel you will get. It is the reason that one survey is the number I trust most in this entire guide.
The built statistics: synthetic, and still worth keeping
Now the contested family. A visibility score is constructed: you cannot watch real travellers talk to private assistants, so you build a panel of prompts, run each one repeatedly per engine, and measure three things.
Are you named in the answer a guest reads — and is the name linked, or just text? A single run is close to a coin toss; repeated and averaged it becomes a stable rate — the same hotel holds the top slot in 50.5% of identical reruns, up to 96.1% in tight markets.
Is your hotel linked — and to your own site (a direct booking path) or to an OTA listing of you? The clickable, traffic-bearing version of a mention, and the built statistic that maps most directly onto a referral session in your logs.
Who gets cited instead of you — the OTAs, the review sites, the rival hotels the answer leaned on. The most useful by-product of the whole exercise: a per-engine list that doubles as your to-do list.
The fair objection is that you wrote the questions, so you are grading yourself on a test of your own making. The defence is method, not denial. A panel modelled on how the engines actually behave — location queries trigger a real web search 98% of the time against 8% for definitional ones, then fan out into four or five sub-queries — and anchored to the destinations you genuinely compete in is sampling real intent, the way survey research samples a population. The same method replicates on niches I have no stake in, which is what you would expect from a measurement and not from a flattering coincidence.
Getting better isn’t just “more content” — it’s controlled experiments
Ask any dashboard how to improve and it says: publish more. It is not wrong — more good, structured content genuinely helps, the way eating vegetables helps. It is just unfalsifiable advice. You can always be told to write another page, and you can never tell which page did anything. That is the lazy version, and it is the version the critics are right to mock.
The version that produces knowledge is a controlled experiment, where the two families of number finally meet:
- 1Baseline both
Track the built score and the real measures for two or three weeks untouched, so you know the natural wobble in each.
- 2Change exactly one thing
One structured page, one schema fix, one review push, one third-party placement — timestamped. One lever, or you learn nothing.
- 3Watch the per-engine built statistic
Look at the engine you targeted, not a blend. A change that wins ChatGPT may do nothing on Gemini, because they ground on different sources.
- 4Confirm in the real measures
Did ChatGPT-User log hits and branded search move in the same window? Two independent families agreeing is as close to proof as this gets.
- 5Claim timing, not cause
Report “I changed X on this date and both lines moved over the next runs.” Show the sequence; let it persuade. Overclaiming is the original sin here.
Some grubby things genuinely work
Here is the part nobody likes to put on a slide. The single biggest lever is often not your beautiful website at all, because the engines read about you more than they read you. The source mix makes that embarrassingly clear: Grok grounds its answers on Reddit 54.5% of the time and Facebook 63.5%, and before its March cull ChatGPT pulled 14% of its hotel sources straight from Reddit. So being talked about in the right forum thread is, annoyingly, real leverage — frequently more than a month of on-site copy.
The catch is the other edge of the same blade: you do not control it, and it is fragile. When ChatGPT moved to 5.3 on 5 March 2026, the share of hotel answers leaning on user-generated platforms collapsed from 21% to 2% overnight — Reddit alone down 71%. A tactic that was carrying your visibility one week was switched off by a model update the next. Grubby tactics are real, but treat them as weather, never as foundation.
This is also why the built statistics earn their keep defensively. The competition column is what tells you a third of your visibility was riding on one platform before a model update pulls the rug — a concentration risk you would never see in a booking report.
Where the methodology still has to go
None of this is finished. The gap between a vanity score and a real instrument is mostly method, and there are three fronts where the work is still live.
Away from amenity-keyword panels that flatter you, toward need-based, geography-anchored prompts that follow a guest across the whole booking conversation rather than scoring turn one. The mechanics are their own measurement guide.
Treating server-log parsing of AI user-agents as the primary instrument rather than a footnote, because GA4 structurally undercounts and the ground moves fast — the 7 May link rollout reshaped the numbers in a single day. The traffic study is the worked example.
Making “how did you hear about us?” a deliberate instrument, and learning to read the AI-driven slice hiding inside branded search — the blind spot where most of the value sits because 94% of users loop back to Google before they book.
When both families line up: a hotel built on the score
The cleanest test I could run was to build a hotel’s demand from nothing — Hotel Ranque, no OTA contracts, no ad budget — with AI visibility as effectively the only channel, and watch whether the built statistic and the real measures told the same story. (The name is a pun on rank. I am not above that.)
The built score moved first, in the order you would predict: a long-tail Perplexity appearance around week four, consistent presence across the big engines by week twelve, generic “best boutique hotel in Paris” placements past week twenty. Then the real measures followed the same curve — enquiries from a trickle to more than the rooms could hold, which is when I stopped. And the survey on the booking form closed the loop in the guests’ own words:

One property is one data point, and I would not generalise an industry from it. But it is the clean case the whole argument needs: a built statistic that led, real measures that confirmed, and guests who said so at check-in. You do not get that alignment if the score is measuring nothing.
FAQ
Further reading
This guide is the “what counts as measuring” layer. For the mechanics of building the built statistics without fooling yourself, the prompt-tracking guide is the companion; the figures above come from my own research library.
How to Measure AI Hotel Traffic
The real measures: server logs as gold standard, the attribution gap, and why analytics undercounts AI.
The ChatGPT Direct-Traffic Explosion
What happened to 17,000 hotels the day ChatGPT started embedding links: +62% overnight.
ChatGPT 5.3 Halved Its Hotel Sources
The March 2026 cutover that collapsed user-generated sources from 21% to 2% — why dirty tactics are fragile.
AI Rankings Consistency Study
Why the built score is signal not noise — 50.5% to 96.1% top-spot stability.