Measurement Guide · June 2026

Does an AI Visibility Score Actually Mean Anything?

Half the fights about AI-search metrics come from filing two different things under one word. There are real measures — server logs, referral traffic, bookings, the “how did you hear about us?” line on your booking form — and there are built statistics: the prompt-panel scores that count mentions, citations and sources. One is ground truth. The other is a constructed proxy. Both are useful; confusing them is where the trouble starts.

I run both on hotels I built myself, so this is the honest version: which numbers to trust as outcomes, which to treat as early signals, what actually moves them (controlled experiments — and yes, a few grubby tactics), and how far the measurement still has to go.

21 / 52
guests who named AI search on a from-scratch hotel’s booking form
+62%
overnight jump in AI referral sessions when links went live
50.5%
top-spot stability in the built score — structure, not dice

Two kinds of number, routinely confused

When someone says “AI visibility is unmeasurable” and someone else says “our score went up 30%,” they are usually both right, because they are talking about different instruments. It is worth separating them on the table before arguing about either.

Real measures

Ground truth. Things that happened in the world: a log line, a session, a guest typing your name, a booked room.

  • · Trust them as outcomes.
  • · Lagging, and often under-counted.
  • · You cannot fake them; you also cannot rush them.
Built statistics

Constructed proxies. You write a panel of prompts, run them, and count how often you are mentioned, cited, and beside whom.

  • · Trust them as leading indicators.
  • · Early, and movable.
  • · Synthetic by construction — their honesty is in the method.

The useful question is never “is the score a booking?” — of course it is not. It is “does the built statistic move before the real measure, and do the two agree when I act?” When they do, you have something close to attribution. The whole rest of this guide is making the two families check each other.

The real measures: what you can take to the bank

These are the numbers a sceptic already trusts, ordered from the cash register backwards. Each tells you something true and hides something important — which is exactly why no single one of them is enough.

Bookings

Tells you: The only number that pays the bills, and the final word on whether any of this worked.

Hides: It is also the quietest about cause. A booked room rarely says which channel sent it, and AI’s fingerprints are usually wiped off by the time the reservation lands.

“How did you hear about us?”

Tells you: One line on the booking form, asked of every guest — the cheapest causal read you will ever get. On Hotel Ranque, 21 of 52 answers said AI search, against 27 for Google.

Hides: Self-reported and lossy: plenty of guests never realise an assistant planted the name, so it undercounts. Treat it as a floor.

Branded & direct search

Tells you: The hidden channel. Roughly 94% of people who ask an assistant still circle back to Google before booking, so demand the AI created shows up as someone typing your name.

Hides: Real and large, but almost always misfiled. Your analytics credits “branded search” for a sale the model actually set up.

Referral sessions

Tells you: The clicks that arrive straight from an answer. When ChatGPT began embedding hotel links on 7 May 2026 these jumped 62% overnight across a 17,000-hotel panel and doubled week over week — net-new, not reshuffled.

Hides: Only the visits where the guest clicked through. The far larger group who read the answer and went straight to Google never appears here.

Server logs

Tells you: The bedrock. The ChatGPT-User, Claude-User and PerplexityBot user-agents in your raw access logs are the only direct proof a model fetched your page — no sampling, no inference.

Hides: A fetch is not a recommendation. It tells you the model read you, not that it told a guest about you. Everything above is downstream of this.

Read the list top to bottom and the blind spot jumps out. The conversion click lands on branded search; the booking credits “direct.” The one place the AI step is unambiguous is the log file, which most hotels never open. Standard analytics watched a single hotel’s AI share crawl from 0.07% to 4.41% over fourteen months — and still filed most of the resulting demand under Google.

Whatever number you land on, it is too low

Notice what every real measure in that list has in common: they all err in the same direction. Logs catch the model’s fetch but not the recommendation it made afterwards. Referral catches the guest who clicked but not the larger crowd who read the answer and opened a fresh tab. Analytics buries the rest in branded and direct, because roughly 94% of people loop back to Google before they book. There is no instrument anywhere in the stack that over-reports AI. When every gauge you own reads low, the true value is not somewhere in the middle — it is past all of them.

And the thing being undercounted is not small. Getting on for a billion people put a question to ChatGPT every week, a base that more than doubled in a year — and Gemini is right behind it, approaching a billion a month, before you even add Perplexity, Copilot and the rest. Travel discovery is happening at that scale whether or not your dashboard has caught up. The expensive mistake is not over-investing in a vanity metric; it is reading “AI is 0.9% of my traffic,” shrugging, and missing that your instrument simply cannot see most of the channel.

~1B
weekly ChatGPT users
~1B
monthly Gemini users
growth in a single year

This is also why, in a black box, the quality of your measurement beats the quantity of it. You cannot instrument the inside of a stranger’s private ChatGPT session, ever. So the highest-fidelity probe you have is the guest’s own memory — and asking “how did you hear about us?” on the booking form is not a soft consolation prize, it is frequently the most direct read of the channel you will get. It is the reason that one survey is the number I trust most in this entire guide.

The mistake is asymmetric, which is what makes it worth naming. Read the channel as bigger than it is and you waste a little effort; read it as smaller — which every default instrument quietly nudges you to do — and you under-invest in front of a billion-user shift. Track AI bookings deliberately, treat the visible number as a floor, and ask the guests directly. The picture that comes back is almost always larger than the dashboard’s.

The built statistics: synthetic, and still worth keeping

Now the contested family. A visibility score is constructed: you cannot watch real travellers talk to private assistants, so you build a panel of prompts, run each one repeatedly per engine, and measure three things.

Mention

Are you named in the answer a guest reads — and is the name linked, or just text? A single run is close to a coin toss; repeated and averaged it becomes a stable rate — the same hotel holds the top slot in 50.5% of identical reruns, up to 96.1% in tight markets.

Citation

Is your hotel linked — and to your own site (a direct booking path) or to an OTA listing of you? The clickable, traffic-bearing version of a mention, and the built statistic that maps most directly onto a referral session in your logs.

Competition

Who gets cited instead of you — the OTAs, the review sites, the rival hotels the answer leaned on. The most useful by-product of the whole exercise: a per-engine list that doubles as your to-do list.

The fair objection is that you wrote the questions, so you are grading yourself on a test of your own making. The defence is method, not denial. A panel modelled on how the engines actually behave — location queries trigger a real web search 98% of the time against 8% for definitional ones, then fan out into four or five sub-queries — and anchored to the destinations you genuinely compete in is sampling real intent, the way survey research samples a population. The same method replicates on niches I have no stake in, which is what you would expect from a measurement and not from a flattering coincidence.

The reason to keep a synthetic number at all: it is early and it is movable. A booking tells you what happened last month; the mention rate tells you what the engines think this week, split by engine, while you can still do something about it. That is the entire job of a leading indicator.

Getting better isn’t just “more content” — it’s controlled experiments

Ask any dashboard how to improve and it says: publish more. It is not wrong — more good, structured content genuinely helps, the way eating vegetables helps. It is just unfalsifiable advice. You can always be told to write another page, and you can never tell which page did anything. That is the lazy version, and it is the version the critics are right to mock.

The version that produces knowledge is a controlled experiment, where the two families of number finally meet:

  1. 1
    Baseline both

    Track the built score and the real measures for two or three weeks untouched, so you know the natural wobble in each.

  2. 2
    Change exactly one thing

    One structured page, one schema fix, one review push, one third-party placement — timestamped. One lever, or you learn nothing.

  3. 3
    Watch the per-engine built statistic

    Look at the engine you targeted, not a blend. A change that wins ChatGPT may do nothing on Gemini, because they ground on different sources.

  4. 4
    Confirm in the real measures

    Did ChatGPT-User log hits and branded search move in the same window? Two independent families agreeing is as close to proof as this gets.

  5. 5
    Claim timing, not cause

    Report “I changed X on this date and both lines moved over the next runs.” Show the sequence; let it persuade. Overclaiming is the original sin here.

How well does this work? A bit. The engines refresh their sources weekly, and much of what an answer cites lives on pages you influence but do not own, so you get timing and correlation, rarely a clean cause. But the built statistic is the closest observable to your action — bookings sit behind price, season and a dozen confounds — so it is where any honest attempt at attribution starts. A bit of real evidence beats a lot of confident hand-waving.

Some grubby things genuinely work

Here is the part nobody likes to put on a slide. The single biggest lever is often not your beautiful website at all, because the engines read about you more than they read you. The source mix makes that embarrassingly clear: Grok grounds its answers on Reddit 54.5% of the time and Facebook 63.5%, and before its March cull ChatGPT pulled 14% of its hotel sources straight from Reddit. So being talked about in the right forum thread is, annoyingly, real leverage — frequently more than a month of on-site copy.

The catch is the other edge of the same blade: you do not control it, and it is fragile. When ChatGPT moved to 5.3 on 5 March 2026, the share of hotel answers leaning on user-generated platforms collapsed from 21% to 2% overnight — Reddit alone down 71%. A tactic that was carrying your visibility one week was switched off by a model update the next. Grubby tactics are real, but treat them as weather, never as foundation.

This is also why the built statistics earn their keep defensively. The competition column is what tells you a third of your visibility was riding on one platform before a model update pulls the rug — a concentration risk you would never see in a booking report.

Where the methodology still has to go

None of this is finished. The gap between a vanity score and a real instrument is mostly method, and there are three fronts where the work is still live.

Prompt methodology

Away from amenity-keyword panels that flatter you, toward need-based, geography-anchored prompts that follow a guest across the whole booking conversation rather than scoring turn one. The mechanics are their own measurement guide.

Traffic measurement

Treating server-log parsing of AI user-agents as the primary instrument rather than a footnote, because GA4 structurally undercounts and the ground moves fast — the 7 May link rollout reshaped the numbers in a single day. The traffic study is the worked example.

AI attribution

Making “how did you hear about us?” a deliberate instrument, and learning to read the AI-driven slice hiding inside branded search — the blind spot where most of the value sits because 94% of users loop back to Google before they book.

When both families line up: a hotel built on the score

The cleanest test I could run was to build a hotel’s demand from nothing — Hotel Ranque, no OTA contracts, no ad budget — with AI visibility as effectively the only channel, and watch whether the built statistic and the real measures told the same story. (The name is a pun on rank. I am not above that.)

The built score moved first, in the order you would predict: a long-tail Perplexity appearance around week four, consistent presence across the big engines by week twelve, generic “best boutique hotel in Paris” placements past week twenty. Then the real measures followed the same curve — enquiries from a trickle to more than the rooms could hold, which is when I stopped. And the survey on the booking form closed the loop in the guests’ own words:

A guest survey titled “How did you hear about us?” with 52 answers: Google Search 27, AI Search 21, Other 4.
Hotel Ranque’s “how did you hear about us?” line, straight from the booking flow: 21 of 52 guests credited AI search, nearly level with Google — on a hotel with no other way of being found. Self-reported, so a floor, not a ceiling.

One property is one data point, and I would not generalise an industry from it. But it is the clean case the whole argument needs: a built statistic that led, real measures that confirmed, and guests who said so at check-in. You do not get that alignment if the score is measuring nothing.

FAQ

A real measure is something that happened in the world and that you can take to the bank: a server-log hit from ChatGPT-User, an AI referral session, a guest typing your name into Google, a booking, a 'how did you hear about us?' answer. A built statistic is a constructed proxy — you write a panel of prompts, run them across the engines, and count mentions, citations and the sources cited alongside you. Real measures are ground truth but lagging and under-counted; built statistics are synthetic but early and movable. You need both, and the point is to make them check each other.

Further reading

This guide is the “what counts as measuring” layer. For the mechanics of building the built statistics without fooling yourself, the prompt-tracking guide is the companion; the figures above come from my own research library.

Summarize with AI

ChatGPTPerplexityClaudeGeminiGrok