Live · updates every Monday · data through July 27, 2026AI Search · Citations · Hotels

Retrieved, rejected, cited:AI's citation throughput in hotel search

TL;DR: Sources are the input; citations are the throughput. Every week our AI Hotel Landscape pipeline logs both what each AI assistant retrieves and what it actually cites. This week ChatGPT cited 18.8% of the hotel sources it pulled; Google AI Mode just 8.7%. And while cross-industry citation data has AI engines rejecting the overwhelming majority of the Reddit pages they retrieve, in hotel search the inversion is total: Reddit has the highest citation throughput of any major domain — once retrieved, a Reddit thread is more likely to be cited than Tripadvisor, Booking.com or any OTA.

NS
Nicolas Sitter
Published July 22, 2026 · refreshed weekly
18.8%
ChatGPT throughput
8.7%
Google AI Mode throughput
52,248
Sources retrieved this week
59%
Reddit throughput on ChatGPT
Read the Report

Executive Summary

Getting retrieved is table stakes. Getting selected is the game.

When an AI assistant answers a hotel question, it first pulls a pool of web sources — the candidate set — then cites a fraction of them in the answer. We call that fraction the citation throughput: citations ÷ retrieved sources — the ratio sometimes called the selection rate. It is the single number that separates “the AI read my page” from “the AI showed my page to a traveller.”

This page computes it live, every week, from the same 616-prompt, 56-destination corpus behind the AI Hotel Landscape. Two engines expose the full funnel (ChatGPT and Google AI Mode), one exposed it historically at a spectacular scale (Grok, 0.3% throughput), and three publish only the survivors (Perplexity, Copilot, Gemini) — the measured-vs-hidden asymmetry every cross-engine citation study runs into.

Section 1

The funnel this week: July 27, 2026

One week of the landscape corpus — 616 hotel prompts across 56 destinations, fired at every engine. “Retrieved” is every URL in the engine's source panel; “cited” is the subset it used inline in the answer.

Retrieved vs. cited hotel sources, week of 2026-07-27. Live data — this table refreshes with each Monday scrape.
EngineRetrievedCitedRejectedThroughputRejection rate
ChatGPT6,0731,1394,93418.8%81.2%
Google AI Mode24,1172,09622,0218.7%91.3%
Engines that only show the survivors. For these, sources = citations by construction, so throughput would read a meaningless 100%.
EngineCitations published this weekWhy no throughput
Perplexity11,074Publishes only the sources it used — the rejected pool is never exposed
Copilot5,702Publishes only the sources it used — the rejected pool is never exposed
Gemini5,282Exposes a cited flag but ~every source carries it — the rejected pool stays hidden
The two measurable engines run very different funnels this week: ChatGPT retrieves ~10 sources per answer and cites roughly a quarter of them; Google AI Mode retrieves ~30 and cites under one in ten — and two-thirds of its pool is google.com itself. Getting into an AI answer is not one problem; it is two problems (get retrieved, then get selected), and each engine weighs them differently.
Section 2

Citation throughput, week by week

The rate is not a constant of the model — it moves when the product changes. Every point below is one Monday's full 616-prompt sweep.

ChatGPTGoogle AI Mode
0%10%20%30%40%50%60%70%80%90%100%Apr 13Apr 27May 11May 25Jun 8Jun 22Jul 6Jul 2018.8%8.7%
View as table
WeekChatGPTGoogle AI Mode
2026-04-1358.1%
2026-04-2062.5%
2026-04-2721.9%
2026-05-0443.3%
2026-05-1197.7%17.5%
2026-05-1894.4%18.9%
2026-05-2582.4%10.7%
2026-06-0165.0%9.3%
2026-06-0868.0%10.1%
2026-06-1566.7%9.4%
2026-06-2251.9%9.4%
2026-06-2958.2%9.5%
2026-07-0659.8%9.6%
2026-07-1325.8%10.4%
2026-07-2026.5%9.4%
2026-07-2718.8%8.7%
The mid-July break: in the week of July 13, 2026, ChatGPT's hotel funnel changed shape. The retrieved pool halved (from ~22 to ~10 sources per answer) and inline citing collapsed (from ~13 to ~3 citations per answer), taking throughput from a steady 52–67% down to ~26%. The raw capture payloads confirm it is a product change, not a pipeline artifact: the same capture method, same prompts, same vendor — but ChatGPT now surfaces fewer sources and commits to far fewer of them. Google AI Mode, by contrast, has held a flat 9–10% for its entire measured history.
Section 3

The Reddit inversion

A popular cross-industry argument says AI doesn't actually prefer Reddit — that at index scale, engines discard nearly every Reddit page they retrieve, and Reddit's ubiquity in AI answers is search visibility, not model preference. Run the retrieved-vs-cited arithmetic inside hotel intent and the ranking flips completely.

Once ChatGPT has pulled a Reddit thread into a hotel answer's source pool, it cites it 59% of the time — the highest throughput of any major domain, double Tripadvisor's and six times Expedia's. Google AI Mode agrees at 57%. The pattern matches our flights study (Reddit cited on 83% of fetches) and price study (100%).

ChatGPT — top retrieved domains (week of 2026-07-27)

Hotel-intent prompts only. Throughput = cited ÷ retrieved per domain.
DomainRetrievedCitedRejectedThroughput
tripadvisor.com759135624
17.8%
booking.com37359314
15.8%
reddit.com228112116
49.1%
oyster.com20744163
21.3%
thehotelguru.com17446128
26.4%
cntraveler.com15448106
31.2%
timeout.com15141110
27.2%
hotelierschoice.com12415109
12.1%

Google AI Mode — top retrieved domains (week of 2026-07-27)

google.com dominates AI Mode's own candidate pool (Maps/Travel inventory) but is selected at only ~7%.
DomainRetrievedCitedRejectedThroughput
google.com14,82017114,649
1.2%
expedia.com728100628
13.7%
tripadvisor.com582302280
51.9%
agoda.com37514361
3.7%
reddit.com356238118
66.9%
booking.com34465279
18.9%
cntraveler.com23413896
59.0%
youtube.com19931168
15.6%
Why hotel-intent throughput runs higher than index-scale estimates: different denominator. An index-scale candidate pool spans every query type; ours is the per-answer source panel on hotel-intent prompts — a pool the engine has already pre-filtered for relevance. The absolute levels aren't comparable; the ranking within the pool is. And within hotel intent, the engines' revealed preference is community and editorial over OTA inventory: three national Tripadvisor TLDs get retrieved constantly and selected at 2–11%, while the .com survives at ~30%. Both things are true at once — search visibility decides who enters the pool, and model preference decides who exits it into the answer (this page).
Section 4

What gets selected: source categories

Same funnel, grouped by source type — live for the current week. OTAs and metasearch fill the retrieved pool; social, editorial and official sources convert into the answer at multiples of their rate.

ChatGPT

Week of 2026-07-27 — categories with ≥20 retrieved sources.
CategoryRetrievedCitedThroughput
Review sites1,437203
14.1%
Editorial & guides1,198323
27.0%
Other1,065131
12.3%
OTAs74284
11.3%
Community & niche guides36876
20.7%
Hotel direct sites25943
16.6%
Hotel chains24235
14.5%
Metasearch24259
24.4%
Social (mostly Reddit)228112
49.1%
Government & tourism boards21666
30.6%

Google AI Mode

Week of 2026-07-27 — categories with ≥20 retrieved sources.
CategoryRetrievedCitedThroughput
Metasearch15,091204
1.4%
OTAs2,519249
9.9%
Other2,207176
8.0%
Editorial & guides1,317464
35.2%
Social (mostly Reddit)908355
39.1%
Review sites666310
46.5%
Hotel direct sites556123
22.1%
Hotel chains34195
27.9%
Community & niche guides31286
27.6%
Government & tourism boards9914
14.1%
The category with the highest throughput on ChatGPT is consistently social — which in this corpus is almost entirely reddit.com. On Google AI Mode the winners are review sites and editorial. On both engines, OTAs convert at roughly half the overall average: heavily retrieved, lightly cited. If your AI strategy is “be on the OTAs,” you are optimizing the doorway, not the room.
Section 5

Grok, the everything-rejector

Before our Grok tracking ended in June 2026, it was hotel search's most extreme funnel — near-total rejection applied to the entire web. Grok retrieved 581,549 hotel sources over its measured lifetime and cited 40,153; in its final measured month its throughput was 0.3% — a 99.7% rejection rate.

Grok, May 4–24, 2026 (final full weeks of tracking): 194,262 retrieved sources, 516 cited.
DomainRetrievedCitedRejectedThroughput
booking.com22,36629122,0751.3%
tripadvisor.com17,813117,8120.0%
travelweekly.com8,10908,1090.0%
expedia.com7,75507,7550.0%
kayak.com5,26605,2660.0%
facebook.com5,16105,1610.0%
hotels.com5,09005,0900.0%
guide.michelin.com3,84303,8430.0%
Look at which 0.3% survived: booking.com and the official sites of luxury chains — fourseasons.com, marriott.com, hyatt.com. Grok rejected Tripadvisor 17,812 times out of 17,813 retrievals. When an engine is this selective, brand-owned pages are the only reliable way through the filter.

Methodology

Corpus. The weekly AI Hotel Landscape sweep: 616 frozen hotel prompts across 56 destinations, fired every Monday at ChatGPT, Google AI Mode, Perplexity, Copilot and Gemini via Bright Data (Grok until June 3, 2026). Running since April 2026; over 1.1 million retrieved sources logged to date.

Retrieved vs. cited. For each capture we store every URL the engine exposes as a consulted source, with a boolean marking whether it was cited inline in the answer. For ChatGPT, “retrieved” merges the visible source panel and its “more sources” overflow; “cited” means the URL appeared as an inline citation pill. For Google AI Mode the structured citation payload carries an explicit cited flag per source.

Citation throughput = cited ÷ retrieved — the ratio sometimes called the selection rate. One scope note: our denominator is the per-answer source pool on hotel-intent prompts, not a cross-intent candidate index — so levels run higher than index-scale estimates, and cross-study comparisons should use rankings, not absolute rates.

Engines without a rate. Perplexity and Copilot expose only cited sources; Gemini exposes a flag that is in practice always true. For those three, this page reports citation volume only — throughput can only be measured where an engine exposes both sides of the funnel.

Live data. Every “this week” number on this page is queried from the same public read-only views that power the open landscape data feed (CC-BY-4.0), and re-renders within an hour of each Monday scrape. The per-domain tables are a July 2026 snapshot computed from the raw capture store until the domain-level view ships in the public feed, at which point they go live automatically.

FAQ

The share of sources an AI assistant actually cites out of everything it retrieved while researching an answer: citations ÷ retrieved sources — the ratio also known as the selection rate. Retrieval means the engine fetched and considered your page; selection means a traveller actually saw it referenced. In hotel search this week, that rate is roughly one in four for ChatGPT and under one in ten for Google AI Mode.

Explore the data behind this page

Every number here comes from the open AI Hotel Landscape feed — weekly, CC-BY-4.0, JSON or CSV.