Retrieved, rejected, cited:AI's citation throughput in hotel search
TL;DR: Sources are the input; citations are the throughput. Every week our AI Hotel Landscape pipeline logs both what each AI assistant retrieves and what it actually cites. This week ChatGPT cited 18.8% of the hotel sources it pulled; Google AI Mode just 8.7%. And while cross-industry citation data has AI engines rejecting the overwhelming majority of the Reddit pages they retrieve, in hotel search the inversion is total: Reddit has the highest citation throughput of any major domain — once retrieved, a Reddit thread is more likely to be cited than Tripadvisor, Booking.com or any OTA.
Executive Summary
Getting retrieved is table stakes. Getting selected is the game.
When an AI assistant answers a hotel question, it first pulls a pool of web sources — the candidate set — then cites a fraction of them in the answer. We call that fraction the citation throughput: citations ÷ retrieved sources — the ratio sometimes called the selection rate. It is the single number that separates “the AI read my page” from “the AI showed my page to a traveller.”
This page computes it live, every week, from the same 616-prompt, 56-destination corpus behind the AI Hotel Landscape. Two engines expose the full funnel (ChatGPT and Google AI Mode), one exposed it historically at a spectacular scale (Grok, 0.3% throughput), and three publish only the survivors (Perplexity, Copilot, Gemini) — the measured-vs-hidden asymmetry every cross-engine citation study runs into.
The funnel this week: July 27, 2026
One week of the landscape corpus — 616 hotel prompts across 56 destinations, fired at every engine. “Retrieved” is every URL in the engine's source panel; “cited” is the subset it used inline in the answer.
| Engine | Retrieved | Cited | Rejected | Throughput | Rejection rate |
|---|---|---|---|---|---|
| ChatGPT | 6,073 | 1,139 | 4,934 | 18.8% | 81.2% |
| Google AI Mode | 24,117 | 2,096 | 22,021 | 8.7% | 91.3% |
| Engine | Citations published this week | Why no throughput |
|---|---|---|
| Perplexity | 11,074 | Publishes only the sources it used — the rejected pool is never exposed |
| Copilot | 5,702 | Publishes only the sources it used — the rejected pool is never exposed |
| Gemini | 5,282 | Exposes a cited flag but ~every source carries it — the rejected pool stays hidden |
Citation throughput, week by week
The rate is not a constant of the model — it moves when the product changes. Every point below is one Monday's full 616-prompt sweep.
View as table
| Week | ChatGPT | Google AI Mode |
|---|---|---|
| 2026-04-13 | 58.1% | — |
| 2026-04-20 | 62.5% | — |
| 2026-04-27 | 21.9% | — |
| 2026-05-04 | 43.3% | — |
| 2026-05-11 | 97.7% | 17.5% |
| 2026-05-18 | 94.4% | 18.9% |
| 2026-05-25 | 82.4% | 10.7% |
| 2026-06-01 | 65.0% | 9.3% |
| 2026-06-08 | 68.0% | 10.1% |
| 2026-06-15 | 66.7% | 9.4% |
| 2026-06-22 | 51.9% | 9.4% |
| 2026-06-29 | 58.2% | 9.5% |
| 2026-07-06 | 59.8% | 9.6% |
| 2026-07-13 | 25.8% | 10.4% |
| 2026-07-20 | 26.5% | 9.4% |
| 2026-07-27 | 18.8% | 8.7% |
The Reddit inversion
A popular cross-industry argument says AI doesn't actually prefer Reddit — that at index scale, engines discard nearly every Reddit page they retrieve, and Reddit's ubiquity in AI answers is search visibility, not model preference. Run the retrieved-vs-cited arithmetic inside hotel intent and the ranking flips completely.
Once ChatGPT has pulled a Reddit thread into a hotel answer's source pool, it cites it 59% of the time — the highest throughput of any major domain, double Tripadvisor's and six times Expedia's. Google AI Mode agrees at 57%. The pattern matches our flights study (Reddit cited on 83% of fetches) and price study (100%).
ChatGPT — top retrieved domains (week of 2026-07-27)
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| tripadvisor.com | 759 | 135 | 624 | 17.8% |
| booking.com | 373 | 59 | 314 | 15.8% |
| reddit.com | 228 | 112 | 116 | 49.1% |
| oyster.com | 207 | 44 | 163 | 21.3% |
| thehotelguru.com | 174 | 46 | 128 | 26.4% |
| cntraveler.com | 154 | 48 | 106 | 31.2% |
| timeout.com | 151 | 41 | 110 | 27.2% |
| hotelierschoice.com | 124 | 15 | 109 | 12.1% |
Google AI Mode — top retrieved domains (week of 2026-07-27)
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| google.com | 14,820 | 171 | 14,649 | 1.2% |
| expedia.com | 728 | 100 | 628 | 13.7% |
| tripadvisor.com | 582 | 302 | 280 | 51.9% |
| agoda.com | 375 | 14 | 361 | 3.7% |
| reddit.com | 356 | 238 | 118 | 66.9% |
| booking.com | 344 | 65 | 279 | 18.9% |
| cntraveler.com | 234 | 138 | 96 | 59.0% |
| youtube.com | 199 | 31 | 168 | 15.6% |
What gets selected: source categories
Same funnel, grouped by source type — live for the current week. OTAs and metasearch fill the retrieved pool; social, editorial and official sources convert into the answer at multiples of their rate.
ChatGPT
| Category | Retrieved | Cited | Throughput |
|---|---|---|---|
| Review sites | 1,437 | 203 | 14.1% |
| Editorial & guides | 1,198 | 323 | 27.0% |
| Other | 1,065 | 131 | 12.3% |
| OTAs | 742 | 84 | 11.3% |
| Community & niche guides | 368 | 76 | 20.7% |
| Hotel direct sites | 259 | 43 | 16.6% |
| Hotel chains | 242 | 35 | 14.5% |
| Metasearch | 242 | 59 | 24.4% |
| Social (mostly Reddit) | 228 | 112 | 49.1% |
| Government & tourism boards | 216 | 66 | 30.6% |
Google AI Mode
| Category | Retrieved | Cited | Throughput |
|---|---|---|---|
| Metasearch | 15,091 | 204 | 1.4% |
| OTAs | 2,519 | 249 | 9.9% |
| Other | 2,207 | 176 | 8.0% |
| Editorial & guides | 1,317 | 464 | 35.2% |
| Social (mostly Reddit) | 908 | 355 | 39.1% |
| Review sites | 666 | 310 | 46.5% |
| Hotel direct sites | 556 | 123 | 22.1% |
| Hotel chains | 341 | 95 | 27.9% |
| Community & niche guides | 312 | 86 | 27.6% |
| Government & tourism boards | 99 | 14 | 14.1% |
Grok, the everything-rejector
Before our Grok tracking ended in June 2026, it was hotel search's most extreme funnel — near-total rejection applied to the entire web. Grok retrieved 581,549 hotel sources over its measured lifetime and cited 40,153; in its final measured month its throughput was 0.3% — a 99.7% rejection rate.
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| booking.com | 22,366 | 291 | 22,075 | 1.3% |
| tripadvisor.com | 17,813 | 1 | 17,812 | 0.0% |
| travelweekly.com | 8,109 | 0 | 8,109 | 0.0% |
| expedia.com | 7,755 | 0 | 7,755 | 0.0% |
| kayak.com | 5,266 | 0 | 5,266 | 0.0% |
| facebook.com | 5,161 | 0 | 5,161 | 0.0% |
| hotels.com | 5,090 | 0 | 5,090 | 0.0% |
| guide.michelin.com | 3,843 | 0 | 3,843 | 0.0% |
Methodology
Corpus. The weekly AI Hotel Landscape sweep: 616 frozen hotel prompts across 56 destinations, fired every Monday at ChatGPT, Google AI Mode, Perplexity, Copilot and Gemini via Bright Data (Grok until June 3, 2026). Running since April 2026; over 1.1 million retrieved sources logged to date.
Retrieved vs. cited. For each capture we store every URL the engine exposes as a consulted source, with a boolean marking whether it was cited inline in the answer. For ChatGPT, “retrieved” merges the visible source panel and its “more sources” overflow; “cited” means the URL appeared as an inline citation pill. For Google AI Mode the structured citation payload carries an explicit cited flag per source.
Citation throughput = cited ÷ retrieved — the ratio sometimes called the selection rate. One scope note: our denominator is the per-answer source pool on hotel-intent prompts, not a cross-intent candidate index — so levels run higher than index-scale estimates, and cross-study comparisons should use rankings, not absolute rates.
Engines without a rate. Perplexity and Copilot expose only cited sources; Gemini exposes a flag that is in practice always true. For those three, this page reports citation volume only — throughput can only be measured where an engine exposes both sides of the funnel.
Live data. Every “this week” number on this page is queried from the same public read-only views that power the open landscape data feed (CC-BY-4.0), and re-renders within an hour of each Monday scrape. The per-domain tables are a July 2026 snapshot computed from the raw capture store until the domain-level view ships in the public feed, at which point they go live automatically.
FAQ
Explore the data behind this page
Every number here comes from the open AI Hotel Landscape feed — weekly, CC-BY-4.0, JSON or CSV.