Retrieved, rejected, cited:AI's citation throughput in hotel search
TL;DR: Sources are the input; citations are the throughput. Our AI Hotel Landscape pipeline logs both what each AI assistant retrieves and what it actually cites. This week ChatGPT cited 24.8% of the hotel sources it pulled; Google AI Mode just 14.1%. And while cross-industry citation data has AI engines rejecting the overwhelming majority of the Reddit pages they retrieve, in hotel search the inversion is total: Reddit has the highest citation throughput of any major domain — once retrieved, a Reddit thread is more likely to be cited than Tripadvisor, Booking.com or any OTA.
Executive Summary
Getting retrieved is table stakes. Getting selected is the game.
When an AI assistant answers a hotel question, it first pulls a pool of web sources — the candidate set — then cites a fraction of them in the answer. We call that fraction the citation throughput: citations ÷ retrieved sources — the ratio sometimes called the selection rate. It is the single number that separates “the AI read my page” from “the AI showed my page to a traveller.”
We measure it over the July 13–26, 2026 window, from the same 616-prompt, 56-destination corpus behind the AI Hotel Landscape. Two engines expose the full funnel (ChatGPT and Google AI Mode), one exposed it historically at a spectacular scale (Grok, 0.3% throughput), and three publish only the survivors (Perplexity, Copilot, Gemini) — the measured-vs-hidden asymmetry every cross-engine citation study runs into.
Read the numbers as directional. Each engine exposes its retrieved and cited lists differently, and our capture isn't perfectly identical across them — so the exact percentages move week to week and method to method, and I wouldn't bank on any single decimal. What doesn't move is the shape, and the gaps are wide enough to state plainly: ChatGPT cites roughly one source in five, Google AI Mode fewer than one in ten, and a Reddit thread converts several times better than any OTA. Directional, yes — but the difference is stunning.
The funnel: week of August 10, 2026
The landscape corpus — 616 hotel prompts across 56 destinations, fired at every engine. “Retrieved” is every URL in the engine's source panel; “cited” is the subset it used inline in the answer.
The gap between the two bars is the throughput. AI Mode pulls ~3.5× more sources per answer than ChatGPT but cites barely more of them.
| Engine | Retrieved | Cited | Rejected | Throughput | Rejection rate |
|---|---|---|---|---|---|
| ChatGPT | 6,331 | 1,570 | 4,761 | 24.8% | 75.2% |
| Google AI Mode | 19,967 | 2,820 | 17,147 | 14.1% | 85.9% |
| Engine | Citations published | Why no throughput |
|---|---|---|
| Perplexity | 11,003 | Publishes only the sources it used — the rejected pool is never exposed |
| Copilot | 4,513 | Publishes only the sources it used — the rejected pool is never exposed |
| Gemini | 4,994 | Exposes a cited flag but ~every source carries it — the rejected pool stays hidden |
Citation throughput, week by week
The rate is not a constant of the model — it moves when the product changes. Every point below is one Monday's full 616-prompt sweep.
View as table
| Week | ChatGPT | Google AI Mode |
|---|---|---|
| 2026-04-13 | 58.1% | — |
| 2026-04-20 | 62.5% | — |
| 2026-04-27 | 21.9% | — |
| 2026-05-04 | 43.3% | — |
| 2026-05-11 | 97.7% | 17.5% |
| 2026-05-18 | 94.4% | 18.9% |
| 2026-05-25 | 82.4% | 10.7% |
| 2026-06-01 | 65.0% | 9.3% |
| 2026-06-08 | 68.0% | 10.1% |
| 2026-06-15 | 66.7% | 9.4% |
| 2026-06-22 | 51.9% | 9.4% |
| 2026-06-29 | 58.2% | 9.5% |
| 2026-07-06 | 59.8% | 9.6% |
| 2026-07-13 | 25.8% | 10.4% |
| 2026-07-20 | 26.5% | 9.4% |
| 2026-07-27 | 18.8% | 8.7% |
| 2026-08-03 | 18.2% | 11.0% |
| 2026-08-10 | 24.8% | 14.1% |
result_source tag naming the retrieval tier that fetched it — for hotels, 99.85% one licensed tier (see our result_source study). The tag has a known history of being dialled in and out: in this corpus it peaked at ~89% of retrieved hotel sources in early June 2026, ran at 35–75% through July — and in the July 27 scrape it disappeared entirely, zero occurrences across all 616 captures' raw streams. A halved source panel, collapsed inline citing, and now a withdrawn provenance tag: ChatGPT's retrieval surface is being actively reworked. If the tag returns, our pipeline picks it up automatically and this note will be updated.The Reddit inversion
A popular cross-industry argument — made well in Dejan's “Reddit & AI” analysis — says AI doesn't actually prefer Reddit: that at index scale, engines discard nearly every Reddit page they retrieve, and Reddit's ubiquity in AI answers is search visibility, not model preference. Run the retrieved-vs-cited arithmetic inside hotel intent and the ranking flips completely.
Once ChatGPT has pulled a Reddit thread into a hotel answer's source pool, it cites it 59% of the time — the highest throughput of any major domain, double Tripadvisor's and six times Expedia's. Google AI Mode agrees at 57%. The pattern matches our flights study (Reddit cited on 83% of fetches) and price study (100%).
ChatGPT — top retrieved domains (week of 2026-08-10)
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| tripadvisor.com | 816 | 286 | 530 | 35.0% |
| booking.com | 303 | 63 | 240 | 20.8% |
| timeout.com | 196 | 80 | 116 | 40.8% |
| tripadvisor.ca | 191 | 16 | 175 | 8.4% |
| thehotelguru.com | 184 | 65 | 119 | 35.3% |
| reddit.com | 183 | 126 | 57 | 68.9% |
| oyster.com | 171 | 42 | 129 | 24.6% |
| tripadvisor.co.uk | 169 | 13 | 156 | 7.7% |
Google AI Mode — top retrieved domains (week of 2026-08-10)
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| google.com | 13,557 | 623 | 12,934 | 4.6% |
| expedia.com | 327 | 103 | 224 | 31.5% |
| youtube.com | 322 | 148 | 174 | 46.0% |
| tripadvisor.com | 303 | 136 | 167 | 44.9% |
| booking.com | 255 | 136 | 119 | 53.3% |
| reddit.com | 223 | 130 | 93 | 58.3% |
| cntraveler.com | 209 | 119 | 90 | 56.9% |
| agoda.com | 205 | 15 | 190 | 7.3% |
Domain throughput over time (ChatGPT)
The same cited ÷ retrieved ratio, tracked per domain. Every top domain rode the spring “cite-almost-everything” stretch up toward 100%; the mid-July funnel change then pulled them apart — and reddit.com held the top of the pack while the OTAs fell hardest.
View as table
| Week | reddit.com | tripadvisor.com | booking.com | expedia.com |
|---|---|---|---|---|
| 2026-04-13 | 94.6% | 50.3% | 21.9% | 51.2% |
| 2026-04-20 | 96.2% | 50.0% | 34.2% | 68.6% |
| 2026-04-27 | 100.0% | 3.9% | 1.6% | 4.7% |
| 2026-05-04 | 99.0% | 54.4% | 11.1% | 29.1% |
| 2026-05-18 | 100.0% | 92.5% | — | 94.4% |
| 2026-05-25 | 100.0% | 74.5% | 81.4% | 61.6% |
| 2026-06-01 | 100.0% | 64.7% | 65.3% | 58.6% |
| 2026-06-08 | 100.0% | 68.5% | 79.2% | 53.8% |
| 2026-06-15 | 100.0% | 65.4% | 78.9% | 45.6% |
| 2026-06-22 | 100.0% | 37.2% | 35.4% | 36.9% |
| 2026-06-29 | 49.7% | 28.6% | 29.1% | 22.3% |
| 2026-07-06 | 47.2% | 34.2% | 33.2% | 28.5% |
| 2026-07-13 | 59.2% | 28.8% | 23.8% | 8.6% |
| 2026-07-20 | 59.3% | 30.1% | 23.9% | 11.9% |
| 2026-07-27 | 49.1% | 17.8% | 15.8% | 6.9% |
ChatGPT, hotel-intent prompts, weeks with ≥15 retrieved sources for the domain. Snapshot through the week of July 27, 2026.
Where Reddit converts — by destination and by prompt
Reddit's edge isn't a quirk of one city or one query. Split ChatGPT's Reddit citations by destination and by prompt type and it stays high across the board — cited 40–80% of the time it's pulled, everywhere.
By destination (top cities)
By prompt type
ChatGPT, reddit.com throughput (cited ÷ retrieved), July 13–26 window. Destinations shown are those with the most Reddit pulls; prompt types are the 11 corpus intents.
What gets selected: source categories
Same funnel, grouped by source type, with ChatGPT and Google AI Mode on the same rows so you can compare them directly. OTAs and metasearch fill the retrieved pool; social, editorial and review sources convert into the answer at multiples of their rate.
| Category | Top domains | ChatGPT cited / retrieved | throughput | AI Mode cited / retrieved | throughput |
|---|---|---|---|---|---|
| Social (mostly Reddit) | reddit.com, youtube.com, facebook.com | 112 / 228 | 49.1% | 355 / 907 | 39.1% |
| Review sites | tripadvisor.com, oyster.com, tripadvisor.ca | 200 / 1,426 | 14.0% | 307 / 644 | 47.7% |
| Editorial & guides | cntraveler.com, thehotelguru.com, timeout.com | 293 / 1,068 | 27.4% | 419 / 1,122 | 37.3% |
| Metasearch | google.com, kayak.com, hotelscombined.com | 57 / 229 | 24.9% | 199 / 15,045 | 1.3% |
| OTAs | booking.com, expedia.com, agoda.com | 78 / 654 | 11.9% | 241 / 2,374 | 10.2% |
| Community & niche guides | santorinidave.com, budgetyourtrip.com, hotelierschoice.com | 70 / 311 | 22.5% | 75 / 261 | 28.7% |
| Hotel chains | marriott.com, all.accor.com, fourseasons.com | 30 / 196 | 15.3% | 53 / 191 | 27.7% |
| Other | destination.com, foratravel.com, budgetyourtrip.com | 38 / 245 | 15.5% | 6 / 131 | 4.6% |
Snapshot: week of July 27, 2026, hotel-intent prompts. Throughput = cited ÷ retrieved. Rows ordered by ChatGPT citation throughput; the three most-retrieved domains shown per category.
Grok, the everything-rejector
Before our Grok tracking ended in June 2026, it was hotel search's most extreme funnel — near-total rejection applied to the entire web. Grok retrieved 581,549 hotel sources over its measured lifetime and cited 40,153; in its final measured month its throughput was 0.3% — a 99.7% rejection rate.
| Domain | Retrieved | Cited | Rejected | Throughput |
|---|---|---|---|---|
| booking.com | 22,366 | 291 | 22,075 | 1.3% |
| tripadvisor.com | 17,813 | 1 | 17,812 | 0.0% |
| travelweekly.com | 8,109 | 0 | 8,109 | 0.0% |
| expedia.com | 7,755 | 0 | 7,755 | 0.0% |
| kayak.com | 5,266 | 0 | 5,266 | 0.0% |
| facebook.com | 5,161 | 0 | 5,161 | 0.0% |
| hotels.com | 5,090 | 0 | 5,090 | 0.0% |
| guide.michelin.com | 3,843 | 0 | 3,843 | 0.0% |
Methodology
Corpus. The weekly AI Hotel Landscape sweep: 616 frozen hotel prompts across 56 destinations, fired every Monday at ChatGPT, Google AI Mode, Perplexity, Copilot and Gemini via Bright Data (Grok until June 3, 2026). Running since April 2026; over 1.1 million retrieved sources logged to date.
The prompts. Eleven intents per destination — one control plus ten shopper variations (budget, persona, neighborhood, landmark, amenity) — filled in for all 56 cities, so 11 × 56 = 616. The templates:
Full corpus, per-city, with the destination map on the AI Hotel Landscape.
Retrieved vs. cited. For each capture we store every URL the engine exposes as a consulted source, with a boolean marking whether it was cited inline in the answer. For ChatGPT, “retrieved” merges the visible source panel and its “more sources” overflow; “cited” means the URL appeared as an inline citation pill. For Google AI Mode the structured citation payload carries an explicit cited flag per source.
Citation throughput = cited ÷ retrieved — the ratio sometimes called the selection rate. One scope note: our denominator is the per-answer source pool on hotel-intent prompts, not a cross-intent candidate index — so levels run higher than index-scale estimates, and cross-study comparisons should use rankings, not absolute rates.
Engines without a rate. Perplexity and Copilot expose only cited sources; Gemini exposes a flag that is in practice always true. For those three, this page reports citation volume only — throughput can only be measured where an engine exposes both sides of the funnel.
Data window. This is a point-in-time study of the July 13–26, 2026 window, not a weekly-updating dashboard. The headline funnel figures are read from the same public read-only views that power the open landscape data feed (CC-BY-4.0); the per-domain and per-category tables are a fixed snapshot computed from the raw capture store for that window.
FAQ
Explore the data behind this page
Every number here comes from the open AI Hotel Landscape feed — weekly, CC-BY-4.0, JSON or CSV.