Snapshot · July 13–26, 2026AI Search · Citations · Hotels

Retrieved, rejected, cited:AI's citation throughput in hotel search

TL;DR: Sources are the input; citations are the throughput. Our AI Hotel Landscape pipeline logs both what each AI assistant retrieves and what it actually cites. This week ChatGPT cited 24.8% of the hotel sources it pulled; Google AI Mode just 14.1%. And while cross-industry citation data has AI engines rejecting the overwhelming majority of the Reddit pages they retrieve, in hotel search the inversion is total: Reddit has the highest citation throughput of any major domain — once retrieved, a Reddit thread is more likely to be cited than Tripadvisor, Booking.com or any OTA.

NS
Nicolas Sitter
Published July 2026 · data window July 13–26, 2026
24.8%
ChatGPT throughput
14.1%
Google AI Mode throughput
46,808
Sources retrieved (this week)
59%
Reddit throughput on ChatGPT
Read the Report

Executive Summary

Getting retrieved is table stakes. Getting selected is the game.

When an AI assistant answers a hotel question, it first pulls a pool of web sources — the candidate set — then cites a fraction of them in the answer. We call that fraction the citation throughput: citations ÷ retrieved sources — the ratio sometimes called the selection rate. It is the single number that separates “the AI read my page” from “the AI showed my page to a traveller.”

We measure it over the July 13–26, 2026 window, from the same 616-prompt, 56-destination corpus behind the AI Hotel Landscape. Two engines expose the full funnel (ChatGPT and Google AI Mode), one exposed it historically at a spectacular scale (Grok, 0.3% throughput), and three publish only the survivors (Perplexity, Copilot, Gemini) — the measured-vs-hidden asymmetry every cross-engine citation study runs into.

Read the numbers as directional. Each engine exposes its retrieved and cited lists differently, and our capture isn't perfectly identical across them — so the exact percentages move week to week and method to method, and I wouldn't bank on any single decimal. What doesn't move is the shape, and the gaps are wide enough to state plainly: ChatGPT cites roughly one source in five, Google AI Mode fewer than one in ten, and a Reddit thread converts several times better than any OTA. Directional, yes — but the difference is stunning.

Section 1

The funnel: week of August 10, 2026

The landscape corpus — 616 hotel prompts across 56 destinations, fired at every engine. “Retrieved” is every URL in the engine's source panel; “cited” is the subset it used inline in the answer.

6,331
ChatGPT sources retrieved
1,570
ChatGPT sources cited
19,967
AI Mode sources retrieved
2,820
AI Mode sources cited
Retrieved vs. cited, per answer

The gap between the two bars is the throughput. AI Mode pulls ~3.5× more sources per answer than ChatGPT but cites barely more of them.

0510152010.32.5ChatGPT24.8% cited16.22.3Google AI Mode14.1% citedretrieved / answercited / answer
Retrieved vs. cited hotel sources, week of 2026-08-10.
EngineRetrievedCitedRejectedThroughputRejection rate
ChatGPT6,3311,5704,76124.8%75.2%
Google AI Mode19,9672,82017,14714.1%85.9%
Engines that only show the survivors. For these, sources = citations by construction, so throughput would read a meaningless 100%.
EngineCitations publishedWhy no throughput
Perplexity11,003Publishes only the sources it used — the rejected pool is never exposed
Copilot4,513Publishes only the sources it used — the rejected pool is never exposed
Gemini4,994Exposes a cited flag but ~every source carries it — the rejected pool stays hidden
The two measurable engines run very different funnels this week: ChatGPT retrieves ~10 sources per answer and cites roughly a quarter of them; Google AI Mode retrieves ~30 and cites under one in ten — and two-thirds of its pool is google.com itself. Getting into an AI answer is not one problem; it is two problems (get retrieved, then get selected), and each engine weighs them differently.
Section 2

Citation throughput, week by week

The rate is not a constant of the model — it moves when the product changes. Every point below is one Monday's full 616-prompt sweep.

ChatGPTGoogle AI Mode
0%10%20%30%40%50%60%70%80%90%100%Apr 13May 4May 25Jun 15Jul 6Jul 2724.8%14.1%
View as table
WeekChatGPTGoogle AI Mode
2026-04-1358.1%
2026-04-2062.5%
2026-04-2721.9%
2026-05-0443.3%
2026-05-1197.7%17.5%
2026-05-1894.4%18.9%
2026-05-2582.4%10.7%
2026-06-0165.0%9.3%
2026-06-0868.0%10.1%
2026-06-1566.7%9.4%
2026-06-2251.9%9.4%
2026-06-2958.2%9.5%
2026-07-0659.8%9.6%
2026-07-1325.8%10.4%
2026-07-2026.5%9.4%
2026-07-2718.8%8.7%
2026-08-0318.2%11.0%
2026-08-1024.8%14.1%
The mid-July break: in the week of July 13, 2026, ChatGPT's hotel funnel changed shape. The retrieved pool halved (from ~22 to ~10 sources per answer) and inline citing collapsed (from ~13 to ~3 citations per answer), taking throughput from a steady 52–67% down to ~26%. The raw capture payloads confirm it is a product change, not a pipeline artifact: the same capture method, same prompts, same vendor — but ChatGPT now surfaces fewer sources and commits to far fewer of them. Google AI Mode, by contrast, has held a flat 9–10% for its entire measured history.
The retrieval-tier tag vanished the same month. ChatGPT stamps each retrieved page with an undocumented result_source tag naming the retrieval tier that fetched it — for hotels, 99.85% one licensed tier (see our result_source study). The tag has a known history of being dialled in and out: in this corpus it peaked at ~89% of retrieved hotel sources in early June 2026, ran at 35–75% through July — and in the July 27 scrape it disappeared entirely, zero occurrences across all 616 captures' raw streams. A halved source panel, collapsed inline citing, and now a withdrawn provenance tag: ChatGPT's retrieval surface is being actively reworked. If the tag returns, our pipeline picks it up automatically and this note will be updated.
Section 3

The Reddit inversion

A popular cross-industry argument — made well in Dejan's “Reddit & AI” analysis — says AI doesn't actually prefer Reddit: that at index scale, engines discard nearly every Reddit page they retrieve, and Reddit's ubiquity in AI answers is search visibility, not model preference. Run the retrieved-vs-cited arithmetic inside hotel intent and the ranking flips completely.

Once ChatGPT has pulled a Reddit thread into a hotel answer's source pool, it cites it 59% of the time — the highest throughput of any major domain, double Tripadvisor's and six times Expedia's. Google AI Mode agrees at 57%. The pattern matches our flights study (Reddit cited on 83% of fetches) and price study (100%).

ChatGPT — top retrieved domains (week of 2026-08-10)

Hotel-intent prompts only. Throughput = cited ÷ retrieved per domain.
DomainRetrievedCitedRejectedThroughput
tripadvisor.com816286530
35.0%
booking.com30363240
20.8%
timeout.com19680116
40.8%
tripadvisor.ca19116175
8.4%
thehotelguru.com18465119
35.3%
reddit.com18312657
68.9%
oyster.com17142129
24.6%
tripadvisor.co.uk16913156
7.7%

Google AI Mode — top retrieved domains (week of 2026-08-10)

google.com dominates AI Mode's own candidate pool (Maps/Travel inventory) but is selected at only ~7%.
DomainRetrievedCitedRejectedThroughput
google.com13,55762312,934
4.6%
expedia.com327103224
31.5%
youtube.com322148174
46.0%
tripadvisor.com303136167
44.9%
booking.com255136119
53.3%
reddit.com22313093
58.3%
cntraveler.com20911990
56.9%
agoda.com20515190
7.3%

Domain throughput over time (ChatGPT)

The same cited ÷ retrieved ratio, tracked per domain. Every top domain rode the spring “cite-almost-everything” stretch up toward 100%; the mid-July funnel change then pulled them apart — and reddit.com held the top of the pack while the OTAs fell hardest.

reddit.comtripadvisor.combooking.comexpedia.com
0%10%20%30%40%50%60%70%80%90%100%Apr 13Apr 27May 18Jun 1Jun 15Jun 29Jul 13Jul 2749.1%17.8%15.8%6.9%
View as table
Weekreddit.comtripadvisor.combooking.comexpedia.com
2026-04-1394.6%50.3%21.9%51.2%
2026-04-2096.2%50.0%34.2%68.6%
2026-04-27100.0%3.9%1.6%4.7%
2026-05-0499.0%54.4%11.1%29.1%
2026-05-18100.0%92.5%94.4%
2026-05-25100.0%74.5%81.4%61.6%
2026-06-01100.0%64.7%65.3%58.6%
2026-06-08100.0%68.5%79.2%53.8%
2026-06-15100.0%65.4%78.9%45.6%
2026-06-22100.0%37.2%35.4%36.9%
2026-06-2949.7%28.6%29.1%22.3%
2026-07-0647.2%34.2%33.2%28.5%
2026-07-1359.2%28.8%23.8%8.6%
2026-07-2059.3%30.1%23.9%11.9%
2026-07-2749.1%17.8%15.8%6.9%

ChatGPT, hotel-intent prompts, weeks with ≥15 retrieved sources for the domain. Snapshot through the week of July 27, 2026.

Where Reddit converts — by destination and by prompt

Reddit's edge isn't a quirk of one city or one query. Split ChatGPT's Reddit citations by destination and by prompt type and it stays high across the board — cited 40–80% of the time it's pulled, everywhere.

By destination (top cities)

Sydney
84%
Hanoi
73%
Tokyo
71%
New Orleans
69%
Istanbul
68%
San Francisco
65%
Rio de Janeiro
63%
Amalfi Coast
61%
Beijing
60%
Nairobi
58%
Cape Town
56%
Shanghai
50%
Seoul
48%
Las Vegas
47%

By prompt type

best hotels (control)
70%
boutique / neighborhood
69%
rooftop pool
68%
family-friendly
65%
luxury
65%
business
62%
couples
59%
affordable / budget
57%
solo leisure
52%
neighborhood
51%
near a landmark
39%

ChatGPT, reddit.com throughput (cited ÷ retrieved), July 13–26 window. Destinations shown are those with the most Reddit pulls; prompt types are the 11 corpus intents.

Why hotel-intent throughput runs higher than index-scale estimates: different denominator. An index-scale candidate pool spans every query type; ours is the per-answer source panel on hotel-intent prompts — a pool the engine has already pre-filtered for relevance. The absolute levels aren't comparable; the ranking within the pool is. And within hotel intent, the engines' revealed preference is community and editorial over OTA inventory: three national Tripadvisor TLDs get retrieved constantly and selected at 2–11%, while the .com survives at ~30%. Both things are true at once — search visibility decides who enters the pool, and model preference decides who exits it into the answer (this page).
Section 4

What gets selected: source categories

Same funnel, grouped by source type, with ChatGPT and Google AI Mode on the same rows so you can compare them directly. OTAs and metasearch fill the retrieved pool; social, editorial and review sources convert into the answer at multiples of their rate.

CategoryTop domainsChatGPT
cited / retrieved
throughputAI Mode
cited / retrieved
throughput
Social (mostly Reddit)reddit.com, youtube.com, facebook.com112 / 22849.1%355 / 90739.1%
Review sitestripadvisor.com, oyster.com, tripadvisor.ca200 / 1,42614.0%307 / 64447.7%
Editorial & guidescntraveler.com, thehotelguru.com, timeout.com293 / 1,06827.4%419 / 1,12237.3%
Metasearchgoogle.com, kayak.com, hotelscombined.com57 / 22924.9%199 / 15,0451.3%
OTAsbooking.com, expedia.com, agoda.com78 / 65411.9%241 / 2,37410.2%
Community & niche guidessantorinidave.com, budgetyourtrip.com, hotelierschoice.com70 / 31122.5%75 / 26128.7%
Hotel chainsmarriott.com, all.accor.com, fourseasons.com30 / 19615.3%53 / 19127.7%
Otherdestination.com, foratravel.com, budgetyourtrip.com38 / 24515.5%6 / 1314.6%

Snapshot: week of July 27, 2026, hotel-intent prompts. Throughput = cited ÷ retrieved. Rows ordered by ChatGPT citation throughput; the three most-retrieved domains shown per category.

The category with the highest throughput on ChatGPT is consistently social — which in this corpus is almost entirely reddit.com. On Google AI Mode the winners are review sites and editorial. On both engines, OTAs convert at roughly half the overall average: heavily retrieved, lightly cited. If your AI strategy is “be on the OTAs,” you are optimizing the doorway, not the room.
Section 5

Grok, the everything-rejector

Before our Grok tracking ended in June 2026, it was hotel search's most extreme funnel — near-total rejection applied to the entire web. Grok retrieved 581,549 hotel sources over its measured lifetime and cited 40,153; in its final measured month its throughput was 0.3% — a 99.7% rejection rate.

Grok, May 4–24, 2026 (final full weeks of tracking): 194,262 retrieved sources, 516 cited.
DomainRetrievedCitedRejectedThroughput
booking.com22,36629122,0751.3%
tripadvisor.com17,813117,8120.0%
travelweekly.com8,10908,1090.0%
expedia.com7,75507,7550.0%
kayak.com5,26605,2660.0%
facebook.com5,16105,1610.0%
hotels.com5,09005,0900.0%
guide.michelin.com3,84303,8430.0%
Look at which 0.3% survived: booking.com and the official sites of luxury chains — fourseasons.com, marriott.com, hyatt.com. Grok rejected Tripadvisor 17,812 times out of 17,813 retrievals. When an engine is this selective, brand-owned pages are the only reliable way through the filter.

Methodology

Corpus. The weekly AI Hotel Landscape sweep: 616 frozen hotel prompts across 56 destinations, fired every Monday at ChatGPT, Google AI Mode, Perplexity, Copilot and Gemini via Bright Data (Grok until June 3, 2026). Running since April 2026; over 1.1 million retrieved sources logged to date.

The prompts. Eleven intents per destination — one control plus ten shopper variations (budget, persona, neighborhood, landmark, amenity) — filled in for all 56 cities, so 11 × 56 = 616. The templates:

best hotels in {city}control
luxury hotels in {city}luxury
affordable hotels in {city} under $200budget
family friendly hotels in {city}families
best hotels in {city} for couplescouples
best hotels in {city} for solo travelerssolo
best hotels in {city} for business travelersbusiness
hotels near {landmark}, {city}landmark
best hotels in {neighborhood}, {city}neighborhood
boutique hotels in {neighborhood}, {city}boutique
hotels in {city} with rooftop poolamenity

Full corpus, per-city, with the destination map on the AI Hotel Landscape.

Retrieved vs. cited. For each capture we store every URL the engine exposes as a consulted source, with a boolean marking whether it was cited inline in the answer. For ChatGPT, “retrieved” merges the visible source panel and its “more sources” overflow; “cited” means the URL appeared as an inline citation pill. For Google AI Mode the structured citation payload carries an explicit cited flag per source.

Citation throughput = cited ÷ retrieved — the ratio sometimes called the selection rate. One scope note: our denominator is the per-answer source pool on hotel-intent prompts, not a cross-intent candidate index — so levels run higher than index-scale estimates, and cross-study comparisons should use rankings, not absolute rates.

Engines without a rate. Perplexity and Copilot expose only cited sources; Gemini exposes a flag that is in practice always true. For those three, this page reports citation volume only — throughput can only be measured where an engine exposes both sides of the funnel.

Data window. This is a point-in-time study of the July 13–26, 2026 window, not a weekly-updating dashboard. The headline funnel figures are read from the same public read-only views that power the open landscape data feed (CC-BY-4.0); the per-domain and per-category tables are a fixed snapshot computed from the raw capture store for that window.

FAQ

The share of sources an AI assistant actually cites out of everything it retrieved while researching an answer: citations ÷ retrieved sources — the ratio also known as the selection rate. Retrieval means the engine fetched and considered your page; selection means a traveller actually saw it referenced. In hotel search this week, that rate is roughly one in four for ChatGPT and under one in ten for Google AI Mode.

Explore the data behind this page

Every number here comes from the open AI Hotel Landscape feed — weekly, CC-BY-4.0, JSON or CSV.