July 2026AI Search · Retrieval · Flights

ChatGPT's result_source, part two:flights don't fly on the hotel stack

TL;DR: We re-ran our hidden result_source study on flights: 2,007 fresh ChatGPT captures from the US and the UK. The hotel-style tier monopoly is gone — the licensed labrador tier falls from 99.85% to 74.6%, Bright Data's tier carries 23%, and the open-web serp tier appears for the first time. Only 8.6% of flight questions skip the web (hotels: 37.3%). And the most-trusted flight source isn't an airline, an OTA or a metasearch — it's Reddit, cited on 83% of its fetches.

NS
Nicolas Sitter
Published July 21, 2026
2,007
ChatGPT captures
23,308
Fetched documents
74.6%
labrador (was 99.85%)
83%
Reddit cite rate
Read the Report

Executive Summary

The retrieval stack behind ChatGPT is vertical-specific — and flights use all of it.

Our hotel study found ChatGPT sourcing hotel answers almost entirely (99.85%) from one licensed retrieval tier. To test whether that generalises, we ran 67 frozen flight prompts × US + UK × 15 iterations = 2,007 captures over four days in July 2026 and parsed the same hidden fields from the raw stream.

It doesn't generalise. Flights run on a genuinely mixed stack (labrador 74.6%, bright 23.1%, oxylabs 1.7%, serp 0.6%), almost every flight question triggers a live search, half of the documents ChatGPT actually cites bypass the tagged retrieval layer entirely — and citation trust looks nothing like hotels: flat across airlines, OTAs and metasearch, with one glaring exception. Reddit.

Section 1

The tier monopoly is gone

Every page ChatGPT retrieves carries the undocumented result_source tag naming the pipeline that fetched it. For hotels, one licensed tier had a 99.85% monopoly. For flights, 82.8% of the 23,308 fetched documents carry a tag — and the mix is completely different.

chatgpt-flights-tier-mix-2026
result_source distribution among tier-tagged documents (flights: 19,305 tagged fetches; hotels: 30,002 tagged citations).
result_sourceFlights (this study)Hotels (Dec 2025–Jun 2026)What it carries for flights
labrador74.62%99.85%Licensed tier — Google, OTAs, metasearch, editorial
bright23.05%0.14%Bright Data datasets — skyscanner.net is its #1 domain
oxylabs1.70%0.01%Scraped open web — the points/deals blogosphere
serp0.63%0% (never appeared)Live SERP — fare-comparison and metasearch pages
The serp tier — zero appearances in 30,002 hotel citations — shows up for flights, and Bright Data's bright tier goes from a rounding error to nearly a quarter of tagged fetches. Notably, skyscanner.net is the single biggest domain in the bright tier (388 documents, ahead of google.com and expedia.com): metasearch inventory reaches ChatGPT through the scraper pipeline, not the licensed feed.
Section 2

The search gate barely exists for flights

Before any retrieval, turn_use_case classifies the query and decides whether the web is touched at all. For hotels, 37.3% of questions were answered from training data. Flights are the opposite: prices move too fast to memorise.

turn_use_case distribution across 2,007 flight captures vs the hotel study.
turn_use_caseShare of flight turnsHotelsHits the web?
instant search77.7%3.3%Yes
instant answers8.7%1.3%Partial
text8.6%37.3%No — training data only
search5.0%31.8%Yes
local (maps)0%26.3%
Two pipelines that dominate hotels are simply absent for flights: the local/maps bucket (26.3% of hotel turns) never fires, and in 2,007 captures we saw zero shopping or booking widgets — no Google-Flights-style panel exists inside ChatGPT. Flight answers are prose plus citations, which makes who gets cited the entire visibility game.
Section 3

Citations come through two doors — and half bypass the tagged layer

A parsing detail with a real finding inside. The URLs ChatGPT actually cites are stamped ?utm_source=chatgpt.com, so they never string-match the tagged fetch layer. Joining them back by canonical URL splits the 4,003 cited documents into three groups:

Provenance of 4,003 cited documents relative to the tier-tagged fetch layer.
Cited-document provenanceShareWhat it means
Footnote-only50.2%No tagged fetch counterpart at all — enters via a separate, untagged browse path
Domain-only match35.1%Same domain was fetched with a tier tag, but a different page
Canonical-URL join14.6%The exact tagged page — overwhelmingly google.com via labrador
Half of what ChatGPT cites for flights never appears in the tagged retrieval layer. The tiered system supplies the bulk fetch (Google, OTA and metasearch pages); a second, untagged path supplies the long tail that actually gets footnoted — route utilities like flightsfrom.com, Reddit threads, editorial, airline pages, Skyscanner and Kayak. The hotel data never showed this split.
Section 4

Retrieved ≠ cited: Reddit is the only source ChatGPT treats like a brand site

For hotels, official brand sites were cited on 77–86% of their retrievals. For flights the citation curve is flat — almost everything sits between 3% and 36% — with one outlier at hotel-brand-site altitude.

chatgpt-flights-retrieved-vs-cited-2026
Top flight domains: times retrieved vs share actually cited, all 2,007 captures.
DomainRetrievedCited %Type
google.com2,02321%Metasearch
flightsfrom.com1,12526%Route utility
expedia.com1,04010%OTA
cheapflights.co.uk68622%Metasearch
flightconnections.com67328%Route utility
kayak.com66915%Metasearch
skyscanner.net53510%Metasearch
skyscanner.com42122%Metasearch
Reddit is cited on 83% of the fetches where it appears — four times the rate of Google, five times Kayak, eight times Expedia. This mirrors our hotel price study, where Reddit hit a 100% cite rate: when ChatGPT wants a “real person” answer about travel, it reads the aggregators and quotes the forum.
Section 5

Brand authority is query-scoped: name the airline and its site becomes the answer

The hotel study's headline — official brand sites get cited when retrieved — does replicate for airlines, but only when the traveller names the brand. On “What is {airline}'s checked baggage allowance?” the airline's own site is fetched in 68–83% of captures and cited in essentially every capture where it was fetched:

Named-airline prompts (baggage / economy experience): share of captures where the airline's own website was fetched and cited.
AirlineCapturesOwn site fetchedOwn site cited
Ryanair6083%83%
Singapore Airlines6078%78%
Emirates6078%73%
British Airways6077%77%
Air France6075%75%
Delta6075%75%
The same airline sites sit at 16–28% cited on anonymous route queries. ChatGPT doesn't “trust airlines” globally — it trusts the official source for questions about that brand. For airlines, owning your baggage/fees/policy pages is the highest-certainty AI visibility asset; for route and price queries, the battle happens on metasearch, utilities and Reddit.
Section 6

“Where should I book?” — the answer set ChatGPT actually cites

Seven advice prompts (“best website for cheap flights”, “book direct or through a third party?”…) probe which booking brands ChatGPT puts in front of travellers. The cited set is deal newsletters, consumer editorial, Reddit — and the metasearch engines:

Top cited domains on the booking-advice prompt family (420 captures).
DomainCitations on advice promptsType
going.com54Deal alerts
kiplinger.com49Finance editorial
reddit.com48Forum / UGC
kayak.com42Metasearch
skyscanner.com42Metasearch
moneysavingexpert.com38Consumer editorial (UK)
which.co.uk34Consumer editorial (UK)
kayak.co.uk32Metasearch
help.skyscanner.net29Metasearch help pages
travelandleisure.com25Editorial
Skyscanner and Kayak are squarely inside ChatGPT's booking-advice answer set — including, in Skyscanner's case, via its help-centre pages (help.skyscanner.net, 29 citations), which ChatGPT uses to explain how the product works. Support content is citation inventory.
Section 7

US vs UK: same engine, different storefronts

Every prompt ran from both a US and a UK vantage point. The sourcing strategy barely moves — the tier mix and class mix are near-identical — but the domains localise.

Share of captures where the brand appears among fetched documents, by capture country.
BrandUS presence (share of captures)UK presenceLocalisation
Google49.2%45.8%google.com both
Kayak33.0%31.9%kayak.com vs kayak.co.uk
Expedia29.9%29.3%expedia.com vs expedia.co.uk
Skyscanner24.1%28.0%skyscanner.com vs skyscanner.net / .gg / .ie
The hotel-study null result — IP doesn't change sourcing — mostly holds for flights at the strategy level. But ChatGPT picks the country storefront: UK captures fetch skyscanner.net, kayak.co.uk and expedia.co.uk, and UK-specific editorial (moneysavingexpert.com, which.co.uk) enters the cited set. Multi-domain brands should treat every country TLD as its own AI surface.
Methodology

Study Design

Data Collection

  • 67 frozen prompts × 2 countries (US, GB) × 15 iterations = 2,007 captures (3 lost to timeouts), 18–21 July 2026, via Bright Data.
  • Four prompt families: 10 routes × 4 conditions (cheapest / round-trip price / best airline / nonstop, live August dates), 8 named airlines × 2, 7 booking-advice prompts, 4 no-search-bait questions.
  • result_source and turn_use_case parsed from the raw SSE stream: 23,308 fetched documents (82.8% tier-tagged), 4,003 cited documents joined back by canonical URL (footnote URLs carry ?utm_source=chatgpt.com).

Caveats

  • One platform (ChatGPT), one collection vendor, a 4-day window — tier shares are a snapshot; the hotel study showed the tag is dialled in and out over weeks.
  • 10 routes and 8 airlines; a different panel shifts the domain table (not the structural findings).
  • serp (n=122) and oxylabs (n=329) are small samples — directional.
  • Fresh matrix vs the hotels' 6-month backfill: comparisons are fetched-docs vs tagged-citations where noted.

Summarize with AI

ChatGPTPerplexityClaudeGeminiGrok
FAQ

Frequently Asked Questions

Continue Reading

More field tests of how AI engines find and cite sources.

All Research