July 2026AI Search Studies

AI Search for Specialty Coffee in Seoul (2026):four engines, three different social webs

TL;DR: Seoul’s specialty cafés barely publish their own websites — 64 of the 494 registry cafés list an own-domain site — and each AI engine fills that hole from a different platform. ChatGPT routes through tourism guides (35.5% of its citations) and Reddit (17.4%), with zero Instagram. Copilot’s partial batch (27 of 92 captures landed) goes majority-Instagram at 53.2%. Perplexity gives 23.6% to Korean platforms, mostly the Naver family. Café-owned sites collect 2.3% of ChatGPT’s citations — under even Marseille’s 10% — so the coffee-vertical break we set out to test is confirmed. And Copilot’s steadiest trait in the series, a 74–97% entity-website share, collapses to 25.7%.

Published July 29, 2026
2.3%
ChatGPT citations to café-owned websites
53%
Copilot citations from Instagram (n=27 captures)
409
Instagram citations — the top non-Google domain
Read the Report

Executive Summary

Seoul is the series’ third coffee city, chosen to answer a question Marseille left open — and to see what AI engines do in a market whose local web runs on Naver instead of standalone sites.

When Marseille put ChatGPT’s own-website share at 10% against a ~32% norm from the service studies, two explanations fit equally well: coffee as a vertical, or Marseille as a city. Seoul’s answer is decisive. ChatGPT cited café-owned websites 2.3% of the time (7 of its 299 citations) — and the reason is visible in the registry itself. Of 494 real cafés, just 64 list a website on their own domain; another 83 put an Instagram, Naver or Kakao page in the website field. There is close to nothing for a citation to land on.

What fills the vacuum became the study’s headline. Each engine substitutes a different platform: ChatGPT goes to Seoul’s tourism-guide layer (35.5% local editorial, led by seoultourism.org) and Reddit (17.4%), Perplexity to the Naver ecosystem (219 of its 975 citations), Copilot to café Instagram profiles (53.2% of its partial batch), and AI Mode mostly back to Google’s own surfaces (56.8%). Ask four engines the same Seoul café question and the supporting sources come from four barely overlapping corners of the internet.

The collateral finding: Copilot, whose entity-website share had stayed between 74% and 97% through every earlier study, drops to 25.7% here. Its batch is partial (27 of 92 captures — section 3 explains why), so we hold the claim loosely, but the direction is hard to dismiss: with no café websites to anchor on, the entity engine grabs the closest thing a Seoul café has to an entity page, which is its Instagram.

Terminology used throughout: a citation is a URL an engine attached as support; a mention is a café named in the visible answer prose. Source sections count citations, the leaderboard counts mentions, and section 5 shows why conflating them would wreck this particular dataset.

Section 1

Where the citations go when the websites are missing

All 4,153 cited URLs, bucketed into an eleven-part taxonomy per engine. The rose slice — café-owned websites — is the thinnest entity layer this series has measured, and a new bucket (Korean platforms: Naver, Kakao, Daum, Tistory) was needed to describe where the local web actually lives.

source-mix-by-platform-seoul-coffee

Integer percentages from raw bucket counts; each column sums to 100 ± 1 from rounding. Copilot’s column rests on 27 captures / 171 citations — the partial batch, flagged wherever it appears.

ChatGPT

A tourism-guide engine in Seoul

35.5% of ChatGPT’s 299 citations are local editorial — its largest single bucket here by a wide margin — headed by seoultourism.org (46 cites). Cafés with no site of their own get represented by whoever wrote the guidebook.

Perplexity

The one engine reading Naver

23.6% of Perplexity’s citations sit in the Korean-platform bucket — more than double any other engine — and it also keeps a real OTA layer alive (8.7%). Of the four engines it is the only one whose source mix would look plausible to a Korean user.

AI Mode

Business as usual, Korean accent

56.8% of AI Mode’s 2,708 citations point back at Google surfaces, inside its established 53–80% series range. The Korean twist is in its social slice: Instagram (6.5%) is chased by Lemon8, a ByteDance lifestyle app, at 44 citations.

The registry explains the chart before any engine enters the picture: 64 own-domain websites across 494 cafés (13%), versus 41% in Marseille, 67% for Paris bistros and 73% for Berlin tattoo studios. Seoul cafés maintain a Naver Place listing and an Instagram grid instead of a homepage — so the citation graph routes around the missing layer, each engine in its own way.
Section 2

Three engines, three different social webs

In earlier studies the engines disagreed about how much social to cite. In Seoul they disagree about which social network exists. ChatGPT’s social layer is Reddit and only Reddit; Copilot’s is Instagram; Perplexity’s is the Naver blogosphere. The same city, scraped the same day, through four engines that have apparently indexed four different internets.

social-surface-by-engine-seoul-coffee

ChatGPT: Reddit, zero Instagram

52 of ChatGPT’s 299 citations (17.4%) are Reddit threads. Its Instagram count, in a city whose café scene is arguably the most photographed on earth: zero. That repeats the Berlin tattoo result, where instagram.com also never appeared in a ChatGPT citation, and it now looks like a firm property of the engine.

Perplexity: the Naver family

blog.naver.com, cafe.naver.com and related Naver properties account for 219 of Perplexity’s 975 citations (22.5%). Korean café discovery genuinely happens on Naver blogs, and Perplexity is the only engine whose citations reflect that.

Copilot: majority Instagram

91 of Copilot’s 171 citations (53.2%) are Instagram URLs — the first majority-social column any engine has produced in this series. Partial batch (n=27), so treat the decimal with suspicion; the majority itself would need a very strange missing 65 captures to undo.

A Seoul café owner asking “where should I be visible for AI?” gets three different true answers depending on the engine: Reddit threads for ChatGPT, Naver blog coverage for Perplexity, an active Instagram for Copilot and AI Mode’s social slice. In this market there is no single social channel that covers AI search.
Section 3

Copilot loses its anchor: 74–97% → 25.7%

Through seven prior studies, Copilot was the engine you could predict blind: it cites the business’s own website, overwhelmingly. The yoga and bike-shop articles reported 95–97%; recomputed values run 89% (Tokyo), 83% (Marseille), 95% (Berlin tattoo), 74% (Paris bistros). Seoul is the first market that denies it the raw material, and the habit gives way.

25.7%
of Copilot’s Seoul citations go to café-owned websites (44 of 171).
Its majority bucket is now Instagram: 53.2%.
copilot-entity-share-across-series

The mechanism seems mundane once stated. Copilot’s personality was never “prefers websites” in the abstract — it resolves an entity, then cites that entity’s canonical web location. In every earlier market the canonical location was a homepage. For a Seoul café the canonical location is an Instagram profile, so the same resolution step now emits instagram.com URLs. The engine may well be running unchanged; the ground underneath it moved.

The honesty caveat, stated as loudly as the finding: this rests on 27 captures. Both proxy countries’ 46-item Copilot snapshots came back dominated by error items, and one US-side refire clawed back 9 more captures before we stopped (both languages are present in what landed). We imputed nothing and dropped nothing silently — every Copilot figure in this article carries its n.

Small n, big signal. A 25.7% entity share from 27 captures could drift by several points with a full batch — but it cannot drift back to 74–97%, and the Instagram majority (91 of 171 citations) would survive any plausible completion of the grid. We publish it with the caveat attached rather than sit on it.
Section 4

The “~32% baseline” was always a service-sector number

Eight cases now sit on this chart, and they no longer read as a baseline with exceptions. They read as two populations. Verticals where the business explains and books its service on its own site — yoga, bike shops, tattoo — cluster at 32–42%. Food and retail scenes sit at 8, 10, 0.9 and now 2.3.

chatgpt-own-website-share-eight-cases

Seoul was picked to settle exactly this. Marseille’s 10% could have been a coffee thing or a Marseille thing; running coffee through a third, structurally different city was the test. At 2.3%, the verdict is in: the break is a vertical trait, not a city artifact. Wherever the businesses themselves publish thin owned webs, ChatGPT builds its answer from third parties — guides, aggregators, Reddit — and the city only decides which third parties those are.

In Seoul the substitute of choice is the tourism-guide layer: 35.5% of ChatGPT’s citations are local editorial, with seoultourism.org alone at 46 of 299 — a bigger single share of the engine’s pool than all 494 cafés’ websites combined (7).

Comparisons that survive scrutiny

The service-side values come from Paris yoga (32), Berlin yoga (32), Amsterdam bikes (42) and Berlin tattoo (35); the floor from Tokyo (8), Marseille (10) and Paris bistros (0.9). Every value was re-read from the sibling article’s data arrays before publishing this one.

Why 2.3 beats even the bistros’ 0.9 barely

Paris bistros at 0.9% had a large registry of restaurants with websites that ChatGPT simply ignored in favour of the food-guide press. Seoul’s 2.3% has the opposite anatomy: the engine shows no aversion to café sites — there are only 64 of them to find. Same chart position, different failure mode; the fix for a café owner (build the site) differs from the bistro owner’s (get the site cited).

Section 5

The consensus cafés

Ranked by how many of the 303 captured answers name each brand in prose. Fritz Coffee Company leads at 63 mentions and appears on every engine that fired, with Coffee Libre (60) and Anthracite Hapjeong (56) close behind — the canonical first-wave roasters, which, notably, the original Apify grid scrape never returned. All three entered the registry through Google Places recovery.

seoul-coffee-leaderboard-answer-mentions

Per-engine mention matrix

#CaféAI Mode92 capturesChatGPT92 capturesPerplexity92 capturesCopilot27 captures
1Fritz Coffee Company22.8%(21)28.3%(26)15.2%(14)7.4%(2)
2Coffee Libre16.3%(15)30.4%(28)13%(12)18.5%(5)
3Anthracite Coffee Hapjeong.17.4%(16)27.2%(25)13%(12)11.1%(3)
4ACR15.2%(14)18.5%(17)14.1%(13)22.2%(6)
5Namusairo Coffee23.9%(22)14.1%(13)10.9%(10)18.5%(5)
6Coffee Hanyakbang23.9%(22)12%(11)10.9%(10)7.4%(2)
7Leesar Coffee13%(12)15.2%(14)8.7%(8)3.7%(1)
8Rewire Coffee13%(12)17.4%(16)4.3%(4)11.1%(3)
9Cafe Onion Anguk12%(11)8.7%(8)16.3%(15)3.7%(1)
10LowKey Seongsu12%(11)18.5%(17)2.2%(2)3.7%(1)
11LEEDORIM Coffee & Vegan Bakery Gyeongbokgung Cafe20.7%(19)8.7%(8)0(0)11.1%(3)
12Center Coffee Gwanghwamun5.4%(5)9.8%(9)13%(12)0(0)
Top 12 cafés — mention ranking with the (broken) citation metric alongside
#CaféAnswer mentionsCitation scoreEngines (of 4)
1Fritz Coffee Company6304
2Coffee Libre6024
3Anthracite Coffee Hapjeong.5604
4ACR50344
5Namusairo Coffee5034
6Coffee Hanyakbang45164
7Leesar Coffee35104
8Rewire Coffee3564
9Cafe Onion Anguk3504
10LowKey Seongsu3104
11LEEDORIM Coffee & Vegan Bakery Gyeongbokgung Cafe30103
12Center Coffee Gwanghwamun2663

The citation-score column deserves its own autopsy, because in Seoul the dual-metric design produced its widest split yet — in both directions at once. Direction one: the consensus #1, Fritz, scores 0 on citations. Its registry row carries no own-domain website for the domain matcher to catch (fritz.co.kr shows up inside captures but never in the registry’s website field), so the counter is blind to it — the same failure class as NAJS in the Berlin tattoo study, except now it sits on the very top row. Cafe Onion Anguk (#9, 35 mentions) and LowKey Seongsu (#10, a registered lowkeycoffee.com) also score 0.

Direction two: the pipeline’s cite-counted CSV crowns a different #1 entirely — 카페 공동, a Nowon café with a score of 101 from a single platform. Its listed “website” is its Instagram page, so it absorbs instagram.com citations that have nothing to do with it. The mapdata generator’s non-identifying-domain guard excludes it from the board you see above; the raw CSV retains it as a cautionary exhibit.

In a market where 64 of 494 venues have an own-domain website, domain-matched citation counting is structurally blind — it erases the real winner (Fritz, 63 mentions, cite score 0) and crowns an Instagram-keyed artifact (카페 공동, score 101). The mention leaderboard is the only ranking we trust for Seoul, and any “AI visibility score” built on citation-domain matching would be unusable here.
Section 6

Instagram tops the table; ChatGPT never touches it

The cross-engine domain table makes the three-way social split concrete. instagram.com leads all non-Google domains at 409 citations, blog.naver.com takes second at 192, and reddit.com — the top community domain in the Tokyo, bistros and tattoo studies — is third at 174. Marseille had already put Instagram in front (237 cites); Seoul’s new twist is that Naver blogs pass Reddit too.

DomainEnginesCitesWhat it is
instagram.com3/4409Café profiles — top non-Google domain; none of these come from ChatGPT
blog.naver.com2/4192Naver blogs — Korea’s dominant blogging platform
reddit.com3/4174Community threads; ChatGPT’s preferred social surface
youtube.com2/455Café-tour and vlog content
lemon8-app.com2/448ByteDance’s lifestyle app — AI Mode’s #2 social source
seoultourism.org2/448Seoul tourism guide — ChatGPT’s single biggest source (46 of its 299)
facebook.com2/435Café pages
brewatlas.co4/433Specialty-coffee directory — cited by every engine that fired
diningcode.com4/431Korean restaurant/café aggregator — also all four engines
perfectdailygrind.com2/430Global specialty-coffee press
v.daum.net3/429Daum content portal

google.com is excluded from the table: its 1,536 citations all come from AI Mode referencing Google’s own surfaces, and would drown every row above (see section 10).

Only two domains reach all four engines

brewatlas.co (33 cites), a specialty-coffee directory, and diningcode.com (31), a Korean dining aggregator, are the sole domains cited by every engine that fired. Neither is a café. For a scene without websites, the closest thing to engine-consensus infrastructure is a pair of third-party directories.

The ChatGPT asymmetry, again

All 409 Instagram citations come from Copilot, AI Mode and Perplexity; ChatGPT contributes none while holding Reddit at 17.4% of its pool. Two data points (Berlin tattoo, Seoul) now show ChatGPT citing zero Instagram in Instagram-native verticals — whatever the cause, Instagram investment currently buys café visibility in three engines out of four.

Section 7

English Seoul and Korean Seoul barely intersect

On the control prompt, ChatGPT’s English and Korean top-5 lists share 2 cafés of 8 distinct — a 25% overlap, landing exactly where Tokyo (25%) and Paris bistros (25%) did and leaving Berlin tattoo’s 67% as the series outlier. Below the consensus brands, the two language worlds diverge completely: every district prompt overlaps at 0%.

EN vs KO top-5 overlap by prompt (ChatGPT, local proxy)
Prompt templateEN vs KO top-5 overlap
vibe: aesthetic cafés43% (3/7)
control (best specialty coffee shops)25% (2/8)
matcha25% (2/8)
vegan-friendly25% (2/8)
roasters25% (2/8)
cold brew25% (2/8)
tearoom crossover25% (2/8)
vibe: brunch25% (2/8)
vibe: work-friendly25% (2/8)
pour-over11% (1/9)
price: cheap11% (1/9)
district: Ikseon-dong0% (0/10)
district: Seongsu-dong0% (0/10)
district: Yeonnam-dong0% (0/10)
pastry / bakery cafés0% (0/10)
vibe: quiet0% (0/10)

Overlap = shared entities ÷ distinct entities across both top-5 lists (Jaccard). KO mentions were bridged to the registry through a 49-entry hand-checked Hangul→Latin alias table (커피리브레 → Coffee Libre).

The one prompt where the languages agree most is aesthetic cafés (43%) — the photogenic flagships are famous in both webs. The prompts where they agree least are the local ones: ask about Ikseon-dong, Seongsu-dong or Yeonnam-dong and the Korean answer draws on neighbourhood Naver-blog knowledge while the English answer recycles the international guide circuit. Zero shared names, in all three districts.

The series’ usual companion metric — how strongly the prompt language couples to country-code domains — simply cannot be computed here, and that absence is itself the result. Tokyo coupled Japanese prompts to .jp at 5.0×; Berlin yoga measured .de at 1.5×. Seoul’s 64 own-domain cafés (a mix of .com, .co.kr and .kr) are too few for any trustworthy ratio, because the Korean-language layer engines actually cite lives on platform subdomains: blog.naver.com, v.daum.net, *.tistory.com. Korean web identity is platform-hosted, so a country-TLD metric has nothing to grip.

One genuinely new language behaviour surfaced: 14 of the 303 captures recommended no venue at all, and 13 of those are Perplexity answering in Korean (the fourteenth is Perplexity in English). That is 13 of its 46 Korean captures — 28% — returning generic how-to-find-a-café advice with zero names. No EN-market run in this series has produced refusals at that rate, from any engine.

Section 8

494 cafés on the map, 12 in the answers

The dots are the 499-row registry (494 real cafés after cleanup), built in two Apify passes plus 187 Google Places recoveries; the numbered markers are the top 12 by answer mentions. The favourites hug the centre-north arc — Jung-gu’s Euljiro corridor, Jongno, Yeonnam/Hapjeong and Seongsu — while the grid’s outer-district cafés (Dobong, Nowon, Gwanak) never break into an AI answer.

All 494 registry cafés across Seoul

Top 12 cafés by answer mentions — click a marker for per-engine counts

1Fritz Coffee Company2Coffee Libre3Anthracite Coffee Hapjeong.4ACR5Namusairo Coffee6Coffee Hanyakbang7Leesar Coffee8Rewire Coffee9Cafe Onion Anguk10LowKey Seongsu11LEEDORIM Coffee & Vegan Bakery Gyeongbokgung Cafe12Center Coffee Gwanghwamun

Does asking for a district get you that district?

For the six district templates, we checked whether ChatGPT’s resolved venues actually sit in the asked-for neighbourhood. The spread is enormous:

District promptEN accuracyKO accuracy
Ikseon-dong100% (5/5)100% (2/2)
Mangwon-dong75% (3/4)— (0 resolved)
Seongsu-dong57% (4/7)67% (4/6)
Yeonnam-dong38% (3/8)— (0 resolved)
Euljiro31% (5/16)0% (0/4)
Hannam-dong8% (1/13)17% (1/6)

Ikseon-dong — small, named, walled — scores 100% in both languages; Hannam-dong bottoms out at 8–17% because the engine answers with Itaewon/Yongsan-wide picks. Caveat before quoting these: Places-recovered venues carry no district field, so denominators are small and this is a best-effort measure, an improvement on the Berlin/Amsterdam registries where district accuracy could never register at all.

Section 9

Eight studies, one scoreboard

Seoul against the running series metrics. Prior values were recomputed from the data arrays inside Marseille coffee, Berlin tattoo, Paris bistros and Tokyo bookstores as published on this site.

MetricPrior studiesSeoul coffeeReading
ChatGPT own-website %Services 32–42 (Paris yoga 32 · Berlin yoga 32 · bikes 42 · tattoo 35) · food/retail: Tokyo 8 · Marseille 10 · bistros 0.92.3%Food/retail floor confirmed as a vertical trait
Copilot entity-website %Tattoo 95 · Tokyo 89 · Marseille 83 · bistros 74 (yoga/bikes articles: 95–97)25.7% (n=27)First break of the series’ steadiest metric
AI Mode google.com %Range 53–80 (tattoo 53 · Tokyo 61 · bistros 70 · Marseille 80)56.8%Inside range — untouched by the Korean web
Top social domainMarseille: Instagram 237 over Reddit 175 · tattoo: reddit 237 · Tokyo: reddit 104 · bistros: reddit 74Instagram 409 · Naver blogs 192 · Reddit 174Naver blogs pass Reddit too; ChatGPT still cites zero Instagram
Language→TLD couplingTokyo .jp 5.0× · Berlin yoga .de 1.5× · tattoo .de 0.87×UnmeasurableKorean web identity is platform-hosted — no TLD to couple to
EN vs local top-5 overlap (control)Tattoo 67 · bistros 25 · Tokyo 25 · Marseille 1125%Tattoo remains the outlier
Mentions #1 vs cites #1Tattoo: OMEN both · Marseille: Deep both · bistros: splitFritz 63 mentions / cite score 0 · CSV crowns an Instagram-keyed artifactWidest dual-metric divergence yet, in both directions
Section 10

Engine-level field notes

ChatGPT: live web on 87 of 92

95% of ChatGPT captures triggered web search, and the run produced 800 venue-panel entities overall. The five no-search answers leaned on training-data memory of the famous roasters — which happens to be the one tier of Seoul coffee that memory covers adequately.

AI Mode × Korea: the batch that landed

After AI Mode rejected FR-proxy prompts in two French studies, the KR-proxy batch went through 92/92 with no trigger issues. Its Korean flavour shows up in sources rather than availability: Lemon8 at 44 citations is its second social domain, a platform no other study in this series ever surfaced.

Perplexity’s Korean-language refusals

The 14 zero-recommendation captures are a one-engine, one-language phenomenon: 13 are Perplexity in Korean, one is Perplexity in English, and the other three engines produced none at all. An engine that names venues fluently in English going name-free in 28% of its Korean answers is a localisation gap that visibility tooling built on EN prompts would never detect.

Copilot: how a batch dies

Both the US and KR 46-item Copilot snapshots returned payloads made mostly of error items; a targeted US refire recovered 9 captures, bringing the total to 27 of 92. We chose to publish the partial column with its n on every figure — the alternative, quietly leaving Copilot out, would have hidden the series’ most interesting break.

For cafés & roasters

If you pull shots in Seoul

  • Your Instagram is your citable entity page. With 409 citations, instagram.com carries more AI-search weight here than every café website combined — Copilot and AI Mode resolve cafés to their profiles directly. Keep the handle, hours and location current; that profile is what three of four engines link.
  • Naver blog coverage feeds Perplexity. The Naver family took 219 of Perplexity’s 975 citations. Blogger visits and Naver Place upkeep already matter for Korean customers; they now also decide whether an AI engine can source you.
  • ChatGPT reaches you through guides and Reddit. seoultourism.org alone supplied 46 of its 299 citations, and Reddit threads another 17.4% — while Instagram supplied zero. Placement in the English guide layer and presence in r/seoul-style threads are the only levers this engine responds to.
  • An own-domain site is a cheap differentiator, at low current stakes. Only 64 of 494 cafés have one, and the engines cite what exists (2.3% for ChatGPT). If the entity-citation channel grows, the site also fixes the measurement blindness that gives even Fritz a citation score of 0.
  • Work both languages, because the answers don’t mix. EN/KO top-5 overlap is 25% on the control prompt and 0% on every district prompt. English guide coverage and Korean Naver coverage are two separate campaigns reaching two separate answer sets.
Conclusion

What AI search does when the open web is somewhere else

Seven earlier studies ran in markets where local businesses at least plausibly lived on the open web. Seoul is the first where they don’t — café identity here is a Naver Place listing and an Instagram grid — and none of the four engines failed for it. Each one simply re-routed to whichever platform its retrieval already favoured: ChatGPT to guides and Reddit, Perplexity to Naver, Copilot to Instagram, AI Mode to itself. The answers stayed fluent and confident throughout; only the plumbing underneath changed completely.

That re-routing had a casualty and a beneficiary. The casualty is measurement: citation-domain matching, the standard mechanic of AI-visibility tooling, returns garbage in a market without entity domains — scoring the consensus favourite at zero while a mis-keyed Instagram artifact tops the raw CSV. The beneficiary is anyone who understands the split: with the engines this divided, a Seoul café can rank in one AI assistant and be invisible in the next, and the difference is traceable to specific, workable surfaces.

The open question Seoul sets up: is this the profile of every platform-first web, or a Korea-specific shape? A Tokyo café run, or a Jakarta or São Paulo one, would say whether “the engines each pick a platform” generalises. Marseille earned its follow-up; this result deserves one too.

Methodology

Study design

Data collection

  • 23 prompt templates × 2 languages (EN/KO) × 2 proxy countries (US/KR) × 4 AI engines = 368 theoretical items; 303 captured. Per engine: ChatGPT 92/92, Perplexity 92/92, Google AI Mode 92/92 — the AI Mode × KR batch landed cleanly, unlike the FR-proxy rejections of the French studies — and Copilot 27/92 (see below).
  • Gemini was not fired this run. It was dropped before capture to keep the run inside its scrape cost cap, after the seed required a second, district-targeted Apify pass. Its absence is disclosed everywhere it matters; no Gemini values are imputed anywhere in this article.
  • Copilot partial batch: both proxies’ 46-item snapshots returned mostly error items; a US refire recovered 9 captures for a final 27/92 with both languages represented. Every Copilot statistic carries the n=27 / 171-citation caveat.
  • Totals: 303 captures · 4,153 cited URLs · 800 map entities · 2,285 extracted café mentions (71.9% resolved; the bistros study shipped at 70%). Captured 2026-07-29 via Bright Data.
  • Registry: 314 Apify Google Maps rows in two passes (200 city-grid — which skewed to outer residential districts — plus 114 targeted at Seongsu/Yeonnam/Hannam/Ikseon/Mangwon/Euljiro/Gangnam/Hongdae). One Texas geocoding leak deleted and one brunch spot classified out → 312. Google Places recovery then added 187 venues the grid missed — including Fritz, Coffee Libre, Anthracite and Onion, the exact cafés the engines recommend most → 499-row registry, 494 real cafés.
  • Scrapes ran through the pipeline’s in-process fallback runner (Modal was unreachable from this sandbox); identical payloads, tables and parsing as the standard path.

What we measured

  • Cafés named per answer (brand-aggregated; ranked by mentions, with the citation score shown for contrast)
  • Cited URLs bucketed into an eleven-part taxonomy, including a Korean-platform bucket new to this study
  • Per-engine social-surface concentration (Reddit vs Instagram vs Naver)
  • ChatGPT own-website share vs the seven prior cases
  • EN vs KO top-5 overlap per template, via a 49-entry Hangul→Latin alias table
  • District-targeting accuracy for six named neighbourhoods
  • Zero-recommendation (refusal) captures by engine and language

From prose answers to countable cafés

Engines reply in running text, and in two scripts — a Korean answer will recommend 커피리브레 where the English one says Coffee Libre. The pipeline’s NER pass reads each of the 303 answers with an LLM under fixed extraction rules (named venues only, recommendation context required, no inference beyond the text), then a deterministic resolver matches names to the registry: normalised exact match, then fuzzy similarity, with a hand-checked 49-entry Hangul→Latin alias table bridging the script gap the fuzzy matcher cannot cross. Frequently recommended names that still failed to resolve were checked against Google Places, which is how 187 canonical-scene venues entered the registry. Final resolution: 71.9% of 2,285 mention rows.

Extractor disclosure: this run’s extraction was performed by a Claude model rather than the pipeline’s usual Gemini extractor (the sandbox could not reach the Gemini key) — same prompt rules, same deterministic resolution path. As a QA gate, all 2,285 extracted names were mechanically verified to appear in their source answer text; zero came back suspect. The KO prompt translations are competent but were not reviewed by a native speaker, the same limitation the Tokyo study carried for Japanese.

Caveats

  • Copilot is a partial batch (27 of 92 captures, 171 citations). Its 25.7% entity share and 53.2% Instagram majority are directionally strong and numerically soft.
  • Gemini is absent by design (cost cap), so cross-engine claims cover four engines, and series comparisons involving Gemini stop at the prior studies.
  • The website-field undercount cuts both ways: Places-recovered rows carry no website field, so “64 of 494 with own-domain sites” is an upper-bound picture built on incomplete data — the layer is genuinely thin, but the exact figure is soft.
  • Citation scores are structurally unreliable in this market (section 5); we publish them only to document the failure mode.
  • District accuracy is best-effort: recovered venues lack a district field, leaving small denominators.
  • NER resolution is precision-first; unresolved mentions are dropped, so brand counts are conservative lower bounds.
  • Single-city, single-run snapshot (July 2026); engine behaviour moves with model updates.
  • Disclosure: no personal affiliation with any Seoul café.
FAQ

Frequently Asked Questions

Summarize with AI

ChatGPTPerplexityClaudeGeminiGrok