Paper deep dive
Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Qiaoni Shi, Kai Zhu, Kai Gu
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/15/2026, 8:16:19 AM
Summary
This paper examines how AI search interfaces, particularly ChatGPT, reshape web traffic allocation and digital intermediation compared to traditional search engines like Google. Using URL-level Comscore U.S. desktop clickstream data, the authors demonstrate that AI search retains the majority of information needs internally, producing outbound clicks in only 5.2% of sessions versus 31.1% for Google. Residual AI search traffic skews toward specialized, non-ad-supported destinations (e.g., academic, developer, SaaS) and causally displaces traditional search queries by 9.4% as access expands. The study identifies a fundamental economic shift where intermediaries resolve queries directly, weakening the traditional referral bargain that historically linked search traffic, content production, and digital advertising.
Entities (10)
Relation Signals (8)
ChatGPT ā hasoutboundclickrate ā 52%
confidence 96% Ā· ChatGPT produces outbound clicks in only 5.2% of conversation sessions
Google ā hasoutboundclickrate ā 31.1%
confidence 96% Ā· compared with 31.1% of Google queries
AI search ā displaces ā Traditional search
confidence 92% Ā· Wider access cuts search use by 9.4%
OpenAI ā released ā ChatGPT Search
confidence 90% Ā· OpenAI extended ChatGPT Search to paid subscribers on October 31 2024
AI search ā resolvesneedsinside ā Intermediary
confidence 90% Ā· information needs can be resolved inside the intermediary
ChatGPT ā routestrafficto ā Specialized destinations
confidence 88% Ā· they skew toward specialized destinations and away from ad-supported sites
Digital intermediation ā undergoesshiftdueto ā AI search adoption
confidence 87% Ā· Our findings identify a central economic shift in digital intermediation: AI search might satisfy information needs inside the intermediary
Comscore U.S. desktop clickstream ā usedforanalysis ā AI search displacement
confidence 85% Ā· Using URL-level Comscore U.S. desktop clickstream... estimate traditional search displacement
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream, we compare ChatGPT and Google information-seeking occasions and exploit ChatGPT Search access expansions to estimate traditional search displacement. ChatGPT produces outbound clicks in only 5.2% of conversation sessions, far below Google's referral ratio. The remaining clicks are not a scaled-down Google stream: they skew toward specialized destinations and away from ad-supported sites. Wider access cuts search use by 9.4%, with search-referral losses largest for informational categories. Our findings identify a central economic shift in digital intermediation: AI search might satisfy information needs inside the intermediary while weakening the referral bargain that has linked search, traffic, and content production on the open web.
Tags
Links
- Source: https://arxiv.org/abs/2607.07652v1
- Canonical: https://arxiv.org/abs/2607.07652v1
Trouble viewing inline? Open PDF directly ā
Full Text
198,077 characters extracted from source content.
Expand or collapse full text
Answering Without Referring: How AI Search Rewrites the Webās Economic Bargain Qiaoni Shi, Kai Zhu, and Kai Gu Bocconi University, Milan, Italy This version: July 2026 Abstract Search engines have long allocated attention on the web by routing users from queries to websites. AI search changes this arrangement because information needs can be resolved inside the intermediary. Using URL-level Comscore U.S. desktop clickstream, we compare ChatGPT and Google information-seeking occasions and exploit ChatGPT Search access expansions to estimate traditional search displacement. ChatGPT pro- duces outbound clicks in only 5.2% of conversation sessions, far below Googleās referral ratio. The remaining clicks are not a scaled-down Google stream: they skew toward specialized destinations and away from ad-supported sites. Wider access cuts search use by 9.4%, with search-referral losses largest for informational categories. Our find- ings identify a central economic shift in digital intermediation: AI search might satisfy information needs inside the intermediary while weakening the referral bargain that has linked search, traffic, and content production on the open web. Keywords: AI search; web traffic; digital intermediation; attention economy; search adver- tising. 1 arXiv:2607.07652v1 [cs.CY] 8 Jul 2026 1 Introduction Search engines lowered the cost of finding information (Bakos, 1997; Goldfarb and Tucker, 2019) and became the webās dominant mechanism for allocating attention (Simon, 1971; Athey and Ellison, 2011). For two decades, the economic bargain was simple: users expressed information needs, search engines ranked and monetized queries, and destination websites received visits they could convert into advertising impressions, subscriptions, purchases, or brand equity (Athey and Ellison, 2011; Edelman et al., 2007; Varian, 2007; Ghose and Yang, 2009). The bargain was implicit rather than contractual: search engines captured value at the query, while content producers supplied information and captured value after the click. AI search, an answer interface that combines web retrieval with language-model synthesis, changes this bargain by moving resolution inside the intermediary. Instead of returning only ranked links, the interface can retrieve from websites, synthesize a response, and satisfy the information need before the user visits a source. Figure 1 contrasts this architecture with Google search. Google monetizes the query and routes attention through a click; AI search may end with an in-chat answer, leaving websites only an optional residual click-out stream. This is a paradigm shift in the webās traffic economy: discovery no longer has to cul- minate in a visit, because the intermediary can resolve the need before the user reaches a source. Citations and links can preserve attribution, but they do not recreate the atten- tion, ad impression, conversion opportunity, or subscriber relationship attached to a visit. That difference is economically consequential because traffic remains central to content pub- lishing, digital advertising, and online commerce (Chiou and Tucker, 2017; Jeon and Nasr, 2016; Zhao and Berman, 2026). The empirical question is therefore how often AI search retains information needs, where the residual referrals go, and whether wider access reduces traditional search. 1 Figure 1. AI search changes web intermediation from routed visits to residual referrals. Notes: The left column summarizes traditional search: the intermediary ranks links, monetizes the query, and routes the user to a destination website. The right column summarizes AI search: retrieval and synthesis can resolve the task inside the intermediary, so websites receive only optional residual click-out traffic. A growing academic literature documents pieces of this transition. Recent work has been valuable in measuring how LLM adoption changes online behavior, whether LLM use sub- stitutes for or complements traditional search, and how publishers respond to generative AI (Padilla et al., 2025; Gholami et al., 2026; Zhao and Berman, 2026). A closely related paper documents ChatGPT referrals in destination-side e-commerce traffic (Kaiser and Schulze, 2026). We build on these studies by examining user-side information seeking under AI searchāthat is, language-model interfaces with web retrievalāincluding absorbed sessions that leave no destination-side trace, and by comparing those occasions with traditional search across the open web. 2 We provide that user-side view using URL-level Comscore U.S. desktop clickstream from October 2024 through July 2025, a window that spans the initial ChatGPT Search rollout. The data record the same householdās page loads, timestamps, and HTTP referrers, allowing us to reconstruct ChatGPT conversation sessions, traditional search queries, and outbound referrals. We compare ChatGPT sessions and Google queries within the same household- week and estimate within-household, within-week differences in routing. We then exploit three expansions of ChatGPT Search access: October 31 2024 for paid subscribers, Decem- ber 16 2024 for free logged-in users, and February 5 2025 for anonymous browsers (OpenAI, 2024). A stacked difference-in-differences design pools cohorts newly eligible at each date and compares them with reweighted households without pre-expansion ChatGPT or Claude activity to estimate the effect of wider access on traditional search use (Cengiz et al., 2019; Callaway and SantāAnna, 2021). Three results characterize how AI search changes the webās traffic allocation. First, ChatGPT retains most expressed information needs. It produces a clean outbound click in only 5.2% of conversation sessions, compared with 31.1% of Google queries, and the gap persists within the same household and week. Second, the residual traffic is not a proportional sample of traditional search traffic. Users invoke ChatGPT relatively more in reference and tool-oriented browsing contexts, but ChatGPT clicks out most often in technical and e-commerce contexts; its referrals favor reference, academic, developer, and tools/SaaS destinations and avoid much of the ad-supported web. Third, wider access to ChatGPT Search reduces traditional search queries by 9.4% on average and 17.0% after twenty weeks, with the loss concentrated in informational categories. The paperās primary contribution is to identify the routing margin through which AI search reallocates attention. Traditional search reduced search costs while usually routing users onward; AI search can reduce those costs by resolving the task inside the intermediary. By measuring sessions that do and do not exit the AI interface, comparing ChatGPT and Google within household-week, and using access expansions to estimate search displacement, we connect three margins that are usually observed separately: retention inside AI search, 3 the composition of residual referrals, and downstream substitution away from traditional search. Our claim is deliberately narrower than a welfare claim: we measure a change in observable traffic allocation, not consumer surplus, publisher revenue, or long-run content production. The findings contribute to three literatures. First, we extend research on attention, search, and digital intermediation by distinguishing intermediaries that route filtered infor- mation needs from answer interfaces that can resolve those needs themselves (Simon, 1971; Goldfarb and Tucker, 2019; Athey and Ellison, 2011; Chiou and Tucker, 2017; Jeon and Nasr, 2016). Second, we complement research on generative-AI adoption, productivity, and online participation by tracing how use inside the AI interface changes behavior outside it (Noy and Zhang, 2023; Brynjolfsson et al., 2025; Bick et al., 2024; Chatterji et al., 2025; Burtch et al., 2024; del Rio-Chanona et al., 2024). Third, we extend emerging evidence on AI search and website traffic by observing absorbed information-seeking occasions across the web, compar- ing intermediaries within households, and identifying traditional search displacement after access expansions (Padilla et al., 2025; Kaiser and Schulze, 2026; Gholami et al., 2026; Zhao and Berman, 2026). The paper proceeds as follows. Section 2 describes the data and methods. Section 3 presents the results. Section 4 discusses implications and future research. 2 Data and Method URL-level Comscore U.S. desktop clickstream data are well suited to the routing questions we study because they follow the same households across intermediaries and over time. The raw data contain foreground and background requests; we construct a foreground-visit layer for user-facing browsing and retain timestamps, HTTP referrers, search-query URLs, ChatGPT conversation endpoints, and surrounding browsing contexts. This user-side view lets us observe both referrals to destination websites and information-seeking occasions that leave no destination-side trace. The rest of this section defines the panel, user-side units, referral 4 measures, domain taxonomy, access-expansion cohorts, and adoption patterns that support these tests. Data and units. The full Comscore sample contains between 168,467 and 238,315 active U.S. desktop households per month. A balanced sub-panel of 45,386 households appears with at least four foreground page loads in every month; we use it only for the traditional search displacement analysis. A foreground visit is an analytic page-load layer with a valid HTML MIME type and response code. This restriction removes assets, background requests, and redirects from the raw clickstream. A search query is a foreground Google, Bing, or Yahoo load whose URL matches the platformās search pattern. Online Appendix Sections OA.1.1ā OA.1.3 document panel construction, traffic cleaning, and endpoint coverage. Conversation sessions. Raw ChatGPT records do not map one-to-one to occasions of user information seeking. ChatGPT is a single-page application, so one conversation can generate repeated endpoint loads as the user sends messages, revisits the tab, or continues the exchange. Each conversation-endpoint URL contains a conversation identifier. We define a conversation session as the maximal sequence of loads on one machine that share this identifier, opening a new session when the same identifier reappears after more than one hour. A different identifier always begins a new session. This construction deduplicates repeated loads without combining distinct conversations. It also gives ChatGPT a user- side unit comparable to a Google query: one occasion on which a household brought an information need to an intermediary. Online Appendix Section OA.1.4 reports the threshold sensitivity. Referrals. We observe an outbound referral when a foreground visit to a third-party web- site carries ChatGPT or a search engine in its HTTP referrer; a visit with no observable referrer is treated as a direct visit rather than a referral from search or another external source. A clean referral further removes self-referential traffic, platform-internal navigation, and search-results pages reached from another source. We keep genuine user clicks and 5 exclude machine-initiated programmatic requests, following the URL patterns in Online Ap- pendix Section OA.2.1. These user-side events provide a common denominator for comparing where attention ends. For platform p, we define the referral ratio as the share of conversation sessions or search queries that produce at least one clean outbound referral. Its complement is absorptiveness, the share that ends without a clean referral. We compare one ChatGPT conversation session with one Google query because each represents an occasion on which a household brought an information need to an intermediary. The measures do not condition on a results-page impression or destination visit. They describe traffic rather than consumer welfare: an absorbed session may be valuable to the user even though it creates no observable visit for a website. Domain classification and category shares. We classify 4,266 matched-support des- tination domains into content-type and monetization categories. A high-confidence sub- set of 3,245 domains supports content-type comparisons, surrounding-context labels, and displacement-by-category results (Online Appendix Appendix OA.3). For category C and event set E, we report the difference in category shares: ā E (C) = Share CG,E (C)ā Share G,E (C). Positive values mean that category C is more common for ChatGPT than for Google in the event set being analyzed. In the user-intent analysis, E is classifiable surrounding-context anchors; in the destination-composition analysis, E is clean outgoing referrals. Access expansions and design. OpenAI extended ChatGPT Search to paid subscribers on October 31, 2024, free logged-in users on December 16, and anonymous browsers on February 5, 2025. We define cohorts using pre-expansion signals that remain fixed afterward. The paid-subscriber group contains 84 households (P); the logged-in group contains 2,440 API-dominant households (L); and the anonymous-browser group contains 1,358 anonymous- dominant households (A). The pooled treated sample contains 3,882 households. The pre- 6 ferred N itt control contains households with no pre-expansion ChatGPT or Claude activity (Claude being the closest competing chat-search assistant). We reweight treated observations toward the American Community Survey distribution using age by household-income cells. Each treated cohort appears only in its own stack and is matched to its own N itt control, de- fined as the households with no ChatGPT or Claude activity as of that expansionās date; the control pool therefore shrinks for later expansions. Online Appendix Sections OA.6.1āOA.6.2 give the endpoint classifier, cohort-threshold sensitivity, and balance diagnostics. Adoption. ChatGPTās monthly active-household reach rises from about 4.6% in Octo- ber 2024 to a peak near 7.9% in mid-2025 (Figure 2). Gemini, Claude, and Perplexity remain below 1.5%, while Googleās reach stays near 49%. We infer search-enabled ChatGPT conversation sessions from an audited endpoint signal, after subtracting contemporaneous non-search surfaces within the session. 1 This inferred search-use series rises by approximately a factor of forty-five per active household from November 2024 to its April 2025 peak. Be- cause adoption trends alone cannot show whether this activity adds new tasks or displaces traditional search, the displacement analysis below relies on the access-expansion design. 3 Results We organize the findings in three steps. We first measure how often information-seeking occasions exit each intermediary. We then examine the browsing contexts in which the residual referral traffic occurs and which destinations receive it. Finally, we use the three access expansions to estimate whether AI search displaces traditional search. This sequence moves from user-side retention, to differential routing, to causal substitution. 1 The signal is the paragen_submission endpoint. It is shared with Canvas, file uploads, and other paragen-routed A/B surfaces, so the measure is an imperfect proxy and likely an upper bound. It remains the best panel-wide search signal before July 2025, when message-level f/conversation coverage reaches analytic scale; Online Appendix Section OA.1.3 details the endpoint audit. 7 Figure 2. ChatGPT reach rises after access expansions while Google reach remains stable. Paid Logged-in No Login ChatGPT Claude Gemini Perplexity 0.0% 2.0% 4.0% 6.0% 8.0% 2024-11 2025-01 2025-03 2025-05 2025-07 Share of households (a) LLM visit users. Paid Logged-in No Login Yahoo Bing Google 0% 20% 40% 60% 2024-11 2025-01 2025-03 2025-05 2025-07 Share of households (b) Search-engine query users. Notes: Each series reports the monthly share of active Comscore U.S. desktop households that visits an LLM platform (panel a) or issues a search-engine query (panel b). Dashed vertical lines mark the October 31, December 16, and February 5 ChatGPT Search access expansions. ChatGPT reach rises from 4.6% in October 2024 to a peak near 7.9%; Google remains near 49%. 3.1 AI search retains most information-seeking occasions Measuring routing from the intermediary. The central measurement decision is the denominator. A click-through rate begins with a results-page impression or a displayed link. We instead begin with an information-seeking occasion observed at the intermediary and ask whether it produces at least one clean referral. For ChatGPT, that occasion is a conversation session; for Google, it is a query. A conversation can contain several messages and repeated page loads before any click occurs. We therefore count the whole exchange once, using the conversation identifier, and ask whether that exchange ever generated a clean referral. Figure 3 varies the session boundary from thirty minutes to five hours. The estimates move little, showing that the comparison does not depend on the one-hour boundary. ChatGPT routes substantially fewer information-seeking occasions than Google. Its monthly referral ratio rises from approximately 2.5% to a peak near 6.5%, but remains far below Googleās per-query ratio (Figure 3). Pooled across the sample, ChatGPT pro- duces a clean referral in 5.2% of sessions, compared with 31.1% of Google queries (Table 1). The household-level pattern is starker: 74.4% of 56,578 ChatGPT-active households never produce a clean ChatGPT referral during the ten-month window, compared with 9.6% of 8 291,747 Google-querying households. Figure 3. ChatGPT routes far fewer information-seeking occasions than Google under every session definition. Paid Logged-inNo Login 1 h (primary) 0.5 h 5 h 0.0% 2.0% 4.0% 6.0% 2024-11 2025-01 2025-03 2025-05 2025-07 Referral ratio (a) ChatGPT conversation sessions. Paid Logged-in No Login Bing Google Yahoo 0% 10% 20% 30% 40% 2024-11 2025-01 2025-03 2025-05 2025-07 Referral ratio (b) Search-engine queries. Notes: Points report the monthly share of user-side information-seeking occasions that produce at least one clean outbound referral. Panel (a) plots ChatGPT conversation sessions: the red series uses the preferred one-hour boundary, and gray series use thirty-minute and five-hour boundaries. Panel (b) plots Google, Bing, and Yahoo search queries, with Google highlighted. Dashed lines mark the three access expansions. The session rule changes ChatGPTās estimates little: ChatGPT rises from approximately 2.5% to a peak near 6.5%, compared with Googleās pooled per-query benchmark of 31.1% in Table 1. These measures describe observable traffic, not whether a task was completed or whether the intermediary created value for the user. A session without a clean referral may end with a satisfactory answer, an abandoned task, or a destination visit whose referrer was stripped. A session with a referral may contain several messages but enters the ratio only once. Absorptiveness therefore identifies where observable attention ends while remaining agnostic about the userās welfare and the reason no click occurred. Public benchmarks support the scale of the result but also show why definitions matter. 2 Industry reports using per-query or bounded-session denominators generally place ChatGPT outbound routing between approximately 3% and 7%, which contains our pooled estimate and overlaps most of the monthly series. Estimates based on clicks per continuous visit can be 2 Online Appendix Section OA.2.3 reconciles our estimates with public benchmarks. The appendix records each benchmarkās numerator, denominator, attribution window, population frame, and traffic filters. 9 Table 1. ChatGPT retains far more information seeking than Google. MetricGoogleChatGPT User-side eventSearch query Conversation session Total events61,453,859409,133 Total clean referrals31,345,265166,915 Referral ratio31.1%5.2% Active households291,74756,578 Zero-referral household share9.6%74.4% Notes: The sample covers Comscore U.S. desktop activity from October 2024 through July 2025. A user-side event is one Google query or one ChatGPT conversation session. The total-clean-referrals row counts qualifying third-party destination visits; the referral ratio is the share of user-side events producing at least one such visit. The zero-referral share is the percentage of active households that never produce a clean referral from the named intermediary during the sample window. much larger because one visit can contain several information-seeking occasions and several clicks. Server-side estimates, by contrast, can be lower when browsers strip AI referrers, while our clean-referral filter excludes authentication, infrastructure, and automated requests that broader click counts may retain. Comparing the same household across intermediaries. The raw gap can combine persistent differences between the households that use each intermediary, common changes across weeks, task selection across intermediaries, and differences created by the interfaces themselves. We make progress on the first two factors by restricting the analysis to household- weeks active on both ChatGPT and Google and estimating Referral ratio itq = β 1q = ChatGPT + α i + γ t + ε itq ,(1) where the dependent variable is the share of intermediary-q information-seeking occasions for household i in week t that produce a clean referral. The dual-active restriction pairs ob- servations from the same household and week. Household fixed effects remove time-invariant 10 differences across panelists, while week fixed effects absorb common calendar shocks. Stan- dard errors are clustered by household. In this paired sample, the raw referral ratios are 8.1% for ChatGPT and 39.6% for Google, a 31.5 percentage-point (p) difference. These means differ from the aggregate rates because the regression weights household-week-intermediary cells rather than individual information-seeking events. Adding household and week fixed effects yields Ė Ī² = ā0.290 (SE = 0.002; Online Appendix Section OA.2.2). The modest change from the raw gap indicates that persistent household composition and common weekly conditions explain little of the difference. It does not, however, hold the underlying task fixed: even in the same week, a household may bring different information needs to ChatGPT and Google. This comparison does not separate differences in the tasks brought to each intermediary from differences created by the intermediaries. Section 3.2 examines those information needs using surrounding browsing. 3.2 AI search is selective in intent and destination Section 3.1 shows that most information-seeking occasions do not leave ChatGPT. We now examine the traffic that remains. We first use surrounding browsing to characterize the information needs households bring to each intermediary and the contexts in which ChatGPT routes. We then examine which websites receive those residual clicks and whether traffic concentrates on the same destinations across households. User intent and conditional routing. We do not observe prompt text, so we use the householdās other foreground browsing within fifteen minutes of each ChatGPT session or Google query as a behavioral proxy for intent, which we call the surrounding context. Using this taxonomy, we exclude self-referential traffic and search-results pages, then assign each anchor to the dominant content category in its surrounding context. Panel (a) of Figure 4 compares the context mix surrounding ChatGPT and Google use. Panel (b) asks, conditional on a ChatGPT session occurring in a given context, whether that session produces a referral. 11 Online Appendix Appendix OA.4 tests alternative windows, contamination rules, session gaps, and dominance thresholds. Households use ChatGPT for a distinct mix of tasks. The largest difference is that ChatGPT sessions are far more likely to occur soloāwith no other eligible browsing in the surrounding windowāthan Google queries (+23.1 p), so the chat itself is often the householdās only observed task environment. ChatGPT use is also relatively concentrated in reference/knowledge (+5.2 p) and tools/SaaS (+1.6 p), and less concentrated in social media (ā8.7 p), adult (ā7.7 p), entertainment/gaming (ā5.5 p), and e-commerce mar- ketplace (ā4.7 p) contexts. Conditional routing follows a different ordering. ChatGPTās referral ratio is highest in developer/technical contexts (13.4%), academic research (12.1%), and e-commerce brands (9.4%), and lowest in adult, social, and reference/knowledge con- texts. The tasks that draw households to ChatGPT are therefore not necessarily the tasks that send them back to the web: reference needs are prevalent but often retained, whereas technical and e-commerce needs are more likely to produce a click. The surrounding-context proxy above also lets us probe task selection more directly. Adding a surrounding-context fixed effect to the same household-week comparison on the context-classifiable sample barely moves the ChatGPT coefficient, from ā0.302 to ā0.300 (SE = 0.002; Online Appendix Table OA.2.4). This check does not hold prompts or exact tasks fixed, but it shows that the broad browsing context around the occasion does not explain the routing gap. Destination composition. Outgoing referrals still matter because they are the observable visits that websites can attribute, monetize, and convert into audience relationships. Destina- tion composition applies the category-share difference above to clean outgoing referral clicks. It therefore describes the traffic that exits each intermediary, not all information-seeking occasions that begin there. Relative to Google referrals, ChatGPTās residual referrals con- tain more reference/knowledge websites by 13.1 p, tools/SaaS by 8.7 p, academic research by 5.4 p, and developer/technical websites by 4.7 p (Figure 5). They contain less social 12 Figure 4. ChatGPT attracts reference and tool tasks but routes technical and e-commerce tasks. G 31.1 | CG 22.4 G 9.0 | CG 1.3 G 7.3 | CG 1.8 G 9.3 | CG 4.6 G 3.6 | CG 1.5 G 2.7 | CG 1.0 G 0.5 | CG 0.3 G 0.8 | CG 1.2 G 0.1 | CG 0.5 G 10.5 | CG 12.1 G 10.5 | CG 15.7 G 14.6 | CG 37.7 social media adult entertainment gaming ecommerce marketplace news journalism ecommerce brand government institutional developer technical academic research tools SaaS reference knowledge solo -10 p+0 p +10 p +20 p +30 p +40 p ChatGPT - Google share (p) (a) ChatGPT minus Google surrounding-context share. 13.4% 12.1% 9.4% 8.5% 8.1% 7.8% 6.4% 6.2% 5.5% 5.0% 4.3% 3.8% adult solo social media reference knowledge entertainment gaming tools SaaS news journalism ecommerce marketplace government institutional ecommerce brand academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral (b) Referral ratio by context. Notes: Surrounding context is each householdās other eligible foreground browsing within ±15 minutes of a ChatGPT session or Google query. We remove platform, competing search/LLM, and search-results pages; label anchors by the dominant high-confidence content category when it reaches 50%; and place anchors with no classifiable third-party browsing in the solo bucket. Mixed anchors are dropped and remaining shares are renormalized. Panel (a) plots ChatGPT minus Google category shares; panel (b) plots ChatGPT referral ratios by context. Error bars are 95% confidence intervals. media by 15.8 p and fewer ad-supported websites by 27.6 p. The monetization tilt favors non-profit/public, freemium SaaS, and subscription websites over ad-supported and transac- tional destinations. This is especially consequential for visit-funded sites: AI searchās smaller referral pool also bypasses the ad-supported destinations most dependent on routed atten- tion. AI search therefore does not distribute its smaller referral pool proportionally across the websites that traditional search served; this ordering is robust to the classifier-confidence and taxonomy choices (Online Appendix Section OA.3.3). Concentration across the market and within households. Concentration has differ- ent meanings at the aggregate and household levels. Aggregate concentration pools referrals across all households and asks whether total traffic converges on the same destinations. By this measure, ChatGPTās referral pool is less concentrated than Googleās: the floor-corrected normalized Herfindahl indexāwhich equals one when all referrals converge on a single desti- nation and zero when they spread as evenly as the destination count allows (Online Appendix 13 Figure 5. ChatGPTās residual referrals favor tools and knowledge over social and ad-supported sites. G 12.7 | CG 25.8 G 6.0 | CG 14.7 G 0.3 | CG 5.7 G 0.7 | CG 5.4 G 4.3 | CG 7.4 G 0.5 | CG 3.3 G 3.6 | CG 5.3 G 15.1 | CG 8.4 G 12.2 | CG 4.3 G 9.6 | CG 0.1 G 35.2 | CG 19.4 social media adult entertainment gaming ecommerce marketplace ecommerce brand government institutional news journalism developer technical academic research tools SaaS reference knowledge -10+0+10+20+30 ChatGPT - Google share (p) (a) Content type. G 6.8 | CG 24.8 G 8.0 | CG 17.6 G 2.7 | CG 4.8 G 2.0 | CG 2.8 G 20.3 | CG 17.2 G 60.3 | CG 32.7 ad supported transactional mixed other subscription paywall freemium SaaS nonprofit public -20+0+20+40 ChatGPT - Google share (p) (b) Monetization. Notes: The sample is the high-confidence (ā„ 0.90) subset of the 4,266 matched-support destination domains classified for ChatGPT and Google (3,245 domains). Bars report each categoryās share of ChatGPT referrals minus its share of Google referrals, in percentage points. Positive values indicate categories more common in ChatGPTās residual referral traffic than in Googleās. Panel (a) groups desti- nations by content type; panel (b) groups them by monetization model. Error bars are 95% confidence intervals. Section OA.5.1)āis between 1.87 and 3.47 times higher for Google across five support cut- offs (Table 2). Google referrals converge on large destinations such as YouTube, Reddit, and Wikipedia. ChatGPT instead sends relatively more traffic to lower-volume academic, reference, and tool-oriented websites (Figure 6). Household concentration asks a different question: within a household-week, how evenly does each intermediary distribute that householdās referrals across destinations? Among household-weeks with at least one or two referrals from each intermediary, the ChatGPTā Google difference is statistically indistinguishable from zero (Table 3). At the five- and ten-referral thresholds, the difference becomes positive, indicating greater concentration among heavier ChatGPT users. The two levels reconcile through cross-household hetero- geneity. Google appears concentrated in aggregate because many households converge on the same large websites. ChatGPT appears dispersed in aggregate because different house- holds reach different specialty websites, not because each household distributes its attention 14 Figure 6. ChatGPT-favored destinations are smaller and more specialized. python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org python.org mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com mdpi.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com oup.com frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org frontiersin.org wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com wired.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com biomedcentral.com invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io invideo.io springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com springer.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com visualstudio.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com wiley.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com pinterest.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com steampowered.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com ebay.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com homedepot.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com tiktok.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com twitter.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com amazon.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com x.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com reddit.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com fandom.com ChatGPT over-represents Google over-represents -0.5 +0.0 +0.5 +1.0 +1.5 1,00010,000100,0001,000,000 Total referrals to domain (ChatGPT + Google, log scale) log 10 (ChatGPT share / Google share) Notes: Each point represents one destination domain. The horizontal axis reports total ChatGPT and Google referrals to the domain on a logarithmic scale. The vertical axis reports the base-ten logarithm of the domainās ChatGPT referral share divided by its Google referral share. Values above zero favor ChatGPT; values below zero favor Google. ChatGPT-favored domains tend to be lower- volume academic, reference, and tool websites, whereas Google-favored domains include large social, marketplace, and encyclopedia websites. more evenly. Aggregate diversification is therefore a between-household pattern rather than a within-household one. Online Appendix Appendix OA.5 reports support-cutoff, long-tail, and robots.txt analysis. Together, the surrounding-context and destination results explain how ChatGPTās refer- ral stream differs from Googleās. Households arrive at the AI intermediary with a different mix of information needs; click-out occurs in only some contexts; and the outgoing traffic that remains reaches a different set of websites. Different households then reach different specialty destinations, making the aggregate referral pool appear broad even when household- level attention is not. The surrounding context provides a behavioral, context-based reading of these patterns, not a prompt classification or an identified mechanism. The evidence remains consistent with both intent selection and intermediary-driven retention. 15 Table 2. Google referrals converge more strongly on the same destinations. Support cutoffChatGPT Google Google/ChatGPT Top 1000.0360 0.06741.87 Top 1,0000.0148 0.03962.67 Top 10,0000.0078 0.02703.47 Matched support0.0131 0.03522.68 Full domain set0.0055 0.01913.47 Notes: Entries are floor-corrected normalized Herfindahl indices calculated from each intermediaryās aggregate referral shares; higher values indicate that more traffic converges on fewer destinations. Rows vary the destination-support cutoff to ensure that the comparison is not driven by the long tail. The final column divides Googleās index by ChatGPTās. Googleās referral pool is more concentrated under every cutoff and is 3.47 times as concentrated on the full domain set. The five cutoffs shown are a representative excerpt; Online Appendix Table OA.5.1 reports the full cutoff grid across the percentile, absolute-rank, cumulative-share, and minimum-referral families. 3.3 Wider AI search access reduces traditional search The preceding results show that ChatGPT retains many information-seeking occasions and routes a different residual stream. The next question is whether wider access changes tra- ditional search use itself. This margin matters because search queries are the source of search-advertising inventory and the starting point for many website referrals. If AI search merely adds a new channel for separate tasks, traditional search need not fall. If households substitute toward AI search, fewer information needs enter the search engines that previously routed traffic onward. We use two complementary access-based comparisons. The preferred design pools the three access expansions and compares each treated cohort with its own reweighted N itt control: paid subscribers gaining access on October 31, 2024 (P = 84), free logged-in users on December 16 (L = 2,440), and anonymous browsers on February 5, 2025 (A = 1,358). A narrower December 16 comparison contrasts logged-in users with anonymous users who gain access later; it has less power but compares households already using ChatGPT. For 16 Table 3. ChatGPT concentration emerges only among heavier referring households. ā„ 1 ref ā„ 2 refs ā„ 5 refs ā„ 10 refs ChatGPT indicatorā0.0010.0000.035 ā 0.056 ā (0.002) (0.002) (0.003) (0.006) Mean normalized HHI, Google0.0720.0710.0750.077 Mean normalized HHI, ChatGPT 0.0690.0700.1070.130 Household FEYesYesYesYes Week FEYesYesYesYes Observations48,641 35,644 12,6843,721 Notes: The dependent variable is the household-weekās floor-corrected normalized Herfindahl index, a [0, 1] concentration index equal to one when all referrals fall on a single destination and zero when they spread as evenly as the destination count allows. Columns retain household-weeks meeting the indicated minimum referral count for each intermediary. The index is defined only for household-weeks with at least two distinct destinations for the measured intermediary, so single-destination weeks are dropped. The ChatGPT-indicator coefficient compares ChatGPT with Google within the same household and week; positive values indicate greater concentration through ChatGPT. The mean-normalized-HHI rows report each intermediaryās average household-week index as a level benchmark for the indicator. Standard errors clustered by household appear in parentheses. The difference is indistinguishable from zero at the one- and two-referral thresholds and positive among heavier referring households. The four thresholds shown are an excerpt; Online Appendix Table OA.5.2 reports the full sweep from ā„ 1 to ā„ 20 referrals. ā p < 0.10, ā p < 0.05, ā p < 0.01. household i, stack s, and calendar week t, let W ist = tā t ā s denote relative week, with t indexed in calendar weeks. We estimate Y ist = X wĢø=ā1 β w 1W ist = w Treated is +α si + γ st + ε ist ,(2) with stack-by-household and stack-by-week fixed effects and standard errors clustered by household. The outcome is weekly Google, Bing, and Yahoo search-query loads. The pooled treated sample contains 3,882 households before demographic-completeness restric- tions. Each cohort enters only its corresponding expansion stack against its own reweighted 17 N itt control, defined from households clean of ChatGPT and Claude as of that expansionās date. Online Appendix Section OA.6.1 gives the endpoint rules; Online Appendix Sec- tion OA.6.3 reports the full causal design and inference. Wider access reduces weekly traditional search queries by 3.14 per treated household, or 9.4% of the pre-expansion mean of 33.51 (Figure 7). The event-study path is flat before access and turns negative afterward, with larger declines after sustained exposure; joint pre-trend tests fail to reject parallel trends for both the stacked design and the December 16 within- adopter comparison (Online Appendix Section OA.6.4). The cleanest single comparisonā December 16 logged-in users against not-yet-treated anonymous usersāshows the same neg- ative displacement, strengthening with exposure (Online Appendix Figure OA.6.2). Figure 7. Wider ChatGPT Search access reduces traditional search queries. -10 -5 0 5 -11 -7 -315913 17 21 2529 Weeks relative to ChatGPT-Search rollout ATT per user-week Notes: The outcome is weekly Google, Bing, and Yahoo query loads per household. The stacked difference-in-differences design pools cohorts gaining access on October 31, December 16, and February 5 and compares each with the ACS-reweighted N itt control. Event week ā1 is omitted. Points are event- time coefficients; ribbons are 95% confidence intervals with standard errors clustered by household. The pooled post-expansion effect is ā3.14 queries, or ā9.4% of the pre-expansion mean of 33.51. Matched-window averages show how this displacement grows with exposure (Table 4). The preferred three-event specification reaches 17.0% after twenty weeks. The October cohort is small, and adding it barely moves the two-event estimate. A cleaner December-only comparison between two groups already using ChatGPT yields a smaller 8.2% decline after 18 twenty weeks. We read the difference as evidence that comparisons with nonusers may retain some selection on AI adoption; the direction of the effect survives both designs, alternative controls, and an independent individual-adoption design (Online Appendix Section OA.6.7). Table 4. Traditional search displacement grows with exposure. Design / controlw ā„ 0 w ā„ 5 w ā„ 10 w ā„ 15 w ā„ 20 Dec. 16, L vs. A, reweightedā4.9% ā3.8% ā5.5% ā5.4% ā8.2% Three-event N itt , reweighted (preferred) ā9.4% ā7.9% ā11.3% ā12.7% ā17.0% Three-event N itt , unweightedā10.9% ā9.5% ā12.6% ā14.2% ā18.4% Notes: Entries are matched-window average effects on weekly Google, Bing, and Yahoo query loads, expressed as percentages of the pre-expansion mean. A column labeled w ā„ h averages event-time effects from week h through the end of the common support. The preferred row pools the three access expansions, uses the N itt control, and reweights observations to American Community Survey age-by- income cells. The December 16 comparison instead contrasts logged-in cohort L with anonymous cohort A. The preferred estimate grows from ā9.4% after access to ā17.0% after twenty weeks. Online Appendix Table OA.6.5 reports full inference and estimator checks. The loss in referral traffic from traditional search engines falls disproportionately on in- formational websites. Search-engine referrals decline by 32.8% for academic research, 26.5% for reference/knowledge, 15.1% for developer/technical, and 13.4% for news/journalism des- tinations (Figure 8). Transactional and recreational categories show smaller or statistically indistinguishable referral changes. Total visits also decline for academic research, refer- ence/knowledge, and news/journalism, indicating that lost search-engine referrals are not fully replaced by other observed channels. This category-level pattern is descriptive hetero- geneity within the preferred design, not a separately identified mechanism. 19 Figure 8. Traditional search losses are largest for informational destinations. -33% -15% -7% -1% -1% -11% -13% -27% -8% -9% entertainment gaming ecommerce marketplace ecommerce brand social media tools SaaS government institutional news journalism developer technical reference knowledge academic research -60%-40%-20%+0% ATT (% of pre-mean per user-week) (a) Search-engine referral visits. -29% -12% -4% -6% -1% -7% -19% -21% -5% -4% entertainment gaming tools SaaS ecommerce brand social media ecommerce marketplace government institutional developer technical news journalism reference knowledge academic research -40%-20%+0% ATT (% of pre-mean per user-week) (b) Total destination visits. Notes: Points report destination-category effects from the preferred weighted stacked design, scaled by each categoryās pre-expansion mean; bars are 95% confidence intervals. Panel (a) uses downstream visits attributable to Google, Bing, or Yahoo search referrals. Panel (b) uses total destination visits from all sources. Search-engine referral losses are largest for academic research, reference/knowledge, news/journalism, and developer/technical destinations, while marketplace and entertainment referral effects are small and statistically indistinguishable from zero. This heterogeneity connects the displacement effect to the intent results. The website cat- egories that lose search-engine referral traffic are also the categories where households bring more informational tasks to ChatGPT and where many sessions end without a click. The pattern links retention inside ChatGPT to downstream losses in routed traffic: households use ChatGPT relatively more for informational tasks; many of those tasks remain inside ChatGPT; and wider access reduces the traditional search paths that had sent users from search engines to websites. 4 Discussion AI search shifts the web from routing information needs to resolving them inside the inter- mediary. ChatGPT sends outbound referrals in only 5.2% of conversation sessions, versus 31.1% of Google queries; three-quarters of ChatGPT-active households never click out dur- 20 ing the panel window. The remaining traffic is selective: reference and tool-oriented use is common, outbound clicks are most likely in technical and e-commerce contexts, and refer- rals tilt toward reference, academic, developer, and tool/SaaS destinations. Wider access to ChatGPT Search reduces traditional search queries by 9.4% on average and 17.0% after twenty weeks, with the largest referral losses in informational categories. These patterns identify the routing margin. Routed traffic has been the webās practical attribution system. A query becomes a visit, and that visit creates exposure, monetization opportunities, and an audience relationship that websites, advertisers, and regulators can observe. AI search weakens this record when it uses web information but satisfies the user without a referral. The shift matters for publishers, advertisers, and policy because each group relies on routed visits to value, allocate, or govern attention. For publishers and content producers, the results show where attribution and licensing debates are most exposed. Informational content is valuable in the tasks households bring to ChatGPT, yet many of those tasks end without referrals and show the largest search-referral losses. Our estimates give these negotiations traffic benchmarks: how often sessions route, which destinations receive residual referrals, and which content categories lose traditional search traffic. For advertisers and search marketers, the results show that AI search changes both query inventory and the locations where attention remains reachable. The access expansions reduce traditional search use most clearly for research, reference, and how-to tasks, while transactional and recreational categories change less. This evidence maps the contraction in available queries and routed attention, the margin that precedes any equilibrium adjustment in bids, prices, or channel mix. For measurement and policy, the results show why destination-side traffic records are incomplete when an intermediary can answer without routing. User-side measures of ab- sorbed sessions, residual referrals, and displaced search queries help discipline attribution, compensation, and traffic-sharing proposals. They clarify the economic object at stake: the 21 redistribution of attention and commercial opportunity when an intermediary uses web in- formation without routinely sending users onward. The main limitations point to the next research frontier. We observe U.S. desktop be- havior, not consumer surplus, publisher revenue, or long-run content investment. Future work should connect user-side behavior to platform logs, revenue records, and supply-side responses. The central dynamic question is whether content producers continue to create high-quality information at the same scale when AI search uses web content while returning fewer visits. Answering it requires measuring user value and producer incentives jointly as AI search becomes a primary gateway to online information. 22 Online Appendix This online appendix documents the complete empirical pipeline behind the main paper. Ap- pendix OA.1 covers panel construction, foreground-traffic cleaning, endpoint coverage, conversation sessions, the three access-expansion cohorts, ACS reweighting, and the intensiveāextensive adoption decomposition. Appendix OA.2 defines user-side referral measurement and reconciles our referral ratios with external benchmarks. Appendix OA.3 documents the domain classification, destina- tion taxonomy, confidence thresholds, manual validation, and the verbatim classification prompt. Appendix OA.4 tests the surrounding-browsing intent proxy across contamination rules, session gaps, window lengths, and dominance thresholds. Appendix OA.5 reports destination-composition and concentration robustness, including the long tail and robots.txt blocking. Appendix OA.6 presents the complete three-shock displacement design, balance and weighting diagnostics, alterna- tive controls and thresholds, estimator checks, matched-window dynamics, category heterogeneity, and external validation. 23 OA.1 Data construction and validation This section documents how we convert raw Comscore URL records into the household-level panel and user-side demand events used in the paper. It proceeds from panel inclusion and foreground- traffic filters to endpoint coverage, conversation construction, and adoption margins. The treated and control cohorts and the population reweighting are inputs to the causal design and are docu- mented with it in Appendix OA.6. OA.1.1 Panel definition and threshold sensitivity The balanced sub-panel is the set of households, indexed by machine ID, with at least Ļ fg =4 foreground records (text/html content type, http_rc in the valid set, after applying the outflow- exclusion list) in every one of the ten months. The threshold Ļ fg =4 trades off coverage against engagement: relaxing to Ļ fg =1 admits +4,822 low-activity panelists; tightening to Ļ fg =50 removes 15,009 panelists with month-to-month coefficient of variation approximately 0.92. Table OA.1.1 reports the marginal-user counts and cross-month coefficient of variation at each threshold. Table OA.1.1. Panel-threshold sensitivity. Sample: Comscore US Desktop active-household panel, all 50,208 households with at least one foreground record in every month of the ten-month panel window October 2024 ā July 2025. n_panelists is the count of households with at least Ļ fg valid foreground records in every month of the ten-month window. āavg fg/moā is the mean per-month foreground activity of the marginal users introduced or dropped relative to the Preferred Ļ fg =4 panel. CV is the cross-month coefficient of variation in foreground activity for the marginal set. Ļ fg n_panelists+marginal āmarginal avg fg/moCV 150,208+4,8220577.1 1.084 247,967+2,5810562.8 1.061 346,477+1,0910606.3 1.042 4 (Preferred)45,38600ā 544,4030 ā983512.3 1.036 1041,1050 ā4,281571.9 1.010 2037,0570 ā8,329601.4 0.973 5030,3770 ā15,009651.9 0.922 The marginal users at Ļ fg =1 have CV approximately 1.08 across monthsāthey are bursty rather than habitualāwhereas those at Ļ fg =50 have CV approximately 0.92, indicating very stable usage. Choosing Ļ fg =4 balances inclusion breadth against the cross-month-stability requirement that supports within-user identification. OA.1.2 Foreground filter and HTTP response codes The raw Comscore stream is filtered to foreground traffic before any analysis. A row counts as foreground if (a) its mimetype equals text/html and (b) its HTTP response code falls in the success-or-redirect set VALID_HTTP_RC = 200, 201, 202, 203, 204, 206, 301, 302, 303, 304, 307, 308. The mimetype filter discards 71 other distinct mimetype values observed in the panel window, grouped broadly into scripts (application/javascript, text/x-python), images (image/png, image/webp, image/svg+xml), fonts (font/woff2, font/ttf), audio/video (audio/mpeg, 24 video/mp4), JSON / API responses (application/json, application/json+protobuf, application/problem+json), server-sent-events (text/event-stream), document binaries (application/pdf, application/zip, DOCX/XLSX/PPTX OOXML strings), form-submission types (multipart/form-data, application/x-w-form-urlencoded), header-leak artefacts (cross-origin, sameorigin, report-uri /_/bardchatui/cspreport), and a long tail of mal- formed strings (1, gfet4t7, tk/lws, and base64-mangled blobs). The HTTP-RC filter excludes nonstandard codes (e.g., 299, 999) and all 4x/5x errors. Both filters are applied in tandem: a record is retained only if its content type is text/html and its HTTP response code is in the valid set. OA.1.3 ChatGPT endpoints Comscore records the path of each network request a panelistās browser issuesāhost, directory, and pageābut not request bodies or authorization headers. We therefore identify ChatGPT activity by matching these paths to the platformās internal endpoints, whose semantics are established through public reverse-engineering of the ChatGPT web app. 3 Table OA.1.2 defines each endpoint the paper usesāthe user action that issues it and its role in our measurement. Two prefixes encode login state: backend-api/* requests carry a bearer token and so fire only for logged-in users, whereas backend-anon/* requests use a request-level sentinel token with a CSRF cookie and identify pre- authenticated or anonymous use. 4 3 Endpoint behavior is documented by the community-maintained everything-chatgpt catalogue of Chat- GPTās internal API (https://github.com/terminalcommandnewsletter/everything- chatgpt, ac- cessed 2026-06-24) and by independent reverse-engineering write-ups; see, e.g., How I Successfully Reverse- Engineered ChatGPT to Create an Unofficial API Wrapper, HackerNoon (https://hackernoon.com/how- i-successfully-reverse-engineered-chatgpt-to-create-an-unofficial-api-wrapper, accessed 2026-06-24). 4 The anonymous flow requests CSRF cookies, then posts to backend-anon/sentinel/chat-requirements for a sentinel token, then posts to backend-anon/conversation; see the archived anonymous-chatgpt client (https://github.com/Mr-Destructive/anonymous-chatgpt, accessed 2026-06-24). 25 Table OA.1.2. ChatGPT endpoints used in the paper: definition and role. Comscore observes the URL path of each request, not its body or headers; endpoint semantics are established from public reverse-engineering of the ChatGPT web app. backend-api/* implies a logged-in (bearer-token) request and backend-anon/* a pre-authenticated or anonymous (sentinel-token) request. EndpointDefinition and role in the paper chatgpt.com foreground page loadGET. A user-visible page load (homepage or /c/uuid); the broadest but shallowest signal. Visit denominator. backend-api/conversation/uuid GET. The browser opening or reloading one conversation; each distinct UUID is one conversation. Conversation denominator. backend-anon/conversation/uuid GET. The same for anonymous sessions; negligible volume. backend-api/f/conversation 5 POST. One logged-in message send (a newer endpoint variant that ramps MarāJul 2025). Message-level referral attribution. backend-anon/f/conversationPOST. One anonymous message send. paragen_submission 6 POST. A front-end submission to a paragen-routed surface (web search, Canvas, image generation, or file upload). Search-enabled proxy after filtering the contaminating non-search surfaces. These endpoints partition ChatGPT activity into three observability tiersāforeground page vis- its (broadest coverage, shallowest semantics), background backend-api endpoints, and background backend-anon endpointsāwith mostly disjoint populations. Table OA.1.3 reports the panel-user union over the ten-month window for each. Table OA.1.3. Panel-user endpoint coverage (10-month union, panel N =45,386). ā% panelā is the share of all panelists; ā% visitā uses the foreground-visit user set (N =9,481) as denominator. Background endpoints are identified by URL-pattern matching. EndpointPanel users (10-mo) % panel% visit chatgpt.com foreground visits9,48120.9100.0 backend-api/conversation/uuid5,26511.655.5 backend-anon/conversation/uuid3480.83.7 backend-api/f/conversation (Jul 2025)2,3615.2 65.7 (Jul) backend-anon/f/conversation (Jul 2025)9362.1 26.1 (Jul) Two features of these endpoints bear on the downstream identification. First, paragen_submission is a client-side A/B-test harness that tags outputs across several surfacesāimage generation, file handling, Canvas, autocompletion, and web searchāso a raw count overstates true search use. The first surfaces carry clean URL markers but web search does not, so we keep a paragen_submission hit as search-paragen only when none of these contaminating markers fires within ±60 s on 5 Unlike the other endpoints here, the f/conversation variant is not separately documented in public reverse-engineering sources. Direct inspection of the ChatGPT web appās network activity shows that sending a single message issues exactly one f/conversation POST. As an undocumented internal endpoint it may change without notice. 6 The paragen submission surface was identified through front-end reverse-engineering by Tibor Blaho (https://x.com/btibor91, accessed 2026-06-24). 26 the same household (the non-search paragen markers: images/bootstrap, image-generation bootstrap; backend-api/files, file upload and attachments; /textdocs, Canvas text documents; and generate_autocompletions, autocompletions). The residual remains an upper bound on search use. Second, message-level attribution via f/conversation reaches analytic scaleā6,056 message-send users on the full active-household sample, against 184K July sendsāonly in July 2025; pre-July message coverage is sparse, so the message-level referral ratio (Section OA.2.1) is reported on July 2025 alone. (The panel-level coverage table above counts send-or-prepare events on the 45,386-household balanced panel, a narrower and prepare-inclusive population, so its f/conversation cells are not directly comparable to this all-user send count.) OA.1.4 Definition of a conversation session A ChatGPT conversation session is defined by grouping all backend-api/conversation/uuid and backend-anon/conversation/uuid loadsāso that both logged-in and anonymous users are includedāon the same machine_id that share a UUID, with a same-UUID gap exceeding 3,600 seconds opening a new session. We call the surviving unit a conversation session (not a raw āloadā) to distinguish it from individual page loads. The number of distinct conversations (unique UUIDs, 263,008) is invariant to this threshold by construction; only the merging of adjacent same-UUID loads into sessions depends on it, and that count varies only modestly with the gap (Table OA.1.4). Table OA.1.4. Conversation-session count by same-UUID gap threshold. The number of distinct conversations (unique UUIDs) is invariant to the threshold; only the merging of adjacent same- UUID loads into sessions depends on it, and the resulting session count varies only modestly around the 1 h Preferred. Same-UUID gapSessionsvs. 1 h 0.25 h (900 s)449,391+9.8% 0.5 h (1,800 s)429,074+4.9% 1 h (3,600 s, Preferred)409,133ā 2 h (7,200 s)391,282 ā4.4% Distinct conversations (any gap) 263,008 invariant Per-UUID gap distribution. The per-UUID gap distributionāthe time between two con- secutive loads of the same conversation UUID by the same userāis sharply bimodal. Consecu- tive same-UUID pairs, where no other UUID intervenes, are browser auto-refreshes and tab-focus reloads, with a median gap of 13 s; non-consecutive pairs, where loads of other UUIDs intervene, are genuine returns to a prior conversation, with a median gap of 4,946 s (approximately 82 min). The 1 h threshold sits in the valley between the two modes. OA.1.5 Intensive vs. extensive decomposition ChatGPT growth across the panel window can be decomposed into an extensive marginārising reach, a larger share of households activeāan intensive marginārising amount, more activity per active householdāand a cross term (new households that arrive already highly active, the empirical signature of a feature rollout among an expanding eligibility frontier). 27 Formally, express volume per household. Let a t = N t /M t be the share of the M t active house- holds reached in month tāthe reach, or extensive margināand Ģq t the mean count per reached householdāthe amount, or intensive margināwhere a unit is a conversation session or a search- enabled conversation. Per-household volume is v t = a t Ģq t = V t /M t , and between a baseline month b and an endpoint month e its change decomposes exactly as āv = (a e ā a b ) Ģq b |z extensive (reach) + a b ( Ģq e ā Ģq b ) | z intensive (amount) + (a e ā a b )( Ģq e ā Ģq b ) | z cross . The extensive term holds the amount fixed at its baseline level and counts only the rise in reach; the intensive term holds reach fixed and counts only the rise in amount; the cross term is the interaction. Because the active-household base M t grows month to month, normalizing by it nets out volume growth that merely reflects an expanding panel, so these per-household shares isolate genuine adoption and differ from a raw-count decomposition. Table OA.1.5 reports them, dividing each term by āv, with per-household growth multiple v e /v b . Table OA.1.5. KitagawaāOaxaca decomposition of per-household volume change between baseline and endpoint months. Sample: Comscore US Desktop active-household panel, October 2024 ā July 2025. The change splits into an extensive margin (rising reachāa larger active-household share), an intensive margin (rising amountāmore use per reached household), and a cross term (new households that arrive already using the feature heavily), following the identity above. The two columns are nested units: cleaned conversation sessions (Section OA.1.4) and search-enabled conversations, defined as conversation sessions with at least one paragen_submission event and no contaminating non-search paragen surface within±60 s (Section OA.1.3). Volume is normalized by the month-specific active-household base M t . Search-enabled growth is reported at the April 2025 peak endpoint; from May the f/conversation send endpoint begins appearing and may disrupt the paragen signal, so later endpoints are unreliable (Section OA.1.3). Conv. sessions Search-enabled conv. Baseline / endpoint2024-10 / 2025-072024-11 / 2025-04 Growth multiple (per household)3.0 times45.0 times Extensive share65.5%66.2% Intensive share14.9%1.1% Cross share19.6%32.7% Per-household conversation-session growth between October 2024 and July 2025 splits 65.5% extensive (reach), 14.9% intensive (amount), and 19.6% crossānew households drive most of the rise, but reached households also deepen their use. Search-enabled conversation growth is more rollout-shaped: 66.2% extensive, 1.1% intensive, and 32.7% cross at the April 2025 peak endpoint. The cross term is the second-largest componentānew users running the search feature heavily upon arrivalāthe empirical signature of a feature rollout among a sticky audience. The decline after April has two non-exclusive explanations. On the measurement side, the f/conversation send endpointānegligible through Aprilābegins appearing in May and ramps through July, reaching analytic scale only in July (Section OA.1.3); the front-end changes accompa- nying that rollout plausibly disrupt the paragen_submission surface from which the search-enabled signal is inferred, so part of the measured decline from May onward may be an endpoint-migration artefact rather than a genuine drop in use. Behaviorally, the later JuneāJuly fall is also consis- tent with seasonal disengagement. We therefore anchor the decomposition at the April peakāthe 28 last month before the f/conversation rollout. Figure OA.1.1 plots the ChatGPT conversation- volume and search-enabled trajectories (panels (a) and (b))āeach rising into the spring, with search-enabled use peaking in April and then softening. Figure OA.1.1. ChatGPT conversation volume and search-enabled use per active household, October 2024 ā July 2025. PaidLogged-in No Login 0.000 0.100 0.200 0.300 2024-11 2025-01 2025-03 2025-05 2025-07 Sessions / household (a) All conversation volume. Paid Logged-in No Login 0.0000 0.0050 0.0100 0.0150 0.0200 2024-11 2025-01 2025-03 2025-05 2025-07 AI-search sessions per household (b) Search-enabled conversations. Notes: Monthly per-active-household ChatGPT conversation volume, October 2024 through July 2025. Panel (a) shows all conversations; panel (b) shows search-enabled conversations (sessions with a paragen_submission event surviving the ±60 s contamination filter). Both rise through the spring, with search-enabled use peaking in April and softening from May (see text). Table OA.1.5 decomposes the changes into reach and amount. OA.2 Referral measurement and industry benchmarks This section defines a referral and the clean-referral filter (Section OA.2.1), shows that the within- household referral-ratio gap is robust to household and week fixed effects (Section OA.2.2), and situates the two main ratios against published industry benchmarks (Section OA.2.3). OA.2.1 Definition of a clean referral A referral is an outbound foreground visit to a third-party website whose HTTP referrer identifies ChatGPT or a search engine as the source. A clean referral additionally survives the exclusion filters below, which remove self-referential traffic, platform-internal navigation, search-engine in- termediation, and non-content endpoints. The referral ratio for an intermediary is the share of its user-side unitsāChatGPT conversation sessions or Google /search queriesāthat produce at least one clean referral; a session or query that fires several clean clicks contributes once to the numerator, so the ratio counts denominator units that produce any referral, not the total number of click events. We report this ratio at three levels: per ChatGPT conversation session (5.2%, the main user-side unit), per Google /search query (31.1%, the symmetric search-engine denominator), andāas a finer cross-checkāper ChatGPT message send (Table OA.2.2). Outflow exclusion list. A curated allow-list excludes 147 destination domains that are not user-navigation clicks. Five categories of exclusion: ⢠Auth/OAuth surfaces (microsoftonline.com, accounts.google.com)ānot a destination, just an identity-provider redirect step. 29 ⢠Platform-internal CDN/asset hosts (oaiusercontent.com, oaistatic.com, chatgptusercontent.com, characterai.io, sora.com)āstatic-asset URLs internal to the LLM platform itself. ⢠Infrastructure / CDN (browser-intake-datadoghq.com, stripecdn.com, wp.com, prodregistryv2.org, featureassets.org)ātelemetry and CDN endpoints, not content. ⢠AI-tool extensions and detection-bypass utilities (gptzero.me, quillbot.com, undetectable.ai, zerogpt.com, humbot.ai, writehuman.ai, stealthwriter.ai, originality.ai, approx- imately 30 in total)ābrowser-extension noise unrelated to the AI search substitute under study. ⢠Adware / tracker / parking clusters (approximately 70 entries; e.g., topodat.info, earthview3dmaps.com, onlinemanualspdf.co, adnxs-simple.com, webpkgcache.com, secure-check.co, linewize.net)āgibberish URL patterns, no real content destinations. Search-ecosystem and same-source exclusions. On top of the foreground filter and the OUTFLOW exclusion list, the referral filter also drops (i) the search-engine ecosystem (Google, Bing, Yahoo, and all their subdomains), including ancillary authentication and file-service endpoints such as OAuth, asset-upload, drive.google.com, and docs.google.comāthese reflect search in- termediation or file transfer into the LLM rather than genuine human information-seeking; and (i) same-source self-referrals (LLM ā LLM, search ā search). ChatGPT referral attribution. A foreground click is attributed to a ChatGPT conversa- tion if its raw load timestamp falls within the inter-load window [load_ts, next_load_ts) on the same machine. We do not require the conversation UUID to match the destination URLāmany chat.openai.com post-redirect URLs lose UUID context as users navigate, and the load-window approximation captures click-out behavior at the user level rather than the link level. The window is not a tunable threshold: it spans the full interval between consecutive conversation loads, so every clean referral between sessions is attributed. Search-query referral attribution. A Google, Bing, or Yahoo /search query is the search- engine user-side unit, identified by matching the /search path on the full URL string rather than on Comscoreās parsed path components, which split the same path inconsistently across hosts. A foreground outflow within the inter-visit window [visit_ts, next_visit_ts) counts as a search referralāthe symmetric counterpart to the conversation-load window above. Referral validation: referrer host and landing-page depth. Two URL-pattern diagnostics support reading the surviving clean referrals as genuine in-interface click-outs rather than machine-initiated requests.First, the 99.7% chatgpt.com referrer share is produced by our filters, not by the raw data (Panel A): the unfiltered ChatGPT- referred outflow is 84.7% application/json machine fetches and splits its referrer host chatgpt.com/auth.openai.com/openai.com at roughly 69/15/13 percent; the foreground text/html filter removes the JSON fetchesācitation prefetches, metadata and favicon loadsāand the clean-referral definition removes the authentication and marketing surfaces. This establishes interface-origin (the surviving click was issued from the conversation interface), not human-origin on its own. Second, landing-page depthāthe number of nonempty slash-delimited path segments below the host, so example.com/ has depth 0 and example.com/watch depth 1ātracks Google closely (Panel B): 75.6% of ChatGPT referrals reach a deep (depth ā„ 1) content page versus 30 81.4% for Google, with an identical median depth of one. A homepage-heavy distribution would signal navigational or brand traffic; the deep-page share instead resembles a search referral. The human-click reading rests on the conjunction of the two diagnostics, not on either alone. Table OA.2.1. Referral validation: ChatGPT referrer host by filter stage and landing-page depth. Panel A. Referrer host by filter stage (all-user, ten months) Filter stage chatgpt.com Other OpenAI (such as auth.openai.com) Unfiltered (any MIME)69.1%30.9% Clean referral99.7%0.3% Panel B. Landing-page depth Intermediary Homepage Deep content Median depth ChatGPT24.4%75.6%1.0 Google18.6%81.4%1.0 Notes: Panel A: ChatGPT referrer host over the full ten-month all-user outflow at two filter stagesā unfiltered (any MIME type) and the paperās clean-referral definition; āOther OpenAIā aggregates auth.openai.com, openai.com, and the remaining OpenAI subdomains. In the unfiltered data 84.7% of these rows are application/json machine fetches, removed by the foreground text/html filter. Panel B: share of clean referrals landing on a site homepage versus a deeper content page, for ChatGPT (N =5,318 classifiable referrals) and Google (N =874,266); ādeep contentā is any non-homepage path (depthā„ 1). A high deep-content share is consistent with answer- or results-driven click-out rather than navigational traffic. ChatGPT message-level referral ratio. At the finer message-send granularity, the ratio is computed on individual f/conversation sends. Message coverage reaches analytic scale only in July 2025 (Table OA.1.3); pre-July sends are too sparse to yield a stable ratio, so we report it for July 2025 alone. Table OA.2.2 shows that 1.2% of logged-in sends are followed by a referral (0.9% within five minutes, the click-like window) and 1.1% of anonymous sends are. The message-level ratio is thus even lower than the session-level ratio and an order of magnitude below the search- engine rate, consistent with absorption of the information-seeking occasion at the message level rather than routing it onward. Table OA.2.2. Message-level referral ratio, July 2025. Message sends (Jul 2025)Sends With referral Ratio Logged-in (backend-api)184,0302,231 1.2% within 5 min1,623 0.9% Anonymous (backend-anon) 31,320360 1.1% Notes: Share of ChatGPT message sends followed by an outbound referral, July 2025āthe only month in which f/conversation message coverage reaches analytic scale (cf. Table OA.1.3). A message is a POST to backend-api/f/conversation (logged-in) or backend-anon/f/conversation (anonymous), with pre-flight prepare events excluded. The within-five-minute row restricts to referrals firing in the click-like window after a send. The message-level ratio sits below the 5.2% session-level ratio and an order of magnitude below Googleās 31.1% per-query rate. 31 OA.2.2 Within-user referral-ratio regressions The body compares ChatGPT and Google within the same household and week using Equation 1, estimated on household-weeks active on both intermediaries. Table OA.2.3 reports the full specifi- cation ladder. The raw cross-intermediary gap of roughly 31 percentage points (p) is essentially unchanged by household and week fixed effects: the coefficient on the ChatGPT indicator isā0.289 under household fixed effects alone and ā0.290 under household-plus-week fixed effects. Persistent differences in who uses each intermediary and common weekly conditions therefore explain little of the routing gap. Table OA.2.3. The within-user referral-ratio gap is robust to household and week fixed effects. SpecificationChatGPT indicatorSE Observations Pooled OLSā0.315 ā 0.0012,790,136 Household FEā0.289 ā 0.0022,790,136 Household + week FEā0.290 ā 0.0022,790,136 Notes: The dependent variable is the share of a householdās intermediary-q information-seeking occasions in a week that produce a clean referral, stacked across intermediaries so each dual-active household-week contributes one ChatGPT cell and one Google cell. The ChatGPT-indicator coefficient is Ė Ī² in Equation 1. The paired raw means are 0.081 for ChatGPT and 0.396 for Google. Standard errors are clustered by household. ā p < 0.01. A remaining concern is that households may bring systematically different tasks to each inter- mediary even within the same week. We address the observable part of this concern by adding the surrounding-context category (Online Appendix Appendix OA.4) as a third fixed effect, which holds the dominant content type of the householdās nearby browsing fixed. Table OA.2.4 reports the result. Adding the context fixed effect moves the coefficient only fromā0.302 toā0.300, so dif- ferences in the broad type of task households bring to each intermediary do not account for the gap. The household-plus-week level differs slightly from theā0.290 above for two reasons tied to the con- text dimension itself. First, the observation grain differs: each household-week-intermediary cell of Table OA.2.3 is decomposed into one cell per dominant surrounding-content type, and because cells enter the regression unweighted by activity volume, the single week-level referral share is replaced by several context-level shares whose unweighted mean need not equal itāso β targets a different population-weighted average even when the underlying events are identical. Second, the samples differ in support: Table OA.2.4 retains only activity assignable to a well-defined surrounding con- text, dropping search visits without an enclosing browsing session and all cells whose surrounding content is ambiguous (mixed, solo, or other), so the within-user variation that identifies β is drawn from a selected subset of the household-weeks in Table OA.2.3. 32 Table OA.2.4. Adding a surrounding-context fixed effect barely changes the gap. SpecificationChatGPT indicatorSE Observations Pooled OLSā0.329 ā 0.0014,618,397 Household + week FEā0.302 ā 0.0024,618,397 Household + week + context FEā0.300 ā 0.0024,618,397 Notes: The dependent variable and intermediary stacking match Table OA.2.3, but the sample is the mixed-1h surrounding-context variant (Online Appendix Appendix OA.4), which assigns each anchor the dominant nearby content category; the context fixed effect holds that category fixed. The household-plus- week coefficient (ā0.302) differs from theā0.290 of Table OA.2.3 because the context dimension changes both the estimand and the support. (i) Grain: each household-week-intermediary cell is decomposed into one cell per dominant context type and cells enter unweighted by activity volume, so the unweighted mean of the context-level shares need not equal the week-level share and β targets a different population- weighted average. (i) Support: only activity assignable to a well-defined context is retainedāsearch visits without an enclosing session and ambiguous-context cells (mixed, solo, other) are droppedāso identification draws on a selected subset of household-weeks. Within this sample the context fixed effect itself moves the coefficient only from ā0.302 to ā0.300. The paired raw means are 0.085 for ChatGPT and 0.414 for Google. Standard errors are clustered by household. ā p < 0.01. OA.2.3 Industry benchmarks Industry sources report a wide range of āCTR-likeā numbers for ChatGPT, between 3% and 7% on per-query or bounded-session denominators, with denominator units, click definitions, and popula- tion frames that are not directly comparable. Table OA.2.5 collates our Preferred values alongside the closest published benchmarks. Table OA.2.5. ChatGPT and Google referral-ratio benchmarks across industry sources and our Preferred. Sample (our rows): Comscore US Desktop active-household panel, October 2024 ā July 2025. DV: share of denominator units producing ā„ 1 clean outbound click. āPer queryā = denominator is one search query; āper sessionā = one ChatGPT engagement window; āper visitā = one continuous on-platform visit. Industry rows reproduce the publication-period scope of the cited source; peer-reviewed sources are cited as bibliography keys and grey-literature reports in footnotes with access dates. SourceMetricValue Google Ours (per query, foreground click) Preferred, all-user31.1% SparkToro / Datos 7 (Jul 2024) Clicks to open web / queries36% Pew Research Center 8 (Jul 2025) Per query, ā„ 1 trad-link click15% Datos 9 (Mar 2025)Organic CTR, US40.3% ChatGPT Ours (per session, all-user)Preferred5.2% Ours (per message, Jul 2025 pooled) Robustness, message denominator1.2% Peec AI 10 Per search3ā5% Similarweb 11 Conversion rate (transactional AI refs) approximately 7% Ahrefs 12 Outbound links / continuous visit1.4 33 The Ahrefs 1.4-links-per-visit number on AI is not an outlier reading of the same construct: it counts outbound links per continuous on-platform visit rather than the share of sessions producing any click-out, and counts AI-initiated outbound as āclicks.ā Read against their denominators, the benchmarks bracket our two main ratios rather than contradicting them. On the Google side, our per-query 31.1% sits inside the published per-query band: SparkToro/Datos route roughly 36% of US-query clicks to the open web, Datosā organic CTR is 40.3%, and Pew records a traditional-link click on 15% of result pages. On the ChatGPT side, our per-session 5.2% is of the same order as Peec AIās 3ā5% per-search synthesis and Similarwebās approximately 7% transactional-referral conversion. The closest methodological analogāSemrushās U.S. clickstream panelāreports that ChatGPTās web-search feature is active on only 34.5% of queries as of February 2026 (down from 46% in late 2024), which caps how often an outbound referral is even possible and is consistent with a single-digit per-session click-out rate. 13 The robust pattern is the cross-intermediary contrast: on every comparable denominator, ChatGPTās propensity to send a click onward is several times lower than Googleās. OA.3 Domain classification The surrounding-context, destination-composition, and displacement-by-category results rest on one measurement stepāassigning each destination domain an economic label. We describe how we select the matched-support universe of 4,266 domains, what taxonomy we apply, how the labels distribute, how well the classifierās confidence calibrates against ground truth, how the labels hold up against blind human coding, and the verbatim prompt we send to GPT-4o. Domains enter the universe by their own referral coverage rather than by editorial judgment: the Preferred universe is the union of the top-2,500 LLM-referred and top-2,500 search-referred domains, which captures roughly 71% of referral mass on each side. We assign each domain two economic labels with a grounded GPT-4o classifier: ⢠content type (12 values, the last an explicit abstention): reference knowledge, news journal- 7 Rand Fishkin (SparkToro, on Datos clickstream data), 2024 Zero-Click Search Study, 1 July 2024, https: //sparktoro.com/blog/2024-zero-click-search-study-for-every-1000-us-google-searches- only-374-clicks-go-to-the-open-web-in-the-eu-its-360/ (accessed 2026-06-24). 8 A. Chapekis and A. Lieb, Google Users Are Less Likely to Click on Links When an AI Summary Appears in the Results, Pew Research Center, 22 July 2025, https://w.pewresearch.org/short-reads/2025/ 07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the- results/ (accessed 2026-06-24). 9 Datos, State of Search Q2 2025: Behaviors, Trends, and Clicks Across the US & Europe, https:// datos.live/report/state-of-search-q2-2025/ (accessed 2026-06-24). 10 Malte Landwehr (Peec AI), The Real Search Engine Market Share of ChatGPT, 23 June 2026, https:// peec.ai/blog/the-real-search-engine-market-share-of-chatgpt (accessed 2026-06-24). Synthesizes an approximately 3ā5% ChatGPT click-through ratio per search (iPullrank 3.8ā5.4%; Semrush approximately 3% per search) against Googleās approximately 40%. 11 Similarweb, 2025 Generative AI Landscape: From Platforms to Pathways (press release, 2 December 2025), https://ir.similarweb.com/news-events/press-releases/detail/138/ai-discovery-surges- similarwebs-2025-generative-ai-report-says (accessed 2026-06-24). 12 Ahrefs, AI Makes Up 0.1% of Traffic, but Clicks Arenāt Everything, https://ahrefs.com/blog/ai- traffic-research/ (accessed 2026-06-24): ChatGPT users click 1.4 external links per visit versus 0.6 for Google search users. 13 Semrush, ChatGPT Traffic Analysis: Insights from 17 Months of Clickstream Data, updated 7 April 2026, https://w.semrush.com/blog/chatgpt-search-insights/ (accessed 2026-06-24). 34 ism, academic research, social media, ecommerce brand, ecommerce marketplace, tools SaaS, developer technical, government institutional, entertainment gaming, adult, and other. ⢠monetization (6 values, the last an explicit abstention): ad supported, subscription paywall, freemium SaaS, transactional, nonprofit public, and mixed other. The classifier always returns one of these values; a domain it cannot place is labelled other or mixed_other at low confidence rather than left unlabelled. A separate unknown bucket appears in the distribution below: it is not a model label but an out-of-band residual stamped by the post- processing code when a batch fails to parse or the API call errors, so every domain in that batch receives unknown at confidence zero. Preferred composition uses the confidence ā„ 0.90 subset (76.1% of classified domains, i.e. 3,245 domains); the relaxation to ā„ 0.60 shifts content-type magnitudes by ⤠1.5 p at every category without changing direction (see Section OA.3.3 below). We run the classifier as gpt-4o (OpenAI EU endpoint eu.api.openai.com) at temperature=0, in batches of five domains, with a single-pass call per batch through the OpenAI Responses API; Section OA.3.4 reproduces the full prompt verbatim. The classifier may invoke a web-search tool to fetch unrecognized destinations before labelling them. Per-dimension distribution. Table OA.3.1 reports the marginal distribution of content type and monetization on the 4,266-domain matched-support universe. The unknown bucket is the out- of-band parse/API-failure residual described above, not a model label; the modelās own abstention is other. Table OA.3.1. Per-dimension classifier distribution on the 4,266-domain matched-support uni- verse. Sample: union of the top-2,500 LLM-referred and top-2,500 search-referred domains over October 2024 ā July 2025. Cell entries are domain counts N and within-dimension percentages. content typemonetization CategoryN % CategoryN % reference knowledge803 18.8ad supported1,708 40.0 tools SaaS686 16.1 transactional890 20.9 other490 11.5 nonprofit public680 15.9 ecommerce brand425 10.0 freemium SaaS578 13.5 entertainment gaming348 8.2 subscription paywall 248 5.8 government institutional 345 8.1mixed other137 3.2 news journalism327 7.7 unknown25 0.6 ecommerce marketplace 259 6.1 adult250 5.9 social media166 3.9 academic research82 1.9 developer technical80 1.9 unknown5 0.1 Manual validation. We assess whether the machine labels are trustworthy by hand-validating them against a blind human coding of a 100-domain stratified sample on the same two dimensions. The design and results are reported in Section OA.3.1 below: overall agreement is substantial-to- almost-perfect (Cohenās Īŗ = 0.76 on content type, 0.82 on monetization). 35 Sample sensitivity. The rank cutoff is a choice, and a reader should know what is lost by stopping at the top 2,500 per side. Table OA.3.2 traces that sensitivity: as the cutoff K rises, the share of ChatGPT and Google referral mass the sample covers rises with sharply diminishing returns. Table OA.3.2. Sample sensitivity to the rank cutoff. Sample: Comscore US Desktop active- household panel, October 2024 ā July 2025. K is the rank cutoff applied to each side (LLM-referred top-K and search-referred top-K); intersections, sizes, and own-side and union-side coverage (share of platform referral mass on the cutoff support) are reported. The Preferred universe takes K=2,500 on each side. K |LLM only| |both| |union| LLM cov. Search cov. Union (LLM) 5003441568440.5150.5780.530 1,000697303 1,6970.5930.6350.608 2,0001,399601 3,3990.6800.6900.696 2,500 (Preferred)1,766734 4,2660.7110.7080.726 3,0002,154846 5,1540.7350.7220.751 5,0003,658 1,342 8,6580.8210.7600.834 10,0007,545 2,455 17,5450.9410.8070.948 Past K=3,000 additional domains add negligible referral mass: lifting LLM coverage from 0.73 to 0.95 takes Kā10,000, and the extra domains are tail-rank destinations that contribute negligibly to either platformās referral mass. The K=2,500 cutoff sits at the joint Pareto optimum, covering approximately 70% of referrals on both sides while leaving the long, low-value tail unclassified. OA.3.1 Validation against human coding To assess whether the labels behind the composition results agree with human judgment, we validate them against a blind human reference. One of the authors hand-labelled a 100-domain sample on the same two dimensions, using the same codebook, blind to the LLM labels. We draw the sample by rank-stratified random sampling (fixed seed) rather than as a top-100: the 4,266 destinations are ordered by referral count and split into a head (1ā200), a middle (201ā1,000), and a tail (1,001+), with 40, 30, and 30 domains drawn at random within each. By spreading validation across the whole size distribution rather than concentrating it in the head, the agreement statistics gauge real classifier performance across the full range of destinations rather than a handful of high- traffic, easily classified platforms. The 100 sampled domains span ranks 10ā4,082 and carry only approximately 7.5% of total referral volume; because the head is sampled at random (40 of 200), the traffic-dominant mega-platforms (YouTube, Reddit, Amazon) are mostly absent. Because the sample is not self-weighting, population figures are reweighted by the known sampling fractions (reported below). Overall agreement. Table OA.3.3 reports exact agreement, referral-volume-weighted agree- ment, and Cohenās Īŗ with a percentile-bootstrap 95% confidence interval (5,000 resamples, fixed seed). Both dimensions reach substantial-to-almost-perfect chance-corrected agreement overall (Īŗ = 0.76 for content type, 0.82 for monetization, reading Īŗ ā (0.61, 0.80) as substantial and Īŗā„ 0.81 as almost perfect per Landis and Koch (1977)). Volume-weighted agreement is uniformly higher than unweighted (0.81ā0.91), because disagreements concentrate on lower-traffic, ambiguous sites. This within-sample figure is not a population traffic-weighted estimate; the design-reweighted 36 estimates appear below. In total 30 of 100 domains disagree on at least one dimension (17 on content type only, 9 on monetization only, 4 on both), i.e. 21 content-type and 13 monetization disagreements. Table OA.3.3. Agreement between the GPT-4o classifier and a blind human coder on the 100- domain validation sample. āDetectableā excludes the 12 domains on which the LLM abstained (content type = other, see below). Exact and volume-weighted agreement are shares; Īŗ is Cohenās chance-corrected agreement with a percentile-bootstrap 95% interval. Cats is the number of cate- gories present in the sample. SubsetDimensionn Cats Exact Vol-wtd Īŗ 95% CI Overallcontent type 10011 0.790.81 0.76 [0.67, 0.85] Overallmonetization 1006 0.870.91 0.82 [0.72, 0.90] Detectable (LLM Ģø= other) content type 8811 0.840.84 0.82 [0.72, 0.90] Detectable (LLM Ģø= other) monetization 886 0.880.90 0.82 [0.72, 0.92] The abstention bucket. The codebook instructs the model to emit content type = other (confidence < 0.6) when it cannot determine what a homepage shows ā an explicit abstention rather than a substantive category. The LLM abstained on 12 of 100 domains; the human independently coded 6 as other, and the two agreed on 5 (genuine agreement). A salient slice of the disagreeing abstentions is small, non-mainstream adult sites: the model recognizes and correctly labels large adult brands (adult precision = 1.00) but has no training knowledge of obscure adult properties and its grounding tool will not crawl them, so it falls back to other. The remaining abstentions are the deliberately quarantined frontier-LLM domains and assorted small/ungroundable portals. These abstention-driven misses account for 7 of the 21 content-type disagreements; because they reflect missing information rather than mis-classification, the detectable subset (n = 88) is the cleaner measure of label quality where the model commits, and there both dimensions reach almost perfect agreement (Īŗ = 0.82). Where the labels diverge. Disagreement is systematic and semantically adjacent, not ran- dom noise (Table OA.3.4). On content type the LLM is highly precise on identifiable categories ā adult, academic research, and government institutional (precision 1.00), entertainment gaming (0.94), ecommerce brand (0.93), tools SaaS (0.92). On monetization, ad supported and transac- tional are recovered very well (F1 0.90 and 0.96); the one weak spot is mixed other (recall 0.29), where the LLM resolves āmixedā sites to a single dominant model ā coder conservatism, not an error of kind. 37 Table OA.3.4. Per-category precision, recall, and F 1 of the LLM against the human coder on the validation sample, treating the human label as ground truth. āSupportā is the human count; āAssignedā is the LLM count. Both dimensions, full 100-domain sample. CategorySupport Assigned Precision Recall F 1 content type entertainment gaming25170.940.64 0.76 tools SaaS17130.920.71 0.80 ecommerce brand13140.931.00 0.96 adult12101.000.83 0.91 reference knowledge7110.550.86 0.67 other6120.420.83 0.56 ecommerce marketplace560.831.00 0.91 news journalism590.561.00 0.71 social media540.750.60 0.67 government institutional431.000.75 0.86 academic research111.001.00 1.00 monetization ad supported43480.850.95 0.90 transactional25231.000.92 0.96 freemium SaaS11120.750.82 0.78 nonprofit public971.000.78 0.88 mixed other730.670.29 0.40 subscription paywall570.711.00 0.83 By layer and population estimates. Because the sample is a disproportionate stratified draw, the raw averages are not population figures. Table OA.3.5 gives two design-aware viewsāi.e., estimates that correct for the disproportionate stratified sampling rather than averaging the raw sample. By rank stratum (Panel A), content-type agreement is stable down the traffic distribution (Īŗ ā 0.73ā0.77 everywhere), while monetization is almost perfect in the head (Īŗ = 0.92, where revenue models are obvious) and softens to substantial in the tail (Īŗ = 0.71). Panel B reweights to all 4,266 domains: a per-domain HorvitzāThompson estimator (inverse sampling fraction) gives 0.77 / 0.79 for content type and monetization, and a post-stratified traffic-weighted estimator gives 0.79 / 0.88. However it is weighted, content-type agreement lands at 0.77ā0.79 and monetization at 0.79ā0.88 across the full universe. 38 Table OA.3.5. Design-aware agreement. Panel A reports exact agreement and Cohenās Īŗ within each rank stratum. Panel B reweights the sample to the 4,266-domain universe: the per-domain estimator weights each sampled domain by its stratumās inverse sampling fraction; the traffic- weighted estimator weights predicted categories by their universe referral share. Coverage is the share of the universe the estimator spans. content typemonetization Panel A: by rank stratumExact ĪŗExactĪŗ Head (rank 1ā200)n=400.80 0.760.950.92 Middle (rank 201ā1,000)n=300.80 0.770.870.80 Tail (rank 1,001+)n=300.77 0.730.770.71 Panel B: population estimatescontent typemonetizationCoverage Design-weighted, per-domain (HorvitzāThompson)0.770.79100% Post-stratified, traffic-weighted0.790.88 approximately 99% Interpretation. The composition facts do not hinge on the adjacent distinctions that drive the residual disagreement. The groupings used in the text ā e-commerce as a whole; ad-supported ver- sus transactional; the share of referrals to UGC/social media and reference knowledge ā are exactly the level at which human and machine agree most, and the disagreements cluster on within-family confusions (gaming-editorial vs. news, brand vs. marketplace, single-model vs. mixed monetiza- tion) that collapse away at that grouping. We therefore treat the LLM classification as a valid measurement instrument. OA.3.2 Classifier confidence: distribution and calibration Every content-type result in the paper restricts to high-confidence labels; this section documents the confidence score and tests whether it separates correct labels from incorrect ones. The classifier emits a per-domain confidence score alongside each pair of labels. Across the 4,266 classified domains the score has mean 0.870 and median 0.90 (range 0.0ā0.99), and 76.1% (3,245 domains) clear the ā„ 0.90 high-confidence cutoff; Table OA.3.6 reports the full distribution. Confidence tracks groundability: on the detectable subset (content type /ā other, unknown, n = 3,771) the model is confident ā median 0.90, with 82.3% at ā„ 0.90 ā whereas on the abstention/residual bucket (content type ā other, unknown, n = 495) confidence drops to median 0.65, and the 5 genuine unknown rows carry confidence 0.0. The low-confidence tail is therefore overwhelmingly the same small/ungroundable-site population that drives the abstention misses documented in Section OA.3.1. 39 Table OA.3.6. Classifier-confidence distribution on the 4,266-domain matched-support universe. Cell entries are domain counts N within each confidence bin and the corresponding bin share; the final column is the cumulative share at or above the binās lower edge. Confidence bin N% Cumulative ā„ ā„ 0.903,245 76.176.1 0.80ā0.90487 11.487.5 0.70ā0.80199 4.792.1 0.60ā0.70108 2.594.7 < 0.60227 5.3100.0 Notes: Bins are right-open below 0.90 and closed at the top (ā„ 0.90). The cumulative column reads downward: e.g. 87.5% of classified domains carry confidenceā„ 0.80 and 94.7% carryā„ 0.60. The < 0.60 bin is the abstention/residual tail. Confidence calibration. A high score is only useful if it predicts a correct label. We test this directly by joining the confidence column to the 100-domain blind-human validation sample of Section OA.3.1, which lets us ask whether the score flags the labels the model gets wrong. The sample mirrors the universe share: 79 of the 100 sampled domains are high-confidence (ā„ 0.90), 87 at ā„ 0.80, and 97 at ā„ 0.60 (median 0.90, in line with the 76.1% universe share). Table OA.3.7 splits exact agreement and Cohenās Īŗ against the human coder into the high-confidence (ā„ 0.90, n = 79) and low-confidence (< 0.90, n = 21) bins. Table OA.3.7. Classifierāhuman agreement split by confidence bin on the 100-domain validation sample. Exact is the share of domains on which the LLM and human labels match; Īŗ is Cohenās chance-corrected agreement. content type monetization Subsetn ExactĪŗ ExactĪŗ Full sample100 0.790 0.761 0.870 0.817 High-conf (ā„ 0.90) 79 0.835 0.813 0.861 0.813 Low-conf (< 0.90) 21 0.619 0.525 0.905 0.768 The score is well calibrated for content type: the ā„ 0.90 bin reaches almost-perfect agreement (exact 0.835, Īŗ = 0.813) ā essentially the detectable-subset quality ā while the low-confidence tail drops to moderate (exact 0.619, Īŗ = 0.525). The high-confidence bucket is precisely where the labels are most trustworthy, and the score correctly flags the harder approximately 21%. The score does not discriminate monetization quality: agreement is essentially flat across bins (the low-conf cell is even slightly higher on exact, lower on Īŗ, on n = 21). This is expected ā the single per-domain confidence reflects whether the model can identify what the site is (the abstention mechanism), not how readable its revenue model is. Because confidence discriminates content-type quality, all three content-type analyses ā sur- rounding context, destination composition, and displacement-by-content-type ā restrict to the confidence ā„ 0.90 subset (3,245 domains), giving a common high-quality labeled domain set across exhibits. No monetization-specific confidence filter is applied, because confidence does not track monetization quality. 40 OA.3.3 Classifier confidence and taxonomy robustness This subsection tests whether the composition tilt is an artifact of the high-confidence filter. The Preferred content-type composition uses the confidence ā„ 0.90 subset of the matched-support clas- sifier output, and relaxing the threshold to ā„ 0.60 admits an additional approximately 19% (about 790 domains), mostly tail destinations, without changing the direction or shape of the composition tilt. Figure OA.3.1. Destination-composition results survive lower classifier-confidence thresholds. G 12.9 | CG 25.9 G 6.4 | CG 15.6 G 0.3 | CG 5.7 G 0.8 | CG 5.5 G 0.6 | CG 3.4 G 4.5 | CG 7.2 G 3.7 | CG 5.6 G 14.6 | CG 8.1 G 12.6 | CG 4.5 G 9.8 | CG 0.1 G 33.8 | CG 18.3 social media adult entertainment gaming ecommerce marketplace ecommerce brand news journalism government institutional developer technical academic research tools SaaS reference knowledge -10+0+10+20+30 ChatGPT - Google share (p) (a) Content type. G 6.6 | CG 23.5 G 8.0 | CG 17.0 G 2.7 | CG 5.1 G 1.9 | CG 2.8 G 19.9 | CG 17.5 G 60.9 | CG 34.1 ad supported transactional mixed other subscription paywall freemium SaaS nonprofit public -20+0+20+40 ChatGPT - Google share (p) (b) Monetization. Notes: Bars report ChatGPTās referral share minus Googleās by destination category under alterna- tive classifier-confidence thresholds. Panel (a) varies the content-type threshold; panel (b) varies the monetization threshold. The ordering remains stable: ChatGPT favors tools, reference, academic, and developer destinations and routes relatively less traffic to social and ad-supported websites. OA.3.4 Verbatim domain-classification prompt We reproduce below the literal prompt sent to GPT-4o for the domain classification of Ap- pendix OA.3, exactly as sent, so the labels can be regenerated. At run time we fill the block placeholders DOMAIN_1. . . DOMAIN_5 with the five domains in each batch and call the OpenAI Responses API with a web_search tool enabled and temperature=0; the model may call web_search 0+ times per batch before emitting the final JSON array. The per-domain web_search flag in the output is recovered post-hoc from the APIās tool-call events (substring matching the queried URL against batch domains, with ambiguous calls attributed conservatively). <instructions> You are classifying website domains along two dimensions for an academic study of how AI search redistributes web traffic.,ā The taxonomy is organized around AI-search displacement risk -- how easily an LLM can answer the user's underlying need without sending the user to the destination. Groupings (reference vs. social, brand ecommerce vs. marketplace ecommerce, adult as a policy-filtered bucket, frontier LLM chatbots quarantined into "other") are chosen because they isolate quantitatively different displacement patterns in the downstream regressions. ,ā ,ā ,ā ,ā Classify each domain along these two dimensions: content_type, monetization. Provide a one-sentence rationale and a float confidence score in [0.0, 1.0].,ā 41 Label each domain based on its homepage and the primary user activity the site supports. Cite the homepage evidence behind your label in the rationale (e.g., "homepage is a feed of subreddit post cards with vote counts and comment counts"). If you cannot determine the site's function -- domain is unrecognized, parked, or multi-mode with no clear primary activity -- set content_type to "other", monetization to "mixed_other", confidence <= 0.3, and say so explicitly in the rationale. Do not guess from domain name or TLD alone. ,ā ,ā ,ā ,ā ,ā </instructions> <dimensions> <dimension name="content_type"> 12 values, listed alphabetically (the ordering is intentional and not a ranking). - "academic_research": Peer-reviewed journals, preprint servers, research organizations, academic databases. Examples: nih.gov, arxiv.org, nature.com, sciencedirect.com, pubmed.ncbi.nlm.nih.gov.,ā - "adult": Explicit-content sites. Quarantined into its own category because the AI-Google gap on adult content is a policy artifact (LLMs filter by design) not an organic compositional difference; mixing it with other buckets contaminates the substantive findings. Examples: pornhub.com, xvideos.com, xhamster.com, onlyfans.com. ,ā ,ā ,ā - "developer_technical": Code repositories, package registries, developer documentation, technical reference for programmers. Examples: github.com, npmjs.com, developer.mozilla.org, stackoverflow-for-teams-style docs. ,ā ,ā - "ecommerce_brand": Manufacturer / direct-to-consumer sites selling their own goods or services. The site's catalog is its own brand. Examples: nike.com, apple.com, patagonia.com, tesla.com.,ā - "ecommerce_marketplace": Multi-seller marketplaces and broad-line retailers that aggregate third-party or multi-brand catalogs. Examples: amazon.com, ebay.com, etsy.com, walmart.com, target.com.,ā - "entertainment_gaming": Gaming sites, mod communities, fan-fiction archives, play-focused destinations. **Also includes any wiki, encyclopedia, or database whose entire scope is one game, game franchise, or game series -- even when the homepage looks like a general reference index. A "wiki" or "db" or ".wiki" / "*pedia.net" / "*db.net" domain devoted to a single game belongs here, NOT in reference_knowledge.** Examples: steamcommunity.com, fandom.com, nexusmods.com, gamebanana.com, gta5-mods.com, ign.com when gaming-only, runescape.wiki, stardewvalleywiki.com, pokemondb.net, bulbagarden.net (Pokemon), paradoxwikis.com, minecraft.wiki, terrariawiki.org, oldschool.runescape.wiki. ,ā ,ā ,ā ,ā ,ā ,ā - "government_institutional": .gov sites, public-service portals, international institutions, authoritative policy/regulatory sources. Examples: irs.gov, cdc.gov, usa.gov, europa.eu, who.int.,ā - "news_journalism": Original reporting, current news, opinion, and editorial commentary consumed as timely/topical content. Examples: nytimes.com, reuters.com, apnews.com, bbc.com, forbes.com.,ā - "other": Three distinct uses. (a) **Frontier general-purpose LLM chatbots, AI-search assistants, multi-LLM aggregator chatbots, and persona-based / companion AI chatbots** whose primary product IS the conversational AI. Covers (i) general-purpose assistants (chatgpt.com, openai.com, claude.ai, gemini.google, perplexity.ai, pplx.ai, deepseek.com, ollama.com), (i) multi-LLM aggregator wrappers (chaton.ai, monica.im, iask.ai, toolbaz.com), and (i) persona / companion / roleplay AI chatbots where users converse with AI-generated characters even when the homepage shows a UGC-character feed (character.ai, polybuzz.ai, spicychat.ai, replika.ai, dream.ai, woebothealth.com, and similar). These are all the AI-search substitute being studied; bucketing them across "tools_saas" / "social_media" / "entertainment_gaming" would scatter the AI substitute across multiple buckets and dilute the very category the analysis treats as the substitute. Use confidence >= 0.9 here. (b) **AI-only single-purpose generators** for image / video / voice / music / avatars whose entire product is the AI generation step (midjourney.com, runwayml.com, suno.com, udio.com, elevenlabs.io, openart.ai, leonardo.ai, ideogram.ai, invideo.io, synthesia.io, heygen.com, descript.com, fliki.ai, pika.art, lalal.ai, soundraw.io, craiyon.com, aiva.ai, magicstudio.com, aistudios.com, maxstudio.ai, and similar). They are AI-substitution-adjacent destinations whose displacement dynamics belong with category (a), not with general productivity SaaS. Use confidence >= 0.9 here. Note: domain-specific tools that *embed* AI features but whose primary product is the underlying tool (Grammarly, Canva, Notion, Adobe Photoshop, SodaPDF, Zoom) stay in "tools_saas". (c) Domains the homepage cannot pin down (unrecognized, parked, multi-mode with no clear primary activity). Use confidence < 0.6 here. ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā 42 - "reference_knowledge": Encyclopedic, how-to, dictionary, tutorial, or general reference content consumed as timeless lookup. **Also includes education-purpose platforms even if their business model is SaaS** -- learning management systems (Canvas/instructure.com, Schoology, Blackboard, Ellucian, Clever, ClassLink, PowerSchool, Jupitered), homework / study tools (Chegg, Quizlet, Quizizz, Mathway, Symbolab, WolframAlpha, Desmos, GeoGebra, Turnitin, Duolingo), K-12 adaptive-learning platforms (IXL, Edgenuity, DreamBox, Prodigy, Kahoot), tutoring marketplaces (Preply, Wyzant, italki, Cambly, Outschool), exam proctoring (Honorlock, ProctorU, Respondus), MOOCs and online courses (Coursera, edX, Udemy, Cengage, McGraw-Hill), education publishers (Pearson), and student career platforms (Handshake). The user's purpose at all of these is learning / coursework / education-knowledge lookup, not productivity tooling. **Also includes hospitals and health publishers** (mayoclinic.org, clevelandclinic.org, webmd.com, healthline.com, drugs.com, medlineplus.gov, kidshealth.org): their homepages present a health-reference index (symptoms, conditions, drugs, tests), not an institutional/government portal -- bucket them with reference_knowledge, NOT government_institutional, even when the .org TLD and clinic branding suggest "institution". Examples: wikipedia.org, wikihow.com, britannica.com, dictionary.com, instructure.com, khanacademy.org, chegg.com, coursera.org, edgenuity.com, ixl.com, preply.com, mayoclinic.org, webmd.com. ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā ,ā - "social_media": All social and consumption-on-platform UGC and streaming sites -- text-first social/Q&A/threaded discussion *and* rich-media (video, audio, image-centric) UGC and professional streaming libraries. The user is on the platform either to read/post text or to watch/listen; consumption typically requires the platform. Examples: reddit.com, quora.com, stackoverflow.com, x.com, twitter.com, facebook.com, linkedin.com, youtube.com, instagram.com, tiktok.com, spotify.com, netflix.com, twitch.tv. ,ā ,ā ,ā ,ā ,ā - "tools_saas": Productivity tools, web apps, utilities, design tools, SaaS dashboards used for general work -- *not* education or learning, *not* frontier LLM chatbots, and *not* AI-only single-purpose generators. The user uses or logs into a tool. Examples: canva.com, grammarly.com, notion.so, zoom.us, docs.google.com. Note: education-SaaS (LMS, study tools, MOOCs, homework helpers, K-12 learning platforms, tutor marketplaces, exam proctoring) goes to "reference_knowledge"; frontier general-purpose LLM chatbots and AI-only single-purpose generators go to "other" (see below); Shopify storefronts and the merchant tooling at shopify.com / *.myshopify.com go to "ecommerce_marketplace". ,ā ,ā ,ā ,ā ,ā ,ā </dimension> <dimension name="monetization"> 6 values, listed alphabetically (the ordering is intentional and not a ranking). - "ad_supported": Default for any free-to-access site whose dominant homepage revenue signal is advertising (display ads, video pre-rolls, sponsored content units, ad slots in feeds). **A low-uptake premium tier -- "remove ads", YouTube Premium, Quora+ -- does NOT promote the site to mixed_other.** The premium tier has to materially co-fund the site, not just exist as an optional upsell. Examples: cnn.com, weather.com, youtube.com, quora.com, fandom.com, allrecipes.com, reddit.com. NOTE: if the site is a SaaS app with visible pricing tiers (spotify, dropbox, notion), use freemium_saas / subscription_paywall instead -- those are matched earlier in the monetization decision tree below. ,ā ,ā ,ā ,ā ,ā ,ā - "freemium_saas": Functional free tier plus paid upgrade (SaaS or app). Examples: spotify.com, notion.so, grammarly.com, chatgpt.com, dropbox.com.,ā - "mixed_other": Reserved for sites with TWO OR MORE roughly equal revenue streams visible on the homepage. Use only when the secondary stream is substantial -- a metered paywall most readers hit, marketplace commissions on a multi-revenue platform, in-app commerce comparable to ad revenue. NOT a fallback for "doesn't fit elsewhere". Examples: nytimes.com (metered paywall reaches a large share of readers), twitch.tv (ads + subs + bits + commerce), bandcamp.com (commerce + ad-free streaming). ,ā ,ā ,ā ,ā - "non_profit_public": Foundation-funded, government-funded, .edu, .gov, hospital/clinic non-profits. Examples: wikipedia.org, nih.gov, usa.gov, mayoclinic.org, clevelandclinic.org, hopkinsmedicine.org. **NOT for commercial health publishers (webmd.com, healthline.com, drugs.com) -- those run on ad revenue; bucket those as monetization=ad_supported even though their content_type is reference_knowledge under the hospital/health exception.** ,ā ,ā ,ā ,ā - "subscription_paywall": Recurring fee for access, whether metered or hard paywall. Examples: wsj.com, ft.com, netflix.com, disneyplus.com.,ā - "transactional": Revenue from on-site goods or services sales (shopping cart, booking flow, checkout). Examples: amazon.com, nike.com, walmart.com, expedia.com, doordash.com.,ā Monetization decision tree -- apply top-to-bottom and stop at the first match: 1. Shopping cart / booking / checkout on homepage -> "transactional" 2. Hard paywall (no free articles visible) -> "subscription_paywall" 3. SaaS pricing tiers visible, free tier exists -> "freemium_saas" 4. SaaS pricing tiers visible, no free tier -> "subscription_paywall" 5. Foundation/non-profit/.gov/.edu/hospital footer with no ads or commerce -> "non_profit_public" 6. Visible ads + free content (any "premium" tier is optional/secondary) -> "ad_supported" (default) 7. Metered paywall OR substantial second stream comparable to the first -> "mixed_other" </dimension> </dimensions> 43 <decision_tree> Apply these steps in order. STEP 1 -- Evidence about the homepage. What does the homepage show -- navigation, headlines, product tiles, login/paywall copy, footer, about? If the domain is unrecognized, parked, empty, or redirects to an unrelated placeholder -> content_type = "other", confidence <= 0.3, note it in rationale. STOP the content_type decision here. ,ā ,ā STEP 2 -- Apply the adult filter first. If the homepage displays explicit adult content -> content_type = "adult". Do not reach for any other content_type. Continue to monetization normally.,ā STEP 3 -- What is the primary user activity on the site? 3a. Looking up timeless reference, how-to, dictionary, or encyclopedic content -> "reference_knowledge". 3b. Reading current news, original reporting, opinion, or editorial -> "news_journalism". 3c. Reading peer-reviewed research, preprints, or academic databases -> "academic_research". 3d. Browsing/posting on social, Q&A, threaded discussion, broadcast feeds, professional networks, or watching/listening on UGC/streaming media platforms (reddit, quora, stackoverflow, x/twitter, facebook, linkedin, youtube, instagram, tiktok, spotify, netflix, twitch) -> "social_media". ,ā ,ā 3e. Buying from a single brand's own catalog -> "ecommerce_brand". 3f. Buying across a multi-seller marketplace or broad-line retailer -> "ecommerce_marketplace". 3g. Using a web app, productivity tool, or logging into a SaaS dashboard -> "tools_saas". 3h. Visiting a code repo, package registry, or developer documentation -> "developer_technical". 3i. Accessing a government/.gov/public-service portal or institutional authority -> "government_institutional".,ā 3j. Browsing gaming, fandom wiki, fan-fiction, or mod communities -> "entertainment_gaming". 3k. Site's primary product is a frontier / general-purpose LLM chatbot or AI-search assistant (ChatGPT, Claude, Gemini, Perplexity, DeepSeek, Ollama, and similar) -> "other" (confidence >= 0.9, study-design quarantine). ,ā ,ā 3l. Nothing above fits, or evidence is insufficient -> "other" (confidence < 0.6). STEP 4 -- Resolve disambiguation with the rules below. STEP 5 -- Pick monetization using the monetization decision tree above (shopping cart -> transactional -> hard paywall -> subscription_paywall -> SaaS tiers -> freemium_saas / subscription_paywall -> non-profit footer -> non_profit_public -> free + ads -> ad_supported -> multiple substantial streams -> mixed_other). ,ā ,ā ,ā STEP 6 -- Calibrate confidence. 0.9-1.0: homepage evidence is unambiguous (clear masthead, clear product catalog, clear login/paywall, clear .gov header).,ā 0.6-0.9: homepage evidence is suggestive but requires inference across multiple signals, or the domain is multi-mode and you had to pick a primary mode.,ā <0.6: homepage was thin, ambiguous, behind auth, parked, or contradictory. Use content_type = "other" rather than guessing at a specific label.,ā </decision_tree> <disambiguation_rules> The most confusable boundaries, all resolved from observed homepage evidence. tools_saas vs. other (frontier LLM and generative-AI exceptions): - Frontier general-purpose LLM chatbots and AI-search assistants whose primary product IS the conversational AI -- chatgpt.com, openai.com, claude.ai, gemini.google, perplexity.ai, pplx.ai, deepseek.com, ollama.com -> "other". These are the AI-search substitute being studied; quarantining them keeps tools_saas from absorbing the category the analysis treats as the substitute. Confidence >= 0.9 -- this is a study-design rule, not an evidence-thin guess. ,ā ,ā ,ā ,ā - Multi-LLM aggregator chatbots (chaton.ai, monica.im, iask.ai, toolbaz.com) -> "other". Same logic: primary product is the conversational AI, just routed through multiple back-end models.,ā - Persona-based / companion / roleplay AI chatbots where users converse with AI-generated characters (character.ai, polybuzz.ai, spicychat.ai, replika.ai, dream.ai, woebothealth.com) -> "other". The homepage may show a UGC-style character gallery, but the primary user activity is still chatting with an LLM-generated agent -- they are AI-substitution-adjacent and belong with the other conversational chatbots, not with social_media or tools_saas. ,ā ,ā ,ā ,ā - AI-only single-purpose generators for image / video / voice / music / avatars (Midjourney, Runway, Suno, Udio, ElevenLabs, OpenArt, Leonardo, Ideogram, InVideo, Synthesia, HeyGen, Descript, Fliki, Pika, Lalal, Soundraw, Craiyon, AIVA, MagicStudio, AIStudios, MaxStudio) -> "other". Their entire product is the AI-generation step, so they are AI-substitution-adjacent and belong with the conversational chatbots, not with productivity SaaS. ,ā ,ā ,ā ,ā 44 - Domain-specific tools that *embed* AI features but whose primary product is the underlying tool (Grammarly, Canva, Notion, Adobe Photoshop, SodaPDF, Zoom) -> "tools_saas".,ā ecommerce_brand vs. ecommerce_marketplace: - Single-brand catalog selling its own goods (homepage shows one brand, one product line) -> "ecommerce_brand". Examples: nike.com, apple.com, patagonia.com.,ā - Multi-seller marketplace or broad-line retailer (homepage shows many brands, category navigation across unrelated product classes) -> "ecommerce_marketplace". Examples: amazon.com, ebay.com, walmart.com, etsy.com, target.com. ,ā ,ā news_journalism vs. reference_knowledge: - Read as current/topical (articles with datelines, running headlines, opinion pieces, reviews) -> "news_journalism".,ā - Looked up as timeless reference (dictionary entries, how-to articles, wiki pages, encyclopedic entries, health-condition pages) -> "reference_knowledge".,ā - Health/finance reference with paid writers (webmd.com, healthline.com, investopedia.com) -> "reference_knowledge" when homepage is a reference index, not a news feed.,ā reference_knowledge vs. academic_research: - General reference or encyclopedic (wikipedia, britannica, wikihow) -> "reference_knowledge". - Peer-reviewed, preprint, or research-org primary sources (arxiv, nature, nih.gov research pages) -> "academic_research".,ā - Edge: pubmed, sciencedirect -> "academic_research". reference_knowledge vs. tools_saas (education exception): - LMS (Canvas/instructure, Schoology, Blackboard, Ellucian, Clever, ClassLink, PowerSchool, Jupitered), study tools (Chegg, Quizlet, Quizizz, Mathway, Symbolab, WolframAlpha, Desmos, GeoGebra, Turnitin, Duolingo), K-12 adaptive-learning platforms (IXL, Edgenuity, DreamBox, Prodigy, Kahoot), tutor marketplaces (Preply, Wyzant, italki, Cambly, Outschool), exam proctoring (Honorlock, ProctorU, Respondus), MOOCs (Coursera, edX, Udemy, Cengage, McGraw-Hill), education publishers (Pearson), student career platforms (Handshake) -> "reference_knowledge". The user's purpose at these is learning or coursework, not productivity tooling -- even though the business model is SaaS. ,ā ,ā ,ā ,ā ,ā ,ā - General productivity SaaS / design tools (Canva, Notion, Grammarly, Zoom, Slack, Figma) -> "tools_saas". social_media (text-first social and rich-media UGC/streaming, merged): - Both text-first social (reddit, quora, stackoverflow, twitter/x, facebook, linkedin) and rich-media UGC/streaming (youtube, tiktok, instagram, spotify, netflix, twitch) -> "social_media". The previous split between "social_platform" and "media_content" has been merged: both buckets describe consumption-on-platform UGC/streaming, and the analytical contrast that matters now is social_media vs. reference_knowledge / news_journalism, not text-vs-media within social_media. ,ā ,ā ,ā ,ā - Adult sites still go to "adult", not "social_media" (adult filter applied first per Step 2). developer_technical vs. tools_saas: - Code repos, package registries, API docs, developer-specific reference (github, npmjs, pypi, MDN, docs.python.org) -> "developer_technical".,ā - General SaaS dashboards, productivity web apps, design tools (canva, notion, grammarly, zoom) -> "tools_saas".,ā government_institutional vs. reference_knowledge: - .gov, .mil, .int, public-service portals, regulatory/tax agencies -> "government_institutional". - Encyclopedic or how-to content on a .org or .com that happens to be non-profit -> "reference_knowledge". - **Hospitals and health publishers (mayoclinic.org, clevelandclinic.org, webmd.com, healthline.com, drugs.com, medlineplus.gov, kidshealth.org) -> "reference_knowledge", NEVER government_institutional.** Their homepages present a health-reference index (symptoms, conditions, drugs, tests), not an institutional service portal. The .org TLD and clinic branding are not sufficient evidence for government_institutional. Research arms of NIH etc. -> "government_institutional" or "academic_research" depending on homepage. ,ā ,ā ,ā ,ā ,ā ecommerce_marketplace vs. tools_saas: - Homepage shows shopping cart, product grid, prices -> "ecommerce_marketplace" (or "ecommerce_brand"). - Homepage shows "sign in", tool interface, or SaaS pricing tiers -> "tools_saas". - **Shopify exception**: shopify.com (the merchant SaaS landing page) and *.myshopify.com (the per-merchant storefront subdomain) -> "ecommerce_marketplace". Even though shopify.com markets a SaaS to merchants, end-user panel traffic to shopify.com is dominated by buyer-side checkout flows, and *.myshopify.com tenant subdomains ARE storefronts. Bucketing these as marketplace destinations matches what panelists actually do at the URL. ,ā ,ā ,ā ,ā entertainment_gaming vs. social_media: - Gaming storefront / fandom wiki / mod community / fan-fiction archive -> "entertainment_gaming". 45 - Video/audio/image streaming or text/Q&A social on a general consumer platform (including Twitch live-streaming and gaming-adjacent text discussion on twitter/reddit) -> "social_media".,ā entertainment_gaming vs. reference_knowledge (game-specific wikis/databases): - General-purpose reference, encyclopedia, or how-to (wikipedia, britannica, wikihow, dictionary.com) -> "reference_knowledge".,ā - **Wiki, encyclopedia, or database whose entire scope is a single game, franchise, or game series (runescape.wiki, stardewvalleywiki.com, pokemondb.net, bulbagarden.net, paradoxwikis.com, minecraft.wiki) -> "entertainment_gaming", NOT reference_knowledge.** The site is reference-shaped but the topic is play, not general knowledge. The "*.wiki" / "*db.net" / "*pedia.net" pattern does not save it for reference_knowledge if the scope is a single game. ,ā ,ā ,ā ,ā </disambiguation_rules> <examples> 14 reference classifications covering the 12 categories and the major split cases (hospital exception, Shopify exception, generative-AI quarantine, education exception, mixed_other for metered paywalls). Use them as calibration anchors. Each rationale is phrased as though you had just viewed the homepage. ,ā ,ā [ "domain": "reddit.com", "content_type": "social_media", "monetization": "ad_supported", "confidence": 0.95, "rationale": "Homepage is a feed of subreddit post cards with vote counts and comment counts -- canonical threaded discussion. Bucketed under the merged social_media category (text-first social and rich-media UGC/streaming combined). Ads dominate the feed; Reddit Premium and awards are secondary upsells, so this stays ad_supported, not mixed_other.", ,ā ,ā ,ā ,ā "domain": "chatgpt.com", "content_type": "other", "monetization": "freemium_saas", "confidence": 0.99, "rationale": "Homepage is the ChatGPT conversational interface -- a frontier general-purpose LLM chatbot. Quarantined into'other' per study design (the AI-search substitute being studied, not productivity tooling). Freemium with paid Plus/Pro tiers.", ,ā ,ā ,ā "domain": "quora.com", "content_type": "social_media", "monetization": "ad_supported", "confidence": 0.95, "rationale": "Homepage shows a threaded Q&A feed with answer cards and'Follow' buttons -- text-first social bucketed under social_media. Ads dominate the homepage; Quora+ exists but is a secondary upsell, so this stays ad_supported not mixed_other.", ,ā ,ā ,ā "domain": "nike.com", "content_type": "ecommerce_brand", "monetization": "transactional", "confidence": 0.99, "rationale": "Homepage is Nike's own product catalog (shoes, apparel) with shopping-cart CTA. Single-brand DTC storefront, not a multi-seller marketplace. Revenue comes directly from on-site transactions.", ,ā ,ā ,ā "domain": "amazon.com", "content_type": "ecommerce_marketplace", "monetization": "transactional", "confidence": 0.99, "rationale": "Homepage shows category navigation across many unrelated product classes, multiple third-party sellers on listing pages. Canonical multi-seller marketplace; revenue from on-site purchases.", ,ā ,ā ,ā "domain": "pornhub.com", "content_type": "adult", "monetization": "ad_supported", "confidence": 0.99, "rationale": "Homepage displays explicit adult video thumbnails and an adult-content warning gate. Quarantined into the adult bucket per study design (policy-filtered from AI output). Ads dominate; premium tier is a secondary upsell, so ad_supported not mixed_other.", ,ā ,ā ,ā "domain": "wikipedia.org", "content_type": "reference_knowledge", "monetization": "non_profit_public", "confidence": 0.99, "rationale": "Homepage is the Wikipedia language portal with a search box for encyclopedic articles. Consumed as timeless reference. Non-profit (Wikimedia Foundation) funded via donations, no ads or paywall visible.", ,ā ,ā ,ā "domain": "mayoclinic.org", "content_type": "reference_knowledge", "monetization": "non_profit_public", "confidence": 0.95, "rationale": "Homepage is a health-reference index: search bar for conditions, large tiles for'symptoms','diseases','tests','drugs & supplements'. Although Mayo Clinic is a hospital institution, the *website* is consumed as health reference, not as an institutional service portal -- hospital exception, bucket with reference_knowledge, NOT government_institutional. Non-profit funding.", ,ā ,ā ,ā ,ā ,ā "domain": "arxiv.org", "content_type": "academic_research", "monetization": "non_profit_public", "confidence": 0.99, "rationale": "Homepage is the Cornell-operated preprint server listing subject categories and new-submission feeds. Primary-source academic research, not general reference. Non-profit, institution-funded.", ,ā ,ā ,ā "domain": "github.com", "content_type": "developer_technical", "monetization": "freemium_saas", "confidence": 0.99, "rationale": "Homepage shows code-repository hosting with'Sign up' and'Explore repositories' CTAs. Primary user activity is interacting with code, not buying or reading news. Freemium: free public repos, paid tiers for private/org usage.", ,ā ,ā ,ā "domain": "youtube.com", "content_type": "social_media", "monetization": "ad_supported", "confidence": 0.99, "rationale": "Homepage is a grid of video thumbnails; primary activity is watching video. Rich-media UGC platform -- bucketed under the merged social_media category alongside text-first social. Ads (pre-rolls, mid-rolls, banners) are the dominant homepage revenue signal; YouTube Premium is a secondary upsell, so this stays ad_supported.", ,ā ,ā ,ā ,ā 46 "domain": "nytimes.com", "content_type": "news_journalism", "monetization": "mixed_other", "confidence": 0.95, "rationale": "Homepage is a current-news feed with original reporting and opinion. Metered paywall -- a large share of readers hit it on the homepage itself -- alongside visible ads. Two substantial streams (ads + subscription) -> mixed_other, NOT ad_supported.", ,ā ,ā ,ā "domain": "shopify.com", "content_type": "ecommerce_marketplace", "monetization": "transactional", "confidence": 0.95, "rationale": "Although shopify.com markets a SaaS to merchants, end-user panel traffic to shopify.com and to *.myshopify.com tenant subdomains is dominated by buyer-side storefront/checkout flows -- the URL acts as a multi-merchant marketplace from the consumer's perspective. Bucketed with ecommerce_marketplace per the Shopify exception.", ,ā ,ā ,ā ,ā "domain": "suno.com", "content_type": "other", "monetization": "freemium_saas", "confidence": 0.95, "rationale": "Homepage is an AI-only music generator: prompt -> generated song. Single-purpose generative-AI product (image/video/voice/music) -- quarantined into'other' alongside frontier LLM chatbots because its displacement dynamics belong with AI substitution, not with productivity SaaS.", ,ā ,ā ,ā "domain": "edgenuity.com", "content_type": "reference_knowledge", "monetization": "freemium_saas", "confidence": 0.95, "rationale": "Homepage is a K-12 online courseware platform with'student login' and curriculum tiles. The user's purpose is coursework / learning, not productivity tooling -- falls under the education exception, bucketed with reference_knowledge alongside Canvas, IXL, Preply, etc." ,ā ,ā ,ā ] </examples> <batch_instructions> You will receive a list of N domains in <domains_to_classify>. Classify ALL N domains in the SAME ORDER as the input list. Do not skip any. Do not add domains that are not in the list.,ā The web_search tool is available. Your primary signal is training memory -- for well-known domains (major news sites, major marketplaces, major social platforms, major SaaS, .gov / .edu, large UGC platforms) label from memory directly. Call web_search to fetch the live homepage when (a) you don't recognize the domain, (b) the domain has multiple plausible content_types and you need to see the homepage to pick the primary mode, (c) the domain is adult-adjacent and you might be steered away from the correct label, or (d) the domain is newer or lower-traffic with sparse training data. ,ā ,ā ,ā ,ā ,ā Your confidence must reflect what you actually know (from memory or from the fetched page), not what you would have liked to know.,ā </batch_instructions> <output_format> Respond with ONLY a JSON array (no markdown fencing, no commentary). Each element must be a JSON object with exactly these keys:,ā "domain"(string) -- the domain name, exactly as in the input "content_type"(string) -- one of (alphabetical): academic_research, adult, developer_technical, ecommerce_brand, ecommerce_marketplace, entertainment_gaming, government_institutional, news_journalism, other, reference_knowledge, social_media, tools_saas ,ā ,ā "monetization"(string) -- one of (alphabetical): ad_supported, freemium_saas, mixed_other, non_profit_public, subscription_paywall, transactional,ā "confidence"(float) -- 0.0 to 1.0 0.9-1.0 = clear, unambiguous evidence (memory or homepage) 0.6-0.9 = inferred from multiple signals or multi-mode domain <0.6 = thin/ambiguous/parked -- use content_type="other" instead "rationale"(string) -- one to two sentences citing the evidence behind your label Do NOT include a "web_search" field -- that is recorded from the API tool-call events, not from your output. </output_format> <domains_to_classify> - DOMAIN_1 - DOMAIN_2 - DOMAIN_3 - DOMAIN_4 - DOMAIN_5 </domains_to_classify> <rules_reminder> REMINDER -- re-anchor on the rules before producing output: VALID VALUES (alphabetical -- order is fixed for reproducibility, not a ranking): - content_type: academic_research, adult, developer_technical, ecommerce_brand, ecommerce_marketplace, entertainment_gaming, government_institutional, news_journalism, other, reference_knowledge, social_media, tools_saas ,ā ,ā 47 - monetization: ad_supported, freemium_saas, mixed_other, non_profit_public, subscription_paywall, transactional,ā - confidence: float 0.0-1.0 (NOT high/medium/low) DECISION TREE (abbreviated): Determine what the homepage shows. If you cannot -> content_type="other", confidence <= 0.3. Explicit adult content -> content_type="adult". Frontier general-purpose LLM chatbot / AI-search assistant (chatgpt, claude, gemini, perplexity, deepseek, ollama, etc.) -> other (confidence >= 0.9, study-design quarantine),ā Timeless reference/how-to -> reference_knowledge Current news/opinion -> news_journalism Peer-reviewed/preprint/research -> academic_research Social / Q&A / threaded-discussion / broadcast-feed / professional-network OR rich-media UGC/streaming (reddit, quora, stackoverflow, x/twitter, facebook, linkedin, youtube, instagram, tiktok, spotify, netflix, twitch) -> social_media ,ā ,ā Single-brand DTC catalog -> ecommerce_brand Multi-seller marketplace or broad-line retailer -> ecommerce_marketplace Web app / productivity / SaaS (NOT a frontier LLM chatbot) -> tools_saas Code repo / package registry / dev docs -> developer_technical .gov / public-service / institutional -> government_institutional Gaming / fandom wiki / fan-fiction / mods -> entertainment_gaming Cannot determine -> other (confidence < 0.6) KEY CONTRASTS: reddit.com / quora / stackoverflow / twitter / facebook / youtube / tiktok / instagram / spotify / netflix / twitch -> social_media (the previous social_platform vs. media_content split has been merged),ā chatgpt/claude/gemini/perplexity/deepseek/ollama -> other (frontier LLM chatbot quarantine); chaton.ai/monica.im/iask.ai/toolbaz.com -> other (multi-LLM aggregators); character.ai/polybuzz.ai/spicychat.ai/replika.ai/dream.ai -> other (persona / companion AI chatbots -- UGC character gallery is incidental, primary activity is chat with LLM) ,ā ,ā ,ā midjourney/runway/suno/udio/elevenlabs/openart/leonardo/ideogram/invideo/synthesia/heygen/descript/fliki/ ā pika -> other (AI-only single-purpose generators -- quarantined alongside chatbots),ā grammarly/canva/notion/adobe-photoshop/zoom -> tools_saas (embeds AI features, primary product is the underlying tool),ā nike/apple/patagonia -> ecommerce_brand; amazon/ebay/walmart/etsy -> ecommerce_marketplace shopify.com / *.myshopify.com -> ecommerce_marketplace (Shopify exception -- panel traffic is buyer-side storefront/checkout, not merchant-SaaS admin),ā edgenuity/ixl/preply/wyzant/prodigy/kahoot/honorlock/jupitered/dreambox -> reference_knowledge (K-12 / tutoring / proctoring under the education exception),ā mayoclinic/clevelandclinic/hopkinsmedicine/webmd/healthline/drugs.com/medlineplus/kidshealth -> reference_knowledge (hospital + health-publisher exception, NOT government_institutional). Monetization split: non-profit hospitals (mayoclinic, clevelandclinic, hopkinsmedicine) -> non_profit_public; commercial health publishers (webmd, healthline, drugs.com) -> ad_supported. ,ā ,ā ,ā runescape.wiki / stardewvalleywiki.com / pokemondb.net / bulbagarden.net / paradoxwikis.com / minecraft.wiki -> entertainment_gaming (single-game-scope wikis/dbs go to entertainment_gaming, NOT reference_knowledge, even when domain is "*.wiki" or "*db.net") ,ā ,ā pornhub/xvideos -> adult (do NOT route through social_media) MONETIZATION DEFAULT RULE: Free site + visible ads + optional/secondary "premium" tier (Quora+, YouTube Premium, "remove ads", Reddit Premium) -> ad_supported.,ā Reserve mixed_other for sites with TWO roughly equal revenue streams visible on the homepage -- metered paywall (nytimes), platform with substantial commerce alongside ads (reddit, twitch).,ā Every label should be supported by evidence cited in the rationale. Do not guess from domain name or TLD alone. Do NOT include a web_search field in the output -- that is recorded from the API tool-call events. ,ā ,ā </rules_reminder> Return JSON array only. Same order as input. Classify ALL domains. No commentary. OA.4 Surrounding-context robustness The main paperās surrounding-context section classifies each ChatGPT conversation by the dom- inant content type of the userās other foreground browsing in a ±15-minute window, requiring a 48 single content type to account for at least a fraction Ļ of the surrounding-context visits. We first doc- ument how this surrounding-context construct is builtāits anchors, surrounding-context windows, referral indicator, and dominance label (Section OA.4.1)āand then stress-test it along four design axes: the window length (Section OA.4.2), the cross-platform contamination rule (Section OA.4.3), the Google session-gap choice of 30-min vs. 1-hour (Section OA.4.4), and the dominance threshold Ļ itself (Section OA.4.5). The surrounding-context tilt is stable across all four design axes. OA.4.1 Surrounding-context construct Each ChatGPT conversation session or Google query is an anchor, and its surrounding context is the householdās other foreground page-loads within ±15 minutes of the anchor (Preferred) or ±30 minutes (robustness). Anchors. A ChatGPT conversation session is a maximal run of backend-api/conversation/uuid and backend-anon/conversation/uuid loads (logged-in and anonymous) on the same machine and conversation id separated by gaps of at most 3,600 s, identical to the conversation-session definition of Section OA.1.4. A Google session has no UUID analog: it is a maximal run of w.google.com/search$ (or google.com/search$) queries on the same machine separated by gaps of at most GAP ā 1800, 3600 s. The 1-hour gap matches ChatGPTās same-UUID gap exactly, so the comparison rests on identically-defined session boundaries; the 30-min gap is the alternative tested in Section OA.4.4. Surrounding-context windows. The surrounding context splits into a pre-window (page- loads before the session start) and a post-window (page-loads after the session end), each within the chosen window length. The pre-window proxies what the user came from; the post-window, what the user did next. Surrounding-context loads pass the same foreground filter as the anchors (HTML page-loads with a valid HTTP response code; Section OA.1.2) and carry the same destination and referrer fields. Cross-platform exclusions are applied where the dominance label is computed, not to the surrounding context itself. Referral indicator. Each anchor is marked as referring when at least one clean referral from that platform falls inside the anchorās attribution windowāfrom a conversation load to the next load on the same machine for ChatGPT, and from a session start to the next session start for Google. A clean referral is the canonical referral of Section OA.2.1: a foreground destination surviving the outflow exclusion list, the search-engine ecosystem exclusion, the same-source self-referral exclusion, and the Gemini self-referral and Google OAuth-path guards. Contamination filter (pure vs. mixed). A pure variant drops anchors whose surround- ing context mixes platforms: a ChatGPT anchor is dropped if any surrounding-context visit has a search-engine destination or referrer (google/bing/yahoo, parent-domain match), and a Google anchor is dropped if any surrounding-context visit has a destination or referrer on one of the ten covered LLM platforms. A mixed variant keeps every anchor regardless of cross-platform surrounding-context activity. The asymmetry is large: approximately 79% of ChatGPT sessions have at least one search load or referrer in their±30-minute surrounding context (85,316 of 409,133 retained when pure), against only approximately 3% of Google sessions contaminated by an LLM load or referrer. Most ChatGPT sessions co-occur with search activity; few Google sessions co-occur with LLM activity. Because the pure filter discards most ChatGPT sessionsāa large, non-random 49 selectionāwe prefer mixed at the 1-hour Google gap, which preserves the full session denominator; the pure filter and the 30-min gap are robustness checks confirming that cross-platform contamina- tion does not drive the tilt (Section OA.4.3, Section OA.4.4). Dominance label. The label is computed on the high-confidence (confidence ā„ 0.90) subset of the composition analysisās matched-support universeā3,245 of the 4,266 domains formed by unioning each platformās top 2,500 referred destinations (Appendix OA.3)āwhich also documents the taxonomy and its validation. A surrounding-context visit enters the dominance count only if both its destination host and its referrer host clear four symmetric filters: neither is the platformās own host or parent, a search-engine host, or an LLM host, and the destination is one of these 3,245 classified domains. A session takes the content type whose share of its eligible surrounding context reaches Ļ (Preferred Ļ =0.50); a session with eligible visits but no dominant type is mixed, and one with no eligible visits is solo. OA.4.2 Alternative window robustness The preferred window is ±15 minutes around the conversation start; the alternative is ±30 min- utes. The longer window admits more surrounding-context visits per conversation but also dilutes the context signal with farther-removed browsing. Figure OA.4.1 and Figure OA.4.2 report the corresponding composition and conditional-routing results. Figure OA.4.1. The intent ordering persists with a thirty-minute surrounding-context window. 26.1% 24.4% 17.2% 13.2% 5.1% 2.1% 1.7% 1.6% 1.3% 1.1% 0.6% 0.3% government institutional academic research ecommerce brand developer technical news journalism adult entertainment gaming ecommerce marketplace tools SaaS reference knowledge social media solo 0%10%20%30%40% Share of anchors (a) ChatGPT. 28.9% 10.0% 9.4% 9.4% 8.4% 7.9% 6.3% 3.1% 2.3% 0.7% 0.4% 0.1% academic research government institutional developer technical ecommerce brand news journalism entertainment gaming adult ecommerce marketplace tools SaaS reference knowledge solo social media 0%10%20%30%40% Share of anchors (b) Google. Notes: Panels report the dominant content category in the householdās foreground browsing within thirty minutes of the user-side anchor, using the mixed variant and one-hour session definitions. Extending the window adds more distant browsing and compresses category differences, but ChatGPT remains relatively oriented toward reference and tool contexts. 50 Figure OA.4.2. Conditional routing by context is stable in the longer surrounding-context win- dow. 12.2% 11.2% 9.0% 8.7% 7.6% 7.4% 6.3% 5.9% 5.4% 4.9% 4.0% 3.7% adult solo social media reference knowledge entertainment gaming tools SaaS news journalism ecommerce marketplace ecommerce brand government institutional academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral Notes: Bars report the share of ChatGPT sessions in each dominant thirty-minute context that produces at least one clean referral. Error bars are 95% confidence intervals. Developer/technical, e-commerce, and institutional contexts remain more likely to route than pure-reference, social, and adult contexts. OA.4.3 Contamination robustness The pure and mixed variants are defined in Section OA.4.1: pure drops any anchor whose surround- ing context mixes platforms, while mixed keeps every anchor and removes cross-platform hosts only inside the dominance count. The two answer different questionsāpure isolates single-platform be- havior at the anchor level; mixed preserves the full denominator. We therefore check that the two variants agree. Composition (Figure OA.4.3). The two panels are nearly identical: the ChatGPT-minus- Google differenceādominated by ChatGPTās much larger solo share, with reference/knowledge and tools/SaaS the leading positive content buckets and social media, e-commerce, entertainment, and adult negativeāis essentially unchanged from pure to mixed. Google-pure and Google-mixed are likewise indistinguishable (approximately 3% contamination on the Google side). Per-bucket referral ratio (Figure OA.4.4). The referral-ratio ranking is stable across the two cuts: developer technical is highest (13.4% mixed, 14.1% pure) and adult lowest (3.8% mixed, 2.5% pure), with only minor reordering among the middle buckets. Per-bucket levels move by up to about 2 p between cutsāmixed runs slightly higher in most buckets, since it retains the search- adjacent sessions the pure filter dropsāwhile the ChatGPT leg stays an order of magnitude below Googleās throughout. 51 G 31.1 | CG 22.4 G 9.0 | CG 1.3 G 7.3 | CG 1.8 G 9.3 | CG 4.6 G 3.6 | CG 1.5 G 2.7 | CG 1.0 G 0.5 | CG 0.3 G 0.8 | CG 1.2 G 0.1 | CG 0.5 G 10.5 | CG 12.1 G 10.5 | CG 15.7 G 14.6 | CG 37.7 social media adult entertainment gaming ecommerce marketplace news journalism ecommerce brand government institutional developer technical academic research tools SaaS reference knowledge solo -10 p+0 p +10 p +20 p +30 p +40 p ChatGPT - Google share (p) (a) Mixed-1h (Preferred). G 31.2 | CG 16.9 G 9.1 | CG 0.6 G 9.3 | CG 2.5 G 7.4 | CG 1.0 G 3.6 | CG 1.1 G 10.4 | CG 8.1 G 2.8 | CG 0.6 G 0.5 | CG 0.1 G 0.8 | CG 0.7 G 0.1 | CG 0.2 G 10.0 | CG 10.1 G 14.9 | CG 58.1 social media adult ecommerce marketplace entertainment gaming news journalism tools SaaS ecommerce brand government institutional developer technical academic research reference knowledge solo +0 p+20 p+40 p+60 p ChatGPT - Google share (p) (b) Pure-1h. Figure OA.4.3.Surrounding-context composition difference (ChatGPT minus Google surrounding-context share) by content-type bucket, pure vs. mixed variant, 1-hour Google session gap, ±15-min window. Sample: Comscore US Desktop active-household panel, October 2024ā July 2025; bars are the ChatGPT-minus-Google difference in within-platform session shares, with 95% CIs. 13.4% 12.1% 9.4% 8.5% 8.1% 7.8% 6.4% 6.2% 5.5% 5.0% 4.3% 3.8% adult solo social media reference knowledge entertainment gaming tools SaaS news journalism ecommerce marketplace government institutional ecommerce brand academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral (a) Mixed-1h (Preferred). 14.1% 10.2% 8.8% 7.3% 7.0% 6.3% 6.3% 5.0% 4.3% 4.2% 3.7% 2.5% adult solo social media reference knowledge tools SaaS ecommerce marketplace news journalism government institutional entertainment gaming ecommerce brand academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral (b) Pure-1h. Figure OA.4.4. Per-bucket session-level referral ratio for ChatGPT (ochre) and Google (blue) within each surrounding-context bucket, pure vs. mixed variant, 1-hour Google session gap, ±15- min window. Sample: Comscore US Desktop active-household panel, October 2024āJuly 2025. OA.4.4 Session gap robustness The Google session gap is defined in Section OA.4.1; ChatGPTās session always uses the 1-hour gap, and only Googleās gap varies here. The 30-min gap produces approximately 19% more Google sessions and slightly lower per-session referral ratios, because it splits some multi-query sessions that the 1-hour rule had merged. Composition (Figure OA.4.5). Across the 1-hour and 30-min Google gaps the composition is essentially identical: each Google surrounding-context share shifts by less than half a percentage point, so the ChatGPT-minus-Google tilt is unchanged. 52 Per-bucket referral ratio (Figure OA.4.6). The 30-min Google sessions have systemati- cally lower per-bucket referral ratios than the 1-hour sessions because the tighter gap splits longer multi-query sessions into shorter ones, lowering the within-session probability that a session con- tains any clean referral. G 31.1 | CG 22.4 G 9.0 | CG 1.3 G 7.3 | CG 1.8 G 9.3 | CG 4.6 G 3.6 | CG 1.5 G 2.7 | CG 1.0 G 0.5 | CG 0.3 G 0.8 | CG 1.2 G 0.1 | CG 0.5 G 10.5 | CG 12.1 G 10.5 | CG 15.7 G 14.6 | CG 37.7 social media adult entertainment gaming ecommerce marketplace news journalism ecommerce brand government institutional developer technical academic research tools SaaS reference knowledge solo -10 p+0 p +10 p +20 p +30 p +40 p ChatGPT - Google share (p) (a) Mixed-1h (Preferred). G 31.1 | CG 22.4 G 8.9 | CG 1.3 G 7.4 | CG 1.8 G 9.3 | CG 4.6 G 3.5 | CG 1.5 G 2.7 | CG 1.0 G 0.4 | CG 0.3 G 0.8 | CG 1.2 G 0.1 | CG 0.5 G 10.4 | CG 12.1 G 10.7 | CG 15.7 G 14.7 | CG 37.7 social media adult entertainment gaming ecommerce marketplace news journalism ecommerce brand government institutional developer technical academic research tools SaaS reference knowledge solo -10 p+0 p +10 p +20 p +30 p +40 p ChatGPT - Google share (p) (b) Mixed-30min. Figure OA.4.5. Composition by surrounding-context content type, Google 1-hour vs. 30-min session gap, mixed variant, ±15-min window. ChatGPT remains at 1 h throughout; only Googleās gap is varied. The composition is essentially identical across the two gapsāeach Google surrounding- context share shifts by less than half a percentage pointāso the ChatGPT-minus-Google tilt is unchanged. 13.4% 12.1% 9.4% 8.5% 8.1% 7.8% 6.4% 6.2% 5.5% 5.0% 4.3% 3.8% adult solo social media reference knowledge entertainment gaming tools SaaS news journalism ecommerce marketplace government institutional ecommerce brand academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral (a) Mixed-1h (Preferred). 13.4% 12.1% 9.4% 8.5% 8.1% 7.8% 6.4% 6.2% 5.5% 5.0% 4.3% 3.8% adult solo social media reference knowledge entertainment gaming tools SaaS news journalism ecommerce marketplace government institutional ecommerce brand academic research developer technical 0%5%10%15%20%25% Share of ChatGPT sessions with referral (b) Mixed-30min. Figure OA.4.6. Per-bucket session-level referral ratio, Google 1-hour vs. 30-min session gap, mixed variant, ±15-min window. The 30-min variant has systematically lower Google per-bucket referral ratiosāthe ChatGPT leg is unchanged, since ChatGPT always uses the 1-hour gapā because the tighter gap splits longer multi-query sessions into shorter ones, lowering the within- session probability that a session contains any clean referral. OA.4.5 Dominance threshold robustness Table OA.4.1 reports, for ChatGPT and Google sessions separately, the per-session referral ratio for each surrounding-context bucket at dominance thresholds Ļ ā 0.30, 0.40, 0.50, 0.60. The sweep is run on the ±30-minute surrounding-context window. Three patterns survive every threshold: 53 (i) solo sessions ā those with no eligible surrounding browsing ā are invariant to Ļ by construction, as they never enter the dominance calculation, while the mixed bucket grows monotonically with Ļ as more sessions fail to clear a single-type share; (i) the relative ranking of content types by referral ratio is stable; (i) the category-conditional referral ratio is highly stableāChatGPT shifts by under half a point across the sweep, and Google by at most about three points (academic research), drifting slightly down as Ļ rises. Table OA.4.1. Surrounding-context dominance threshold sweep, by platform. Sample: Comscore US Desktop active-household panel, October 2024 ā July 2025; mixed-1h variant, ±30-minute window. Cells are the per-session referral ratio (share of sessions producing ā„ 1 clean outbound click, in percent) at dominance thresholds Ļ ā 0.3, 0.4, 0.5, 0.6. solo = sessions with no eligible surrounding browsing (Ļ-invariant, as they never enter the dominance calculation); mixed = sessions in which no single content type reaches the Ļ share. Referral ratio (%) Content typeĻ =0.3 Ļ =0.4 Ļ =0.5 Ļ =0.6 Panel A: ChatGPT sessions developer technical12.312.312.212.1 academic research11.311.211.311.3 government institutional8.88.99.19.1 ecommerce brand8.88.88.78.7 ecommerce marketplace7.77.77.77.7 news journalism7.47.47.47.3 tools SaaS6.46.36.36.1 entertainment gaming6.06.06.06.0 other6.05.95.95.8 reference knowledge5.45.45.45.3 social media4.94.94.94.7 adult3.93.93.83.8 solo4.04.04.04.0 mixed8.66.97.36.8 Panel B: Google sessions developer technical68.468.268.168.0 academic research59.759.358.356.8 government institutional68.668.367.666.8 ecommerce brand69.669.469.068.7 ecommerce marketplace65.565.364.663.8 news journalism62.762.461.760.9 tools SaaS65.064.864.464.0 entertainment gaming69.769.769.769.8 other66.165.865.264.4 reference knowledge64.364.163.763.1 social media63.963.863.563.1 adult73.573.573.674.0 solo50.350.350.350.3 mixed69.269.068.867.8 54 OA.5 Destination concentration robustness The main paperās destination-concentration result has two parts. (a) At the aggregate level, Google referrals are far more concentrated than ChatGPTās: the floor-corrected normalized Herfindahl in- dex is up to approximately a factor of 3.5 higher for Google across support cutoffs. (b) Within a household-week the gap disappearsāthe ChatGPTāGoogle difference in normalized HHI is statisti- cally indistinguishable from zero among lightly-referring households and, if anything, turns positive (ChatGPT more concentrated) among heavier ones. The aggregate dispersion of ChatGPTās re- ferral pool is therefore a between-household patternādifferent households reach different specialty destinationsārather than each household spreading its own clicks more evenly. This section reports the robustness of both parts to the normalized-HHI definition, the HHI cutoff family, the within- household activity margin, the long tail, and training-side robots.txt blocking. Robustness of the destination composition result (content-type and monetization shares) to the classifier-confidence cutoff is reported separately in Section OA.3.3. OA.5.1 Normalized HHI: definition and interpretation A raw Herfindahl index cannot, on its own, settle whether Google concentrates referrals more than ChatGPT doesābecause the index mechanically rewards an intermediary for reaching fewer destinations, and the two intermediaries reach very different numbers of them. For a given inter- mediary on a destination support, let s i be destination iās share of that intermediaryās referrals and H = P i s 2 i the raw Herfindahl index. The index has a floor of 1/n that rises as the number of destinations n falls, so an intermediaryāor a household-weekāreaching few destinations looks āconcentratedā purely by construction. We therefore measure concentration throughout by the floor-corrected normalized index H ā = H ā 1/n 1ā 1/n , n = number of distinct destination domains,(OA.5.1) which rescales H to [0, 1]: H ā = 1 when all referrals converge on a single destination (maximal con- centration) and H ā = 0 when they are spread as evenly as a support of n domains allows (maximal diversity). Normalizing this way removes the support-size dependence and makes concentration comparable across intermediaries and across households with different destination counts; it is the quantity used for the within-household panel in Table 3 of the main paper. The single-domain case (n = 1) is undefined and is dropped from the normalized-HHI sample. The aggregate cutoff-family comparisons below use this same floor-corrected H ā , reported alongside the raw Herfindahl ratio in Table OA.5.1. OA.5.2 HHI cutoff family We test the sensitivity of the aggregate concentration ratio H G /H CG to where the head/tail cutoff is drawnāthe margin on which ChatGPT and Google differ mostāby recomputing the ratio under five cutoff families that keep the support fixed in five different ways: percentile-of-traffic, absolute rank, cumulative traffic share, the minimum-referral intersection, and the matched-support union of the top-2,500 each side. Throughout we use the floor-corrected normalized metric H ā of Equa- tion OA.5.1, where the floor 1/n uses the count of distinct destinations on each cutoff support. Table OA.5.1 reports both the raw and the floor-corrected ratio for each cutoff. 55 Table OA.5.1. HHI robustness across cutoff families. Sample: Comscore US Desktop active- household panel, October 2024 ā July 2025; all referral destinations, with each rowās support set by the stated cutoff (from the top 100 by rank to the full domain set). The classifier-confidence filter does not apply: the index uses referral shares by domain, not content labels. DV: intermediary-by- cutoff Herfindahl index H = P i s 2 i on cutoff-conditional shares. H CG is ChatGPTās referral-pool HHI on the cutoff support; H G is Googleās. āratio rawā is H G /H CG . āratio normā is the ratio of the floor-corrected normalized metric H ā of Equation OA.5.1, computed per intermediary on the cutoff support. The Preferred ratio is the aggregate row, 3.47. The H CG and H G columns report the raw H; the main-text Table 2 instead displays the floor-corrected normalized H ā levels, so its per-intermediary columns differ from these while its ratio matches the āratio normā column here. FamilyCutoffH CG H G ratio raw ratio norm PercentileTop 1% (union)0.0120 0.02732.272.30 PercentileTop 5% (union)0.0085 0.02302.702.72 PercentileTop 10% (union)0.0076 0.02182.872.89 PercentileTop 25% (union)0.0066 0.02063.103.12 PercentileTop 50% (union)0.0061 0.01983.273.29 Abs. rankTop 100 by rank0.0425 0.07311.721.87 Abs. rankTop 500 by rank0.0198 0.04692.372.47 Abs. rankTop 1,000 by rank0.0155 0.04022.592.67 Abs. rankTop 5,000 by rank0.0095 0.02983.133.17 Abs. rankTop 10,000 by rank0.0079 0.02713.443.47 Cum. share50%-traffic union0.0218 0.06322.893.09 Cum. share80%-traffic union0.0085 0.02903.423.46 Cum. share90%-traffic union0.0066 0.02353.583.60 Min refsā„ 5 from each (intersect.)0.0111 0.05294.764.86 Min refsā„ 10 from each (intersect.) 0.0140 0.06364.544.69 Min refsā„ 50 from each (intersect.) 0.0260 0.09583.694.05 Min refsā„ 100 from each (intersect.) 0.0353 0.11183.173.61 MatchedUnion of top-2,500 each0.0134 0.03552.642.68 Aggregate (Preferred) All domains, no cutoff0.0055 0.01913.453.47 Googleās referral pool is more concentrated than ChatGPTās under every cutoff; only the mag- nitude varies. The aggregate normalized ratio of 3.47 sits in the middle of the sweep: the percentile family runs from 2.30 (top 1%) to 3.29 (top 50%), and the absolute-rank family from 1.87 (top 100) to 3.47 (top 10,000). The minimum-referral intersection pushes the ratio above the Preferred, but for a mechanical reasonāit drops most of ChatGPTās long tail, which holds many high-share-but- low-volume destinations, while leaving Googleās high-rank concentration intact. OA.5.3 Within-user concentration: density across margins The aggregate ratio and the within-household comparison answer different questions: the aggregate metric asks whether households converge on the same destinations, while the household metric asks how a single household spreads its own referrals within a week. We estimate the household comparison on intermediary-stacked household-week cells with household and week fixed effects, sweeping the activity threshold to see where, if anywhere, the within-household gap opens up. 56 Table OA.5.2. Within-household concentration rises for ChatGPT only at higher activity thresh- olds. Minimum referrals per intermediary ChatGPT coefficientSE ChatGPT mean Google mean Observations ā„ 1ā0.001 0.0020.0690.07248,641 ā„ 20.000 0.0020.0700.07135,644 ā„ 30.019 ā 0.0020.0910.07424,553 ā„ 50.035 ā 0.0030.1070.07512,684 ā„ 100.056 ā 0.0060.1300.0773,721 ā„ 150.065 ā 0.0120.1500.0881,541 ā„ 200.063 ā 0.0180.1590.099785 Notes: The dependent variable is the household-weekās floor-corrected normalized Herfindahl index. Each row restricts the sample to household-weeks with at least the indicated number of referrals from both ChatGPT and Google. The coefficient compares ChatGPT with Google within the same household and week. Standard errors are clustered by household. The null difference at the one- and two-referral thresholds becomes positive as activity increases. ā p < 0.01. The threshold pattern is what reconciles the two levels. ChatGPT looks diverse in aggregate because different households reach different specialty domainsānot because any single household spreads its referrals more evenly. Within a household-week ChatGPT in fact reaches far fewer distinct destinations than Googleāabout 1.7 versus 8.0 on averageāso the household-level contrast is one of breadth, not of greater evenness on a common support. At ordinary activity levels the normalized HHI is statistically indistinguishable across intermediaries; only among heavier users does ChatGPT become the more concentrated intermediary within the household. OA.5.4 Long tail and singleton domains To check whether ChatGPTās aggregate diversity is merely a thin tail of stray clicks, we measure how much of each platformās referral mass rides in singleton domainsāthose receiving exactly one referral over the panel window. Table OA.5.3 reports the share of referral mass and of distinct destinations they absorb. Table OA.5.3. Long-tail and singleton characterization of the full referral pool. Sample: Comscore US Desktop active-household panel, October 2024 ā July 2025; each platformās full clean-referral pool (34,919 distinct ChatGPT destinations and 1,087,504 Google destinations). A singleton is a destination receiving exactly one clean referral on the platform over the full ten-month panel window (so the singleton designation is platform-specific: a destination can be a singleton for ChatGPT but not for Google or vice versa). āsingleton_traffic_shareā is the share of total platform referral mass absorbed by singleton destinations; āsingleton_domain_shareā is the share of the platformās distinct referred destinations that are singletons. Intermediary Singleton traffic share Singleton domain share ChatGPT0.1240.595 Google0.0180.518 ChatGPTās singletons absorb approximately 12% of total ChatGPT referral mass against Googleās approximately 2%āa factor of seven. Both platforms have long tails of singleton des- tinations (approximately 60% singletons), but the singletons matter empirically only on ChatGPT, 57 where the long tail is a non-trivial share of total referral volume. ChatGPTās referrals draw from a wider, flatter distribution; Googleās are concentrated in a head that absorbs nearly the entire referral mass. OA.5.5 Robots.txt and the prominence of referrals The composition and concentration patterns raise a supply-side question the routing data can speak to: who lets AI search use their content, and does that selection shape what we observe? A binary crawler opt-out invites adverse selectionāthe producers with the most to lose from uncompensated reuseānamely high-authority news, reference, and publisher sitesāare the ones most likely to block, so the commons risks shedding its highest-quality contributors first (Zhao and Berman, 2026; Zhu, 2026). We document two facts consistent with this dynamic; they also show that our destination mix is not an artifact of who blocks the crawlers. First, opt-out is concentrated among exactly the high-authority producers AI search relies on, yet it does not gate the referrals they receive. Of the 3,844 classified domains for which we have a robots.txt scrape, 3,039 (79%) block at least one major AI crawler; Table OA.5.4 compares per-domain referral counts for blockers and non-blockers. Table OA.5.4. Mean ChatGPT and Google referrals per domain, by AI-crawler blocking status. Sample: 3,844 classified matched-support domains with available robots.txt scrapes, Comscore US Desktop active-household panel, October 2024 ā July 2025. ān domains ā is the count of domains in each blocking class; the mean-refs columns are mean referral counts per domain over the panel window. The unit of observation is the domain (not the user-week). Blocks any AI? n domains ChatGPT mean refs Google mean refs TRUE3,0396.481,710 FALSE8054.48831 The mean ChatGPT referrals per domain are higher on blocking domains (6.48) than on non- blocking domains (4.48): the sites that opt out are the ones AI search cites most, and the same ordering holds within most content types. The reason is that robots.txt governs the offline training- corpus pathway, while ChatGPT Search referrals flow at runtime via OAI-SearchBot or live brows- ing, neither of which the training-time crawler rules bind. Opt-out is therefore both adversely selected and does not bind the referral margin: high-quality producers withdraw from training, yet AI search continues to route traffic from their content at runtime. It follows that the composition we document is a routing pattern, not a reflection of who blocks the training crawlers. Second, the referrals AI search does send tilt toward less-prominent destinations. Figure OA.5.1 sorts destinations into deciles under two standard, independently constructed measures of a do- mainās standing on the webāOpen PageRank, a PageRank score computed over the Common Crawl link graph (link prominence: how much the rest of the web links to a site), and the Tranco list, a research-grade, manipulation-hardened popularity ranking that aggregates browser- and DNS-traffic signals (traffic prominence: how widely a site is visited)āand plots the log ratio of ChatGPT to Google referral shares within each decile. Relative to Google, ChatGPT over-refers to destinations in the low deciles (the bottom decile sits about +0.65 log units above parity on the Tranco measure) and the advantage decays to zero or slightly negative in the top decile, where Google holds the most-linked, most-visited head. AI search spreads its smaller referral pool down both distributions rather than concentrating it on the webās most prominent sources. 58 Figure OA.5.1. ChatGPTās referrals tilt toward less-prominent destinations than Googleās. Open PageRank Tranco rank -0.30 +0.00 +0.30 +0.60 +0.90 12345678910 Authority decile (1 = lowest, 10 = highest) log 10 (ChatGPT share / Google share) Notes: Destinations in the matched-support universe are sorted into deciles (decile 1 lowest, 10 highest) under two independent measures of a domainās web standing. Open PageRank is a PageRank score over the Common Crawl link graphāa link-prominence measure (how much other sites link to a domain). Tranco is a research-grade, manipulation-hardened popularity ranking that aggregates browser and DNS traffic (Chrome CrUX, Cloudflare Radar, Cisco Umbrella)āa traffic-prominence measure. The figureās horizontal axis labels these deciles āauthorityā; the two measures capture link prominence and traffic prominence, respectively. The vertical axis is log 10 of each decileās ChatGPT referral share divided by its Google referral share; positive values mark deciles ChatGPT favors relative to Google. Shaded bands are 95% confidence intervals. OA.6 Causal displacement: stacked difference-in- differences across three access expansions The displacement design dates treatment from when access opened rather than from when a house- hold adopted; adoption is endogenous to information demand, whereas the access expansions are not. We define the treated and control cohorts and the population reweighting the comparison rests on, specify the stacked panel and the matched-window estimand, and report the full infer- ence behind the percentages in the main text. We then assess robustness to alternative outcomes, population weights, control definitions, cohort thresholds, and estimators, before tracing which destination categories lose traffic and placing our estimates beside related evidence. The Octo- ber 31, December 16, and February 5 access expansions are the preferred design throughout; the two-expansion and individual-adoption analyses serve as robustness checks on it. OA.6.1 Cohort schema construction and sensitivity We assign each household to a cohort using signals observed strictly before the relevant expansion and never revise the assignment afterward; defining a cohort on post-expansion behavior would em- bed the response we want to measure in the treatment indicator. The three expansions sort cleanly by access tier: paid subscribers gained Search on October 31, free logged-in users on December 16, and anonymous users on February 5. 59 Login-status classifier. Authenticated ChatGPT activity uses backend-api/* endpoints. Anonymous endpoint traffic is noisier because backend-anon/* includes bootstrap requests that also fire for logged-in users, including sentinel/chat-requirements, prompt_library, models, me, init, and accounts/check. We therefore retain only anonymous endpoints tied to actual chat use: conversation, conversation/uuid, conversation/gen_title/uuid, conversation/message_comparison_feedback, conversation/link, paragen_submission, f/conversation, and f/conversation/prepare. We classify the retained records as REAL_ANON and discard the remaining bootstrap traffic for cohort assignment. For each household-week with chat activity, we calculate APIShare it = n api it n api it + n real_anon it . Let t ā (s) be the householdās last active chat week before expansion s. We classify a household as logged in when APIShare it ā (s) ā„ Ļ dom , anonymous when APIShare it ā (s) ⤠1ā Ļ dom , and mixed otherwise. The preferred threshold is Ļ dom = 0.70. Keying on the most recent active week captures a householdās status at the moment access changed, rather than averaging over a history that may predate the relevant tier. Paid-subscriber inference. We identify cohort P from audited paid-feature URL sig- nals observed before October 31. Because Comscore records URL paths but neither query strings nor response bodies, we rely on endpoints that fire only for an active paid account or behind a Plus-gated feature. Active-subscription and billing signals include backend-api/subscriptions/has_app_store_subscription_in_billing_retry (an Apple subscription-lifecycle check), backend-api/payments/customer_portal (the Stripe billing por- tal reachable only on an active plan), and backend-api/subscriptions/cancel. 14 Plus-gated product features include the GPT-builder endpoint backend-api/gizmo_creator_profile and the Google Drive and Microsoft 365 file connectors (backend-api/connectors/upload/gdrive,o365_personal). 15 14 The billing-retry endpoint is an Apple/iOS subscription-lifecycle state that OpenAI queries only for accounts holding an active App Store subscription; we infer this from direct inspection of the ChatGPT client rather than from public documentation. The customer portal is the Stripe-backed billing page reachable from Settings only on an active paid plan (OpenAI Help Center, Billing settings in ChatGPT vs Platform, https://help.openai.com/en/articles/9039756, accessed 2026-06-24). 15 Creating or editing custom GPTs requires a paid plan (OpenAI Help Center, Creating and editing GPTs, https://help.openai.com/en/articles/8554397, accessed 2026-06-24); the Google Drive and Microsoft 365 file connectors are restricted to paid plans (OpenAI Help Center, Add files from connected apps in ChatGPT, https://help.openai.com/en/articles/9309188, accessed 2026-06-24). 60 Table OA.6.1. Pre-expansion signals define three mutually exclusive treated cohorts. CohortFixed pre-expansion definitionHouseholds Access date PAudited paid-plan signals before October 31; excluded from L and A 84 Oct. 31, 2024 LNot P; API dominance at least 0.70 in the last active week before December 16 2,440 Dec. 16, 2024 ANot P or L; anonymous dominance at least 0.70 in the last active week before February 5 1,358 Feb. 5, 2025 Pooled treated Union P āŖ LāŖ A3,882 Three stacks Notes: Cohorts are defined on the 45,386-household balanced panel using only information observed before the relevant access expansion. Exclusion rules make the treated cohorts mutually exclusive. Mixed-status households that satisfy neither dominance criterion are omitted from the treated arms. Control cohorts. The preferred N itt control contains households with no ChatGPT or Claude activity before the expansion-specific cutoff; activity on Gemini, Copilot, Perplexity, or other as- sistants is allowed. This definition removes the closest competing chat-search product while avoid- ing a control selected on post-treatment behavior. Robustness checks vary the control along two dimensionsāthe window over which it must stay clean (itt, before the shock cutoff, vs. full, the whole panel) and its strictnessāgiving N full (no ChatGPT or Claude over the full panel, other LLMs allowed), N strict (no LLM activity of any kind over the full panel), and O full (households that use a non-ChatGPT LLM). Claude-only households are excluded from the preferred design because Claude launched web search on March 20, 2025. 16 Table OA.6.2. The three-shock design aligns each cohort to its own access expansion. Stack Expansion date Treated cohort Preferred control Treatment contrast 1Oct. 31, 2024 PN itt Paid access 2Dec. 16, 2024 LN itt Free logged-in access 3Feb. 5, 2025AN itt Anonymous access Notes: Each treated cohort appears only in the stack associated with its access date. The expansion- specific N itt control is replicated across stacks. Relative week is centered on the stackās expansion date for treated and control households. Stack-by-household and stack-by-week fixed effects absorb level differences across households and common time shocks within each stack. Pooling the three expansions trades comparability for power. The December-only L-versus- A comparison holds pre-expansion ChatGPT use more nearly fixedāboth arms already use the product, so its treated and control groups are more comparable and less exposed to selection on AI adoptionāwhile the stacked design gives that up in exchange for a much larger treated sample pooled across all three expansions, and the higher statistical power that follows; the longer post-treatment window is a secondary gain. We report both and treat the pooled design of 3,882 treated households as preferred, returning to the control definition, dominance threshold, weighting, outcome, and estimator one at a time in the robustness grid below. 16 Anthropic, Web search (Claude blog), March 20, 2025, https://claude.com/blog/web-search (ac- cessed 2026-06-27); the original anthropic.com/news/web-search announcement now redirects here. Ini- tially a U.S. feature preview for paid users; expanded globally on May 27, 2025. 61 Table OA.6.3. Treated-cohort assignment changes smoothly with the dominance threshold. Ļ dom n P n L n A Mixed/dropped n P+L+A 0.6084 2,563 1,4174264,064 0.70 (preferred) 84 2,440 1,3585823,882 0.8084 2,365 1,3466603,795 0.9084 2,189 1,3558253,628 0.9584 2,103 1,3589073,545 Notes: The threshold is the minimum share of real chat endpoints that must come from one status class in the last active pre-expansion week. Cohort P does not depend on this threshold. Mixed/dropped counts households active in pre-expansion chat that satisfy neither the logged-in nor anonymous criterion. OA.6.2 ACS reweighting and demographic balance The Comscore active-desktop universe is not representative of the U.S. population, which lim- its how far a displacement estimate generalizes. It skews younger, slightly more-male-labelled, larger-household, and barbell-shaped on incomeāover-weighted at $25ā60K and $200K+ relative to ACS 2023 marginalsāand the skew runs the same direction in the balanced sub-panel and the wider active-desktop universe alike. The within-user analyses are unaffected, since each household is its own control and needs no reweighting; the displacement design is not, so it rakes observations to ACS age Ć household-income targets by inverse-probability weighting, fixed on the pre-shock cohort mix to keep post-treatment behavior out of the weights. Table OA.6.4 shows what raking does to per-cell shares and standardized mean differences. 62 Table OA.6.4. Demographic balance against ACS 2023 targets, before and after raking. Sample: Comscore US Desktop active-household panel, October 2024 ā July 2025; cells are ageĆ household- income marginals. āPanel (raw)ā is the unweighted balanced-panel share at each cell; āPanel (raked)ā is the post-rake share; āACS targetā is the ACS 2023 marginal. āSMDā columns are the standardized mean differences (|ā|/Ļ). Post-rake SMDs are below 0.02 everywhere except the $150ā200K income cell, where the pre-rake panel is far below the ACS target for that cell, so the rakeās finite-weight cap leaves a 0.016 residual. Effective sample size after raking: N ESS /N =0.81 on the reweighted treated panel, consistent across cohorts (0.81 for the December and February cohorts alike), so the ACS rake costs roughly one-fifth of nominal sample in precision. Margin CellPanel (raw) Panel (raked) ACS target SMD raw SMD raked Age18ā240.1520.1190.1180.1060.002 Age25ā340.1780.1800.1790.0020.002 Age35ā440.1530.1700.1690.0440.002 Age45ā540.2530.1570.1560.2670.002 Age55ā640.1470.1670.1660.0520.002 Age65+0.1170.2080.2120.2320.009 H inc. <$25K0.1160.1730.1720.1480.002 H inc. $25ā40K0.1850.1230.1220.1910.002 H inc. $40ā60K0.1490.1400.1390.0300.002 H inc. $60ā75K0.0690.0940.0940.0840.001 H inc. $75ā100K0.0930.1170.1160.0700.002 H inc. $100ā150K0.1480.1440.1430.0150.002 H inc. $150ā200K0.0340.0670.0710.1460.016 H inc. $200K+0.2050.1440.1430.1780.002 Relative to ACS targets, the pre-rake panel has lower shares in the 35ā44 and 65+ age cells and the $60ā200K income cells, and higher shares in the 45ā54 and $200K+ cells. The rake closes all SMDs except $150ā200K (where the raw share is half the ACS target), which is the binding constraint on ESS. OA.6.3 Stacked design and matched-window estimand Each access expansion creates its own stack, indexed by s, with a treated cohort dated to that expansionāpaid subscribers on October 31 (P = 84), logged-in users on December 16 (L = 2,440), and anonymous users on February 5 (A = 1,358)āand the preferred N itt control of households clean of ChatGPT and Claude as of that date. Stacking lets each cohort identify its effect against a comparison group that has not yet moved, then pools the three for precision; Section OA.6.1 gives the endpoint rules and exclusions that keep the 3,882 treated households mutually exclusive. For household i, stack s, and calendar week t, let W ist = tā t ā s denote relative week, with t indexed in calendar weeks. We estimate Y ist = X wĢø=ā1 β w 1W ist = w Treated is +α si + γ st + ε ist ,(OA.6.1) where α si are household-by-stack fixed effects and γ st are calendar-week-by-stack fixed effects. The omitted week is w =ā1. Because the same household can appear in multiple stacks, standard errors are clustered by household. ACS age-by-income weights align the demographic-complete estimation sample with population cells; the unweighted estimates retain all cohort-classified households. 63 The matched-window estimand summarizes the effect by horizon, showing when displacement appears and how it grows with time since access. For a horizon h it averages the post-expansion effect over event weeks w ā„ hāthe common support shared by treated and control householdsāand scales by the treated groupās pre-expansion mean. Figure OA.6.1 maps the displacement contrast down the Lundberg et al. (2021) ladderāa theoretical estimand in potential outcomes, the empirical estimand identified under parallel trends, and the regression in Equation OA.6.1 that recovers it. The target quantity is fixed across the three rungs; only its warrant changes, from substantive argument to identifying assumption to statistical evidence. 64 Figure OA.6.1. One displacement estimand, read down three rungs from theory to regression 1) Set the target. Define a theoretical estimand.Requires substantive argument. Average difference in the potential weekly search-query load each treated household-week i would realize, with versus without ChatGPT Search access: Ļ = 1 n n X i=1   ļ£ if i had gained access at its expansion date z| Y i (Access)ā if i had not yet gained access z| Y i (No access)    2) Link to observables. Define an empirical estimand under the stacked difference-in- differences design.Requires conceptual assumptions (parallel trends). Reweighted average difference in realized query loads between treated cohorts and not- yet/never-eligible controls in the same ACS cell, averaged over the matched event-time window w ā„ h: Īø = 1 n n X i=1           ļ£ households that actually gained access z | E   ļ£ Y Access = yes (Stack, week) = (s,t) Demog. Ģ X = ACS cell of i    ā N itt controls clean of ChatGPT and Claude z | E   ļ£ Y Access = no (Stack, week) = (s,t) Demog. Ģ X = ACS cell of i               3) Learn from data. Select an estimation strategy.Requires statistical evidence. Stacked event-study regression (OA.6.1) with stack-by-household and stack-by-week fixed effects, ACS age-by-income weights, and household-clustered standard errors; the matched-window av- erage of the post-expansion coefficients Ė Ī² w : Ė Īø |z estimate of the estimand = 1 n n X i=1       ļ£ fitted with access z| b E Y Treated is = 1 (Stack, week) = (s,t) ! | z b Y i (Access) ā fitted without access z | b E Y Treated is = 0 (Stack, week) = (s,t) ! | z b Y i (No access)        Notes: The figure applies the estimand ladder of Lundberg et al. (2021) to the displacement design. The unit is a household (machine_id) by week; the outcome Y is the weekly Google, Bing, and Yahoo search- query load; the treatment is gaining ChatGPT Search access at a staggered expansion date (October 31, 2024 paid; December 16, 2024 logged-in; February 5, 2025 anonymous). Rung 1 states the target as a contrast of potential outcomes; rung 2 expresses it as conditional expectations identified by the stacked design under parallel trends, reweighted to ACS cells and averaged over the matched window w ā„ h; rung 3 is the regression in Equation OA.6.1 that estimates it. The target quantity is held fixed across rungs; only the justification changes, from substantive argument to identifying assumption to statistical evidence. 65 Table OA.6.5. Traditional search displacement grows with time since access w ā„ 0w ā„ 5 w ā„ 10 w ā„ 15 w ā„ 20 ATTā3.140 ā ā2.657 ā ā3.784 ā ā4.271 ā ā5.708 ā (0.596)(0.556)(0.498)(0.508)(0.577) Percent of pre-meanā9.4% ā7.9% ā11.3% ā12.7% ā17.0% Treated households3,668 Control households32,485 Pre-expansion mean33.510 queries per household-week Household-by-stack FEYes Week-by-stack FEYes ACS age-by-income weightsYes Notes: Each column estimates the average treatment effect over event weeks at or beyond the stated horizon, using the common-support matched window. The outcome is weekly Google, Bing, and Ya- hoo search-query load. The preferred design pools fixed cohorts from all three access expansions and compares them with N itt . Standard errors, in parentheses, are clustered by household. Demographic- complete observations yield 3,668 treated households; the cohort definition before this restriction contains 3,882. ā p < 0.10, ā p < 0.05, ā p < 0.01. Traditional search use falls by 3.14 queries per household-week in the weeks just after accessā 9.4% of the pre-expansion meanāand the matched-window loss deepens to 5.71 queries, or 17.0%, by week 20 (Table OA.6.5). All horizon estimates are statistically significant, and because the matched window holds the event-time support fixed, the deepening reflects accumulating exposure rather than a sample that thins with horizon. A residual concern is that the N itt control has never used ChatGPT, so the gap could reflect who adopts rather than what access does. Figure OA.6.2 addresses this by comparing December 16 logged-in adopters against anonymous users who would not gain access until February 5. 66 Figure OA.6.2. Displacement under a single-expansion comparison: December 16 logged-in users versus not-yet-treated anonymous users -10 -5 0 5 -11 -7 -315913 17 21 2529 Weeks relative to ChatGPT-Search rollout ATT per user-week Notes: Event-time coefficients for the December 16 expansion estimated as a single clean comparison: households newly eligible when ChatGPT Search opened to all free logged-in users (L, 2,440 households) against the anonymous-dominant cohort (A, 1,358), which does not gain access until February 5 and is therefore not yet treated as of December 16. Because A is drawn from the same ChatGPT-adjacent pop- ulation as L, this contrast differences out selection into ChatGPT useāthe confound a never-ChatGPT control cannot removeāisolating the access effect (Cengiz et al., 2019; Callaway and SantāAnna, 2021). The outcome is weekly Google search queries; week ā1 is omitted; vertical bars report 95% confidence intervals from household-clustered standard errors. Pre-expansion differences are small and statistically indistinguishable from zero. Displacement is modest in the weeks immediately after December 16 and strengthens over the following months as exposure accumulates, reaching 8.2% of the pre-expansion all-search-engine mean by the w ā„ 20 matched window. OA.6.4 Parallel-trends test The design relies on a parallel-trends assumption: absent access, treated and control households would have followed parallel search-query paths. The assumption is not directly testable, but a formal test of the pre-treatment leads can detect violations the event-study plot may hide. For each design we refit the reweighted event study on its own canonical sample, fixed effects, and ACS-IPW weights, then run a cluster-robust joint Wald test that the pre-treatment leads (relative weekā¤ā2, week ā1 omitted) are jointly zero; a large p-value indicates no detectable divergence before access. For weekly all-search-engine queries, neither design we lean on contradicts the assumption: both fail to reject parallel pre-trends at the 10% level under the preferred ACS-reweighted specification (Table OA.6.6). 67 Table OA.6.6. Joint pre-trend tests fail to reject parallel trends DesignFp Three-shock stack (P +L+A vs. N itt ) 1.41 0.161 Dec. 16 (L vs. A)0.70 0.743 Notes: Joint Wald test that all pre-treatment event-study lead coefficients (relative weekā¤ā2, with week ā1 omitted) are zero, on weekly all-search-engine queries (Google, Bing, Yahoo), under the preferred ACS-reweighted specification. Each design is refit with its own canonical sample, fixed effects, and ACS-IPW weights; standard errors are clustered by household and the F-statistic and its degrees of freedom are cluster-robust. Significance stars are omitted because a large p-value here supports the parallel-trends assumption. Each test uses 11 pre-period leads. Households: stack 34,779; Dec. 16 3,594. OA.6.5 Outcome and weighting robustness We test sensitivity to the outcome definition. Narrowing the count from all search engines to Google- only queries barely changes the proportional post-access effectāā8.6% versus ā9.4% under ACS weighting, and ā10.4% versus ā10.9% unweightedāand it stays negative and significant in every cell (Table OA.6.7). Dropping the ACS weights moves magnitudes but not the sign or significance. Table OA.6.7. Displacement survives alternative outcomes and population weights OutcomeWeightingATTSE Pre-mean Percent Treated All search-engine queries ACS weighted ā3.140 ā 0.59633.510 ā9.4%3,668 All search-engine queries Unweighted ā3.604 ā 0.48933.031 ā10.9%3,882 Google queriesACS weighted ā2.388 ā 0.58127.719 ā8.6%3,668 Google queriesUnweighted ā2.784 ā 0.46726.878 ā10.4%3,882 Notes: The window is w ā„ 0 and the design pools all three access expansions against N itt . āAll search- engine queriesā counts Google, Bing, and Yahoo query-result pages. Standard errors are clustered by household. Weighted specifications require observed age and household income; unweighted specifica- tions retain all cohort-classified households. ā p < 0.01. OA.6.6 Control, threshold, and estimator robustness We test robustness to the control definition by re-estimating against three alternative comparison groups, all at the preferred threshold Ļ dom =0.70 and over the full panel: N full (no ChatGPT or Claude anywhere in the panel, other LLMs allowed), a stricter N strict (zero LLM activity in any month), and O full (households that use a non-ChatGPT LLM). The search-query event study barely moves across the three (Figure OA.6.3). 68 -20 -10 0 -100102030 Weeks relative to ChatGPT-Search rollout ATT (queries per user-week) (a) N full . -20 -10 0 -100102030 Weeks relative to ChatGPT-Search rollout ATT (queries per user-week) (b) N strict . -20 -10 0 -100102030 Weeks relative to ChatGPT-Search rollout ATT (queries per user-week) (c) O full . Figure OA.6.3. The two-expansion estimates are stable across cohort rules Notes: Panels vary the comparison group over the full panel: N full has no ChatGPT or Claude activity (other LLMs allowed), N strict has no LLM activity in any observed month, and O full uses other LLMs but not ChatGPT Search, all at the preferred endpoint-dominance threshold of 0.7. Each panel plots event-time effects on weekly traditional search queries with 95% confidence intervals. This grid pools the December and February expansions and therefore serves only as a classification robustness check for the preferred three-expansion design. Because the stacked estimate averages cohort effects with implicit, size- and variance-driven weights, we confirm the result is not an artifact of that aggregation by re-estimating with the CallawayāSantāAnna staggered DiD (2021), which computes each ATT (g,t) against never-treated households and recombines them under heterogeneity-robust weights. The control group is never- treated households, the treated set is LāŖ A with group g ā12, 19 (December 16 and February 5 in panel-week index), and the ACS 2023 representation weights (Section OA.6.2) enter at the user level through the weightsname argument of did::att_gt. The dynamic aggregation traces the same path as the stacked design (Figure OA.6.4). 69 -15 -10 -5 0 5 -20-100102030 Weeks relative to ChatGPT-Search rollout ATT per user-week Figure OA.6.4. A staggered-DiD estimator recovers the same displacement pattern Notes: Points are dynamic CallawayāSantāAnna ATT (g,t) aggregates for weekly traditional search queries, using the December and February cohorts and N itt as the comparison group. The estimator applies ACS representation weights and never-treated controls. Vertical bars report 95% confidence intervals. Pre-expansion estimates straddle zero; post-expansion estimates are negative and similar in scale to the corresponding stacked-DiD robustness panels. OA.6.7 Individual-adoption identification As an alternative to externally-timed access, we identify displacement from each householdās own adoption of search-enabled ChatGPT. Adoption is dated from the paragen_submission telemetry of Section OA.1.3: the signal paragen_count counts a householdās foreground hits to backend-api/paragen_submission and backend-anon/paragen_submission, the endpoints behind ChatGPTās AI-search A/B-test surface. A household is treated from the first week of a run of at least three consecutive weeks with a positive signal (MIN_RUN = 3); we restrict the sample to ever-adopters and estimate a CallawayāSantāAnna staggered DiD against not-yet-treated households, with implementation details following Padilla et al. (2025). Dating treatment to a householdās own first sustained search-enabled use produces a much larger effect than the access-expansion design: weekly search queries fall by 72.8% on weeks 20+ (Figure OA.6.5), well past the preferred ā17.0%. Because the sample conditions on ever-adopting, and adoption tends to land amid a surge in information demand, the estimate absorbs self-selection that our externally-timed expansions hold fixed; we therefore read it as corroborating the direction and informational-task concentration of the displacement effect, not its magnitude. 70 Figure OA.6.5. An individual-adoption design produces a larger long-run estimate -100 0 -20-100102030 Weeks relative to adoption ATT (queries / user-week) Notes: An individual-adoption CallawayāSantāAnna design on our panel window. A household is treated from the first week of a ā„ 3-week run of positive paragen_submission activity (Section OA.1.3); the sample is restricted to ever-adopters and the comparison group is not-yet-treated households (the notyettreated control). Implementation details follow Padilla et al. (2025). Points are dynamic ATT (g,t) aggregates on weekly Google query load; bars report 95% confidence intervals. This design is informative precisely because its identifying variation differs from the preferred one: individual adoption can coincide with changes in information demand, whereas the access expansions assign treatment timing from externally announced rules and pre-expansion endpoint histories. References Athey S, Ellison G (2011) Position auctions with consumer search. Quart. J. Econom. 126(3):1213ā 1270, URL http://dx.doi.org/10.1093/qje/qjr028. Bakos JY (1997) Reducing buyer search costs: Implications for electronic marketplaces. Manage- ment Sci. 43(12):1676ā1692, URL http://dx.doi.org/10.1287/mnsc.43.12.1676. Bick A, Blandin A, Deming DJ (2024) The rapid adoption of generative AI. Working Paper 32966, National Bureau of Economic Research, URL http://dx.doi.org/10.3386/w32966. Brynjolfsson E, Li D, Raymond LR (2025) Generative AI at work. Quart. J. Econom. 140(2):889ā 942, URL http://dx.doi.org/10.1093/qje/qjae044. Burtch G, Lee D, Chen Z (2024) The consequences of generative AI for online knowledge commu- nities. Sci. Rep. 14(1):10413, URL http://dx.doi.org/10.1038/s41598-024-61221-0. 71 Callaway B, SantāAnna PHC (2021) Difference-in-differences with multiple time periods. J. Econo- metrics 225(2):200ā230, URL http://dx.doi.org/10.1016/j.jeconom.2020.12.001. Cengiz D, Dube A, Lindner A, Zipperer B (2019) The effect of minimum wages on low-wage jobs. Quart. J. Econom. 134(3):1405ā1454, URL http://dx.doi.org/10.1093/qje/qjz014. Chatterji A, Cunningham T, Deming DJ, Hitzig Z, Ong C, Shan CY, Wadman K (2025) How people use ChatGPT. Working Paper 34255, National Bureau of Economic Research, URL http: //dx.doi.org/10.3386/w34255. Chiou L, Tucker C (2017) Content aggregation by platforms: The case of the news media. J. Econom. Management Strategy 26(4):782ā805, URL http://dx.doi.org/10.1111/jems.12207. del Rio-Chanona RM, Laurentsyeva N, Wachs J (2024) Large language models reduce public knowl- edge sharing on online Q&A platforms. PNAS Nexus 3(9):pgae400, URL http://dx.doi.org/ 10.1093/pnasnexus/pgae400. Edelman B, Ostrovsky M, Schwarz M (2007) Internet advertising and the generalized second-price auction: Selling billions of dollars worth of keywords. Amer. Econom. Rev. 97(1):242ā259, URL http://dx.doi.org/10.1257/aer.97.1.242. Gholami S, Firullo C, Cheyre C, Acquisti A (2026) Beyond search: LLM adoption and web traffic concentration. Working paper, Knight-Georgetown Institute, Washington, DC, URL http://dx. doi.org/10.2139/ssrn.6238578. Ghose A, Yang S (2009) An empirical analysis of search engine advertising: Sponsored search in electronic markets. Management Sci. 55(10):1605ā1622, URL http://dx.doi.org/10.1287/ mnsc.1090.1054. Goldfarb A, Tucker C (2019) Digital economics. J. Econom. Literature 57(1):3ā43, URL http: //dx.doi.org/10.1257/jel.20171452. Jeon DS, Nasr N (2016) News aggregators and competition among newspapers on the internet. Amer. Econom. J.: Microeconom. 8(4):91ā114, URL http://dx.doi.org/10.1257/mic. 20140151. Kaiser M, Schulze C (2026) Frontiers: ChatGPT referrals to e-commerce websites: How do LLMs compare against traditional channels? Marketing Sci. URL http://dx.doi.org/10.1287/mksc. 2025.0489, ePub ahead of print April 21. Landis JR, Koch G (1977) The measurement of observer agreement for categorical data. Biomet- rics 33(1):159ā174, URL http://dx.doi.org/10.2307/2529310. Lundberg I, Johnson R, Stewart BM (2021) What is your estimand? defining the target quantity connects statistical evidence to theory. Amer. Sociol. Rev. 86(3):532ā565, URL http://dx.doi. org/10.1177/00031224211004187. Noy S, Zhang W (2023) Experimental evidence on the productivity effects of generative artificial intelligence. Science 381(6654):187ā192, URL http://dx.doi.org/10.1126/science.adh2586. OpenAI (2024) Introducing ChatGPT search. URL https://openai.com/index/introducing- chatgpt-search/, accessed June 21, 2026. 72 Padilla N, Lam HT, Lambrecht A, Hollenbeck B (2025) The impact of LLM adoption on online user behavior. Working paper, London Business School, London, URL https://ssrn.com/ abstract=5393256. Simon HA (1971) Designing organizations for an information-rich world. Greenberger M, ed., Com- puters, Communications, and the Public Interest, 37ā72 (Baltimore, MD: The Johns Hopkins Press). Varian HR (2007) Position auctions. Internat. J. Indust. Organ. 25(6):1163ā1178, URL http:// dx.doi.org/10.1016/j.ijindorg.2006.10.002. Zhao H, Berman R (2026) Strategic response of news publishers to generative AI. Working paper, University of Pennsylvania, Philadelphia, URL https://ssrn.com/abstract=5992774. Zhu K (2026) Adverse selection in the AI data commons. Working paper, Bocconi University, Milan, URL https://ssrn.com/abstract=6438640. 73