Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain
Source: arXiv:2607.07652 · Published 2026-07-08 · By Qiaoni Shi, Kai Zhu, Kai Gu
TL;DR
This study investigates how AI-driven search interfaces, exemplified by ChatGPT Search, alter the traditional economic model of web traffic allocation dominated by conventional search engines such as Google. Traditional web search engines direct users from queries to external websites, enabling those sites to monetize visits through advertising, subscriptions, or commerce. AI search changes this paradigm by often answering user queries within the intermediary itself, reducing outbound clicks to external destinations. Using detailed URL-level Comscore U.S. desktop clickstream data spanning October 2024 to July 2025, the authors compare ChatGPT and Google search behaviors within the same households and exploit staged access expansions of ChatGPT Search to estimate the degree to which AI search displaces conventional search traffic.
Key findings reveal that ChatGPT generates outbound clicks in only about 5.2% of conversation sessions, dramatically lower than Google's 31.1% referral ratio. The residual outbound traffic from ChatGPT is skewed toward specialized and technical content (e.g., academic, reference, developer tools) and away from ad-supported general-interest websites. Moreover, wider ChatGPT access causally reduces traditional search query volumes by approximately 9.4% on average and up to 17% after 20 weeks, with the largest displacement occurring in informational search categories. These results underscore a fundamental economic shift: AI search intermediaries satisfy many information needs internally, weakening the referral traffic that underpins the longstanding web content ecosystem and advertising-based monetization models.
Key findings
- ChatGPT produces outbound clicks in only 5.2% of conversation sessions versus Google's 31.1% referral ratio (Table 1).
- Within households active on both platforms, ChatGPT referral ratios are 8.1% versus Google’s 39.6%, a 31.5 percentage point difference (Figure 3).
- 74.4% of ChatGPT-active households never produce any outbound referral, compared to only 9.6% for Google-active households.
- ChatGPT's residual referral traffic favors specialized categories like reference/knowledge (+13.1 pp), tools/SaaS (+8.7 pp), academic research (+5.4 pp), and developer/technical (+4.7 pp), while significantly underrepresenting social media (−15.8 pp) and ad-supported sites (−27.6 pp) relative to Google (Figure 5).
- ChatGPT referral traffic is less concentrated in aggregate across domains than Google’s, with Google concentrating on large sites such as YouTube, Reddit, and Wikipedia (Table 2, Figure 6).
- ChatGPT sessions commonly occur without other simultaneous browsing (+23.1 pp solo sessions) and are more frequent in reference/knowledge contexts (+5.2 pp) but less so in social media and e-commerce contexts.
- Wider ChatGPT Search access causally reduces traditional search volume by 9.4% on average, rising to 17.0% after twenty weeks, particularly hitting informational search categories hardest.
- Adding surrounding browsing context as a covariate does not materially reduce the observed ChatGPT vs Google referral gap, implying interface differences drive the routing variation.
Threat model
Not a traditional security threat model; adversaries are conceptualized as website publishers or economic agents affected by the shift in user attention and referral traffic induced by AI search intermediaries altering the routing and monetization on the open web. The study assumes no capability for adversarial manipulation or deception of clickstreams but focuses on the economic displacement and referral dynamics.
Methodology — deep read
Threat model and assumptions: The adversary context here is economic rather than security-focused; the paper studies how AI search changes traffic patterns and referral behaviors compared to traditional search. It assumes users bring information needs to intermediaries (Google, ChatGPT) and measures where attention is routed post-query. The adversary cannot inject fake referrers, and the study only tracks real user browsing behaviors.
Data provenance: The study uses URL-level Comscore U.S. desktop clickstream data from Oct 2024 through July 2025, covering roughly 168k to 238k active households monthly. A balanced sub-panel of 45,386 households with continual activity is used for the displacement analysis. Data include page loads, timestamps, HTTP referrers, Google/Bing/Yahoo search queries, and ChatGPT conversation endpoints.
Defining user side units: One Google search query or one ChatGPT conversation session is treated as one information-seeking occasion. ChatGPT session IDs reconstruct maximal sequences of loads within one hour gaps to deduplicate repeated endpoint hits.
Outbound referrals: Clean outbound referrals are page loads with an HTTP referrer pointing from Google or ChatGPT, excluding self-referrals, platform internal navigation, automated requests, and non-foreground visits.
Domain classification: 4,266 destination domains were classified into content categories (e.g. reference, social media) and monetization models (ad-supported, freemium, subscription). High-confidence subsets (3,245 domains) support category and displacement analyses.
Access expansion design: Three ChatGPT Search rollouts (Oct 31 2024 for paid subscribers, Dec 16 for free logged-in users, Feb 5 2025 for anonymous users) create natural experiments. Cohorts are assembled of households gaining access, with matched controls lacking ChatGPT/Claude activity.
Empirical strategy: Using a stacked difference-in-differences approach comparing treated cohorts to controls, the paper estimates the causal impact of ChatGPT Search access on traditional search query volumes and referral patterns, controlling for household and week fixed effects.
Within-household-week comparisons between Google and ChatGPT measure relative referral rates, controlling for time and household invariant factors.
Task intent proxies: User surrounding browsing within ±15 minutes of the query/session is used as a behavioral proxy for search intent and context.
Concentration analyses evaluate domain referral distribution both in aggregate and within households using normalized Herfindahl indices.
A concrete example: For one household active on both Google and ChatGPT in the same week, the authors count all Google queries and ChatGPT conversation sessions, identify which generated clean referrals, link referrals to classified domains, and measure differences in the rate of outbound clicks and domain types visited. After New Access expansion on December 16, cohorts gaining free logged-in ChatGPT Search access show statistically significant reductions in their traditional search queries versus controls.
Technical innovations
- Using URL-level Comscore clickstream to reconstruct user events on both AI and traditional search interfaces at scale and within-household comparisons.
- Defining ChatGPT conversation sessions as maximal sequences per conversation ID with time thresholds to deduplicate repeated page loads.
- Leveraging staged ChatGPT Search access expansions as natural experiments for causal inference on search displacement.
- Classifying domain-level referral traffic by content type and monetization model to reveal economic implications of AI search’s referral skew.
- Combining user-side proxy for search intent via surrounding browsing context with controlled comparisons to disentangle interface vs task differences.
Datasets
- Comscore U.S. desktop clickstream — ~168k to 238k households monthly — proprietary dataset used with user panel sampling spanning Oct 2024 to Jul 2025
Baselines vs proposed
- Google search queries: 31.1% referral ratio vs ChatGPT sessions: 5.2% referral ratio
- Within household-week fixed effects sample: Google referrals = 39.6% vs ChatGPT referrals = 8.1%
- Wider ChatGPT Search access effect: traditional search queries reduced by 9.4% on average and 17.0% after 20 weeks
Limitations
- Analysis limited to U.S. desktop clickstream; excludes mobile or other geographies, which may differ.
- ChatGPT conversational prompts and user intents are not directly observed, only proxied via surrounding browsing.
- Clean referral detection depends on HTTP referrers which can be stripped or obscured by browsers or privacy features.
- The causal displacement estimates assume treated/control household comparability; unobserved confounders may remain.
- No direct measures of consumer welfare, publisher revenue, or long-run content ecosystem impacts were assessed.
- Focus is on ChatGPT and Google; other AI or search platforms (Claude, Gemini) showed limited reach and were not deeply analyzed.
Open questions / follow-ons
- How do these AI search referral dynamics differ on mobile devices or in other geographic regions?
- What are the long-term impacts on publisher revenue, content creation incentives, and advertising ecosystems?
- How robust are referral patterns and displacement effects when including other AI search competitors or hybrid AI-search models?
- What specific user satisfaction or welfare differences arise from retained vs routed queries in AI search?
Why it matters for bot defense
For bot-defense and CAPTCHA practitioners, this paper sheds light on the evolving digital intermediation landscape where AI search interfaces retain user attention internally and drastically reduce outbound clicks to external sites. This shift implies that traditional CAPTCHA deployment strategies reliant on measuring referral traffic or click-through behaviors may need reconsideration, as fewer user-triggered referrals occur via AI search intermediaries. Moreover, the pronounced decline in traffic to ad-supported and commercial websites could affect the economic incentives underlying CAPTCHA effectiveness—especially in adversarial contexts where bot detection relies on interaction patterns linked to website visits. Understanding AI search’s selective referral skew toward specialized domains may help tailor bot defense measures to new traffic flows rather than legacy referral models. Practitioners should anticipate that AI search intermediaries disrupt not only traffic volumes but also the surface on which bot-detection challenges operate, potentially requiring novel in-intermediary verification mechanisms or analytics.
Cite
@article{arxiv2607_07652,
title={ Answering Without Referring: How AI Search Rewrites the Web's Economic Bargain },
author={ Qiaoni Shi and Kai Zhu and Kai Gu },
journal={arXiv preprint arXiv:2607.07652},
year={ 2026 },
url={https://arxiv.org/abs/2607.07652}
}