Skip to content

Understanding Generative AI-mediated User Engagement with Academic Library Resources

Source: arXiv:2607.20328 · Published 2026-07-22 · By Hae Min Kim, Stacy Stanislaw

TL;DR

This study investigates the emerging role of generative AI platforms as discovery intermediaries for academic library resources, focusing on empirical analysis of referral traffic to Drexel University Libraries from August 2023 to October 2025. It identifies a diverse set of AI platforms—including ChatGPT, Perplexity, Gemini, Copilot, and Claude—as key sources driving user traffic, with ChatGPT dominating. The research documents a notable increase in AI-mediated referrals, especially after the late 2024 introduction of ChatGPT’s 'Sources' feature, which displays clickable citations to external authoritative resources. The study finds that AI-generated outputs preferentially expose library resources that feature structured metadata, stable permalinks, and are openly accessible, primarily directing users to institutional repositories such as electronic theses and dissertations.

The results emphasize a shifting academic information landscape where users increasingly access verified scholarly content via AI-generated pathways rather than traditional, direct search methods. The paper highlights both opportunities and challenges for libraries to adapt service strategies to maintain relevance and visibility amidst evolving AI ecosystems. It further underscores the importance of metadata quality, open access, and transparent linking standards to maximize AI-mediated discoverability.

Key findings

  • AI-mediated referrals to Drexel University Libraries first appeared on August 25, 2023, originating from Perplexity.ai.
  • Between August 2023 and October 2025, 20 distinct generative AI platforms were identified as referral sources, with ChatGPT accounting for 85% of AI-mediated traffic by October 2025.
  • Introduction of ChatGPT's 'Sources' feature in late October 2024 coincided with a 43% month-over-month increase in AI-mediated user traffic.
  • After GPT-5’s release in August 2025, daily ChatGPT-mediated referrals rose by 78% in September and an additional 47% in October 2025 compared to prior months.
  • AI referrals disproportionately direct users to the institutional repository, notably electronic theses and dissertations, signaling AI preference for open access content with rich metadata and stable permalinks.
  • AI-mediated access overall remains a minority proportion of total library traffic but shows a steady upward trend, indicating growing importance of these platforms.
  • Perplexity, Gemini, and Copilot also showed gradual upward referral trends post-2024, consistent with market share estimates of global generative AI usage.
  • Metrics such as user counts were favored over sessions due to bot traffic inflation, aiming for more reliable measures of genuine human engagement.

Threat model

n/a — This study is observational and descriptive, not focused on adversarial threats but on characterizing organic user traffic originating from generative AI referral sources to library systems.

Methodology — deep read

The study conducted an exploratory case study analyzing web referral data from Google Analytics 4 (GA4) for Drexel University Libraries from August 2023 through October 2025. The threat model assumes an academic user seeking library resources via generative AI platforms, with no adversarial manipulation considered.

Referral traffic was classified as AI-mediated if originating from domains affiliated with large language model (LLM)-based generative AI platforms such as chat.openai.com (ChatGPT), perplexity.ai, gemini.google.com, copilot.microsoft.com, or identified through domain patterns ending in .ai or containing 'ai'. This two-step process combined direct domain matching and heuristic filtering. This conservative approach could underestimate total AI referrals due to inconsistent referral signal capture.

The dataset covered user traffic across multiple library-managed platforms: the core library website (Sitecore), discovery service (ExLibris Primo), institutional repository (ExLibris Esploro), library guides, virtual reference, room/event scheduling, and digital exhibits. Drexel’s institutional repository included over 14,000 electronic theses and dissertations and more than 100,000 research items.

GA4 metrics focused primarily on "total users" rather than sessions or page views to minimize distortions from automated bots. Supplementary metrics such as bounce rate and average engagement time contextualized user behavior. Dimensions analyzed included session source (referrer), page path, and hostname.

Data analysis was stratified into three phases: (1) monthly trend analysis tracking the emergence and growth of AI referrals, (2) system-level distribution comparing AI user counts across library platforms, and (3) content-level analysis identifying types of materials accessed (e.g., open access publications, subject guides, theses/dissertations).

Preliminary data processing used Microsoft Excel, while visualization and comparative analyses leveraged Microsoft Power BI. No machine learning or predictive modeling was applied. No code or full raw data release was documented, limiting direct reproducibility.

One concrete example: Post-release of ChatGPT’s Sources feature in October 2024, referral traffic from chat.openai.com to the institutional repository surged by 43% month-over-month, correlating with enhanced citation linking in generative AI responses enabling more direct user navigation to authoritative open access documents. This evidences how technical platform changes altered traffic patterns in a clear cause-effect temporal relationship, though formal causal inference was not established.

Overall, the study leverages longitudinal administrative web analytics to quantitatively characterize ongoing shifts in discovery pathways without experimental intervention. Limitations include reliance on referral signal accuracy and inability to disambiguate underlying user intents.

Technical innovations

  • A novel two-step AI referral identification method combining direct domain matching with heuristic pattern filtering to isolate generative AI platform traffic in web analytics data.
  • Use of comprehensive longitudinal GA4 data across multiple distinct academic library service platforms enabling fine-grained system-level referral attribution.
  • Demonstration of a temporal correlation between generative AI technical feature deployments (e.g., ChatGPT 'Sources') and quantifiable changes in user referral behavior.
  • Integration of user-level engagement metrics to control for bot-induced distortions common in session or page-view counts for library analytics.

Datasets

  • Drexel University Libraries Google Analytics 4 data — approximately 26 months of web referral and engagement logs from August 2023 to October 2025 — internal institutional data

Baselines vs proposed

  • Pre-Sources feature (Oct 2024) ChatGPT referrals: baseline traffic with sporadic monthly growth; post-Sources feature: 43% month-over-month user increase
  • September 2025 ChatGPT referral users increased by 78% vs August 2025 (post GPT-5 release), October 2025 increased by 47% vs September 2025

Figures from the paper

Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.20328.

Fig 1

Fig 1: Monthly Number of Total Users Referred by Generative AI Platforms to All Library Online Systems

Fig 2

Fig 2: Monthly Proportion of AI-Mediated Access Relative to Total Traffic to All Library Online Systems

Fig 3

Fig 3: Monthly Trends in Total Users Across Library Systems

Limitations

  • Referral data relies on accurate and consistent recording of ‘session source’ in GA4, which may omit or misclassify some AI-mediated traffic.
  • No direct measurement of user intent or qualitative assessment of how AI outputs influenced click decisions; only referral counts analyzed.
  • Study focused solely on one institutional library system, limiting generalizability across different academic libraries or disciplines.
  • No adversarial testing or analysis of manipulation / gaming of AI referral pathways was conducted.
  • Did not evaluate downstream user outcomes beyond initial referral (e.g., resource usage quality, satisfaction, or learning impact).
  • No public codebase or data release hampers reproducibility and external validation.

Open questions / follow-ons

  • How do AI-generated citations and links affect user critical evaluation and trust in academic resources?
  • What metadata features most strongly influence AI platform retrieval and citation behavior for scholarly content?
  • How might libraries strategically optimize metadata and access frameworks to maximize AI-mediated discoverability while ensuring ethical data governance?
  • What are the long-term impacts of AI-mediated discovery on traditional library search behaviors and institutional user engagement?

Why it matters for bot defense

For bot-defense and CAPTCHA practitioners focused on bot detection and traffic validation in academic library and scholarly resource contexts, this study highlights the increasing proportion of legitimate user traffic arriving via AI-powered platforms generating synthesized responses with embedded references. AI referrals represent a newer traffic vector distinct from traditional search engine referrals, involving dynamic, conversational interfaces that may issue links programmatically. This evolving discovery pattern implies that library web infrastructure and bot-detection systems need adaptation to differentiate genuine user referrals originating from AI sources versus automated scraping or abuse.

Furthermore, the demonstrated importance of stable permalinks and well-structured metadata for AI indexing suggests that CAPTCHA and bot-defense solutions could consider incorporating external signals from AI referral sources as part of their traffic profiling heuristics. Recognizing the key generative AI platforms as legit referral origins may help reduce false positives in access control systems. However, the study also underscores challenges due to partial or inconsistent referral capture, calling for refined analytics to accurately parse AI referrals. Bot-defense engineers should also be aware of ongoing platform feature updates (e.g., citation linking in ChatGPT) that can abruptly shift traffic patterns, necessitating continuous monitoring to tune defenses accordingly.

Cite

bibtex
@article{arxiv2607_20328,
  title={ Understanding Generative AI-mediated User Engagement with Academic Library Resources },
  author={ Hae Min Kim and Stacy Stanislaw },
  journal={arXiv preprint arXiv:2607.20328},
  year={ 2026 },
  url={https://arxiv.org/abs/2607.20328}
}

Read the full paper

Articles are CC BY 4.0 — feel free to quote with attribution