Skip to content

Generative AI Availability, Grades, and Student Satisfaction at a Large University

Source: arXiv:2607.21534 · Published 2026-07-23 · By James M. Zumel Dumlao, Meng Wang, Zhonghan Xie, Junyao Hu, Ivan Bar, George Chaney et al.

TL;DR

This study rigorously examines the effects of generative AI (GenAI) availability—marked by ChatGPT's public release in late 2022—on student academic performance and satisfaction at a large U.S. university over 2015-2025. Motivated by concerns that students might substitute their cognitive effort by using GenAI on susceptible assessments (e.g., take-home problem sets, essays), potentially inflating grades without learning, the authors developed a human-validated LLM-based method to quantify each course's GenAI susceptibility from syllabi. Employing a difference-in-differences design that compares more to less susceptible courses before and after ChatGPT, while accounting for COVID-19 pandemic disruptions, they analyze 156,135 students, 87,936 offerings, and 1.45 million student-course observations. Contrary to earlier smaller-scale studies suggesting grade inflation in GenAI-susceptible courses, this work finds no significant increase in grades, withdrawal, or failure rates attributable to GenAI availability. Effects on self-reported understanding and interest are negligible or insignificant under most reasonable pandemic-effect models. The findings challenge the “GenAI substitution hypothesis” and suggest that widespread generative AI use has not yet degraded grades or student satisfaction at scale in this setting.

Key findings

  • No significant differential effect of GenAI availability on final grades in highly susceptible courses compared to less susceptible ones, under both transient and persistent COVID-19 effect assumptions.
  • No grade inflation observed among previously lower-performing students; grade effects are null across student prior academic ability terciles measured by residualized GPA rank.
  • No significant changes in course withdrawal or failure rates linked to course GenAI susceptibility post-ChatGPT.
  • Self-reported student understanding and interest in subjects showed no robust declines; only a modest increase in interest is detected assuming transient COVID effects.
  • GenAI susceptibility measure correlates strongly year-to-year (mean correlation 0.73 from 2019), validating the use of 2019 as an anchor pre-exposure year.
  • Prior studies (Hausman et al. 2025, Chirikov 2026a) reported grade increases of ~0.6–1 point out of 100 and ~0.12 GPA points respectively in susceptible courses, findings not replicated here.
  • Observed increase in share of A grades post-ChatGPT in high-susceptibility courses aligns with non-parallel pre-COVID trends and thus cannot be causally attributed to GenAI.
  • Course assessment structures (susceptible vs non-susceptible) extracted at scale via a validated LLM pipeline with human consensus over 525 syllabi, enabling robust treatment assignment.

Threat model

The adversary is a hypothetical student seeking to use generative AI tools to substitute their effort on course assessments that can be completed independently without instructor supervision, to achieve higher grades without proportional learning or mastery. They have access to publicly released GenAI tools such as ChatGPT starting November 2022. They cannot modify course grading policies or circumvent in-person assessments, and their use of GenAI is unobserved by the researchers. The study does not model malicious exploitation beyond cognitive offloading.

Methodology — deep read

  1. Threat Model & Assumptions: The study addresses the hypothesis that students may use generative AI tools such as ChatGPT to automate work on assessments that are susceptible to such tools (e.g., take-home exams, essays), effectively substituting cognitive effort. The implicit adversary is the student seeking to achieve high grades without proportional learning. The study assumes no direct observation of AI use and relies on an ecological natural experiment due to the sudden introduction of ChatGPT in November 2022.

  2. Data: Data encompass syllabi, administrative student records, and course evaluations from a large U.S. university for 2015 through 2025. The dataset covers 156,135 unique students and 87,936 course offerings across 6,836 unique courses. Syllabi (n=36,357 for 2015–2025) were scraped from an institutional archive, containing detailed grading policies enabling extraction of assessment types and their weights. Course evaluation data (43,025 evaluations post-2016) provide student-reported understanding, interest, and workload. Student demographics and transcripts enable calculation of prior academic preparation.

  3. Course GenAI Susceptibility & Treatment Variable: Courses are operationalized as more or less GenAI susceptible based on the proportion of the final grade assigned to assessments feasibly automatable by GenAI (open-book exams, take-home homework, essays, projects). Non-susceptible assessments include closed-book in-class exams, presentations, and live demonstrations. An advanced LLM-based annotation pipeline identifies and weights these from syllabi with human-validated accuracy against 525 labeled syllabi. Susceptibility is fixed using 2019 offerings to avoid confounding from pandemic or GenAI-era instructor policy changes.

  4. Analytical Design: A difference-in-differences (DiD) approach estimates the causal effect of GenAI availability by comparing outcomes in higher vs. lower susceptibility courses, before and after ChatGPT's release. The 'post' period starts Fall 2022. Fixed effects for course, term, and student control for confounders like grade inflation and student composition changes. Two models bound pandemic effects as either persistent or transient during 2020-2022.

  5. Outcomes: Performance metrics include average final GPA on a 4.0 scale, withdrawal rates, failure rates (grade below D-), and distributional analyses by grade thresholds. Satisfaction is captured by median course evaluation responses regarding understanding, interest, and workload.

  6. Heterogeneity & Robustness: Effects are assessed across prior student academic ability, measured via first-term residualized GPA ranks within cohorts. Event-study plots and pre-trend analyses test the parallel trends assumption; however, some pre-trend violations occur, notably for average grades, weakening causal claims there.

  7. Reproducibility: The authors rely on institutional archives and de-identified administrative data not publicly available. The LLM annotation pipeline for syllabi is described in detail with validation metrics, but no public code repository is stated. Fixed weighting and course anchoring choices are transparent and documented.

Concrete example: E.g., a course's 2019 syllabus is parsed with the LLM pipeline to extract 60% of final grade as take-home problem sets and essays (susceptible tasks). Post-ChatGPT, the DiD compares this course's average GPA from Fall 2015–Fall 2021 to Fall 2022–Fall 2025 against a course with 20% susceptibility, controlling fixed effects. No significant positive differential GPA change was found, suggesting no grade inflation from GenAI use in this susceptible course.

Technical innovations

  • Development and human validation of an LLM-based pipeline to infer course assessment types and grading weights from syllabi at scale with high accuracy.
  • Operationalization of course-level GenAI susceptibility using detailed grading policy decomposition rather than proxy features like writing percentage.
  • Anchoring GenAI susceptibility measurement to a pre-pandemic, pre-GenAI year (2019) to avoid contamination from instructor strategic behavior post-shock.
  • A rigorous difference-in-differences empirical design modeling COVID pandemic effects as either transient or persistent to bound estimates of GenAI effects.
  • First university-wide empirical investigation analyzing GenAI effects on student satisfaction measures (self-reported understanding, interest, workload) alongside grades.

Datasets

  • University of Michigan administrative records — 156,135 students, 87,936 course offerings, 2015–2025 — non-public institutional data
  • University syllabi archive — 44,876 syllabi with 36,357 between 2015–2025 — non-public institutional archive
  • Course evaluation survey results — 43,025 course evaluations, 2016–2025 — non-public institutional data

Baselines vs proposed

  • Hausman et al. (2025): AI-compatible courses grade increase ≈ 0.6–1.0 points on a 100-point scale post-ChatGPT — proposed study: no significant effect observed.
  • Chirikov (2026a): Writing/coding exposure courses grade increase ≈ 0.12 GPA points — proposed study: no significant differential GPA increase.
  • For withdrawal rate: baseline and proposed — no significant difference detected.
  • For failure rate: baseline and proposed — no significant difference detected.
  • Student satisfaction (understanding, interest): prior literature limited — proposed study finds mostly null effects except small interest increase assuming transient COVID effects.

Figures from the paper

Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.21534.

Fig 1

Fig 1: plots the event-study dynamics of the grade effect at semester resolution,

Fig 2

Fig 2: Event Study: Average Grade by AbilityRank Tercile. Left boxes: semester-by-

Fig 3

Fig 3: Distribution of GenAI Susceptibility among courses in 2019, 2021, and 2023. Each

Fig 4

Fig 4: Susceptibility Tercile Trajectories, 2019 to 2022 to 2025. Flows trace courses among

Fig 5

Fig 5: Event Study: Grade-Distribution Thresholds. Left boxes: semester-by-semester

Fig 6

Fig 6: Event Study: Withdrawal and Failure Rates. Left boxes: semester-by-semester

Fig 7

Fig 7: Event Study: Course Evaluation Outcomes. Left boxes: semester-by-semester

Fig 8

Fig 8: Raw Dynamics: Student Outcomes by Susceptibility Group. Each point is the

Limitations

  • Study is observational and non-randomized; despite fixed effects and design, some pre-trend violations (e.g., in average grades) may bias causal inference.
  • Results come from one large public U.S. university; findings may not generalize to different types of institutions or international settings.
  • Potential unmeasured instructor or departmental responses to GenAI that alter assessment policies post-ChatGPT may blur treatment contrast.
  • No direct measurement of individual student GenAI usage; susceptibility is a proxy based on course design.
  • Course evaluation data are voluntary and potentially suffer from response biases limiting interpretation of satisfaction results.
  • Limited granularity regarding the evolution of GenAI capabilities and student adoption rates over time, treated as a binary pre/post indicator.

Open questions / follow-ons

  • How do individual-level GenAI usage patterns correlate with learning outcomes and satisfaction beyond ecological course susceptibility?
  • What are longer-term effects of GenAI use on mastery, retention, and downstream labor market outcomes beyond grades?
  • How do instructor and department-level policy adaptations evolve post-GenAI adoption to maintain assessment integrity?
  • Could more granular or direct measures of cheating versus legitimate GenAI assistance clarify substitution impacts?

Why it matters for bot defense

For bot-defense and CAPTCHA practitioners, this study provides indirect but valuable insights into the capability and limits of generative AI tools when used by students in unsupervised assessment contexts. The rigorous ecological analysis shows that, despite concerns, widespread GenAI use does not necessarily produce large-scale output inflation or erosion of user engagement measured by grades and satisfaction proxies. This informs threat modeling around automated content generation and its impact on trustworthiness signals (e.g., grades as a signal of mastery) in systems reliant on human cognitive effort. It also highlights the importance of observable task structures—tasks fully automatable by AI versus those requiring live human demonstration—in assessing vulnerability to AI substitution. Bot-defense systems may similarly need to differentiate between online tasks susceptible to AI-based automation and those that are not. Finally, the methodology of parsing large unstructured policy documents (e.g., syllabi) with LLMs to extract structured susceptibility metrics could inspire analogous approaches for extracting vulnerability signals from complex textual sources in security settings.

Cite

bibtex
@article{arxiv2607_21534,
  title={ Generative AI Availability, Grades, and Student Satisfaction at a Large University },
  author={ James M. Zumel Dumlao and Meng Wang and Zhonghan Xie and Junyao Hu and Ivan Bar and George Chaney and Henry Gold and Misha Teplitskiy },
  journal={arXiv preprint arXiv:2607.21534},
  year={ 2026 },
  url={https://arxiv.org/abs/2607.21534}
}

Read the full paper

Articles are CC BY 4.0 — feel free to quote with attribution