'A bit of chaos and madness': The AI Assessment Scale and the work of assessment reform
Source: arXiv:2606.26729 · Published 2026-06-25 · By Mike Perkins, Darius Postma, Jasper Roe, Susan Sisay, Craig Holdcroft
TL;DR
This qualitative study investigates the real-world implementation of the Artificial Intelligence Assessment Scale (AIAS), a structured framework designed to guide university assessment reform in response to generative AI (GenAI) tools. Conducted across two distinct institutions—a private international university in Vietnam and a public UK university—the study engages 30 academic staff through five focus groups. It explores how faculty perceive, enact, and negotiate the AIAS in their assessment design and governance amid challenges of integrity, equity, and pedagogy. The research identifies six thematic dimensions shaping implementation: recognising and integrating AI, facilitating conditions, building capacity, pathways to adoption, ethics in practice, and reframing pedagogy.
The findings reveal that AIAS provides shared language and legitimacy for incorporating GenAI in assessment, supports reflection on boundaries of appropriate use, and can stimulate authentic assessment redesign. However, its effectiveness is mediated by factors such as governance models, clarity of policies, staff confidence and workload, disciplinary contexts, and alignment with learning outcomes. When disconnected from these institutional conditions and staff capacity, AIAS risks becoming a superficial compliance tool rather than fostering meaningful pedagogical changes. The study thus highlights that assessment reform amidst GenAI-driven disruption requires governance structures and support mechanisms that enable faculty ownership and capacity building, rather than purely top-down mandates or detection-centered approaches.
Key findings
- Faculty valued AIAS as a shared language legitimising GenAI use and clarifying boundaries, easing experimentation (e.g., ‘It’s given me confidence… we’re not going to be reprimanded’).
- Implementation was shaped by institutional governance: rapid top-down rollout (Vietnam) generated momentum but limited faculty ownership, while decentralized department-led adoption (UK) enabled ownership but created fragmented strategy.
- Unclear policies and lack of evaluative mechanisms caused uncertainty and workload burden; faculty worried about fairness when AI use was declared but not verifiable (‘how can we know what they are doing with it?’).
- Capacity building was most effective when practice-based, peer-led, and connected to teaching rather than one-off trainings (‘hands-on peer exchange’ was a ‘big plus’).
- Ethics discussions focused on fairness and critical engagement with AI rather than policing; concerns about inequitable access to AI tools potentially creating a tiered education system.
- AIAS influenced assessment redesign by prompting reflection on aligning AI use with learning outcomes, but risks emerged when scale levels became ends in themselves rather than tools for valid assessment design.
- Faculty described divergent experiences: UK faculty linked AIAS to authentic assessments and student responsibility, while Vietnam faculty emphasized adjustments and the risk of misalignment with outcomes.
- Overall, meaningful GenAI assessment reform depends on institutional conditions: governance clarity, faculty capacity, alignment to outcomes, and attention to equity.
Threat model
The study addresses the challenge of academic staff managing student use of generative AI tools within an institutional environment of ambiguous or evolving policies. The 'adversary' in this context is not a malicious actor per se, but the threat arises from unmanaged or unregulated AI-enabled academic work risking breaches of integrity, equity, and assessment validity. Faculty navigate uncertainties about permissible AI use and lack mechanisms to verify student declarations, thus the threat model involves institutional governance and educator capacity to meaningfully set and enforce assessment boundaries in the face of rapid GenAI adoption.
Methodology — deep read
Threat Model & Assumptions: The study assumes academic staff as primary agents implementing the AIAS within university governance constraints. The adversarial element consists of managing potential misuse of GenAI by students under conditions of uncertain institutional guidance and rapid tech evolution. The focus is on staff experience, not on technical adversaries or direct AI misuse detection.
Data Collection: Data were collected from two diverse institutional contexts: a private international university in Vietnam (I1), which had adopted the first version of AIAS university-wide, and a public UK university (I2), which piloted the second version in a single school. In total, five focus groups involving 30 academic staff were conducted (3 groups at I1 with 17 participants purposively sampled across disciplines, 2 groups at I2 with 13 participants from the Business School via self-selection). Sessions were approximately 90 minutes, audio-recorded, conducted in person at I1 and online at I2.
Data Analysis: A hybrid thematic analysis methodology was applied. Initial deductive coding categories were developed from research questions aligned with a sensitising concept of Critical AI Literacy (CAIL), addressing ethics, equity, and criticality. Codes were refined through iterative discussion and review by an external co-researcher to mitigate bias given that some authors developed AIAS and one institution was the original AIAS site. Thematic synthesis collapsed 10 parent codes into 6 overarching interpretive themes traversing codes and transcripts across both settings.
Research Instruments: A semi-structured focus group protocol balanced consistency with flexibility, crafted by the research team and aligned to institutional contexts. Ethical approval was obtained and participant anonymity preserved.
Analytical Rigor: Reflexivity was maintained through memoing and team discussions to interrogate potential positionality bias. The coding framework underwent external review and triangulation.
Evaluation & Findings: The qualitative findings are based on interpretive thematic constructions, not quantitative metrics or inter-coder reliability. Themes were developed inductively to capture patterns of meaning about AIAS enactment.
Reproducibility: No code or datasets are publicly released due to the qualitative nature and institutional context. The study is transparent about data provenance and analytical procedures but inherently non-replicable.
Example Workflow: A faculty member at I1 participated in a focus group discussing how they integrated AIAS in assessment. Coding categorized their comments under 'Facilitating Conditions' (policy ambiguity), 'Building Capacity' (lack of hands-on training), and 'Reframing Pedagogy' (challenges aligning scale with outcomes). These were synthesized with peer contributions to form themes highlighting institutional barriers and opportunities.
Technical innovations
- Empirical investigation of AIAS implementation through multi-institutional qualitative focus groups, filling a literature gap on lived faculty experience.
- Application of Critical AI Literacy as a sensitising concept in thematic analysis to frame ethical, equitable, and critical engagement dimensions.
- Development of a hybrid thematic coding framework combining deductive and inductive approaches sensitive to institutional context and epistemic nuances.
- Analysis contrasting top-down centralized vs. decentralized department-led governance models for AI assessment reform impact.
- Interpretation of AIAS’s dual role as both a legitimizing framework and a potential compliance mechanism depending on institutional enactment.
Datasets
- Focus group transcripts — 5 focus groups, 30 participants, drawn from a private international university in Vietnam and a UK public university — not publicly available
Limitations
- Study is qualitative and interpretive, limiting generalizability beyond the two institutional contexts studied.
- Participants were self-selected (UK) or purposively sampled (Vietnam), potentially introducing selection bias.
- No quantitative assessment of AIAS effectiveness or impact on student outcomes; findings rely on self-report and perception.
- Potential confirmation bias given involvement of some authors in AIAS development and proximity to one institution where it was developed.
- Data are cross-sectional soon after rollout; longitudinal effects on practice and policy alignment remain unexamined.
- Focus on faculty perspective; student views and actual AI use monitoring were not included.
Open questions / follow-ons
- How do student perspectives on AIAS and GenAI use in assessment compare or contrast with faculty experiences?
- What long-term impacts does AIAS-guided assessment reform have on learning outcomes, academic integrity incidents, and student engagement?
- Can AI detection tools and AIAS frameworks be effectively integrated without undermining equity or adding compliance burdens?
- What models of faculty capacity building sustain effective assessment redesign and academic integrity in fast-evolving AI contexts?
Why it matters for bot defense
For bot-defense and CAPTCHA practitioners, this study offers insight into how academic institutions operationalize frameworks to manage and integrate AI tool use in high-stakes assessment settings. It highlights that effective mitigation of AI-related risks depends not only on detection or outright banning of AI artifacts but on creating shared languages, policies aligned with learning objectives, and staff capacity to interpret and enforce boundaries. The findings emphasize the importance of clarity, governance, and equitable access when deploying AI assessment frameworks—lessons applicable to designing defenses against automated generation or misuse in complex user scenarios. Practitioners can appreciate that technological solutions require embedding within institutional and social contexts to avoid becoming mere compliance checkboxes, paralleling challenges in bot detection systems where policy, user understanding, and infrastructure interplay critically affect enforcement.
Cite
@article{arxiv2606_26729,
title={ 'A bit of chaos and madness': The AI Assessment Scale and the work of assessment reform },
author={ Mike Perkins and Darius Postma and Jasper Roe and Susan Sisay and Craig Holdcroft },
journal={arXiv preprint arXiv:2606.26729},
year={ 2026 },
url={https://arxiv.org/abs/2606.26729}
}