AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge
Source: arXiv:2607.28237 · Published 2026-07-30 · By Muhammad Sajjad Akbar
TL;DR
This study critically evaluates the reliability, authenticity, and citation fidelity of six leading generative AI systems when answering open-ended Islamic knowledge questions covering Qur'anic interpretation, Hadith, Fiqh, ethics, pastoral advice, and Madhhab-sensitive topics. Unlike prior work relying on structured benchmarks or artificial prompt settings, this empirical evaluation uses real-world participant interactions from Australia and the UK with no restrictions on AI platform choice. Using a mixed-method framework, the study assesses response accuracy, hallucination rates, jurisprudential consistency, uncertainty handling, and geographic variability in outputs.
The results show that current generative AI systems demonstrate their highest reliability in domains with broad scholarly consensus, such as Qur’anic interpretation and ethical guidance, producing consistent references and pastoral advice. However, performance substantially degrades in jurisprudentially complex areas like Fiqh and Madhhab-sensitive issues, where hallucinations, incomplete or fabricated citations, and inconsistent handling of school-of-thought disagreements are frequent. Differences in citation completeness and source provenance vary even by geographic access, implying location-sensitive retrieval or moderation effects. Overall, while generative models serve as valuable assistive tools for introductory Islamic education and general learning, they lack authoritative reliability for fatwa issuance or scholarly research without expert verification against primary sources.
Key findings
- AI achieves highest factual and citation reliability in Qur’anic interpretation and ethical guidance domains with high structural consistency across responses (Fig 2).
- Significant hallucination and citation issues occur in Fiqh and Madhhab-sensitive questions, including missing exact hadith numbers and unverifiable religious claims (Table 3).
- Across six AI tools, citation deficiencies such as incomplete references and claims without source verification are frequent, especially in ChatGPT and Copilot outputs.
- AI systems vary in jurisprudential reasoning and uncertainty handling, with some models like DeepSeek showing overconfidence in contentious topics.
- Geographical variability between UK and Australia responses affects source retrieval, citation completeness, and explanatory detail, demonstrating location-based differences.
- Common linguistic and instructional framing patterns are strongly repeated across AI systems, suggesting reliance on overlapping online Islamic educational corpora.
- Performance comparison across 50 open-ended questions indicates domain accuracy significantly drops in complex Fiqh areas versus more straightforward ethical or Qur’anic queries.
- Participants freely using various AI platforms generated realistic response data revealing current generative AI systems are assistive but unsuited as sole authoritative Islamic knowledge sources.
Threat model
Adversary is a typical end-user seeking Islamic religious guidance using publicly accessible generative AI systems without specialized access or insider knowledge. The adversary lacks capability to manipulate underlying AI training or deployment but may be unaware of factual inaccuracies and hallucinations inherent in outputs. The threat is inadvertent misinformation due to AI limitations, not malicious manipulation or direct AI system exploitation.
Methodology — deep read
The study examines the authenticity and reliability of generative AI outputs in Islamic knowledge through a large-scale survey-based empirical evaluation involving 50 open-ended questions divided into 10 thematic sets. These sets cover a spectrum from Qur’anic interpretation, Hadith explanation, and Fiqh jurisprudence to ethical issues and pastoral guidance, intentionally including dilemmas sensitive to Madhhab and scholarly disagreement.
The threat model assumes typical users seeking religious knowledge via public AI systems without constraints on model choice or subscription level, reflecting real-world interaction. Users could select from multiple commercial and research AI platforms, including ChatGPT, Claude, Gemini, Copilot, DeepSeek, and others, enabling analysis of ecosystem and retrieval variations.
Data collection occurred in Australia and the UK to investigate geographic variability in AI responses, motivated by potential location-aware routing, moderation, or infrastructure effects. Responses, totaling five fully completed survey sets analyzed here, were preserved verbatim alongside metadata about AI source and geographic origin.
The evaluation framework combines quantitative and qualitative dimensions: factual correctness; Qur’anic and Hadith citation authenticity (presence, completeness, traceability); hallucination identification, including fabricated or incomplete citations; jurisprudential consistency and handling of Madhhab diversity; recognition of uncertainty and abstention behavior; interpretive reliability; and cross-model and cross-region response consistency.
Analyses included thematic comparison of response styles and linguistic framing, structural similarity measures, and citation verification against authenticated Islamic sources. Comparative assessments across models elucidated differences in stylistic fingerprints, tendency towards overconfidence, and source referencing quality.
No controlled laboratory prompting or structured multiple-choice benchmarks were used; instead, real user-AI interactions simulate practical Islamic information-seeking conditions. This approach prioritizes ecological validity over synthetic evaluation.
Reproducibility is limited as the underlying response dataset and specific AI versions are not publicly released; model heterogeneity and unrestricted user choice introduce naturalistic variability but complicate exact replication.
Technical innovations
- A novel mixed-method empirical evaluation framework combining open-ended user interactions, citation authenticity verification, jurisprudential consistency, and geographic variability analysis in an Islamic knowledge context.
- Use of unrestricted participant choice of multiple commercial and academic AI platforms capturing real-world generative AI religious information-seeking behavior rather than synthetic benchmarks.
- Integration of qualitative thematic analysis with quantitative similarity and citation completeness metrics to measure AI-generated Islamic content authenticity and hallucination tendencies.
- Systematic comparison of AI handling of Madhhab diversity, jurisprudential disagreement, and uncertainty recognition within religiously sensitive queries.
Datasets
- Islamic AI Survey Dataset — 250 responses (5 sets × 50 questions) — collected from participants in Australia and the UK for this study; not publicly released
Baselines vs proposed
- Gemini 2.0 Flash: IslamicMMLU accuracy = 93.8% vs GPT-3.5-turbo: 39.8%
- GPT-4o: IslamicMMLU accuracy = 77.3% vs GPT-4: 59.6%
- Hallucination rate on Islamic citations varied up to 50% across AI tools; DeepSeek showed higher overconfidence than ChatGPT or Claude
- Citation completeness: ChatGPT and Copilot reported more general or incomplete references compared to Claude's academic style
Figures from the paper
Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.28237.

Fig 1: Generative AI adoption and reliability challenges across general factuality bench-

Fig 2: Combined analysis of Set 1 responses, showing structural similarity, common Islamic

Fig 3: Combined analysis of Set 2 responses showing music-related jurisprudential patterns,

Fig 4: Combined analysis of Set 3 responses showing jurisprudential complexity, compara-

Fig 5: Combined analysis of Set 4 responses showing similarity levels, comparative Fajr

Fig 6: Combined analysis of Set 5 responses showing narrative and emotional characteris-

Fig 7: Performance of the evaluated AI systems across six Islamic knowledge domains.

Fig 8: Risk assessment of AI-generated responses across Islamic knowledge domains.
Limitations
- The study’s dataset is limited to 50 open-ended questions and 5 completed survey sets, which may not cover full domain diversity in Islamic knowledge.
- Analysis focuses on English-language AI responses; effects in Arabic or other languages typical for Islamic discourse remain unexamined.
- No adversarial testing of models or controlled evaluation of response correctness with blinded expert verification was conducted.
- Data collected under unconstrained user conditions lacks rigorous control, introducing variability from participant experience and AI version differences.
- Geographic and temporal variability in AI systems may evolve rapidly, limiting the study’s findings to a snapshot in mid-2026.
- Reproducibility is constrained since the response data and exact AI model versions used by participants are not publicly available.
Open questions / follow-ons
- How can generative AI systems be specialized to incorporate authenticated primary Islamic sources and verified scholarly annotations to reduce hallucinations?
- What techniques can be developed to improve certified citation completeness and source traceability specifically tailored for Islamic religious texts?
- How might AI models better recognize and explicitly represent jurisprudential uncertainty and madhhab diversity in responses?
- What impact do language, localization, and interface design have on user trust and verification behaviors when using AI for religious knowledge?
Why it matters for bot defense
For bot-defense and CAPTCHA practitioners, this paper underscores the complexity of evaluating reliability and hallucination in domain-sensitive generative AI outputs, especially where authoritative verification is essential. The study’s mixed-method evaluation framework offers a potential blueprint for assessing trustworthiness beyond typical factuality metrics by emphasizing source citation fidelity and interpretive consistency.
While CAPTCHAs primarily defend against automated abuse, emerging systems increasingly integrate AI-based knowledge retrieval and interaction; understanding AI hallucination risks and geographic variability demonstrated here can inform risk assessment and mitigation strategies for AI-enhanced user engagements. Moreover, the observed behavioral fingerprints and stylistic differentiation among AI platforms suggest nuanced signal sources that could complement behavioral detection of malicious or automated inputs in religious or content-moderation contexts.
Cite
@article{arxiv2607_28237,
title={ AI and Authenticity in Islamic Research: A Critical Evaluation of Generative AI Reliability, Hallucination, and Source Fidelity in Quranic, Hadith, and Fiqh Knowledge },
author={ Muhammad Sajjad Akbar },
journal={arXiv preprint arXiv:2607.28237},
year={ 2026 },
url={https://arxiv.org/abs/2607.28237}
}