Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science
Source: arXiv:2607.22513 · Published 2026-07-24 · By Davide Scarso, Hugo Noronha de Almeida, Joaquim Pina
TL;DR
This study investigates how commercial large language models (LLMs) from four major families—Claude, Grok, GPT, and Gemini—evaluate pseudo-scientific ethnonationalist claims derived from Frank Salter's biosocial framework across multiple time points and interfaces. The key novelty lies in revealing that the epistemic stance these models take is not stable or purely model-inherent but heavily contingent on deployment configurations, such as system prompts, safety layers, interface (API vs web), and silent, undocumented updates. Notably, Grok's Fast versions, powering the default user experience on X, assign credibility scores two to five times higher to pseudo-scientific statements compared to other models, while performing comparably on control evolutionary consensus and rejected Lamarckian statements.
The results demonstrate significant temporal instability and interface divergence: a silent patch shifted Grok's web outputs from chaotic to stable high validation without public notice, and the same Grok model identifier gave radically different scores between API and web endpoints months later. Furthermore, the most epistemically responsible behavior observed—categorical refusal to rate the pseudo-scientific claim—occurred only in select model-interface combinations and eroded in later versions. The study argues this opacity and instability of epistemic mediation in deployed LLMs constitute a serious public concern requiring new forms of continuous auditing and accountability.
Key findings
- Grok's Fast non-reasoning versions assigned mean credibility scores of 70–75 to ethnonationalist pseudo-science in November 2025 (S3 API snapshot), two to five times higher than all other models which scored 15–40.
- In October 2025 (S1), Grok 4 web output for the pseudo-scientific claim was chaotic (range 10–92), but two weeks later (S2) stabilized near 70–75 due to a silent patch, without documented version changes.
- By February 2026 (S4), Grok 4.1 Fast gave stable high scores (mean 75.0, SD=1.6) via API but near-zero mean scores (mean 5.5) via web with unstable bimodal distribution (13 zeros, two values ~40), indicating radical API-web divergence.
- Two model families exhibited refusal to rate pseudo-scientific claims: Claude Opus 4.1 refused categorically via web (all 15 runs) in S3, while GPT-5.1 Chat refused intermittently via API in S3 and S4; both refusals disappeared in successor versions.
- All models scored high control evolutionary consensus prompts near 95–100 and rejected Lamarckian control statements near 0–5 consistently across snapshots and interfaces, confirming calibration for basic scientific facts.
- Within the Grok family, enabling the reasoning variant lowered but did not normalize the elevated pseudo-science scores, and reasoning variants exhibited much higher output variance (SD>10) than non-reasoning variants with near-zero variance.
- Other models also showed API-vs-web scoring gaps of up to 27 points but generally within low score ranges (10–40); Grok Fast showed the largest and most consequential 69.5-point gap between web and API in S4.
- Output variability differed strongly by model: Claude and Gemini 3 Pro returned deterministic scores, GPT models showed moderate variability (SD=1–5), and Grok Fast reasoning variants had very high variability (SD=11–13).
Threat model
The adversary is an ordinary LLM end user relying on deployed models for knowledge evaluation, with no visibility into or control over deployment configuration (system prompts, safety filters, interface routing). The adversary cannot directly inspect or alter model weights or training data but is subject to the unpredictable and opaque epistemic mediation induced by deployment choices, which may propagate pseudo-scientific misinformation or suppress valid refusals without notice.
Methodology — deep read
The study systematically assessed multiple commercial LLM families (Claude by Anthropic, Grok by xAI, GPT by OpenAI, Gemini by Google) across four temporal snapshots: S1 (Oct-Nov 2025), S2 (Nov 2025), S3 (Nov 2025), and S4 (Feb 2026). The threat model assumes that adversaries include end users querying deployed LLMs for knowledge evaluation, with no access to internal configurations or updates; the study focuses on how deployment settings shape outputs rather than direct adversarial input manipulation.
Data were collected via both API and web interfaces, whenever available, for over 25 model versions spanning different releases and configurations (Fast/non-reasoning vs reasoning variants). Each run prompted the model with three fixed statements: a target prompt expressing Salter's ethnonationalist biosocial pseudo-science, a high control prompt conveying standard evolutionary consensus, and a low control prompt expressing a discredited Lamarckian claim. The models were asked to assign a single integer credibility score from 0 (fully unreliable) to 100 (fully supported by scientific consensus), ignoring moral/cultural aspects and providing no explanations.
Prompting was consistent across runs with instructions framing the model as an impartial scientific evaluator. API tests employed automated Python scripts at temperature=0, running 20-30 randomized repetitions per condition per snapshot to measure consistency. Web interface testing was manual, using cleared cache and multiple accounts to reduce context leakage, with archival screenshots for auditing. Model versions tested included multiple Grok Fast and reasoning variants, GPT chat and non-chat versions, Claude Opus and Sonnet variants, and Gemini Flash and Pro variants.
The architecture details of each model are proprietary and not disclosed, but the study distinguishes configuration variants by system prompts, safety filters, and reasoning capabilities imposed on the same underlying model weights. Training regimes are unknown due to commercial nature. Evaluation focused on mean credibility scores, standard deviation, coefficient of variation, and comparison between API and web endpoints, along with detection of outright refusals to rate. The data analysis highlighted temporal stability and interface effects, focusing on Grok’s bifurcated behavior as a central case study.
No official code or weights were released; the authors commit to releasing the full dataset (CSV/JSON) upon publication. The lack of internal deployment transparency limits interpretability of observed silent patches and policy changes, constituting a main challenge.
A concrete example: in the S3 API snapshot, Grok 4.1 Fast non-reasoning scored a mean 74.8 (SD=1.1) on the pseudo-scientific prompt over 20 runs, with tightly clustered values (70-75), whereas Grok 4.1 Fast reasoning gave 50.5 mean (SD=12.8) with high dispersion, and competing GPT-4.1 gave 34.5 mean (SD=2.2). These measurements allow quantifying both high confidence in validation and instability introduced by configuration.
Technical innovations
- Demonstration that a single commercial LLM model can produce diverging epistemic stances depending on deployment configuration variables including system prompts, safety layers, and interface routing.
- Quantitative credibility scoring approach applied to pseudo-scientific claims across multiple LLM versions, timepoints, and interfaces to measure epistemic instability.
- Identification and documentation of silent, undocumented system updates that systematically alter a model’s knowledge validation behavior.
- Revealing interface-dependent divergence of outputs for same model version, highlighting the impact of endpoint deployment on user experience and epistemic mediation.
Datasets
- Custom prompt-response datasets—over 20-30 randomized runs per model-version-prompt combination; collected via API and web interfaces; >25 model versions tested across 4 temporal snapshots from October 2025 to February 2026.
Baselines vs proposed
- Control prompts (evolutionary consensus): All models scored 95–100, establishing scientific calibration baseline.
- Control prompts (Lamarckism): All models scored 0–5, confirming recognition of pseudoscience.
- Grok 4.1 Fast (S3 API): Target prompt mean score 74.8 (SD=1.1) vs GPT-4.1 mean 34.5 (SD=2.2).
- Grok 4.1 Fast (S4 API): 75.0 (SD=1.6) vs Web: mean 5.5 (SD=14.4), 69.5-point gap indicating deployment effect.
- Claude Opus 4.1 (S3 Web): Consistent refusal to rate pseudo-science vs API same version scored flat 25.
- GPT-5.1 Chat (S3 API): Intermittent refusal to rate pseudo-science on some runs vs numerical scores (~17) on others.
Limitations
- Focus limited to a single domain of ethnonationalist pseudo-science; generalizability to other knowledge domains is not tested.
- Temporal snapshots are discrete intervals rather than continuous longitudinal monitoring; some interim changes may be missed.
- Lack of access to internal deployment logs, system prompt content, or safety filter details limits causal explanation of observed behaviors.
- API and web interface testing involved manual web scraping with fewer repetitions, raising possible consistency concerns.
- The refusal behavior is inconsistent within models at temperature=0, complicating interpretation of model intent or policy enforcement.
- The study does not test adversarial perturbations or robustness against deliberate manipulations of model outputs.
Open questions / follow-ons
- What specific deployment configuration elements (system prompts, safety layers, filters) primarily drive the observed epistemic divergence between API and web interfaces?
- How generalizable are these findings across other pseudo-scientific or contested knowledge domains beyond ethnonationalist biosocial claims?
- Can continuous monitoring or real-time auditing tools be developed to detect and signal silent patches or sudden epistemic shifts in deployed LLMs?
- What mechanisms can support more principled and transparent epistemic accountability to users, particularly regarding refusal policies and misinformation endorsement?
Why it matters for bot defense
For bot-defense engineers and CAPTCHA practitioners, this study underscores that LLM outputs are deeply context-dependent not only on input prompts but also on opaque deployment configurations—system prompts, safety layers, versions, and access interfaces all influence epistemic stances unpredictably. This instability complicates efforts to reliably use LLMs as automated knowledge validators or content filters because two identical inputs can produce vastly different outputs depending on endpoint or silent system changes.
In bot or human verification scenarios that employ LLMs to assess content legitimacy or detect misinformation, this research signals caution: validation decisions might reflect deployment artifacts rather than model reasoning. Continuous systematic auditing across interfaces and versions, explicit handling of refusal responses, and awareness of deployment opacity are necessary to maintain trustworthiness. Captcha systems relying on LLMs or similar AI should anticipate these epistemic mediations and embed transparency and fallback checks accordingly.
Cite
@article{arxiv2607_22513,
title={ Opaque Epistemic Mediation: How LLM Deployment Configurations Shape the Validation of Pseudo-Science },
author={ Davide Scarso and Hugo Noronha de Almeida and Joaquim Pina },
journal={arXiv preprint arXiv:2607.22513},
year={ 2026 },
url={https://arxiv.org/abs/2607.22513}
}