Skip to content

AI systems and the reproduction of (standard) language ideologies in World Englishes

Source: arXiv:2607.28528 · Published 2026-07-30 · By Kingsley Ugwuanyi

TL;DR

This paper critically examines how large language models (LLMs) and associated AI systems reproduce, reinforce, and occasionally challenge entrenched standard language ideologies in the context of World Englishes. It argues that AI systems privilege Inner Circle varieties of English (e.g., US and UK English) as normative benchmarks, marginalizing Global South and non-dominant English varieties. This occurs through biases embedded at multiple levels—training data, design choices, evaluation benchmarks, and user feedback—reflecting sociolinguistic hierarchies long studied in language ideology research. The paper uses a prominent public controversy over the lexical item “delve” to illustrate how AI-generated language becomes a site of linguistic policing that mirrors colonial and ideological biases against Global South Englishes. At the same time, the author highlights a "standardisation paradox" whereby AI simultaneously homogenizes English through standard forms while also exposing diverse Englishes through vast corpora and annotator inputs from the Global South.

Drawing on empirical studies, media discourse, social media debates, and AI output analyses, the paper reveals that AI systems are ideological actors shaping what counts as legitimate English—not neutral tools. Public reactions to AI language perpetuate these language hierarchies, but the digital space also enables resistance and contestation. The paper argues for inclusive AI design that embraces English pluricentricity and cautions that ignoring these sociolinguistic factors risks reproducing harmful linguistic inequalities with real-world consequences for speakers of marginalized Englishes.

Key findings

  • Training datasets overrepresent Inner Circle English varieties, embedding a dataist ideology privileging standardized English (citing Erdocia et al. 2024).
  • Common NLP benchmarks such as GLUE and SuperGLUE focus on American English subsets, biasing evaluation outcomes toward one standard (Raji et al. 2021).
  • Fine-tuning with human annotators often promotes standard English norms, causing annotators, even from the Global South, to internalize and reproduce linguistic hierarchies (Erfani 2026).
  • Fleisig et al. (2024) found LLMs like GPT-3.5 and GPT-4 exhibit 19% more stereotyping, 25% more demeaning content, and 9% more comprehension failures on non-standard English varieties compared to Standard American English.
  • Lin et al. (2024) showed widely used LLMs struggle with African American Vernacular English (AAVE), providing irrelevant or incoherent responses despite semantically equivalent prompts.
  • Corpus analysis indicates the frequency of “delve into” usage tripled in 2023–2024 coinciding with generative AI uptake but is unevenly distributed geographically, with higher use in Asian Englishes compared to Nigerian or Ghanaian Englishes (NOW corpus).
  • Public discourse often incorrectly attributes increased AI use of words like “delve” to Global South English influence, reflecting ideological biases rather than empirical linguistic patterns (e.g., Hern 2024, Aggarwal 2024).
  • AI systems embody a 'standardisation paradox' by both homogenizing English through normative training and simultaneously exposing models to diverse Englishes through large-scale data (Mair 2023, 2025).

Threat model

The adversary conceptualized is the embedded language ideology privileging Inner Circle, standardized English varieties within AI systems—arising from training data biases, model design choices, evaluation benchmarks, and social discourse. This ideological adversary marginalizes non-dominant Englishes and enforces linguistic hierarchies not through direct malicious intent but systemic socio-technical decisions and institutional power dynamics.

Methodology — deep read

  1. Threat Model & Assumptions: The paper situates AI systems and LLMs as ideological actors producing language outputs shaped by sociolinguistic biases rather than neutral generators. The 'adversarial' forces examined include embedded language hierarchies privileging Inner Circle Englishes and marginalizing non-dominant varieties, both in training data and public discourse. It assumes AI developers and users often unconsciously replicate ingrained language ideologies.

  2. Data: The author uses publicly available corpora such as the NOW corpus (2010–present) to analyze lexical frequency over time, especially focusing on the usage dynamics of the lexical item “delve.” Empirical findings and references from prior studies (e.g., Fleisig et al. 2024, Lin et al. 2024) about LLM performance on dialectal inputs support claims. Media articles, social media posts, and public commentary provide qualitative discourse data.

  3. Architecture/Algorithm: While no new AI architectures are proposed, the paper synthesizes analyses of current LLM design practices including training on large-scale datasets like Common Crawl and Wikipedia; fine-tuning incorporating reinforcement learning with human feedback (RLHF); and evaluation practices relying on benchmarks like GLUE and SuperGLUE. It discusses how these design components embed ideology by favoring certain language norms.

  4. Training Regime: The paper does not conduct original training but draws on referenced works describing training data sourcing, annotator demographics and instructions, and fine-tuning strategies that enforce linguistic standards and quality controls aligned with Inner Circle norms.

  5. Evaluation Protocol: The paper reviews and interprets findings from multiple empirical evaluation studies assessing fairness, bias, and performance disparities of LLMs on non-standard Englishes and dialects. Metrics include stereotypical and demeaning content rates, comprehension failure rates, and reasoning performance on dialect-specific prompts. It highlights the absence of equitable benchmarking for diverse Englishes.

  6. Reproducibility: As a critical sociolinguistic analysis rather than a technical AI experiment, no code or models are released. The paper relies on a combination of corpus linguistic analysis, secondary empirical AI evaluation studies, media analysis, and public digital discourse. It transparently acknowledges limitations of data provenance and the indirectness of some evidence.

A concrete example discussed is the lexical frequency trajectory of the word “delve” in AI-generated outputs and public usage. The author examines a dataset of 50,000 ChatGPT responses listing common words, then cross-references with frequency data from the NOW corpus revealing a near tripling of use in 2023–2024. Despite this, public commentary misattributes “delve” usage predominantly to African English influence, exposing how language ideology shapes perception beyond linguistic reality.

Technical innovations

  • Conceptualization of LLMs as ideological actors embedding and reproducing standard language ideologies rather than neutral language tools.
  • Identification and analysis of the 'standardisation paradox' in AI language models—simultaneous homogenization and diversification of English varieties.
  • Use of public controversy around a specific lexical item ('delve') as an empirical lens to reveal mechanisms of linguistic policing and ideological reproduction in AI discourse.
  • Highlighting socio-technical AI design decisions (training data curation, RLHF) as active sites where language hierarchies are embedded and maintained.

Datasets

  • NOW corpus — millions of news articles from 2010 to present — public corpus (https://www.english-corpora.org/now/)
  • ChatGPT-generated text dataset (AI Phrase Finder 2024) — 50,000 responses — source not publicly specified

Baselines vs proposed

  • Standard American English (baseline) vs. non-standard English varieties: Fleisig et al. (2024) find 19% more stereotyping, 25% more demeaning content, and 9% more comprehension failures on non-standard inputs.
  • LLM reasoning performance on AAVE prompts: Lin et al. (2024) observe significant brittleness and unfairness compared to Standard American English.

Limitations

  • The study is primarily qualitative and analytical; it does not present new experimental results or quantitative AI model training.
  • Reliance on secondary empirical studies for bias performance metrics means some details about evaluation protocols are not fully explicated.
  • Corpus analysis on lexical frequency (e.g., 'delve') captures written registers but may not represent spoken or informal usage.
  • The critique of AI systems as ideological actors is conceptual; causal links between design decisions and sociolinguistic outcomes could be further empirically validated.
  • Public discourse analysis focuses largely on English varieties from the Global South, which may not generalize to other global language contexts.
  • Data on annotator demographics and labor practices (e.g., pay disparities) is drawn from few studies and may not reflect the full diversity of AI data curation workflows.

Open questions / follow-ons

  • How can AI training datasets be curated or augmented to authentically represent and legitimize diverse English varieties without reinforcing hierarchical bias?
  • What concrete architectural or fine-tuning methodologies can mitigate dialect discrimination in large language models?
  • To what extent can public discourse and user feedback drive changes in AI language norms toward greater pluricentric inclusivity?
  • How do language ideologies embedded in AI affect real-world outcomes like hiring, education, and justice for speakers of marginalized Englishes?

Why it matters for bot defense

For practitioners in bot defense and CAPTCHA design, this paper underscores the importance of recognizing that AI language models and their evaluation are not linguistically neutral—they disproportionately reflect and reinforce dominant English language norms. This can lead to biased system behavior and unfair treatment of users who employ non-standard or Global South English varieties in interactions, including challenge-response tests or language-based authentication methods. Designers should critically assess linguistic assumptions in training and benchmark datasets to avoid excluding or misclassifying legitimate user inputs that diverge from the Inner Circle English norm.

Moreover, the highlighted 'standardisation paradox' implies that while inclusivity efforts face structural constraints, there is potential for adaptive model components (e.g., dialect adapters) to improve equitable performance across Englishes. Bot-defense systems incorporating NLP technologies should consider integrating or fine-tuning for local linguistic varieties to reduce false positives and better reflect the linguistic diversity of users worldwide. Finally, observing how public perception and language policing manifest online can inform how automated challenges are designed to minimize reinforcing stigmatizing attitudes toward marginalized English speakers.

Cite

bibtex
@article{arxiv2607_28528,
  title={ AI systems and the reproduction of (standard) language ideologies in World Englishes },
  author={ Kingsley Ugwuanyi },
  journal={arXiv preprint arXiv:2607.28528},
  year={ 2026 },
  url={https://arxiv.org/abs/2607.28528}
}

Read the full paper

Articles are CC BY 4.0 — feel free to quote with attribution