Identifying a Level-up Pathway for AI-assisted Counterspeech through Elaboration
Source: arXiv:2607.28239 · Published 2026-07-30 · By Han Li, Inhwan Bae, Natalie Bazarova, Drew Margolin
TL;DR
This paper investigates how AI assistance can empower ordinary social media users to craft more effective and authentic counterspeech responses to vaccine-skeptical content, a critical challenge in combatting vaccine misinformation online. The authors designed and tested three generative AI writing support systems varying by assistance stage—co-writing (guided and unguided) versus re-writing—and conducted a randomized controlled trial with 550 participants responding to both statistical and narrative vaccine-skeptical posts. Results showed that AI assistance consistently improved perceived counterspeech effectiveness while only moderately reducing perceived authenticity of tone, with little impact on authenticity of thoughts. Perceived effectiveness strongly predicted willingness to publicly post counterspeech, a key behavioral outcome. AI's key contribution was in facilitating more elaborate writing, leading to longer, more analytically rich, and lexically sophisticated counterspeech messages.
Key findings
- All three AI-assisted conditions increased perceived counterspeech effectiveness by 0.29 to 0.48 points on a 7-point Likert scale compared to human-only condition (Table 1).
- AI-assisted writing reduced perceived authentic tone by about 0.46 to 0.54 points, but had smaller or non-significant effects on perceived authentic thoughts.
- Perceived effectiveness was the strongest predictor of willingness to post counterspeech publicly (β=0.28 to 0.36, p<0.001), while authentic tone did not significantly predict posting intention.
- AI rewriting increased message elaborateness measured by analytical thinking, information density, text complexity, and length compared to both initial drafts and human-only messages (all p<0.05, Fig 4).
- Longer, more complex, and information-dense counterspeech predicted greater perceived effectiveness across both statistical and narrative posts.
- Mediation analysis showed perceived effectiveness partially mediated the effect of message elaborateness on posting intention, especially text complexity and information density (Supplementary S3).
- In the guided AI co-writing condition, 86% of sessions used elaboration-oriented AI functions, producing significantly longer, more complex, and information-rich messages (Cohen’s d=0.45-0.70, Fig 5).
- Prior counterspeech experience and verbal argumentativeness modestly but significantly predicted posting intention (β=0.14 to 0.27, p<0.001).
Threat model
Adversaries are social media actors disseminating vaccine-skeptical narratives and statistics that, while not always factually incorrect, sow doubt and erode vaccine confidence. They can craft ambiguous or persuasive misinformation that is hard for moderation systems to detect or remove. Lay counterspeakers lack expert knowledge and face barriers in crafting credible, effective counterspeech. AI-assisted writing aims to address this imbalance by empowering ordinary users to respond more effectively and thus reduce misinformation’s spread and influence.
Methodology — deep read
Threat Model & Assumptions: The adversary is social media misinformation promoters spreading vaccine-skeptical content via statistical misrepresentation or personal narratives. The study focuses on lay counterspeakers as defenders, assuming they lack statistical literacy or persuasive writing skills and need AI assistance.
Data: Four vaccine-skeptical posts were selected from a Reddit community representing both statistical (2 posts) and narrative (2 posts) evidence types. The data collection involved recruiting 550 participants from Prolific, randomly assigned to one of four writing conditions (guided AI co-writing, unguided AI co-writing, AI rewriting, human-only) in a between-subjects design. Each participant wrote counterspeech responses to two posts (one statistical, one narrative), totaling 1100 message pairs. No prior training was provided to preserve ecological validity of lay users' writing.
Architecture/Algorithm: Three generative AI-assisted writing systems were custom-built: guided AI co-writing (interface with specific functions like “Give me a hint” and “make it substantial”), unguided AI co-writing (more open-ended chatbot style interaction), and AI rewriting (refines existing draft). The guided system encoded message quality knowledge into the interface. The AI models used large language model APIs (details unspecified).
Training Regime: Not applicable to human evaluation study; AI assistance used pretrained LLMs.
Evaluation Protocol: Main metrics were user perceptions via 7-point Likert scales of message effectiveness, authentic tone, and authentic thoughts. Statistical analyses used linear mixed-effects models with participant-level random intercepts to analyze repeated measures. Regression analyses identified predictors of posting intention. Manual qualitative coding of 582 message pairs assessed evidentiary form. Linguistic NLP measures (analytical thinking, information density, text complexity, length) quantified message elaborateness. Mediation models tested indirect effects. Robust statistical significance and effect sizes were reported, including Cohen's d.
Reproducibility: Data was from Reddit and Prolific with manual coding by authors. The paper does not specify public release of code or models. Pretrained LLM APIs were used for AI assistance. The experimental procedure details are clear for replication but exact code/models unavailable.
Example: A participant assigned to guided AI co-writing began by invoking “Give me a hint” to generate relevant arguments, then composed a draft, used “make it substantial” to expand content, and finally submitted a longer, analytically richer counterspeech comment. This final message had higher perceived effectiveness and participant was more willing to post it publicly compared to a human-only drafted comment.
Technical innovations
- Design and evaluation of three distinct AI-assisted counterspeech writing systems varying by assistance stage (co-writing vs rewriting) and mode (guided vs unguided).
- Introduction of elaboration-oriented AI assistance functions (e.g., “Give me a hint” and “make it substantial”) that directly scaffold message length, complexity, and information density.
- Systematic decomposition and measurement of message elaborateness into analytical thinking, information density, text complexity, and length, linked to perceived counterspeech effectiveness.
- Empirical evidence identifying perceived message effectiveness as the dominant psychological driver of counterspeech posting intention, outweighing authenticity concerns.
Datasets
- Vaccine-skeptical social media posts — 4 posts (2 statistical, 2 narrative) — collected and adapted from a vaccine-related Reddit community
- Counterspeech responses — 1100 message pairs from 550 Prolific participants writing responses
Baselines vs proposed
- Human-only condition (control): perceived effectiveness = 5.69 vs guided AI co-writing = 6.07
- Human-only: perceived effectiveness = 5.69 vs unguided AI co-writing = 5.98
- Human-only: perceived effectiveness = 5.69 vs AI rewriting = 6.16
- AI rewriting final submissions vs rewrite drafts: significant increases in analytical thinking, information density, text complexity, and length (all p < 0.05)
Figures from the paper
Figures are reproduced from the source paper for academic discussion. Original copyright: the paper authors. See arXiv:2607.28239.

Fig 1: Visual illustrations of (a) experimental procedure and (b) four

Fig 2: Perceived message effectiveness and authenticity across four

Fig 3: Relative importance of factors predicting counterspeech posting

Fig 4: Group comparisons of four message elaborateness features between

Fig 5: Mean differences in four message elaborateness features between

Fig 6 (page 30).

Fig 7 (page 34).

Fig 8 (page 35).
Limitations
- Limited to vaccine-skeptical content; generalizability to other misinformation topics unclear.
- Use of only four stimulus posts constrains variability and ecological validity of vaccine-skeptical discourse.
- No long-term behavioral follow-up to assess actual posting and public impact of counterspeech messages.
- Authenticity effects measured only via self-report perceived tone/thoughts; external authenticity ratings or objective measures missing.
- AI models used are black-box and not fully described; reproducibility of AI assistance not fully open.
- Participant population from Prolific may not fully represent real-world social media counterspeakers in motivation or skill.
Open questions / follow-ons
- How does AI-assisted counterspeech perform in real social media environments, including actual posting and public reception?
- Can the elaboration-oriented AI assistance functions be optimized to better preserve authenticity while enhancing effectiveness?
- What are the effects of AI-assisted counterspeech on diverse user populations with different backgrounds, beliefs, and digital literacy?
- How might adversaries adapt to or exploit AI-assisted counterspeech tools, potentially generating counter-countermeasures?
Why it matters for bot defense
For bot-defense and CAPTCHA practitioners focused on designing human-AI collaboration tools, this study highlights the nuanced tradeoff between AI augmentation of user-generated text and perceived authenticity, a key consideration for community-driven content moderation or counterspeech systems. The finding that AI support that scaffolds elaboration—providing strategic hints and content expansion—enhances message effectiveness without fully undermining user agency could inform CAPTCHA designs that aim to validate genuine human engagement via more complex, elaborated textual inputs. Additionally, recognizing that increased perceived message effectiveness drives public willingness to post counterspeech suggests that CAPTCHA or human verification mechanisms could prioritize richer linguistic features or cognitive effort as signals of authentic human participation to defend against bots or low-effort automated responses. However, the potential authenticity cost implies defenses should balance enhancement with careful UX design to preserve genuine user expression.
Cite
@article{arxiv2607_28239,
title={ Identifying a Level-up Pathway for AI-assisted Counterspeech through Elaboration },
author={ Han Li and Inhwan Bae and Natalie Bazarova and Drew Margolin },
journal={arXiv preprint arXiv:2607.28239},
year={ 2026 },
url={https://arxiv.org/abs/2607.28239}
}