"Zooming In" on Agentic Web Browsers as Assistive Technologies: A Case Study with a Low-Vision Technology Expert
Source: arXiv:2606.24870 · Published 2026-06-23 · By Laura Colazzo, Giuseppe Anzillotti
TL;DR
This paper investigates the potential of Agentic Web Browsers (AWBs), powered by Large Language Models (LLMs), as assistive technologies for visually-impaired users. AWBs autonomously navigate and interact with web content based on natural language inputs, which could transform traditionally visual web interactions into fluid conversational experiences. The study centers on a detailed case involving a low-vision expert who used Perplexity Comet's voice user interface to perform typical web navigation and form-filling tasks.
The results show the participant valued the conversational fluidity and adaptable response detail of the AWB, indicating strong potential for accessibility improvements. However, limitations such as lack of non-visual feedback during processing, reduced user control over autonomous actions, and transparency issues reduced trust and usability. These challenges highlight the need for inclusive design that enhances user agency and trust while preserving interaction flexibility. The authors position AWBs as promising but requiring further development and broader user evaluations before generalizing as assistive tools for web accessibility.
Key findings
- Participant appreciated the conversational fluidity and the system's ability to adapt response detail to context.
- Lack of non-visual feedback during request processing led to uncertainty about system status, especially when delays occurred.
- Limited user control was observed: the AWB autonomously inserted fabricated data in form fields without confirmation and selected options without informing the user.
- These transparency and control issues undermined the participant's trust in the system.
- Despite these limitations, a technically expert user viewed the AWB's flexible interaction paradigm as highly promising for future web accessibility.
- The case study is limited to a single participant with residual vision and cannot be generalized to all visually-impaired users.
- Future work should focus on inclusivity, user control, trust, and comparative evaluations with existing assistive technologies like screen readers.
Methodology — deep read
The authors conducted an in-depth case study with a single male participant aged 31 who has congenital low vision characterized by complete blindness in one eye and minimal residual vision (<1/20) and visual field (<3%) in the other. The participant is an experienced computer engineer with advanced knowledge of generative AI and familiarity with agentic web browsers.
The study used Perplexity Comet, an AWB with a voice user interface (VUI), which the participant had not previously used. Interaction was through voice commands with real-time auditory cues indicating system readiness. The participant was first given time to familiarize himself with the VUI.
Two task scenarios were performed: (1) an exploratory configuration task on a commercial website and (2) form-filling on a public administration portal. The researchers observed the participant, took notes, and conducted a semi-structured interview afterward to gather subjective feedback.
This qualitative approach focused on user experience—evaluating conversational fluidity, control, transparency, and the overall assistive potential of AWBs. The participant’s expert background enabled detailed reflections on strengths and limitations.
The session was conducted in controlled conditions but only involved one subject, so findings offer preliminary insights rather than statistically generalizable conclusions. No quantitative performance metrics or direct baseline comparisons with other assistive technologies were reported. Rather, the study emphasizes exploratory user-centric evaluation of interaction paradigms.
The paper does not specify hardware or exact speech recognition parameters used, and the VUI’s backend LLM details are referenced from Perplexity Comet’s existing architecture. The observational nature and subjective feedback are core to understanding how AWBs may be shaped for accessibility.
No public code or anonymized datasets were generated as this is a single participant qualitative case study.
Example scenario: The participant used voice commands to navigate a commercial site, locating and configuring products. During this, the AWB’s autonomous actions sometimes executed without user confirmation, inserting default options or fabricated inputs, causing loss of transparency and control. The participant noted appreciation for fluid dialogue but desire for clearer feedback and authority over actions.
Technical innovations
- Preliminary empirical evaluation of an agentic web browser (AWB) as an assistive technology specifically targeting low-vision users.
- Use of voice user interface combined with LLM-powered autonomous web navigation to support visually-impaired web interaction.
- Identification of interaction challenges unique to assistive contexts such as lack of non-visual feedback and trust issues arising from autonomous system actions.
- Framing agentic web browsers through the lenses of inclusivity and accessibility to highlight broader UX implications for agentic interfaces.
Limitations
- Single-subject case study limits generalizability to the broader visually-impaired population.
- Participant has residual vision and is a technology expert, which may bias perceptions positively compared to typical users.
- No quantitative performance metrics or direct comparative evaluation with conventional assistive technologies such as screen readers.
- Focus on one AWB implementation (Perplexity Comet) without exploring other agentic browsers or voice interfaces.
- No adversarial evaluation or stress testing under real-world noisy or multi-tasking conditions.
- Lack of non-visual feedback and limited user control are critical limitations impacting trust and usability.
Open questions / follow-ons
- How do agentic web browsers perform as assistive technologies across a diverse set of visually-impaired users, including fully blind individuals?
- Can novel non-visual feedback mechanisms be developed to improve transparency and situational awareness during autonomous web agent processing?
- What interaction design patterns optimally balance autonomous system action and user control to enhance trust without reducing fluidity?
- How do AWBs compare quantitatively with traditional assistive technologies like screen readers in real task scenarios?
Why it matters for bot defense
For bot-defense and CAPTCHA engineers, this study reveals that the emergent agentic web browsing paradigm could change how assistive users interact with web content, potentially reducing reliance on visual or manual navigation cues that CAPTCHAs often exploit. Implementing AWBs as assistive technologies means that CAPTCHA designs may need reconsideration to accommodate voice-driven, autonomous browsing agents that interpret page semantics and perform interactions.
The transparency and control challenges highlighted suggest that users assisted by AWBs may have difficulty understanding or interrupting automated actions, which could lead to security or usability risks when interacting with anti-bot challenges. CAPTCHAs requiring complex visual interpretation or manual input may remain barriers unless designed with agentic user flows in mind. Hence, bot-defense mechanisms should explore compatibility with conversational and voice-driven agent interfaces, optimizing for accessibility without compromising security.
Cite
@article{arxiv2606_24870,
title={ "Zooming In" on Agentic Web Browsers as Assistive Technologies: A Case Study with a Low-Vision Technology Expert },
author={ Laura Colazzo and Giuseppe Anzillotti },
journal={arXiv preprint arXiv:2606.24870},
year={ 2026 },
url={https://arxiv.org/abs/2606.24870}
}