arXiv · 2503.13843
WebNav: An Intelligent Agent for Voice-Controlled Web Navigation
Abstract
The current state of modern web interfaces, especially in regards to accessibility focused usage is extremely lacking. Traditional methods for web interaction, such as scripting languages and screen readers, often lack the flexibility to handle dynamic content or the intelligence to interpret high-level user goals. To address these limitations, we introduce WebNav, a novel agent for multi-modal web navigation. WebNav leverages a dual Large Language Model (LLM) architecture to translate natural language commands into precise, executable actions on a graphical user interface. The system combines vision-based context from screenshots with a dynamic DOM-labeling browser extension to robustly identify interactive elements. A high-level 'Controller' LLM strategizes the next step toward a user's goal, while a second 'Assistant' LLM generates the exact parameters for execution. This separation of concerns allows for sophisticated task decomposition and action formulation. Our work presents the complete architecture and implementation of WebNav, demonstrating a promising approach to creating more intelligent web automation agents.
Explore related subjects
Keep this discovery
Trisanth Srinivasan, Santosh Patapati. 2025-03-18. WebNav: An Intelligent Agent for Voice-Controlled Web Navigation. https://arxiv.org/abs/2503.13843
Cite the original work for its findings. Save a collection to share your selection of sources.