What problem does it solve?
This Skill removes the friction of manually driving browser tasks by letting an AI navigate pages, inspect UI state, and perform actions through structured commands.
Core Features & Use Cases
- Navigation and Interaction: Open pages, move through browser history, click, type, select, drag, and manage tabs or windows.
- Page Inspection and Extraction: Snapshot accessibility trees, inspect element text, attributes, values, visibility, and page metadata for reliable automation.
- Testing and Workflow Automation: Fill forms, verify UI flows, capture screenshots or PDFs, record sessions, and wait for specific states during web app testing.
- Use Case: Use this Skill to log into a dashboard, fill a form, wait for the destination page to load, and capture the resulting state for validation.
Quick Start
Ask the agent to open the target website, take an interactive snapshot, and then click or fill the referenced elements needed to complete your web task.