
What it can do
Navigate & read
Open pages, follow links, and extract content and data.
Interact
Click buttons, fill forms, select options, submit.
Log in
Sign into sites, with your help for credentials when needed.
Capture
Take screenshots to show you what it sees.
The shared browser
In web chat, the agent’s browser is shared with you. It appears in the Computer panel. Watch it work, or take over whenever a human is faster:- Logins: sign in yourself instead of handing over credentials.
- Captchas & 2FA: solve the human-verification step the agent can’t.
- Payments & confirmations: take over for anything sensitive.
Terminal
The Computer panel also has a Terminal tab. Open it to get a live shell into the agent’s sandbox, useful for inspecting what the agent built, running a quick command, or debugging something without asking the agent to do it for you. The terminal is only shown when you have edit access on the agent.Search vs. browse
The agent uses the lighter tool first. For reading public information it uses web search and fetch. It opens the full browser when it needs to interact: clicking, filling, logging in, or when a page needs a real browser to render.Good uses
- Pulling data from a dashboard that has no API
- Filling out a web form on your behalf
- Checking a page that requires a login
- Capturing a screenshot of how something looks
- Researching across sites that block simple fetching
For anything requiring authentication, a captcha, or a payment, the agent hands off to you in the shared browser.
Connect accounts for repeat access
Authorize a service once so the agent doesn’t need the browser to log in every time.