AI agent browser automation gives an agent a live website it can inspect and operate. The agent can follow a changing page, choose the next action, enter data, work across tabs, and handle files. That turns the browser from a source of information into a tool for completing multi-step work.
What is AI agent browser automation?
AI agent browser automation is a feedback loop: the browser exposes the current page, the model chooses an action, and a browser tool performs it and returns the resulting state for the next decision.
Many conventional browser scripts are built around predetermined steps and locators. An agent can instead choose a new step when a page presents a different control, opens a dialog, or sends the task into another tab. The browser is still only one tool: the caller or system prompt supplies the goal, while the agent chooses actions and evaluates when the work is complete.
The same feedback pattern appears in broader computer-use systems. OpenAI describes its computer-using agent as a loop in which the model sees the current screen, returns an action, and receives an updated screen after execution in its computer-using agent research. Browser agents can also use structured page snapshots, which make text and controls directly addressable without relying on image coordinates for every step.
How AI agents browse and use websites
A browser agent repeats four steps until it reaches the requested outcome. Each step begins from the state returned by the previous one.
| Stage | What happens | Stage result |
|---|---|---|
| Observe | Read the page structure or capture the visible viewport. | Text, controls, element references, or a screenshot |
| Decide | Choose the next action that advances the task. | A target, action, and any required value |
| Act | Click, type, select, scroll, navigate, or operate the mouse. | A changed page, tab, dialog, or downloaded file |
| Verify | Inspect the result before continuing or finishing. | Fresh page state for the next decision |
Structured snapshots are usually the clearest starting point. They return page text and current element references such as @e1. After navigation or another page-changing action, the agent observes again and uses the new references. Screenshots and coordinate-based mouse controls cover visual interfaces that do not expose a useful text control.
What browser-enabled agents can do
When browser automation is the right tool
Browser automation is most useful when the website itself is the available interface: an authenticated portal, a dynamic application, a multi-step form, or a workflow that crosses several pages. A browser is not the default for every web task. Search, page reading, and APIs are often faster and more predictable for narrower jobs.
| Tool | Best fit | Typical result |
|---|---|---|
| Web search | Discovering relevant public pages | Titles, links, and result excerpts |
| Page reading | Extracting readable content from a known public page | Clean text or Markdown |
| API or connector | Stable, structured operations with a supported interface | Validated request and response data |
| Browser automation | Dynamic, authenticated, or UI-only multi-step work | A completed interaction and its resulting state or file |
Connic exposes search, page-reading, and browser tools to the same agent. The web tools documentation covers their inputs and outputs.
How to build a browser agent with Connic
Add web_browser_* to the agent’s tools. This enables the complete browser toolkit, including page interaction, screenshots, tabs, dialogs, and file handling. Connic opens and manages the browser during the run; the agent project needs no separate browser service or browser-control code. The browser tools documentation covers the individual operations.
Browser Automation Demo: log in and create a project
The Browser Automation Demo uses Connic’s own dashboard as its target. It combines browser actions with human input: the person supplies the login details and project title, and the agent completes the form and returns the project’s name and URL.
- Ask for the Connic account email, password, and new project title through three separate human inputs.
- Open the Connic homepage, find the sign-in page, and log in.
- Create a project with the supplied title using Deploy via CLI.
- Return the created project’s name and URL.
The template’s complete agent file defines this task:
version: "1.0"
name: browser-automation-demo
model: connic/gpt-5.6-sol
description: "Browser automation demo that signs in to Connic and creates a project"
system_prompt: |
Ask for the user's Connic email, password, and new project title using the
three human-input tools.
Open https://connic.co, sign in, and create a project with that title using
Deploy via CLI.
Return the created project's name and URL.
tools:
- web_browser_*
approval:
inputs:
- get_connic_email:
prompt: "Ask for the email address used to sign in to Connic."
label: "Connic account email"
- get_connic_password:
prompt: "Ask for the password used to sign in to Connic."
label: "Connic account password"
sensitive: true
- get_project_title:
prompt: "Ask for the title of the new Connic project."
label: "New project title"Each approval.inputs entry creates a tool that pauses the run for a human response. The password uses sensitive: true so its input field is masked and its exact value is redacted from captured traces and logs. The model receives the password to complete the login.
Run the Browser Automation Demo
Install the template, including the complete agent file:
pip install connic-composer-sdk
connic init browser-demo --templates=browser-automation-demoAdd it to a Connic project using the quickstart, then start the agent manually and leave the input empty. Answer the three requests as they appear. The agent returns the created project’s name and URL. Page observations and actions appear in the run trace.
The Browser Automation Demo template includes the full source and setup instructions. To adapt it to another website, change the task in the prompt and the human inputs it needs.
Reuse browser state across runs
The demo asks for login details on each run. For recurring tasks, Connic can also preserve cookies and local storage in a persistent session. A later run can reuse that state and remain signed in, provided the website still accepts the session.
Each run opens a new live browser: open tabs, the current page, element references, and downloads are not carried over. The agent opens the required URL and reads its current state. The session configuration controls which runs share saved browser state.
Connic keeps the page available between tool calls and closes the browser when the run ends. The agent can close it earlier when its task is complete.
Install the Browser Automation Demo, run it with your own inputs, and adapt the workflow to the website you need.
Open the Browser Automation Demo