Your LLM can write the code, summarize the page, and draft the email — but it cannot click the button. browser-use closes that gap. It is an MIT-licensed Python library with nearly 117,000 GitHub stars that turns any LLM into an agent which sees web pages, clicks, types, and scrolls the way a person does. This tutorial installs the open-source library, runs a real task end to end, swaps in the project's own BU2 model, adds a custom tool, and points the agent at your actual Chrome profile.

1. Why this one

Most "browser agents" are closed products with a chat box. browser-use is the opposite: a library you import. The same Agent runs against OpenAI, Anthropic, Google, or Browser Use's own BU2 model, on a local browser or a managed cloud one. It ships three paths — a fully hosted cloud API, a CLI skill that plugs into agents like Claude Code or Codex, and the Python library we use here. The project publishes its own 60-task benchmark (BU Bench v2) and reports 87.4% on the 200-task Odysseys long-horizon suite, ahead of the computer-use agents from the big labs — the rare case where the marketing claim ships with the harness.

2. What you'll need

  • Python 3.11+ and uv — the install path the README documents.
  • An LLM API key: OPENAI_API_KEY covers the quickstart; BROWSER_USE_API_KEY unlocks the BU2 model and cloud browsers.
  • A machine that can launch a browser. The agent drives a real Chromium, headed or headless.
  • No Browser Use account for the local path. The library itself is free and MIT-licensed — you pay only your model provider.

3. Step 1 — Install the library

From your project directory (the uv init line is only needed when starting a fresh project):

uv init --python 3.12   # new project only — skip if you have one
uv add browser-use

4. Step 2 — Add your model key

# .env
OPENAI_API_KEY=...
# BROWSER_USE_API_KEY=...   # only needed for BU2 or cloud browsers

5. Step 3 — Write your first agent

Save this as agent.py. The shape comes straight from the project's quickstart: pick a model, hand the agent a task in plain English, run it, and read history.final_result(). There is no system prompt to write — Agent supplies its own, even when you change models.

import asyncio

from browser_use import Agent, ChatOpenAI
from dotenv import load_dotenv

load_dotenv()


async def main():
    llm = ChatOpenAI(model="gpt-5")

    agent = Agent(
        task="Find the number of stars of the browser-use repo",
        llm=llm,
    )

    history = await agent.run()
    print(history.final_result())


if __name__ == "__main__":
    asyncio.run(main())

6. Step 4 — Run it

uv run agent.py

The agent opens a browser, navigates to GitHub, reads the repository page, and prints the star count. Watch the first run closely: the agent narrates each action it takes — navigate, click, type, scroll, extract — which is the fastest way to learn what it can and cannot see on a page.

A glowing cursor tracing a dotted path of planned clicks across a dark browser window
AI-generated illustration for AI Frontier Post

7. Step 5 — Swap models without touching the agent

BU2 is Browser Use's own model, tuned for browser automation and recommended in the docs. Provider-prefixed IDs route through the Browser Use gateway on the same key; direct wrappers like ChatAnthropic and ChatGoogle use each provider's own key instead.

from browser_use import Agent, ChatBrowserUse

llm = ChatBrowserUse(model='bu-2-0')  # needs BROWSER_USE_API_KEY

# ...or any provider directly, with that provider's key
# from browser_use import ChatAnthropic, ChatGoogle
# llm = ChatAnthropic(model='claude-sonnet-4-6', temperature=0.0)
# llm = ChatGoogle(model="gemini-3-pro-preview")

agent = Agent(task="...", llm=llm)

BU2 pricing is published per million tokens: $0.60 input, $0.06 cached, $3.50 output. The task string and the rest of your code stay identical — only the llm line changes.

8. Step 6 — Teach it your own tools

The agent can call your Python functions mid-task. Register one with Tools and pass it in — the example below is the pattern from the project's own docs:

from datetime import datetime, timezone

from browser_use import ActionResult, Tools

tools = Tools()

@tools.action(description='Get the current date and time in UTC.')
def get_current_time() -> ActionResult:
    return ActionResult(extracted_content=datetime.now(timezone.utc).isoformat())

agent = Agent(task="...", llm=llm, tools=tools)

This is how browser tasks stop being demos and start being pipelines: the agent browses, your tool writes the result to a database, and the next step reads it back.

9. Step 7 — Drive your real browser profile

For sites where you are logged in, reuse your own Chrome profile instead of a fresh browser:

from browser_use import Browser

browser = Browser.from_system_chrome()  # reuses your Chrome profile
agent = Agent(task="...", llm=llm, browser=browser)

Prefer not to run browsers yourself? Browser(use_cloud=True) moves the browser to Browser Use's infrastructure — stealth, proxy rotation, and CAPTCHA handling included at $0.02 per browser-hour. New Google, GitHub, or Microsoft signups get $15 in cloud credit.

A robot eye watching over stacked translucent browser windows, representing cloud browser automation
AI-generated illustration for AI Frontier Post

10. What you built

A script that takes a sentence and drives a real browser to completion, with model choice, custom tools, and browser target each controlled by a single line. From here the natural extensions are scheduled runs (the Python library parallelizes cleanly for scraping and monitoring), the CLI path — browser-use skill install hands browser control to Claude Code, Codex, or OpenClaw — and the Anthropic SDK integration, which exposes all 31 browser actions plus a Bash tool to Claude directly.

11. Honest limitations

  • The library is free; inference and hosted browsers are not. Estimate cost per run before you loop a task overnight.
  • CAPTCHAs are the hard wall. Cloud browsers reduce bot detection, but the docs are explicit: no configuration guarantees every CAPTCHA is avoided or solved.
  • Cloud profile sync moves cookies only — not local storage, IndexedDB, or extensions — so some sites will ask you to sign in again.
  • Browser agents are slow and token-hungry next to API automation. If a site has an API, use the API; save the agent for the sites that don't.
  • An agent driving your logged-in browser can do everything you can. Develop against a throwaway profile first.