Give your AI agent a browser: hands-on with browser-use, the 117K-star library
browser-use (nearly 117K stars, MIT) turns any LLM into an agent that clicks, types, and scrolls like a human. Install the Python library, wire up your model, and drive a real browser from a script — in about 15 minutes.

Your LLM can write the code, summarize the page, and draft the email — but it cannot click the button. browser-use closes that gap. It is an MIT-licensed Python library with nearly 117,000 GitHub stars that turns any LLM into an agent which sees web pages, clicks, types, and scrolls the way a person does. This tutorial installs the open-source library, runs a real task end to end, swaps in the project's own BU2 model, adds a custom tool, and points the agent at your actual Chrome profile.
1. Why this one
Most "browser agents" are closed products with a chat box. browser-use is the opposite: a library you import. The same Agent runs against OpenAI, Anthropic, Google, or Browser Use's own BU2 model, on a local browser or a managed cloud one. It ships three paths — a fully hosted cloud API, a CLI skill that plugs into agents like Claude Code or Codex, and the Python library we use here. The project publishes its own 60-task benchmark (BU Bench v2) and reports 87.4% on the 200-task Odysseys long-horizon suite, ahead of the computer-use agents from the big labs — the rare case where the marketing claim ships with the harness.
2. What you'll need
- Python 3.11+ and
uv— the install path the README documents. - An LLM API key:
OPENAI_API_KEYcovers the quickstart;BROWSER_USE_API_KEYunlocks the BU2 model and cloud browsers. - A machine that can launch a browser. The agent drives a real Chromium, headed or headless.
- No Browser Use account for the local path. The library itself is free and MIT-licensed — you pay only your model provider.
3. Step 1 — Install the library
From your project directory (the uv init line is only needed when starting a fresh project):
uv init --python 3.12 # new project only — skip if you have one
uv add browser-use
4. Step 2 — Add your model key
# .env
OPENAI_API_KEY=...
# BROWSER_USE_API_KEY=... # only needed for BU2 or cloud browsers
5. Step 3 — Write your first agent
Save this as agent.py. The shape comes straight from the project's quickstart: pick a model, hand the agent a task in plain English, run it, and read history.final_result(). There is no system prompt to write — Agent supplies its own, even when you change models.
import asyncio
from browser_use import Agent, ChatOpenAI
from dotenv import load_dotenv
load_dotenv()
async def main():
llm = ChatOpenAI(model="gpt-5")
agent = Agent(
task="Find the number of stars of the browser-use repo",
llm=llm,
)
history = await agent.run()
print(history.final_result())
if __name__ == "__main__":
asyncio.run(main())
6. Step 4 — Run it
uv run agent.py
The agent opens a browser, navigates to GitHub, reads the repository page, and prints the star count. Watch the first run closely: the agent narrates each action it takes — navigate, click, type, scroll, extract — which is the fastest way to learn what it can and cannot see on a page.

7. Step 5 — Swap models without touching the agent
BU2 is Browser Use's own model, tuned for browser automation and recommended in the docs. Provider-prefixed IDs route through the Browser Use gateway on the same key; direct wrappers like ChatAnthropic and ChatGoogle use each provider's own key instead.
from browser_use import Agent, ChatBrowserUse
llm = ChatBrowserUse(model='bu-2-0') # needs BROWSER_USE_API_KEY
# ...or any provider directly, with that provider's key
# from browser_use import ChatAnthropic, ChatGoogle
# llm = ChatAnthropic(model='claude-sonnet-4-6', temperature=0.0)
# llm = ChatGoogle(model="gemini-3-pro-preview")
agent = Agent(task="...", llm=llm)
BU2 pricing is published per million tokens: $0.60 input, $0.06 cached, $3.50 output. The task string and the rest of your code stay identical — only the llm line changes.
8. Step 6 — Teach it your own tools
The agent can call your Python functions mid-task. Register one with Tools and pass it in — the example below is the pattern from the project's own docs:
from datetime import datetime, timezone
from browser_use import ActionResult, Tools
tools = Tools()
@tools.action(description='Get the current date and time in UTC.')
def get_current_time() -> ActionResult:
return ActionResult(extracted_content=datetime.now(timezone.utc).isoformat())
agent = Agent(task="...", llm=llm, tools=tools)
This is how browser tasks stop being demos and start being pipelines: the agent browses, your tool writes the result to a database, and the next step reads it back.
9. Step 7 — Drive your real browser profile
For sites where you are logged in, reuse your own Chrome profile instead of a fresh browser:
from browser_use import Browser
browser = Browser.from_system_chrome() # reuses your Chrome profile
agent = Agent(task="...", llm=llm, browser=browser)
Prefer not to run browsers yourself? Browser(use_cloud=True) moves the browser to Browser Use's infrastructure — stealth, proxy rotation, and CAPTCHA handling included at $0.02 per browser-hour. New Google, GitHub, or Microsoft signups get $15 in cloud credit.

10. What you built
A script that takes a sentence and drives a real browser to completion, with model choice, custom tools, and browser target each controlled by a single line. From here the natural extensions are scheduled runs (the Python library parallelizes cleanly for scraping and monitoring), the CLI path — browser-use skill install hands browser control to Claude Code, Codex, or OpenClaw — and the Anthropic SDK integration, which exposes all 31 browser actions plus a Bash tool to Claude directly.
11. Honest limitations
- The library is free; inference and hosted browsers are not. Estimate cost per run before you loop a task overnight.
- CAPTCHAs are the hard wall. Cloud browsers reduce bot detection, but the docs are explicit: no configuration guarantees every CAPTCHA is avoided or solved.
- Cloud profile sync moves cookies only — not local storage, IndexedDB, or extensions — so some sites will ask you to sign in again.
- Browser agents are slow and token-hungry next to API automation. If a site has an API, use the API; save the agent for the sites that don't.
- An agent driving your logged-in browser can do everything you can. Develop against a throwaway profile first.