The pitch for the AI browser is simple: instead of you clicking through the web, the browser clicks for you. You ask, it books, fills, researches, and reports back. After a year of launches — Perplexity's Comet, The Browser Company's Dia, and OpenAI's ChatGPT Atlas in late 2025 — the marketing has settled into a confident hum. The reality, once you actually hand these browsers errands to run, is more nuanced: they are startlingly good at research, competent at form-filling with a human watching, and still unreliable the moment money, logins, or long multi-step plans enter the picture.

This review synthesizes independent hands-on tests and published benchmarks to map exactly where agent browsers succeed and where they stall.

The contenders#

Three browsers dominate the current conversation, and they take philosophically different approaches:

  • Perplexity Comet — built on Chromium, with Perplexity's retrieval engine at its core. Its assistant has contextual awareness of all open tabs, and it is the only one of the three to offer unlimited agentic actions on its free tier, according to Gadgets 360's head-to-head testing.
  • ChatGPT Atlas — OpenAI's entry, launched October 2025. A polished, minimalist Chromium browser with ChatGPT as the sidebar assistant and an "agent mode" that takes control of web pages. Even paid subscribers get only 40 agentic actions per day.
  • Dia — from The Browser Company (the team behind Arc). The most design-forward of the three: a conversational layer woven through tabs, notes, and chats. Strong on personalization, weakest on actual agentic actions.

Chrome itself, with Gemini integrations, remains the default for most people, and reviewers at Tom's Guide have noted it still beats Atlas on shopping assistance and tab management in places where Gemini is fully available. But the dedicated agent browsers are where the "do the clicking" claim gets tested.

What the hands-on tests actually show#

A ten-task head-to-head between Atlas and Comet (published November 2025) is revealing about the grain of the experience. On finding and applying promo codes, Comet found a working code in under two minutes; Atlas required explicitly activating agent mode first and took longer. In a second round on Overstock, Atlas found a code first but kept trying additional codes even after one worked — wasting time on an already-solved problem — while Comet stopped once it found a working $40 discount.

That pattern — Atlas being capable but less disciplined about knowing when it's done — shows up across tests. Gadgets 360's comparison found Comet could find a product, add it to the cart, add a delivery address, and then hand over the reins at the checkout page rather than completing the purchase itself. That handoff is the honest design choice, and it's where the industry has quietly converged: agents do the browsing, humans do the paying.

On research tasks, the differences sharpen. A 20-task benchmark across five AI browsers (published May 2026) measured hallucination rates with fake URLs counted as automatic failures:

BrowserHallucination rateNotes
Perplexity Comet (Pro Search)4%Best in class; flagged conflicting datasets
Opera Neon11%Strong on PDF summarization
ChatGPT Atlas14%Notable for hallucinated citation URLs
Brave Leo18%Good single-tab summaries, weak multi-tab
Dia~30%Accurate but shallow on deep research

Comet surfaced valid primary sources on 18 of 20 tasks and hyperlinked live URLs 96% of the time. Atlas executed multi-step tasks faster but invented citation URLs under academic pressure. And no browser — none — could retrieve full text from strictly paywalled journals without authenticated institutional access.

What the lab benchmarks say#

Hands-on reviews are anecdotal by nature, so it's worth checking the formal benchmarks, which are sobering:

  • WebArena (2023): the best GPT-4-based agent completed 14.41% of end-to-end web tasks against a human rate of 78.24%. Recent aggregator-reported submissions put top browser agents in the 60–70% range — dramatic progress, though leaderboard figures vary and aren't all third-party verified.
  • VisualWebArena: the original best multimodal agent (GPT-4V with set-of-mark grounding) reached just 16.4% against an 88.7% human baseline, exposing how hard visual grounding — mapping what the agent sees to precise pixel coordinates — really is.
  • WebChoreArena (2025): this newer benchmark extends WebArena into tedious, memory-heavy work, and it's where modern agents fall apart. The best setup (Gemini 2.5 Pro with BrowserGym) dropped from 59.2% on standard WebArena to 44.9% on WebChoreArena. Agents are decent at short, sharp tasks and bad at long, boring ones — the exact opposite of the marketing promise.

The through-line: scoped tasks have gone from mostly failing to mostly succeeding in about three years, but tasks requiring sustained memory, calculation across pages, and judgment still crater performance.

Where they stall: the failure map#

Across tests, the same failure modes recur. Treat this as a checklist before you trust an agent browser with anything important:

  1. Checkout and payments. By design, agents stop at the payment step — Comet literally hands the page back to you. Anything requiring your credit card, or your judgment about spending money, is out of scope. This is a feature, not a bug, but it means "book my flight" really means "fill in my flight search."
  2. Logins and authentication. Agents struggle with multi-factor auth, SSO flows, and session quirks. If a task requires being logged in, expect to babysit.
  3. Knowing when to stop. Atlas's promo-code test is the canonical example: finding a working answer and then continuing to search anyway. Premature stopping is less common than redundant grinding, and both waste your time and your daily action quota.
  4. Hallucinated citations. Atlas's 14% hallucination rate under research pressure means every citation needs a click-through check. A browser that finds real answers but cites fake sources is worse than one that admits ignorance.
  5. Paywalls and CAPTCHAs. No agent browser defeats a paywall without your credentials, and bot-detection systems (CAPTCHAs, Cloudflare challenges) remain a hard wall for autonomous clicking.
  6. Action limits. Atlas caps agentic actions at 40 per day even for paid users — a genuine constraint for anyone planning to run real workflows. Comet's unlimited free-tier actions are a meaningful differentiator here.
  7. Prompt injection and privacy. Berkeley's AgentWatch evaluation scored Atlas at 93.6 and Comet at 86.1 on privacy and safety behavior, with Comet weaker on ambiguous prompts and prompt-injection scenarios. An agent that reads web pages can be manipulated by web pages — this is a live attack surface, and caution around sensitive accounts is warranted.

The practical verdict#

So who should use what? The honest answer depends on the errand:

  • Research and comparison shopping: Comet. The citation transparency, multi-tab awareness, and low hallucination rate make it the strongest research tool of the three, and unlimited free agentic actions mean you can actually use it.
  • General productivity inside the ChatGPT ecosystem: Atlas, but only with a paid subscription — the free tier's locked sidebar and the 40-action daily cap make the free experience frustrating.
  • A conversational, personalized browsing layer: Dia. Just don't expect it to complete errands; it's a companion, not an agent.

And for everyone: the current sweet spot is agent does the legwork, human does the committing. Let the browser gather options, compare prices, fill drafts, and summarize. Take back the wheel for logins, payments, and anything irreversible. Check every citation. Budget for the agent grinding past the point of done.

The trajectory is real — three years ago agents failed almost everything; now they handle scoped tasks well. But "browser that does your errands" is still a co-pilot pitch wearing an autopilot costume. Dress accordingly.

Takeaway#

AI browsers have crossed the line from demo to genuinely useful for research, form prep, and shopping legwork — Comet leads on research accuracy and free agentic actions, Atlas on ChatGPT-integrated workflows (behind a paywall and action cap), Dia on conversational feel. But formal benchmarks show agents still collapse on long, memory-heavy tasks, and every hands-on test converges on the same boundary: agents browse, humans buy. Use them as tireless research assistants with a verification habit, not as autonomous errand-runners.