Browser Automation for Agents arenaBrowser Automation for Agents
The agent-driving-browser layer: frameworks and infrastructure that let AI agents operate real web browsers — natural-language tasks, act/extract/observe primitives, DOM versus vision action modes, persistent logins, stealth and captcha posture, parallel session fleets, and replayable debugging. Boundary note: web-to-structured-data extraction APIs (including Browserbase) are judged in the Web Scraping APIs arena; this arena is about agents ACTING on the web — filling forms, clicking through flows, completing tasks.
52 user stories · 364 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Skyvern ↗ | 58/100 | 15/100 | 46/100 | 53/100 | 7/100 | ★ 7.6k▲ 4.1k/yrnpm 23.7k/wk | 21/28 verified | 73/100 integrity | ||
| 2 | free-tier vs Steel ↗ | 42/100 | 62/100 | 0/100 | 39/100 | 7/100 | ★ 23k▲ 9k/yr | 13/32 verified · 3 disputed | 29/100 integrity | ||
| 3 | free-tier vs Steel ↗ | 49/100 | 60/100 | 10/100 | 14/100 | 9/100 | ★ 789▲ 509/yrnpm 41k/wk | 8/28 verified · 4 disputed | 0/100 integrity | ||
| 4 | usage-based vs Steel ↗ | 51/100 | 69/100 | 0/100 | 14/100 | untested | ★ 24.3k▲ 9.8k/yrnpm 1.2M/wk | 15/24 verified | 42/100 integrity | ||
| 5 | free-tier vs Steel ↗ | 50/100 | 24/100 | 3/100 | 12/100 | 20/100 | ★ 2k▲ 1.1k/yrpypi 3.3k/wk | 15/34 verified · 6 disputed | 5/100 integrity | ||
| 6 | free-tier vs Steel ↗ | 30/100 | 45/100 | 0/100 | 34/100 | 5/100 | ★ 114.6k▲ 61.3k/yrnpm 70.5k/wkpypi 1.7M/wk | 14/27 verified · 1 disputed | 27/100 integrity | ||
| 7 | free-tier vs Steel ↗ | 26/100 | 53/100 | 0/100 | 0/100 | 0/100 | ★ 19▲ 18/yr | 13/20 verified | 9/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Best for automation-engineer
Steel
46/100
Runner-up:
Skyvern (45/100)
8 automation-engineer stories scored
Story matrix — every product × every judged storyStory matrix
Action primitives — stories about action primitives in this arenaAction primitives
Caching
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Action primitives — stories about action primitives in this arenaCache resolved actions or generated code so repeat runs replay deterministically at lower cost and latency than re-prompting the LLM | developer | none 0/10 | fullC 7/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 |
Dom
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Action primitives — stories about action primitives in this arenaDrive the page through DOM-understanding action primitives (act/click/type on described elements) that survive selector and layout changes | developer | partialX 6/10 | fullX 8/10 | disputedD 5/10 | partialC 6/10 | none 0/10 | partialX 6/10 | partialX 5/10 |
Observe
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Action primitives — stories about action primitives in this arenaPreview candidate actions on the current page (observe/plan) before committing the agent to act | developer | none 0/10 | fullT 8/10 | partialX 3/10 | none 0/10 | partialC 4/10 | partialT 6/10 | none 0/10 |
Vision
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Action primitives — stories about action primitives in this arenaSwitch to a vision or computer-use action mode that operates on screenshots for canvases and UIs the DOM path can't handle | developer | none 0/10 | none 0/10 | fullC 7/10 | partialC 3/10 | none 0/10 | partialX 5/10 | none 0/10 |
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 8/10 | fullT 8/10 | fullT 8/10 | fullT 8/10 | fullT 9/10 | fullT 9/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | partialC 6/10 | fullT 8/10 | fullC 7/10 | fullC 8/10 | fullT 9/10 | fullT 8/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | n/a | none 0/10 | none 0/10 | none 0/10 | n/a | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | disputedD 5/10 | fullT 8/10 | fullT 8/10 | fullT 8/10 | partialC 5/10 | fullT 8/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | none 0/10 | none 0/10 | none 0/10 | partialC 5/10 | fullT 8/10 | fullT 8/10 | fullT 7/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | partialT 7/10 | fullT 8/10 | fullT 7/10 | fullT 8/10 | fullT 9/10 | fullT 8/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | fullT 7/10 | fullT 8/10 | fullT 7/10 | fullC 8/10 | fullT 8/10 | partialT 5/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Agentic features
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | partialC 4/10 | n/a | n/a | n/a | n/a | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | partialC 5/10 | partialC 5/10 | partialX 7/10 | partialC 4/10 | none 0/10 | fullX 7/10 | partialC 5/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | n/a | fullX 8/10 | fullC 7/10 | fullC 7/10 | n/a | disputedD 5/10 | fullX 8/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | fullX 8/10 | fullX 9/10 | fullX 7/10 | fullC 8/10 | partialC 5/10 | disputedD 5/10 | partialC 6/10 |
Api quality
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | none 0/10 | none 0/10 | none 0/10 | partialT 3/10 | partialT 6/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | fullT 9/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | none 0/10 | none 0/10 | none 0/10 | partialC 6/10 | fullT 7/10 | partialC 3/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Auth session persistence — stories about auth session persistence in this arenaAuth session persistence
Compat
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Auth session persistence — stories about auth session persistence in this arenaConnect my existing Playwright, Puppeteer, or CDP automation code to the product's browsers instead of rewriting it | developer | fullX 7/10 | partialX 6/10 | partialC 5/10 | fullC 8/10 | fullT 9/10 | fullC 7/10 | none 0/10 |
Credentials
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Auth session persistence — stories about auth session persistence in this arenaStore credentials in a vault and have the agent complete logins including TOTP/2FA challenges without exposing secrets to the model | automation-engineer | partialC 5/10 | none 0/10 | fullX 7/10 | none 0/10 | none 0/10 | partialC 7/10 | none 0/10 |
Profiles
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Auth session persistence — stories about auth session persistence in this arenaPersist logged-in browser state in reusable profiles so agents skip the login wall on every subsequent run | developer | fullX 8/10 | fullC 8/10 | partialC 6/10 | fullC 7/10 | fullT 8/10 | fullC 7/10 | fullC 7/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | partialC 3/10 | none 0/10 | none 0/10 | partialX 6/10 | partialC 4/10 | partialC 4/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | partialC 3/10 | n/a |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | none 0/10 | n/a | none 0/10 | none 0/10 | n/a | none 0/10 | n/a |
Deployment modes — stories about deployment modes in this arenaDeployment modes
Local
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Deployment modes — stories about deployment modes in this arenaRun the agent against a local browser on my own machine for development, without any cloud account | developer | fullX 8/10 | fullC 7/10 | partialC 6/10 | none 0/10 | fullT 9/10 | none 0/10 | none 0/10 |
Framework model support — stories about framework model support in this arenaFramework model support
Frameworks
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Framework model support — stories about framework model support in this arenaPlug the browser layer into agent frameworks (Claude Agent SDK, Vercel AI SDK, LangChain, CrewAI) through documented adapters | developer | none 0/10 | partialT 4/10 | none 0/10 | none 0/10 | partialC 5/10 | partialT 7/10 | none 0/10 |
Models
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Framework model support — stories about framework model support in this arenaBring my own LLM provider — the framework is model-agnostic rather than locked to one vendor's models | developer | fullX 8/10 | none 0/10 | partialC 6/10 | fullC 7/10 | none 0/10 | partialC 5/10 | none 0/10 |
Nl task execution — stories about nl task execution in this arenaNl task execution
Tasks
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Nl task execution — stories about nl task execution in this arenaSubmit a browser task over a hosted HTTP API and receive the result by polling or webhook, without managing any browser myself | ai-native | fullT 8/10 | none 0/10 | partialT 5/10 | partialT 6/10 | none 0/10 | partialT 6/10 | partialT 5/10 |
| Nl task execution — stories about nl task execution in this arenaHand the product a natural-language goal and it completes a multi-step web task end to end — navigating, filling forms, and clicking through flows | developer | fullX 8/10 | partialX 5/10 | disputedD 5/10 | fullC 7/10 | partialT 5/10 | disputedD 5/10 | fullX 7/10 |
Workflows
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Nl task execution — stories about nl task execution in this arenaCompose repeatable multi-step workflows with loops, conditionals, and parameters instead of one-shot prompts | automation-engineer | none 0/10 | partialC 5/10 | partialC 5/10 | none 0/10 | none 0/10 | partialC 5/10 | partialC 4/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 6/10 | n/a | partialT 6/10 | fullT 7/10 | partialT 6/10 | partialT 7/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 | partialT 5/10 | none 0/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | partialT 5/10 | none 0/10 | disputedD 4/10 | none 0/10 | partialT 6/10 | partialX 3/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | fullT 7/10 | partialT 6/10 | fullT 8/10 | none 0/10 | fullT 10/10 | none 0/10 | none 0/10 |
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Pricing
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to paySee transparent per-task or per-browser-hour pricing and documented rate/concurrency limits before committing | developer | partialC 4/10 | none 0/10 | none 0/10 | disputedD 4/10 | none 0/10 | disputedD 3/10 | none 0/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | none 0/10 | partialC 6/10 | partialC 5/10 | none 0/10 | partialT 4/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | none 0/10 | none 0/10 | partialC 4/10 | none 0/10 | none 0/10 | none 0/10 | partialX 4/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | none 0/10 | none 0/10 | partialX 4/10 | none 0/10 | partialT 4/10 | none 0/10 | partialX 5/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Replay debugging — stories about replay debugging in this arenaReplay debugging
Live
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Replay debugging — stories about replay debugging in this arenaWatch a session live and take human control mid-run when the agent gets stuck | automation-engineer | partialC 6/10 | partialC 4/10 | fullC 8/10 | partialC 4/10 | fullT 8/10 | partialC 5/10 | partialC 4/10 |
Replay
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Replay debugging — stories about replay debugging in this arenaDebug a failed agent run from recorded replays — video, screenshots, step-by-step action timelines | automation-engineer | partialC 4/10 | partialC 5/10 | fullX 8/10 | partialC 5/10 | fullC 7/10 | partialC 5/10 | partialC 3/10 |
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Fleets
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Scale parallelism — running many jobs at once — concurrency, fleets, queueingRun a fleet of concurrent browser sessions with documented concurrency limits and programmatic session management | automation-engineer | partialC 4/10 | none 0/10 | none 0/10 | partialX 6/10 | partialT 6/10 | partialX 4/10 | none 0/10 |
Lifecycle
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Scale parallelism — running many jobs at once — concurrency, fleets, queueingGet webhook notifications when tasks and sessions finish instead of polling for status | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Stealth captcha — stories about stealth captcha in this arenaStealth captcha
Captcha
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Stealth captcha — stories about stealth captcha in this arenaRely on a documented captcha stance — automatic solving, human fallback, or explicit non-support — instead of silent task failures | automation-engineer | fullC 8/10 | partialX 3/10 | fullC 8/10 | disputedD 5/10 | fullC 8/10 | disputedD 4/10 | fullX 6/10 |
Posture
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Stealth captcha — stories about stealth captcha in this arenaPoint to the vendor's published acceptable-use and anti-abuse posture governing what its stealth and automation features may be used for | automation-engineer | none 0/10 | none 0/10 | none 0/10 | disputedD 4/10 | none 0/10 | none 0/10 | none 0/10 |
Stealth
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Stealth captcha — stories about stealth captcha in this arenaEnable stealth fingerprinting and residential or geo-targeted proxies so legitimate automations aren't blocked as bots | automation-engineer | fullC 7/10 | partialX 5/10 | none 0/10 | disputedD 5/10 | fullX 8/10 | fullC 7/10 | partialX 5/10 |
Structured extraction — stories about structured extraction in this arenaStructured extraction
Extraction
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaExtract typed, schema-validated data (Zod/Pydantic-style) from pages the agent visits, not just raw text | developer | partialC 4/10 | fullC 7/10 | fullC 7/10 | fullC 8/10 | none 0/10 | disputedD 5/10 | fullC 8/10 |
Files
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaMy agent can download files from and upload files to the sites it operates, with the artifacts retrievable afterwards | developer | none 0/10 | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.