Browser Automation for Agents Arena
Browser Use vs Skyvern
Browser Use
Browser Use, Inc.
Skyvern wins · 13–18 (18 drawn)
Action primitives — stories about action primitives in this arenaAction primitives
Stories about action primitives in this arena
Caching
developerCache resolved actions or generated code so repeat runs replay deterministically at lower cost and latency than re-prompting the LLM
weight 2 · round drawnBrowser Usenone0/10No evidence of caching resolved actions or generated code for deterministic, lower-cost replay; the product is LLM-driven agent automation with sessions/runs but nothing about caching or replay without re-invoking the model. missing for 10: any mention of action/code caching, deterministic replay mechanism, or cost/latency savings from skipping re-prompting.
Skyvernnone0/10No evidence of caching resolved actions or generated code for deterministic, cheaper replay; Skyvern's model is per-run AI-driven navigation via LLM+vision, and community feedback even complains about cost/latency of repeated LLM calls with no mention of a caching mechanism to mitigate this.
- [community] “I tried it out and it's pretty pricey. My OpenAI API bill is $3.20 after using this on a few different pages to test it out... this is alway…”
- [community] “This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…”
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
Dom
developerDrive the page through DOM-understanding action primitives (act/click/type on described elements) that survive selector and layout changes
weight 3 · round to Browser UseBrowser Use's core agent is built around natural-language task instructions (e.g. "Find the top Hacker News story", "Fill in this job application") executed via an LLM-driven agent that perceives the DOM/page state and decides actions, which is the essence of DOM-understanding, layout-resilient automation — this is corroborated by GitHub task examples and community hands-on use (LinkedIn automation, resume filling). However, the evidence pack lacks explicit documentation of discrete act/click/type primitives with described-element targeting or any stated guarantee/mechanism for surviving selector/layout changes; it's inferred from the agent's general design rather than directly documented. missing for 10: explicit API/primitive-level documentation of click/type/act-on-described-element functions, and direct evidence/testing showing resilience to selector or layout changes rather than just general LLM-driven task completion.
- [github] “Task: "Fill in this job application with my resume and information."”
- [github] “Task: "Extract structured data about my followers and export it as a CSV."”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
Skyverndisputedcontradicted5/10Skyvern's docs describe exactly this: natural-language act/click/type primitives with vision+DOM understanding that operate on sites 'never seen before' and fall back to selectors only if needed (skyvern-docs-19, skyvern-gh-1, skyvern-docs-17), positioned explicitly as a replacement for brittle Selenium scripts (skyvern-docs-13). However, a hands-on community test found it worked on the happy path but concretely failed to interact with a layout element (a popup) and struggled to hit a tab on a real site (skyvern-comm-2), contradicting the claim that it robustly survives arbitrary layout changes. Missing for 10: independent benchmark data on selector/layout-change robustness, broader corroboration beyond one hands-on report, and resolution of the observed failure mode.
- [claimed-docs] “Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.”
- [github] “Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…”
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
- [claimed-docs] “You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.”
- [community] “I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…”
Observe
developerPreview candidate actions on the current page (observe/plan) before committing the agent to act
weight 1 · round to SkyvernBrowser Usenone0/10Evidence shows live-preview URLs for human intervention during CAPTCHA/2FA and event polling for observability, but nothing about an explicit preview/plan step where candidate actions are surfaced for developer review before the agent commits to acting.
- [claimed-docs] “If the challenge remains, open the [live preview](/cloud/browser/live-preview) for human control.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event]”
Skyvern's docs mention human-in-the-loop pausing for approval between steps and a VNC stream to watch/take control, which offers some ability to intervene before the agent proceeds, but there is no documented explicit 'plan/preview candidate actions' step (e.g., a dry-run or action list shown before execution). A community comment even notes the absence of assertion/verification-style controls compared to Playwright, suggesting no built-in preview mechanism for validating steps before they run. Missing for 10: an explicit plan/preview UI or API that lists candidate actions before execution, and independent confirmation that the pause-for-approval flow shows planned actions rather than just pausing mid-run.
- [claimed-docs] “Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…”
- [community] “I can't see that option in Skyvern which would have me worrying that process changes would be overlooked and we would unknowingly start ente…”
Vision
developerSwitch to a vision or computer-use action mode that operates on screenshots for canvases and UIs the DOM path can't handle
weight 2 · round to SkyvernBrowser Usenone0/10Browser Use's docs describe DOM-based agent actions, live preview/human handoff for CAPTCHA/2FA, and CDP-based control, but there is no evidence of a dedicated vision or computer-use mode that acts directly on screenshots for canvas/UI elements the DOM can't reach.
Skyvern's core action loop is vision-based: it maps visual elements to actions on pages it has never seen, without custom DOM-specific code (skyvern-gh-1), and separately uses its vision model to detect and solve CAPTCHAs, which are canvas-like elements the DOM can't parse (skyvern-docs-9, skyvern-docs-27). Docs also mention falling back to selectors when useful (skyvern-docs-19), implying vision-first with DOM as a secondary path rather than a purely DOM-based tool needing a special switch. missing for 10: explicit documentation of a discrete 'vision/computer-use mode' toggle, dedicated canvas/non-DOM UI examples (e.g., canvas-drawn widgets, non-HTML apps), and independent benchmarking confirming success on such UIs.
- [github] “Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…”
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…”
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically.”
- [claimed-docs] “Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.”
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round drawnBrowser Use publishes a discoverable llms.txt at docs.browser-use.com/llms.txt confirmed live via probe (HTTP 200), which is exactly the agent-oriented docs entry point an AI-native user could point an agent at, and the broader docs site is structured/markdown-friendly for agent consumption. Missing for 10: no evidence of additional structured formats like llms-full.txt or explicit guidance encouraging agents to consume it, and no independent community confirmation of agents successfully using it.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.browser-use.com/llms.txt # Browser Use > Documentation for Browser Use Cloud Agent and Browser API…”
A probe confirms llms.txt is live and returns a structured summary of Skyvern for agent consumption, and skyvern.com/llms provides an agent-oriented docs page listing features in a scannable format. This directly satisfies pointing an agent at llms.txt or agent-oriented docs. Missing for 10: a docs.md/markdown-mirrored docs endpoint (404) and an accessible OpenAPI spec, which would round out machine-readable documentation.
- [probe] “PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …”
- [claimed-docs] “Visual workflow builder for non-developers — drag-and-drop, no code required”
- [claimed-docs] “Browser recorder that converts manual actions into reusable automations”
- [claimed-docs] “SOP upload — describe a process in plain English and Skyvern builds the workflow”
- [claimed-docs] “Copilot chat for building and debugging workflows interactively”
- [probe] “PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md”
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to SkyvernBrowser Use ships both an open-source Python library and a cloud API (client.runs.create) that are inherently script/automatable, implying headless/CI use, and gh-3 explicitly pitches automating the web 'from your own code, and with any LLM.' However, there is no explicit documentation of headless mode flags, Docker images, or CI pipeline examples/integration guides. Missing for 10: explicit headless-mode configuration docs, CI/CD pipeline examples (e.g., GitHub Actions), and independent confirmation of running unattended in CI.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [claimed-docs] “For a local agent, use the [open-source library](/open-source/quickstart).”
- [claimed-docs] “Launch a browser, connect to its CDP URL, then stop it”
Skyvern ships a code-first SDK/REST API (Python/TypeScript) that connects to a cloud or self-hosted Chromium instance, explicitly positioned as replacing brittle Selenium scripts and integrating browser automation via API into other products, and can run entirely on your own infrastructure with your own LLM keys, supporting headless/scriptable use suitable for CI. missing for 10: explicit CI/CD pipeline documentation or example (e.g. GitHub Actions integration), and independent confirmation of headless execution in automated environments
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [claimed-docs] “You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.”
ai-native userConnect an agent via an official MCP server
weight 3 · round to SkyvernBrowser Usedisputedcontradicted5/10Docs describe an official MCP server enabling Claude, Cursor, Windsurf or any MCP client to run Browser Use tasks (browser-use-docs-12, browser-use-probe-3), but a hands-on community report says the author had to switch tools because Browser Use 'doesn't support MCP integration' in Cursor (browser-use-comm-4), directly contradicting the documented claim. Missing for 10: independent corroboration that MCP connection actually works end-to-end, and resolution of the conflicting user report.
- [claimed-docs] “Run browser automation tasks from your AI coding assistant. Connect to Claude, Cursor, Windsurf, or any MCP client.”
- [probe] “official MCP server documented at https://docs.browser-use.com/cloud/guides/mcp-server”
- [community] “I want to use browser-use in Cursor but I am using another option because it doesn't support MCP integration which is the common language th…”
Skyvern documents an official MCP server (skyvern-docs-7, skyvern-probe-4) that lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser via Skyvern, directly matching the story. Missing for 10: independent/hands-on community verification of the MCP server specifically (community evidence covers other features, not MCP usage) and a clear setup/config example beyond the single doc page.
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
- [probe] “official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp”
ai-native userUse an official CLI
weight 2 · round drawnBrowser Usenone0/10The evidence pack documents a Python SDK, cloud API, MCP server, and web dashboard, but never mentions an official standalone CLI tool for Browser Use. Since a browser-automation product could plausibly ship a CLI, absence of any such evidence means this axis is unmet.
Skyvernnone0/10Evidence covers Skyvern's Python/TypeScript SDKs, REST API, MCP server, and visual dashboard, but nowhere mentions an official CLI tool for AI-native workflows. missing for 10: any documented CLI command, npm/pip CLI package, or terminal-based interface.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
ai-native userDrive the product through a documented public API
weight 3 · round to SkyvernBrowser Use documents a public Cloud API (client.runs.create, sessions, events polling, structured output, CDP connection) across multiple docs pages, indicating a real programmatic interface beyond the UI. However, probes for a formal OpenAPI/swagger spec returned 404s, so there's no machine-readable API contract, only prose docs and SDK examples. Missing for 10: a published OpenAPI/spec artifact, independent third-party confirmation of API usage beyond vendor docs.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.browser-use.com/llms.txt # Browser Use > Documentation for Browser Use Cloud Agent and Browser API…”
Skyvern documents a public API/SDK surface (Python, TypeScript, REST) for creating tasks, running multi-step browser automations, and extracting structured data via JSON schema, matching the ai-native 'drive via documented API' story; it also ships an MCP server for agent control. Missing for 10: a discoverable OpenAPI/swagger spec (probe found 404s for all candidate paths) and independent/hands-on confirmation of API robustness beyond first-party docs.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [claimed-docs] “you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…”
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
- [probe] “official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round drawnBrowser Usenone0/10No evidence of scoped or least-privilege API key/credential issuance (e.g., role-based tokens, permission scopes) for agents; docs mention API keys implicitly via client usage but no mention of scoping, restricted permissions, or credential management features. Community/GitHub evidence also silent on this.
Skyvernnone0/10No evidence of scoped or least-privilege API credential issuance for agents; docs mention API keys and self-hosted LLM keys but nothing about credential scoping, permissions, or restricting agent access levels. Community comments even raise concerns about handling sensitive credentials in plain text with no mitigation shown. Missing for 10: any documentation of scoped API tokens, role-based access control, or least-privilege credential management for agents.
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [community] “you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…”
ai-native userBuild against official SDKs
weight 2 · round drawnBrowser Use ships an official open-source Python library (github, docs-19) and a cloud client SDK with documented usage patterns (client.runs.create, sessions, events polling) shown in docs-1/4/5/14, giving AI-native devs a concrete first-party SDK to build against. Missing for 10: evidence of SDKs beyond Python (e.g. JS/TS), and no OpenAPI spec was found (probe-2 all 404) or independent third-party corroboration of SDK usage.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “For a local agent, use the [open-source library](/open-source/quickstart).”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
Skyvern explicitly documents official Python and TypeScript SDKs plus a REST API for integrating browser automation, with SDK-level primitives like page.extract and data_extraction_schema shown in docs (skyvern-docs-1, skyvern-docs-5, skyvern-docs-19, skyvern-docs-4). Missing for 10: independent/hands-on developer corroboration of SDK usage and a public API reference (OpenAPI spec probe returned 404s, skyvern-probe-3), so quality is capped below full confidence in completeness.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.”
- [claimed-docs] “you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…”
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
ai-native userSubscribe to events via webhooks
weight 2 · round drawnBrowser Usenone0/10The docs describe an event system based on polling `runs.events()` with a cursor, not webhook subscriptions; no evidence anywhere in the pack mentions webhooks, callback URLs, or push notifications.
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
Agentic features
ai-native userSet up automations that run autonomously in the background
weight 2 · round to SkyvernBrowser Use Cloud lets users kick off agent runs via API (client.runs.create) that execute asynchronously in a hosted browser, with session reuse and event polling to monitor progress without keeping a local process open, which supports a form of unattended background execution. However there is no documented scheduling, cron-like triggers, or webhook-based automation setup for recurring/background jobs, so it's unclear whether truly hands-off recurring automations are supported. Missing for 10: explicit scheduling/trigger mechanism, evidence of long-running unattended jobs beyond single API-invoked runs, and independent confirmation of background reliability.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
Skyvern's docs show multi-step, code-first and no-code workflows that run via API or cloud UI, persist browser state, and can pause for human approval while capturing recordings/artifacts — all indicative of autonomous background execution (skyvern-docs-5,6,10,11,21). Zapier integration and API-driven triggering (skyvern-docs-16, skyvern-docs-28) supports running without manual intervention, but there's no explicit documentation of a scheduler, cron-like triggers, or continuous monitoring dashboard for unattended runs. Missing for 10: explicit scheduling/trigger docs, evidence of long-running unattended background jobs, and independent confirmation of reliability at scale (community notes some brittleness, e.g. skyvern-comm-2).
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.”
- [claimed-docs] “Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…”
- [claimed-docs] “Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…”
- [claimed-docs] “Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.”
- [claimed-docs] “Connect to Zapier”
- [claimed-docs] “Skyvern automates browser-based workflows across these platforms — no API keys or custom connectors required.”
- [community] “I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…”
ai-native userOperate the product with natural-language commands
weight 2 · round to Browser UseBrowser Use is fundamentally natural-language driven: tasks are issued as plain-English strings like "Find the top Hacker News story" or "Fill in this job application with my resume and information", with the agent interpreting and executing them autonomously, corroborated by GitHub examples and community hands-on use for LinkedIn automation. missing for 10: independent benchmark of instruction-following accuracy, and clearer docs on limits/failure modes of natural-language parsing.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [github] “Task: "Fill in this job application with my resume and information."”
- [github] “Task: "Extract structured data about my followers and export it as a CSV."”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
Skyvern's core interaction model is natural-language: users provide a prompt describing the goal (docs-18), SOPs in plain English are converted to workflows (docs-24), and a Copilot chat and MCP server let AI assistants/users direct browser actions in natural language (docs-25, docs-7, docs-19). This is corroborated by community reports of using it via prompts on real sites, though with mixed reliability on complex flows. missing for 10: independent benchmarking or hands-on confirmation that natural-language commands reliably handle complex multi-step tasks, and clearer evidence of NL-driven success rates beyond anecdotal HN reports.
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [claimed-docs] “SOP upload — describe a process in plain English and Skyvern builds the workflow”
- [claimed-docs] “Copilot chat for building and debugging workflows interactively”
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
- [claimed-docs] “Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.”
- [community] “I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…”
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round drawnBrowser Usenone0/10Docs show static code snippets (e.g., client.runs.create examples) but no evidence of an interactive, runnable API console/reference; explicit probes for OpenAPI/Swagger specs at standard paths all returned 404, indicating no interactive API explorer exists.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.browser-use.com/llms.txt # Browser Use > Documentation for Browser Use Cloud Agent and Browser API…”
Skyvernnone0/10Evidence shows only static docs describing SDKs and REST usage, with no interactive API reference or runnable-example explorer; probes explicitly found no OpenAPI/Swagger spec at any candidate path (404s), indicating no interactive reference exists.
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
- [probe] “PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md”
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnBrowser Usenone0/10A direct probe for OpenAPI/swagger spec files returned 404 on all candidate paths, and no documentation references a downloadable machine-readable API spec despite having a REST/cloud API.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
Skyvernnone0/10Skyvern offers a REST API (skyvern-docs-1) but a direct probe for OpenAPI/swagger specs at standard paths returned 404 across all candidates, and no docs mention a downloadable machine-readable spec.
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
- [probe] “PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md”
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
ai-native userTest against a sandbox environment without touching production data
weight 1 · round drawnBrowser Usenone0/10The evidence describes cloud browser sessions, live preview, and CDP connections, but nothing indicates a dedicated sandbox/staging mode that isolates test runs from production data or real accounts. Users are shown reusing real logged-in profiles (browser-use-comm-5) rather than isolated test environments, and no docs mention a sandbox distinct from production.
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Log in once, save the profile, then reuse it to start future browsers already logged in.”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
Skyvernnone0/10Skyvern's docs describe cloud or self-hosted execution, credential handling, and observability, but nothing describes a dedicated sandbox/staging mode isolated from production data or systems — missing for 10: any mention of a sandbox environment, test/staging mode, or data isolation guarantees.
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnBrowser Usenone0/10Evidence shows an API version label ("V4") in docs, but there is no documented deprecation policy, versioning changelog, or API stability guarantees anywhere in the pack; the OpenAPI spec probe also 404s, indicating no formal API contract is published.
- [claimed-docs] “V4 returns `run.result` as a string. Ask for JSON only, then validate it client-side”
- [claimed-docs] “Automatic CAPTCHA solving is enabled by default for API V4 Agent runs and standalone Cloud Browser sessions.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
Skyvernnone0/10No evidence of API versioning scheme or a documented deprecation policy; OpenAPI/spec probes all returned 404 and no changelog or versioning docs appear in the pack. Missing for 10: versioned API endpoints (e.g., /v1/), a published deprecation/support policy, and changelog documentation.
Auth session persistence — stories about auth session persistence in this arenaAuth session persistence
Stories about auth session persistence in this arena
Compat
developerConnect my existing Playwright, Puppeteer, or CDP automation code to the product's browsers instead of rewriting it
weight 2 · round to Browser UseDocs explicitly describe connecting existing Playwright/Puppeteer/CDP code to Browser Use's browsers via CDP URL, with a documented choice between Browser Use driving or the developer's own code connecting directly over CDP (docs-2, docs-18), and community reports confirm connecting to a real local Chrome profile via CDP for existing automation. missing for 10: independent hands-on validation specifically with Playwright/Puppeteer libraries (not just CDP raw), and more detail on session/auth persistence when using external code.
- [claimed-docs] “Launch a browser, connect to its CDP URL, then stop it”
- [claimed-docs] “Choose whether Browser Use drives the browser or your Playwright/Puppeteer code connects directly over CDP.”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
Docs show Skyvern's own SDK connects to a cloud Chromium instance over CDP and layers Playwright on top, and describe 'drop-in AI commands on top of Playwright' with fallback to raw selectors, implying some interoperability with existing Playwright code. However there is no explicit guidance or example showing a developer pointing an existing Playwright/Puppeteer/CDP script at Skyvern's managed browser instead of rewriting into Skyvern's task/workflow API, and no independent confirmation of this specific reuse pattern. Missing for 10: explicit BYO-script CDP endpoint docs, Puppeteer-specific support, and hands-on/community verification of dropping in existing automation code unchanged.
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.”
- [claimed-docs] “you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…”
Credentials
automation-engineerStore credentials in a vault and have the agent complete logins including TOTP/2FA challenges without exposing secrets to the model
weight 2 · round to SkyvernBrowser Use supports saved/reused login profiles (docs-10) and a documented 2FA workflow where the run pauses at the challenge and a human enters the code via live view (docs-11/16), but there is no vault/secrets-manager integration for storing credentials and injecting them without model exposure, and TOTP is handled via human-in-the-loop rather than automated secret injection. missing for 10: a credential vault/secrets-manager integration, evidence that passwords/TOTP secrets are injected without ever passing through the model context, fully automated TOTP handling without human intervention, independent confirmation of the 2FA flow working in practice.
- [claimed-docs] “Log in once, save the profile, then reuse it to start future browsers already logged in.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event]”
Skyvern's docs explicitly describe storing credentials via password-manager vault integrations (Bitwarden, 1Password, Azure Key Vault) and automatically handling TOTP/2FA, email, and SMS verification during login flows, matching the story closely [skyvern-docs-8][skyvern-docs-26]. However, there's no independent/hands-on verification that secrets are never exposed to the LLM, and a community comment raises concern about credentials being handled in plain text, so missing for 10: independent security audit or hands-on confirmation of secret-masking from the model, and clarification addressing the community's plaintext-handling concern.
- [claimed-docs] “Skyvern handles logins with stored credentials, TOTP/authenticator codes, email and SMS verification, magic links, and password manager inte…”
- [claimed-docs] “Skyvern handles authentication end-to-end, from simple passwords to multi-factor flows with TOTP codes, email verification, and magic links.”
- [community] “you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…”
Profiles
developerPersist logged-in browser state in reusable profiles so agents skip the login wall on every subsequent run
weight 3 · round to Browser UseDocs explicitly describe saving a login profile once and reusing it to start future browsers already logged in, plus a 2FA guide for handling the initial login flow, and community evidence confirms local profile reuse works for logged-in automation (e.g., LinkedIn). missing for 10: independent/hands-on corroboration specifically of the cloud profile-reuse feature (only local profile reuse is community-validated) and no detail on profile storage/security guarantees.
- [claimed-docs] “Log in once, save the profile, then reuse it to start future browsers already logged in.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event]”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
Skyvern's browser-sessions feature explicitly persists cookies, local storage, and open tabs across operations so 'later operations pick up exactly where the previous one stopped,' and pauses preserve browser state — this directly supports skipping repeated logins. However, the docs don't clearly describe a named 'profile' abstraction, how long sessions persist across truly separate future runs, or how these persisted sessions are managed/reused across different agents or teams. missing for 10: explicit reusable-profile management docs, long-term persistence guarantees across independent runs, independent/hands-on confirmation of skip-login behavior.
- [claimed-docs] “Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.”
- [claimed-docs] “Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…”
- [claimed-docs] “Skyvern handles logins with stored credentials, TOTP/authenticator codes, email and SMS verification, magic links, and password manager inte…”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to Browser UseThe library is pitched for automating the web 'at scale' from custom code (gh-3) and cloud sessions can hold multiple runs, hinting at multi-task orchestration, but there is no explicit documentation of a batch/bulk API, parallel run submission, or looping over many items as a first-class feature. Missing for 10: dedicated bulk/batch endpoint or SDK pattern, concurrency limits/guidance, and hands-on evidence of running many items in one operation.
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to SkyvernBrowser Usenone0/10Browser Use's evidence covers on-demand task runs, sessions, observability polling, CAPTCHA/stealth, and MCP integration, but nothing describes user-defined rules or triggers that fire actions automatically on external events (e.g., webhooks, schedules, conditional triggers). The product is presented as an agent you invoke to perform a task, not an event-driven automation/rules engine.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Run browser automation tasks from your AI coding assistant. Connect to Claude, Cursor, Windsurf, or any MCP client.”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
Skyvern documents a Zapier integration, which could allow external events to trigger Skyvern workflows, but there is no evidence of a native rule/trigger engine, webhooks, or scheduled/event-based automation within Skyvern itself. Missing for 10: documented native event triggers or webhook listeners, schedule-based triggers, and any conditional rule engine inside Skyvern's workflow builder.
- [claimed-docs] “Connect to Zapier”
- [claimed-docs] “Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round drawnBrowser Usenone0/10The evidence pack covers runs, sessions, observability, stealth/proxy/CAPTCHA handling, auth profiles, and MCP integration, but nowhere mentions cron-like scheduling, recurring triggers, or workflow automation over time. No docs or community evidence describe a scheduler or recurring-job feature.
Skyvernnone0/10The evidence pack describes Skyvern's workflow builder, API/SDK, MCP integration, and automation features extensively, but contains no mention of scheduling, cron triggers, or recurring job execution anywhere in the docs, GitHub description, or community discussion. Since Skyvern is a workflow/automation platform, scheduling recurring runs is a fair capability to expect, but it's simply absent from the provided evidence.
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnBrowser Usenone0/10Browser Use documents runs, sessions, event polling, and observability, but there is no evidence of versioning automation definitions, reviewing changes, or rolling back to prior versions of a task/automation. Missing for 10: version history/diffing of automations, review/approval workflow, rollback mechanism.
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
Skyvernnone0/10No evidence in the pack describes version history, change review, or rollback capabilities for Skyvern workflows/automations — only building, running, sharing templates, and artifact capture (recordings/screenshots) are documented. Missing for 10: workflow version history, diff/review UI, rollback-to-previous-version mechanism, any changelog or audit trail for automation edits.
Deployment modes — stories about deployment modes in this arenaDeployment modes
Stories about deployment modes in this arena
Local
developerRun the agent against a local browser on my own machine for development, without any cloud account
weight 2 · round to Browser UseBrowser Use ships an open-source Python library explicitly positioned for local, code-driven automation ('For a local agent, use the open-source library'; 'automate the web at scale, from your own code, and with any LLM'), and community reports confirm running it locally against a real local Chrome browser/profile without a cloud account. missing for 10: no explicit walkthrough showing zero network/account calls during local runs, and no independent benchmark of purely offline/local operation.
- [claimed-docs] “For a local agent, use the [open-source library](/open-source/quickstart).”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
- [claimed-docs] “Launch a browser, connect to its CDP URL, then stop it”
Skyvern is open-source and its self-hosted docs explicitly state it 'runs entirely on your infrastructure: your servers, your browsers, your LLM API keys' (skyvern-docs-12, skyvern-docs-2), which supports running without a cloud account. However, the core SDK/browser-automation flow described elsewhere connects to a 'cloud Chromium instance over CDP' (skyvern-docs-5), suggesting the default path is cloud-based, and no local-machine dev setup details (docker/local browser config, install steps) are shown. Missing for 10: explicit local-browser dev walkthrough, confirmation that the local-first SDK path bypasses cloud Chromium, and independent hands-on confirmation of local-only operation.
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
Framework model support — stories about framework model support in this arenaFramework model support
Stories about framework model support in this arena
Frameworks
developerPlug the browser layer into agent frameworks (Claude Agent SDK, Vercel AI SDK, LangChain, CrewAI) through documented adapters
weight 2 · round drawnBrowser Usenone0/10Evidence documents an MCP server for connecting to Claude, Cursor, or Windsurf, and a Python library for custom code, but there is no mention of documented adapters for Claude Agent SDK, Vercel AI SDK, LangChain, or CrewAI specifically.
- [claimed-docs] “Run browser automation tasks from your AI coding assistant. Connect to Claude, Cursor, Windsurf, or any MCP client.”
- [probe] “official MCP server documented at https://docs.browser-use.com/cloud/guides/mcp-server”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
Skyvernnone0/10Skyvern documents a Python/TypeScript/REST SDK and an MCP server that plugs into Claude Desktop, Claude Code, Codex, Cursor, and Windsurf, but there is no evidence of documented adapters for Claude Agent SDK, Vercel AI SDK, LangChain, or CrewAI. missing for 10: any documented integration guide or adapter for LangChain, CrewAI, Vercel AI SDK, or Claude Agent SDK specifically.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
Models
developerBring my own LLM provider — the framework is model-agnostic rather than locked to one vendor's models
weight 2 · round to Browser UseGitHub docs explicitly market the library as usable with any LLM ('Use the Python library ... with any LLM'), and independent community testing corroborates this by reporting successful use with Gemini models rather than being locked to one vendor. Missing for 10: a dedicated docs page enumerating specific supported providers/configuration examples and broader independent confirmation across multiple providers beyond Gemini.
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [community] “From the first glance, browser-use is compatible with more models, and has (much) more github stars. Coincidentally I played with it over th…”
Docs state self-hosted Skyvern runs with 'your own LLM API keys' on your own infrastructure, implying model-agnosticism rather than lock-in to a single vendor, but there is no explicit list of supported providers/models or first-party guide on swapping LLM backends, and no independent confirmation of multi-provider support. Missing for 10: an explicit supported-providers list/config docs, and community/hands-on evidence of using non-default LLMs.
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
Nl task execution — stories about nl task execution in this arenaNl task execution
Stories about nl task execution in this arena
Tasks
ai agentSubmit a browser task over a hosted HTTP API and receive the result by polling or webhook, without managing any browser myself
weight 2 · round to Browser UseDocs show a hosted Cloud API (client.runs.create) that creates runs, supports polling via runs.events() with cursors, and returns structured results without the caller managing browser infrastructure (stealth, proxies, CAPTCHA solving handled server-side). Missing for 10: no explicit webhook callback mechanism is documented (only polling is shown), and no public OpenAPI spec was found to confirm full REST surface.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “V4 returns `run.result` as a string. Ask for JSON only, then validate it client-side”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [claimed-docs] “Every cloud browser session runs in a hardened Chromium fork with stealth enabled by default — no configuration needed.”
- [claimed-docs] “Residential proxies are enabled by default across 195+ countries.”
- [claimed-docs] “Every Browser Use Cloud browser enables automatic CAPTCHA solving.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
Docs confirm a hosted REST/SDK API where you submit a prompt+URL (optionally a JSON schema) and Skyvern runs the task on cloud Chromium without the caller managing a browser (skyvern-docs-1, skyvern-docs-5, skyvern-docs-18, skyvern-docs-17). However, there is no explicit documentation of a polling endpoint or webhook callback mechanism, and probes found no discoverable OpenAPI spec, so the exact result-retrieval mechanism described in the story is unconfirmed. Missing for 10: explicit webhook/callback docs, explicit polling endpoint docs, and an accessible API reference confirming these mechanics.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
developerHand the product a natural-language goal and it completes a multi-step web task end to end — navigating, filling forms, and clicking through flows
weight 3 · round to Browser UseDocs and GitHub examples show natural-language goals (e.g., "Find the top Hacker News story", "Fill in this job application") driving an agent that navigates, logs in, handles 2FA/CAPTCHA, and completes multi-step flows end-to-end via both cloud API and open-source library; community reports corroborate real-world use (e.g., LinkedIn automation). Missing for 10: independent third-party benchmark of complex multi-step task success rates and more robust evidence of reliability at scale beyond anecdotal community reports.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [github] “Task: "Fill in this job application with my resume and information."”
- [github] “Task: "Extract structured data about my followers and export it as a CSV."”
- [claimed-docs] “Log in once, save the profile, then reuse it to start future browsers already logged in.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Every Browser Use Cloud browser enables automatic CAPTCHA solving.”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
Skyverndisputedcontradicted5/10Skyvern's docs and GitHub strongly claim natural-language goal execution across novel, multi-step web flows (forms, logins, CAPTCHAs) via a single prompt/API (skyvern-docs-17, skyvern-docs-18, skyvern-gh-1), but a hands-on community test found it succeeded only on the 'happy path' and concretely failed on a real multi-step flow (costcotravel.com), struggling to hit a tab and failing to click a popup (skyvern-comm-2). This is a specific documented counter-example contradicting the 'completes multi-step task end to end' claim, not just general skepticism. Missing for 10: independent benchmark results, more hands-on trials showing consistent success on complex/unseen sites, and resolution of the reported failure case.
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [github] “Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…”
- [community] “I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…”
Workflows
automation-engineerCompose repeatable multi-step workflows with loops, conditionals, and parameters instead of one-shot prompts
weight 2 · round to SkyvernBrowser Usenone0/10Evidence shows Browser Use runs are essentially single natural-language task strings within a session/run model (create run, poll events, reuse session) with no documented constructs for loops, conditionals, or parameterized workflow templates; the Python library is described as scriptable but no workflow-composition API (branching, iteration, variables) is shown.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
Skyvern clearly supports multi-step, repeatable workflows via both a code-first SDK and a visual no-code drag-and-drop builder (skyvern-docs-5, skyvern-docs-6, skyvern-docs-22), plus SOP-to-workflow generation and a browser recorder for building reusable automations (skyvern-docs-23, skyvern-docs-24). However, the evidence never explicitly documents loop constructs, conditional branching, or parameterized workflow inputs as first-class workflow-builder features. Missing for 10: explicit documentation of loop/iteration blocks, conditional/branching logic, and named/typed workflow parameters in the workflow builder.
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.”
- [claimed-docs] “Visual workflow builder for non-developers — drag-and-drop, no code required”
- [claimed-docs] “Browser recorder that converts manual actions into reusable automations”
- [claimed-docs] “SOP upload — describe a process in plain English and Skyvern builds the workflow”
- [github] “a no-code workflow builder to help both technical and non-technical users automate manual workflows on any website, replacing brittle or unr…”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round drawnDocs show the API (client.runs.create, sessions, events polling, structured output, live_view_url for 2FA/CAPTCHA handoff, stealth/proxy defaults) mirrors most cloud UI capabilities, and there's an official MCP server for coding-agent access, suggesting broad but not explicitly confirmed feature parity with the dashboard/UI. However there's no discoverable OpenAPI/formal API spec (404s on all candidate paths) and no explicit vendor statement that 100% of UI functionality is API-reachable; missing for 10: a canonical API reference/OpenAPI spec, and explicit parity documentation confirming every UI action (e.g., live preview manual takeover) is independently scriptable via API rather than requiring the UI.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “If the challenge remains, open the [live preview](/cloud/browser/live-preview) for human control.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Run browser automation tasks from your AI coding assistant. Connect to Claude, Cursor, Windsurf, or any MCP client.”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.browser-use.com/openapi.json, https://docs.browser-use.com/swagger.json, https://docs.b…”
- [probe] “official MCP server documented at https://docs.browser-use.com/cloud/guides/mcp-server”
Skyvern's docs show a strong code-first path (Python/TS/REST SDKs, page.extract, workflow creation via API) that covers most core automation tasks also available in the dashboard, and MCP/REST access is documented. However, several UI-only tooling features (drag-and-drop visual builder, browser recorder, SOP upload, copilot chat) are described only as dashboard capabilities with no documented API equivalent, and no public OpenAPI/swagger spec was discoverable to confirm full API-UI parity. Missing for 10: documented API equivalents for recorder/SOP-upload/copilot-chat features, and a discoverable OpenAPI reference confirming complete parity.
- [claimed-docs] “Integrate browser automation into your product with Python, TypeScript, or REST.”
- [claimed-docs] “Use the dashboard to run tasks and build agents visually.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.”
- [claimed-docs] “Visual workflow builder for non-developers — drag-and-drop, no code required”
- [claimed-docs] “Browser recorder that converts manual actions into reusable automations”
- [claimed-docs] “SOP upload — describe a process in plain English and Skyvern builds the workflow”
- [claimed-docs] “Copilot chat for building and debugging workflows interactively”
- [probe] “PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…”
ai-native userExport all of my data in open formats and leave
weight 3 · round to SkyvernBrowser Usenone0/10The evidence pack describes agent runs, sessions, CAPTCHA handling, and pricing, but nothing documents an account/data export feature (e.g., downloading all run history, sessions, or stored data in an open format) that would let a user leave the platform with their data intact; the open-source library allows self-hosting but that's a separate capability from exporting existing cloud account data.
Skyvern is open-source and self-hostable, meaning your data (artifacts, recordings, screenshots, network traffic) stays on your own infrastructure rather than being locked in a vendor's cloud, which implicitly supports data portability. However, there is no explicit documentation of a data export feature, standard open-format export (e.g., JSON/CSV bulk export of run history), or a stated 'leave with your data' workflow. Missing for 10: explicit export functionality/documentation, named open data formats, and any independent confirmation of successful data migration out of the platform.
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [claimed-docs] “Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…”
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
ai-native userRead the product's source under an open license
weight 2 · round to Browser UseThe GitHub repo (browser-use/browser-use) and docs reference an 'open-source library' with a public quickstart, indicating the core Python library's source is publicly viewable, but no evidence explicitly states an open-source license (e.g., MIT/Apache) or shows license text. missing for 10: explicit license file/declaration, confirmation of license type, evidence of full source (vs. cloud API) being open.
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [claimed-docs] “For a local agent, use the [open-source library](/open-source/quickstart).”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.browser-use.com/llms.txt # Browser Use > Documentation for Browser Use Cloud Agent and Browser API…”
Skyverndisputedcontradicted4/10Skyvern is described as open-source with a GitHub repo, and probe evidence confirms it self-identifies as 'open-source' (skyvern-probe-1), but community evidence directly contradicts full open-license access, noting the project is AGPL3 licensed, which is a legally open license but is called out as a practical non-starter/restrictive for many users (skyvern-comm-3). missing for 10: explicit statement of license terms in docs, confirmation of what percentage of the product (cloud vs self-hosted) is actually open-sourced, and independent corroboration that the full source is readable without restriction.
- [github] “Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…”
- [github] “a no-code workflow builder to help both technical and non-technical users automate manual workflows on any website, replacing brittle or unr…”
- [probe] “PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …”
- [community] “Exciting stuff, my employer would be interested but it's AGPL3 licensed so it's a non-starter for them.”
ai-native userSelf-host the core product
weight 3 · round to SkyvernBrowser Use ships an open-source Python library (github.com/browser-use/browser-use) that runs locally and independently of the Cloud API, explicitly positioned as the option for self-hosted/local agents ("For a local agent, use the open-source library"), and community reports confirm running it locally connected to a real browser/profile. missing for 10: no first-party self-hosting guide covering infra/deployment (e.g. Docker, scaling), and no independent audit of parity between self-hosted and cloud feature sets (stealth, CAPTCHA solving, proxies are cloud-only per docs).
- [claimed-docs] “For a local agent, use the [open-source library](/open-source/quickstart).”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
- [community] “If you run it locally, you can connect it to your real browser and user profile where you are already logged in. This works for me for Linke…”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.browser-use.com/llms.txt # Browser Use > Documentation for Browser Use Cloud Agent and Browser API…”
Skyvern has a dedicated self-hosted docs page stating it 'runs entirely on your infrastructure: your servers, your browsers, your LLM API keys' (skyvern-docs-12, skyvern-docs-2), and community evidence confirms it is genuinely open-source (AGPL3) rather than just marketing language (skyvern-comm-3). Missing for 10: independent hands-on confirmation of a successful self-host deployment and clarity on how AGPL licensing affects commercial self-hosting use.
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [community] “Exciting stuff, my employer would be interested but it's AGPL3 licensed so it's a non-starter for them.”
- [probe] “PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Pricing
developerSee transparent per-task or per-browser-hour pricing and documented rate/concurrency limits before committing
weight 2 · round to Browser UsePricing page shows credits-based model ($5+ credits, no subscription, one-time $15 signup credit) but there is no documented per-task or per-browser-hour cost breakdown, and no documented rate/concurrency limits anywhere in the evidence. missing for 10: explicit per-task/per-browser-hour cost figures, documented rate limits, documented concurrency limits, any independent confirmation of pricing transparency.
- [claimed-docs] “One-time $15 credit for eligible Google, GitHub or Microsoft signups ... No card required”
- [claimed-docs] “Credits from $5. No subscription. No expiry.”
Skyvernnone0/10The pricing page is referenced only for its target-audience blurb (skyvern-docs-13); no evidence pack item shows actual per-task or per-browser-hour rates, tiers, or documented rate/concurrency limits. Community comments only express general cost concerns ('pretty pricey', wanting cost down 'at scale') without citing concrete published pricing or limits.
- [claimed-docs] “You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.”
- [community] “I tried it out and it's pretty pricey. My OpenAI API bill is $3.20 after using this on a few different pages to test it out... this is alway…”
- [community] “This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round to SkyvernBrowser Usenone0/10No evidence pack material addresses data residency, region selection, or storage location controls for cloud runs/sessions; the open-source library option avoids the cloud entirely but that's not the same as choosable region/residency within the product. missing for 10: any mention of data center regions, residency options, or storage location controls.
Skyvern offers a self-hosted deployment mode where 'your servers, your browsers, your LLM API keys' run entirely on the user's own infrastructure, which lets an AI-native user control where data resides by choosing their hosting region themselves — but this is achieved only by self-hosting, not via an explicit region/residency selector in the managed cloud product. Missing for 10: documented data residency/region options in the hosted Skyvern Cloud offering, compliance certifications, or explicit multi-region storage controls.
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
ai-native userPrevent my data from being used to train AI models
weight 3 · round to SkyvernBrowser Usenone0/10No evidence in the pack addresses data usage for AI model training, opt-out policies, or privacy/data-retention commitments for Browser Use Cloud or the open-source library.
Skyvern offers self-hosted deployment using your own infrastructure and your own LLM API keys, which implicitly lets users avoid sending data to Skyvern-controlled models/training pipelines, but there is no explicit privacy policy, data-retention statement, or 'we do not train on your data' commitment in the evidence for the hosted/cloud offering. Missing for 10: explicit no-training/data-use policy documentation, opt-out mechanism for the cloud product, and independent confirmation of data handling practices.
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
ai-native userControl data retention and deletion
weight 2 · round to SkyvernBrowser Usenone0/10No evidence of data retention controls, deletion APIs, or privacy/data-lifecycle policy documentation anywhere in the pack; only session/profile reuse and credit pricing are mentioned. missing for 10: retention policy documentation, data deletion API/UI, export/erasure controls, any privacy compliance statement.
Skyvern offers self-hosting (docs-2, docs-12) which gives infrastructure-level control over where data lives, and it captures artifacts (recordings, screenshots, network traffic) per run (docs-11), implying some data exists to manage, but there is no documented retention policy, data deletion API/UI, or export/purge controls for the cloud/hosted product. missing for 10: explicit data retention policy, user-facing deletion/export controls, documentation on how long artifacts/credentials are stored in cloud mode, and independent confirmation that self-hosting actually eliminates vendor-side data retention.
- [claimed-docs] “Run Skyvern on your own infrastructure with your own LLM keys.”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
- [claimed-docs] “Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…”
- [community] “you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…”
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnBrowser Usenone0/10No evidence in the pack mentions telemetry, usage data collection, or any opt-out/privacy setting for Browser Use; the documentation and community items cover unrelated features like stealth browsing, CAPTCHA solving, and MCP integration.
Skyvernnone0/10No evidence pack item discusses telemetry, usage tracking, or an opt-out mechanism; while self-hosting exists, there is no explicit statement about data collection or opt-out controls for the cloud/hosted product. missing for 10: any mention of telemetry collection, privacy policy on usage data, or an opt-out setting/flag.
Replay debugging — stories about replay debugging in this arenaReplay debugging
Stories about replay debugging in this arena
Live
automation-engineerWatch a session live and take human control mid-run when the agent gets stuck
weight 2 · round to SkyvernDocs describe a live_view_url/live preview that lets a human take control mid-run for cases like CAPTCHAs or 2FA, and events can be polled to monitor a run, which supports live-watch-and-intervene workflows. However, this is scoped to specific triggers (CAPTCHA/2FA) rather than a general 'agent gets stuck, operator takes over anytime' workflow, and there's no independent/hands-on evidence confirming smooth mid-run handoff in practice. missing for 10: general-purpose stuck-detection/handoff beyond CAPTCHA/2FA scenarios, independent hands-on confirmation of live takeover working reliably, clear UI/replay-debugging tooling details.
- [claimed-docs] “If the challenge remains, open the [live preview](/cloud/browser/live-preview) for human control.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event]”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
Docs explicitly state a VNC stream lets you watch a live session and take control at any point, plus pause-for-approval human-in-the-loop flows that preserve browser state — directly matching the story. missing for 10: independent/hands-on confirmation of the live takeover UX and details on how control handoff works mid-run beyond the docs description.
- [claimed-docs] “Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…”
- [claimed-docs] “Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.”
Replay
automation-engineerDebug a failed agent run from recorded replays — video, screenshots, step-by-step action timelines
weight 2 · round to SkyvernDocs describe an observability/events stream (runs.events()) for monitoring a run and a live_view_url for real-time human intervention, which could support building a step timeline, but there is no explicit mention of recorded video or screenshot capture for post-hoc replay debugging of failed runs. Missing for 10: documented video recording of sessions, screenshot capture per action, and a dedicated replay/timeline UI for past runs.
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event], and have the user enter the code. The…”
- [claimed-docs] “Ask the first run to stop at the 2FA screen, get its `live_view_url` from the [`browser.ready` event]”
Skyvern docs explicitly state every run captures session recordings, per-step screenshots, AI reasoning traces, and network traffic for debugging, and community feedback corroborates strong observability into workflow steps. missing for 10: independent hands-on verification of the video/timeline UI itself, and no mention of a true step-by-step interactive timeline scrubber beyond artifact capture.
- [claimed-docs] “Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…”
- [community] “This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…”
Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism
Running many jobs at once — concurrency, fleets, queueing
Fleets
automation-engineerRun a fleet of concurrent browser sessions with documented concurrency limits and programmatic session management
weight 2 · round to Browser UseDocs show programmatic session/run creation (client.runs.create, session IDs holding multiple runs) and event polling for observability, implying some ability to manage sessions programmatically, but there is no documented concurrency limit, no fleet/parallel-session guidance, and no scaling architecture described. missing for 10: documented concurrency limits, guidance/examples for running multiple concurrent sessions at scale, rate-limit or quota specs, and independent evidence of parallel session management working in practice.
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
- [claimed-docs] “A **session** holds the agent’s conversation and can reuse its live browser. One session ID can contain multiple runs.”
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
- [github] “Want to automate the web at scale, from your own code, and with any LLM? Use the Python library”
Skyvernnone0/10Evidence covers session persistence, VNC control, self-hosting, and SDK/API access, but nowhere documents concurrency limits, fleet-level session orchestration, or programmatic management of multiple simultaneous browser sessions. Missing for 10: documented concurrency limits, APIs for spinning up/managing many parallel sessions, and any scaling/throughput guidance.
- [claimed-docs] “Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…”
- [claimed-docs] “Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
- [claimed-docs] “Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.”
Lifecycle
developerGet webhook notifications when tasks and sessions finish instead of polling for status
weight 1 · round drawnBrowser Usenone0/10Docs explicitly describe polling patterns for run status (runs.events() with cursor, polling ordered V4 events) but no webhook or callback-based notification mechanism is mentioned anywhere in the evidence pack.
- [claimed-docs] “Poll `runs.events()` with the previous cursor to receive only new events”
- [claimed-docs] “Poll ordered V4 events to monitor a run or build a custom UI.”
Skyvernnone0/10No evidence pack item mentions webhooks, callback URLs, or push notifications for task/session completion; the docs discuss artifacts, VNC streaming, and human-in-the-loop review but nothing about event-driven notification instead of polling. missing for 10: any documentation of webhook/callback support, event subscription API, or notification configuration.
Stealth captcha — stories about stealth captcha in this arenaStealth captcha
Stories about stealth captcha in this arena
Captcha
automation-engineerRely on a documented captcha stance — automatic solving, human fallback, or explicit non-support — instead of silent task failures
weight 2 · round drawnBrowser Use documents a clear captcha stance: automatic CAPTCHA solving is enabled by default for cloud Agent runs and standalone Cloud Browser sessions, with an explicit human-fallback path via the live preview if the challenge persists. This directly matches the story's requirement of a documented stance rather than silent failure. Missing for 10: independent/hands-on verification of captcha-solving success rates and explicit behavior/limits for the self-hosted open-source library (docs focus on Cloud).
- [claimed-docs] “Every Browser Use Cloud browser enables automatic CAPTCHA solving.”
- [claimed-docs] “If the challenge remains, open the [live preview](/cloud/browser/live-preview) for human control.”
- [claimed-docs] “Automatic CAPTCHA solving is enabled by default for API V4 Agent runs and standalone Cloud Browser sessions.”
Skyvern's docs give an explicit, detailed captcha stance: automatic detection and solving via its vision model for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstile, FunCaptcha, MTCaptcha, and text/image captchas, avoiding silent failure ambiguity. Missing for 10: independent/hands-on confirmation that captcha solving works reliably in practice (community evidence discusses pricing, mobile UX, and credential handling but not captcha outcomes specifically), and no documented fallback/human-in-the-loop behavior specifically tied to captcha failures.
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…”
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically.”
Posture
automation-engineerPoint to the vendor's published acceptable-use and anti-abuse posture governing what its stealth and automation features may be used for
weight 1 · round drawnBrowser Usenone0/10The evidence pack documents stealth, proxy, and CAPTCHA-solving features in detail, but there is no published acceptable-use policy, terms of service, or anti-abuse statement governing what these stealth/automation features may or may not be used for; community discussion even raises security/abuse concerns without any vendor policy response cited.
Skyvernnone0/10No evidence in the pack of any published acceptable-use policy, terms of service, or anti-abuse statement covering CAPTCHA-solving/stealth automation features; docs describe capabilities (CAPTCHA bypass, bot bypass) but no governance/AUP language is cited. missing for 10: a published acceptable-use policy, anti-abuse terms, or statement on permitted use of stealth/CAPTCHA-bypass features.
Stealth
automation-engineerEnable stealth fingerprinting and residential or geo-targeted proxies so legitimate automations aren't blocked as bots
weight 2 · round to Browser UseDocs explicitly state cloud browsers run in a hardened Chromium fork with stealth enabled by default and residential proxies across 195+ countries, directly matching the story's stealth+proxy ask, with automatic CAPTCHA solving as a complementary layer. Missing for 10: explicit control/documentation for selecting a specific geo-target rather than automatic 195+ country rotation, and independent/hands-on evidence confirming bot-detection evasion actually works in practice.
- [claimed-docs] “Every cloud browser session runs in a hardened Chromium fork with stealth enabled by default — no configuration needed.”
- [claimed-docs] “Residential proxies are enabled by default across 195+ countries.”
- [claimed-docs] “Every Browser Use Cloud browser enables automatic CAPTCHA solving.”
- [claimed-docs] “Automatic CAPTCHA solving is enabled by default for API V4 Agent runs and standalone Cloud Browser sessions.”
Skyvernnone0/10Evidence covers CAPTCHA solving and authentication/2FA handling, but there is no mention anywhere of stealth fingerprinting, browser fingerprint spoofing, or residential/geo-targeted proxy support. Missing for 10: any documentation of proxy configuration, geo-targeting, or anti-fingerprinting/stealth mode features.
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…”
- [claimed-docs] “Skyvern detects CAPTCHAs using its vision model and solves them automatically.”
Structured extraction — stories about structured extraction in this arenaStructured extraction
Stories about structured extraction in this arena
Extraction
developerExtract typed, schema-validated data (Zod/Pydantic-style) from pages the agent visits, not just raw text
weight 3 · round to SkyvernDocs mention structured output but V4 returns `run.result` as a plain string with a recommendation to 'ask for JSON only, then validate it client-side' — there's no native Zod/Pydantic schema binding or first-party typed-schema extraction feature shown. This is a workaround rather than a built-in schema-validated extraction pipeline. missing for 10: no evidence of a documented schema/type-binding API (e.g., passing a Pydantic/Zod schema directly to the agent), no SDK-level validation helpers, no independent/hands-on confirmation that structured JSON output reliably conforms to a given schema.
- [claimed-docs] “V4 returns `run.result` as a string. Ask for JSON only, then validate it client-side”
- [github] “Task: "Extract structured data about my followers and export it as a CSV."”
Skyvern's docs explicitly support structured, schema-based extraction via `page.extract` with a JSON schema or `data_extraction_schema` param, matching the developer's need for typed output rather than raw text (skyvern-docs-4, skyvern-docs-20, skyvern-docs-18). However, evidence only shows JSON-schema validation, not native Zod/Pydantic model binding, and there's no independent/hands-on confirmation of this specific feature. missing for 10: explicit Zod/Pydantic model integration examples, independent verification of extraction accuracy/schema enforcement.
- [claimed-docs] “you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…”
- [claimed-docs] “you can extract structured data from any page using page.extract with a JSON schema”
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [claimed-docs] “Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…”
Files
developerMy agent can download files from and upload files to the sites it operates, with the artifacts retrievable afterwards
weight 1 · round to SkyvernBrowser Usenone0/10The evidence pack has no documentation of file upload/download handling or of artifacts being stored and retrievable after a run — GH task examples merely reference a resume being filled in and CSV export, but no confirmation these are handled as retrievable files via any API or session mechanism. Missing for 10: explicit file upload API/tooling, download/save-to-cloud-storage feature, and an artifact retrieval endpoint or docs section.
Docs show Skyvern can log into vendor portals and download PDFs (skyvern-docs-15) and captures per-run artifacts like recordings, screenshots, and network traffic retrievable afterward (skyvern-docs-11), implying file download support, but there is no explicit documentation of file upload capability to sites, nor of a dedicated API/UI for retrieving downloaded artifacts as opposed to just run/debug artifacts. missing for 10: explicit upload-to-site capability documentation, a documented file-download/artifact storage API distinct from debugging screenshots, and independent/hands-on confirmation of file transfer working in practice.
- [claimed-docs] “Log into vendor portals, find invoices, download PDFs.”
- [claimed-docs] “Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…”
- [claimed-docs] “Auto-fill and submit applications on Lever, Greenhouse, and more.”
Not comparable on these axes
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · not comparableBrowser Usen/aBrowser Use is itself an agent/automation product; evidence (browser-use-docs-12) shows it ships as an MCP *server* that other clients (Claude, Cursor, Windsurf) connect to, not as an MCP *client* that consumes external MCP servers' tools. Per the agent-role exception, this client-side 'plug in MCP servers' story is out of scope for a product that is itself an agent unless it explicitly runs as an MCP client, which no evidence shows.
- [claimed-docs] “Run browser automation tasks from your AI coding assistant. Connect to Claude, Cursor, Windsurf, or any MCP client.”
- [probe] “official MCP server documented at https://docs.browser-use.com/cloud/guides/mcp-server”
Skyvernnone0/10Evidence only shows Skyvern exposing an MCP *server* so external AI assistants (Claude, Cursor, etc.) can control Skyvern's browser — the reverse of the story, which asks whether Skyvern can consume external MCP servers' tools as a client. No documentation or community evidence shows Skyvern importing or connecting to third-party MCP servers to extend its own toolset.
- [claimed-docs] “The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.”
- [probe] “official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp”
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · not comparableBrowser Use's agent can extract and return structured data/results from web tasks (e.g., extracting follower data to CSV, structured JSON output), which counts as AI-generated output from data it gathers, but there is no evidence of proactive 'insights and suggestions' generated from a user's own stored data inside a product dashboard — it's task-driven extraction, not analytics-style suggestion generation. missing for 10: dedicated insights/suggestions surface, evidence of proactive recommendations, analysis of user's own historical data corpus rather than ad-hoc scraped web data.
- [github] “Task: "Extract structured data about my followers and export it as a CSV."”
- [claimed-docs] “V4 returns `run.result` as a string. Ask for JSON only, then validate it client-side”
- [claimed-docs] “run = client.runs.create("Find the top Hacker News story")”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · not comparableBrowser Usen/aBrowser Use is itself an agent/automation product (the AI acting inside the browser), not a host application that delegates to a separate built-in assistant — this is the agent-role exception where the axis does not apply. It ships as a library/cloud API/MCP server for developers to build agents with, not as an end-user app containing an embedded assistant.
Skyvern's core product IS an AI agent you delegate to via natural-language prompts to complete multi-step browser tasks (skyvern-docs-17, skyvern-docs-18), and it also ships a 'Copilot chat for building and debugging workflows interactively' inside the platform (skyvern-docs-25), plus SOP-to-workflow generation from plain English (skyvern-docs-24). This matches an AI-native user delegating tasks to a built-in assistant. Missing for 10: independent/hands-on validation of the copilot chat feature specifically (community evidence focuses on task execution quality, not the assistant/copilot UX), and no detail on assistant's conversational scope beyond workflow authoring.
- [claimed-docs] “It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.”
- [claimed-docs] “You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.”
- [claimed-docs] “SOP upload — describe a process in plain English and Skyvern builds the workflow”
- [claimed-docs] “Copilot chat for building and debugging workflows interactively”
- [github] “Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…”