Skip to content

Skyvern wins · 228 (19 drawn)

Action primitives — stories about action primitives in this arenaAction primitives

Stories about action primitives in this arena

Caching

  1. developerCache resolved actions or generated code so repeat runs replay deterministically at lower cost and latency than re-prompting the LLM

    weight 2 · round drawn
    Skyvernnone0/10

    No evidence of caching resolved actions or generated code for deterministic, cheaper replay; Skyvern's model is per-run AI-driven navigation via LLM+vision, and community feedback even complains about cost/latency of repeated LLM calls with no mention of a caching mechanism to mitigate this.

    • [community] I tried it out and it's pretty pricey. My OpenAI API bill is $3.20 after using this on a few different pages to test it out... this is alway…
    • [community] This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…
    • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
    Smoothnone0/10

    No evidence of caching resolved actions/generated code for deterministic, cheaper replay; docs mention persistent sessions (auth reuse), structured outputs, and cost efficiency via small models, but nothing about caching or replay of prior task executions to skip re-prompting the LLM. Missing for 10: any mention of action/result caching, replay mechanism, or cost/latency comparison for repeat runs.

    • [claimed-docs] Log in once, then reuse that authentication for future tasks.
    • [claimed-docs] Structured outputs allow you to write deterministic code based on the agent's output. To activate structured outputs, set the `response_mode…
    • [claimed-docs] Smooth uses small and efficient AI models, making it 7x more affordable than browser-use.

Dom

  1. developerDrive the page through DOM-understanding action primitives (act/click/type on described elements) that survive selector and layout changes

    weight 3 · round to Smooth

    Skyvern's docs describe exactly this: natural-language act/click/type primitives with vision+DOM understanding that operate on sites 'never seen before' and fall back to selectors only if needed (skyvern-docs-19, skyvern-gh-1, skyvern-docs-17), positioned explicitly as a replacement for brittle Selenium scripts (skyvern-docs-13). However, a hands-on community test found it worked on the happy path but concretely failed to interact with a layout element (a popup) and struggled to hit a tab on a real site (skyvern-comm-2), contradicting the claim that it robustly survives arbitrary layout changes. Missing for 10: independent benchmark data on selector/layout-change robustness, broader corroboration beyond one hands-on report, and resolution of the observed failure mode.

    • [claimed-docs] Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.
    • [github] Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…
    • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
    • [claimed-docs] You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.
    • [community] I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…

    Smooth's docs describe a 'Session Workflow' that lets you orchestrate smaller tasks, navigate to URLs, and extract data within a persistent session, and the whole product is framed as an AI browser agent that understands pages rather than relying on brittle selectors (smooth-docs-7, smooth-gh-1). However, there is no explicit documentation of discrete act/click/type primitives on described elements, nor any evidence/testing showing these survive selector or layout changes — community comments even note it doesn't fully close the gap versus Playwright-style tools (smooth-comm-8). Missing for 10: explicit act/click/type API reference, documented resilience testing against DOM/selector changes, independent verification of robustness claims.

    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
    • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
    • [community] agent-browser helped a lot over playwright but doesn't completely close the gap.

Observe

  1. developerPreview candidate actions on the current page (observe/plan) before committing the agent to act

    weight 1 · round to Skyvern

    Skyvern's docs mention human-in-the-loop pausing for approval between steps and a VNC stream to watch/take control, which offers some ability to intervene before the agent proceeds, but there is no documented explicit 'plan/preview candidate actions' step (e.g., a dry-run or action list shown before execution). A community comment even notes the absence of assertion/verification-style controls compared to Playwright, suggesting no built-in preview mechanism for validating steps before they run. Missing for 10: an explicit plan/preview UI or API that lists candidate actions before execution, and independent confirmation that the pause-for-approval flow shows planned actions rather than just pausing mid-run.

    • [claimed-docs] Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…
    • [community] I can't see that option in Skyvern which would have me worrying that process changes would be overlooked and we would unknowingly start ente…
    Smoothnone0/10

    Smooth's docs describe live viewing of actions as they execute (live_url) and data extraction, but there is no evidence of a distinct observe/plan step that lets a developer preview candidate actions before committing the agent to act.

    Vision

    1. developerSwitch to a vision or computer-use action mode that operates on screenshots for canvases and UIs the DOM path can't handle

      weight 2 · round to Skyvern
      Skyvernfullclaimed7/10

      Skyvern's core action loop is vision-based: it maps visual elements to actions on pages it has never seen, without custom DOM-specific code (skyvern-gh-1), and separately uses its vision model to detect and solve CAPTCHAs, which are canvas-like elements the DOM can't parse (skyvern-docs-9, skyvern-docs-27). Docs also mention falling back to selectors when useful (skyvern-docs-19), implying vision-first with DOM as a secondary path rather than a purely DOM-based tool needing a special switch. missing for 10: explicit documentation of a discrete 'vision/computer-use mode' toggle, dedicated canvas/non-DOM UI examples (e.g., canvas-drawn widgets, non-HTML apps), and independent benchmarking confirming success on such UIs.

      • [github] Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…
      • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…
      • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically.
      • [claimed-docs] Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.
      Smoothnone0/10

      The evidence pack covers Smooth's session workflows, extraction, custom tools, proxies, and CAPTCHA solving, but nowhere mentions a vision/computer-use mode operating on screenshots for canvases or non-DOM UI elements. This axis is plausible for a browser-automation agent, but no documentation or community report confirms such a capability exists.

      • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
      • [claimed-docs] Extract structured data from the current page by providing a schema.
      • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
      • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.

    Agenticness — how well agents can access and operate the productAgenticness

    How well agents can access and operate the product

    Agent access

    1. ai-native userPoint an agent at llms.txt or agent-oriented docs

      weight 2 · round drawn
      Skyvernfullprobed8/10

      A probe confirms llms.txt is live and returns a structured summary of Skyvern for agent consumption, and skyvern.com/llms provides an agent-oriented docs page listing features in a scannable format. This directly satisfies pointing an agent at llms.txt or agent-oriented docs. Missing for 10: a docs.md/markdown-mirrored docs endpoint (404) and an accessible OpenAPI spec, which would round out machine-readable documentation.

      • [probe] PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …
      • [claimed-docs] Visual workflow builder for non-developers — drag-and-drop, no code required
      • [claimed-docs] Browser recorder that converts manual actions into reusable automations
      • [claimed-docs] SOP upload — describe a process in plain English and Skyvern builds the workflow
      • [claimed-docs] Copilot chat for building and debugging workflows interactively
      • [probe] PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md
      • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
      Smoothfullprobed8/10

      A direct probe confirms docs.smooth.sh/llms.txt returns HTTP 200 with a structured index of docs pages, and the docs themselves are mirrored as .md files (e.g. live-share.md, proxies.md) enabling agent-friendly consumption. One community comment notes the docs pages aren't fully token-efficient, a minor caveat. Missing for 10: independent verification that agents actually consume llms.txt effectively, and no evidence of additional agent-specific doc formats beyond the single llms.txt file.

      • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
      • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
      • [claimed-docs] Log in once, then reuse that authentication for future tasks.
      • [community] Ironically, the landing page and docs pages of Smooth aren't all that token-efficient!
    2. ai-native userRun the product headlessly / in CI for automation

      weight 2 · round to Skyvern
      Skyvernfullclaimed7/10

      Skyvern ships a code-first SDK/REST API (Python/TypeScript) that connects to a cloud or self-hosted Chromium instance, explicitly positioned as replacing brittle Selenium scripts and integrating browser automation via API into other products, and can run entirely on your own infrastructure with your own LLM keys, supporting headless/scriptable use suitable for CI. missing for 10: explicit CI/CD pipeline documentation or example (e.g. GitHub Actions integration), and independent confirmation of headless execution in automated environments

      • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
      • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
      • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
      • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
      • [claimed-docs] You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.
      Smoothpartialprobed6/10

      Smooth is explicitly designed as an SDK/API and CLI for programmatic browser automation, with 'plug-and-play' 4-line-of-code task execution and custom tools/session workflows suited to unattended automation, and a documented CLI (smooth-gh-1, smooth-docs-1, smooth-probe-3). However there is no explicit documentation or example of running it inside a CI pipeline (e.g., GitHub Actions), headless flags, or exit-code/automation-specific guidance. Missing for 10: explicit CI/CD integration docs or examples, confirmation of non-interactive/headless auth flow for pipelines, independent confirmation of CI usage.

      • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
      • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
      • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
      • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
      • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
    3. ai-native userPlug MCP servers into this product so it can use their tools

      weight 3 · round drawn
      Skyvernnone0/10

      Evidence only shows Skyvern exposing an MCP *server* so external AI assistants (Claude, Cursor, etc.) can control Skyvern's browser — the reverse of the story, which asks whether Skyvern can consume external MCP servers' tools as a client. No documentation or community evidence shows Skyvern importing or connecting to third-party MCP servers to extend its own toolset.

      • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
      • [probe] official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp
      Smoothnone0/10

      Smooth documents a 'custom tools' feature for arbitrary functions but there is no mention anywhere in docs, GitHub, or community discussion of MCP server support or integration, so this applicable capability is unevidenced.

      • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
      • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
    4. ai-native userConnect an agent via an official MCP server

      weight 3 · round to Skyvern
      Skyvernfullprobed8/10

      Skyvern documents an official MCP server (skyvern-docs-7, skyvern-probe-4) that lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser via Skyvern, directly matching the story. Missing for 10: independent/hands-on community verification of the MCP server specifically (community evidence covers other features, not MCP usage) and a clear setup/config example beyond the single doc page.

      • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
      • [probe] official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp
      Smoothnone0/10

      Smooth is a browser-automation SDK/CLI used by agents like Claude Code, but the evidence pack contains no mention of an official MCP server or MCP protocol integration anywhere in its docs, GitHub repo, or community discussion — only generic SDK/CLI methods and custom tool functions are documented.

      • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
      • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
      • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
      • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
    5. ai-native userUse an official CLI

      weight 2 · round to Smooth
      Skyvernnone0/10

      Evidence covers Skyvern's Python/TypeScript SDKs, REST API, MCP server, and visual dashboard, but nowhere mentions an official CLI tool for AI-native workflows. missing for 10: any documented CLI command, npm/pip CLI package, or terminal-based interface.

      • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
      • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
      • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
      Smoothfullprobed7/10

      GitHub repo describes Smooth CLI explicitly as 'a browser for AI agents, enabling tools like Claude Code to navigate the web' and a docs probe confirms an official CLI overview page exists, showing a first-party CLI built for AI-agent workflows. Missing for 10: independent hands-on confirmation of the CLI's usage/reliability and more detailed CLI documentation content beyond the overview link.

      • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
      • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
    6. ai-native userDrive the product through a documented public API

      weight 3 · round to Skyvern
      Skyvernfullprobed7/10

      Skyvern documents a public API/SDK surface (Python, TypeScript, REST) for creating tasks, running multi-step browser automations, and extracting structured data via JSON schema, matching the ai-native 'drive via documented API' story; it also ships an MCP server for agent control. Missing for 10: a discoverable OpenAPI/swagger spec (probe found 404s for all candidate paths) and independent/hands-on confirmation of API robustness beyond first-party docs.

      • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
      • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
      • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
      • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
      • [claimed-docs] you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…
      • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
      • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
      • [probe] official MCP server documented at https://skyvern.com/docs/developers/getting-started/mcp
      Smoothpartialprobed6/10

      Smooth provides documented SDK/API methods (task execution, session workflows, structured outputs, custom tools, proxies) and an llms.txt docs index plus a CLI, showing a documented programmatic interface for AI-native use. However, no formal OpenAPI/REST spec was found (404s on all standard paths), and there is no independent corroboration of API robustness beyond vendor docs. missing for 10: a discoverable OpenAPI/REST spec, independent/hands-on verification of API completeness and stability.

      • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
      • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
      • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
      • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
      • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…
      • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
    7. ai-native userIssue scoped/least-privilege API credentials for an agent

      weight 2 · round drawn
      Skyvernnone0/10

      No evidence of scoped or least-privilege API credential issuance for agents; docs mention API keys and self-hosted LLM keys but nothing about credential scoping, permissions, or restricting agent access levels. Community comments even raise concerns about handling sensitive credentials in plain text with no mitigation shown. Missing for 10: any documentation of scoped API tokens, role-based access control, or least-privilege credential management for agents.

      • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
      • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
      • [community] you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…
      Smoothnone0/10

      No evidence of scoped or least-privilege API credential issuance, API key scoping, or permission management for agents; docs cover task execution, sessions, proxies, and privacy features but nothing about credential scoping. missing for 10: scoped API key/token generation, permission/role controls, credential revocation or least-privilege access management.

      • ai-native userBuild against official SDKs

        weight 2 · round to Skyvern
        Skyvernfullprobed7/10

        Skyvern explicitly documents official Python and TypeScript SDKs plus a REST API for integrating browser automation, with SDK-level primitives like page.extract and data_extraction_schema shown in docs (skyvern-docs-1, skyvern-docs-5, skyvern-docs-19, skyvern-docs-4). Missing for 10: independent/hands-on developer corroboration of SDK usage and a public API reference (OpenAPI spec probe returned 404s, skyvern-probe-3), so quality is capped below full confidence in completeness.

        • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
        • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
        • [claimed-docs] Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.
        • [claimed-docs] you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…
        • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
        Smoothpartialprobed6/10

        Smooth documents SDK-style integration (4-line task execution, custom tools, structured outputs, session workflows) and a CLI positioned as "a browser for AI agents" usable with tools like Claude Code, indicating official first-party SDK/CLI support for AI-native workflows. However, there is no OpenAPI spec, no evidence of multi-language SDKs, and no independent/hands-on confirmation of SDK reliability beyond docs and a GitHub repo. missing for 10: OpenAPI/API spec availability, multi-language SDK coverage, independent developer corroboration of SDK usage/quality.

        • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
        • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.
        • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
        • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
        • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
        • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…
      • ai-native userSubscribe to events via webhooks

        weight 2 · round drawn
        Skyvernnone0/10

        No evidence in the pack mentions webhooks or event subscriptions of any kind; Skyvern's documented integration surfaces are REST/SDK APIs, Zapier, and an MCP server, none of which constitute a webhook subscription mechanism.

          Smoothnone0/10

          No evidence pack item mentions webhooks or event subscriptions; Smooth's documented features (live_url, structured outputs, custom tools, sessions) do not include a webhook/event notification mechanism.

          Agentic features

          1. ai-native userSet up automations that run autonomously in the background

            weight 2 · round to Skyvern

            Skyvern's docs show multi-step, code-first and no-code workflows that run via API or cloud UI, persist browser state, and can pause for human approval while capturing recordings/artifacts — all indicative of autonomous background execution (skyvern-docs-5,6,10,11,21). Zapier integration and API-driven triggering (skyvern-docs-16, skyvern-docs-28) supports running without manual intervention, but there's no explicit documentation of a scheduler, cron-like triggers, or continuous monitoring dashboard for unattended runs. Missing for 10: explicit scheduling/trigger docs, evidence of long-running unattended background jobs, and independent confirmation of reliability at scale (community notes some brittleness, e.g. skyvern-comm-2).

            • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
            • [claimed-docs] Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.
            • [claimed-docs] Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…
            • [claimed-docs] Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…
            • [claimed-docs] Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.
            • [claimed-docs] Connect to Zapier
            • [claimed-docs] Skyvern automates browser-based workflows across these platforms — no API keys or custom connectors required.
            • [community] I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…
            Smoothpartialclaimed5/10

            Smooth lets users kick off browser-agent tasks programmatically that run autonomously (navigating, extracting data, solving CAPTCHAs) and provides a live_url to monitor progress, which supports hands-off execution once started. However there's no evidence of scheduling, triggers, webhooks, or persistent 'set it and forget it' background jobs that run without an explicit API call — the model shown is synchronous task invocation, not autonomous background automation setup. Missing for 10: scheduling/cron or event-trigger support, evidence of long-running unattended jobs, and independent confirmation of background execution beyond a single task call.

            • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
            • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
            • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
            • [claimed-docs] Auto-CAPTCHA solvers: Bypass CAPTCHA challenges automatically, allowing for uninterrupted task execution.
          2. ai-native userDelegate tasks to a built-in AI assistant inside the product

            weight 3 · round to Smooth
            Skyvernfullclaimed7/10

            Skyvern's core product IS an AI agent you delegate to via natural-language prompts to complete multi-step browser tasks (skyvern-docs-17, skyvern-docs-18), and it also ships a 'Copilot chat for building and debugging workflows interactively' inside the platform (skyvern-docs-25), plus SOP-to-workflow generation from plain English (skyvern-docs-24). This matches an AI-native user delegating tasks to a built-in assistant. Missing for 10: independent/hands-on validation of the copilot chat feature specifically (community evidence focuses on task execution quality, not the assistant/copilot UX), and no detail on assistant's conversational scope beyond workflow authoring.

            • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
            • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
            • [claimed-docs] SOP upload — describe a process in plain English and Skyvern builds the workflow
            • [claimed-docs] Copilot chat for building and debugging workflows interactively
            • [github] Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…

            Smooth's core product is a task-delegation interface: users hand off a task description (e.g., navigate, extract, multi-step session workflow) to Smooth's built-in AI/browser agent, which executes autonomously and returns live_url and structured outputs (smooth-docs-1, smooth-docs-7, smooth-docs-8, smooth-docs-5). Community hands-on feedback corroborates it executing complex prompts well (smooth-comm-2, smooth-comm-1), though it is agent-facing (tool for other agents like Claude Code) as well as human-facing. Missing for 10: independent/reproducible benchmarks of task success (raised unanswered in smooth-comm-14) and clearer human-only assistant UX beyond API/CLI task calls.

            • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
            • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
            • [claimed-docs] Extract structured data from the current page by providing a schema.
            • [claimed-docs] Structured outputs allow you to write deterministic code based on the agent's output. To activate structured outputs, set the `response_mode…
            • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
            • [community] Super impressive demo. Seems a lot faster than alternatives. How did you achieve that?
            • [community] I just wrote a complex prompt and it did a good job. How do you do evals or testing of your project?
          3. ai-native userOperate the product with natural-language commands

            weight 2 · round to Skyvern
            Skyvernfullcommunity7/10

            Skyvern's core interaction model is natural-language: users provide a prompt describing the goal (docs-18), SOPs in plain English are converted to workflows (docs-24), and a Copilot chat and MCP server let AI assistants/users direct browser actions in natural language (docs-25, docs-7, docs-19). This is corroborated by community reports of using it via prompts on real sites, though with mixed reliability on complex flows. missing for 10: independent benchmarking or hands-on confirmation that natural-language commands reliably handle complex multi-step tasks, and clearer evidence of NL-driven success rates beyond anecdotal HN reports.

            • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
            • [claimed-docs] SOP upload — describe a process in plain English and Skyvern builds the workflow
            • [claimed-docs] Copilot chat for building and debugging workflows interactively
            • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
            • [claimed-docs] Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.
            • [community] I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…
            Smoothpartialclaimed6/10

            Smooth's core interaction model is task-based: you give it a task description that an agent executes in a browser (session workflow, extract, navigate), and it is explicitly positioned as "a browser for AI agents" usable by tools like Claude Code, implying natural-language task instructions. However, no evidence shows an explicit example of a natural-language prompt/command syntax or confirms this is exposed to end-users beyond agent-to-agent orchestration. Missing for 10: explicit example of a natural-language task string/command, confirmation of human-facing NL command interface, independent corroboration of NL usability.

            • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
            • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
            • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
            • [claimed-docs] Extract structured data from the current page by providing a schema.

          Api quality

          1. ai-native userExplore an interactive API reference with runnable examples

            weight 2 · round drawn
            Skyvernnone0/10

            Evidence shows only static docs describing SDKs and REST usage, with no interactive API reference or runnable-example explorer; probes explicitly found no OpenAPI/Swagger spec at any candidate path (404s), indicating no interactive reference exists.

            • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
            • [probe] PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md
            • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
            Smoothnone0/10

            The evidence pack shows only static markdown docs (llms.txt, feature pages) and explicitly shows the openapi.json/swagger endpoints returning 404, indicating no interactive API reference or runnable-example playground exists. No mention of a Swagger UI, Postman collection, or in-browser code runner is present anywhere in docs or community discussion.

            • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
            • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…
          2. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

            weight 2 · round drawn
            Skyvernnone0/10

            Skyvern offers a REST API (skyvern-docs-1) but a direct probe for OpenAPI/swagger specs at standard paths returned 404 across all candidates, and no docs mention a downloadable machine-readable spec.

            • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
            • [probe] PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md
            • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
            Smoothnone0/10

            A direct probe for OpenAPI/Swagger spec files at all standard paths returned 404, and no docs page references a downloadable machine-readable API spec; only an llms.txt (docs index) is available.

            • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…
            • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…
          3. ai-native userTest against a sandbox environment without touching production data

            weight 1 · round drawn
            Skyvernnone0/10

            Skyvern's docs describe cloud or self-hosted execution, credential handling, and observability, but nothing describes a dedicated sandbox/staging mode isolated from production data or systems — missing for 10: any mention of a sandbox environment, test/staging mode, or data isolation guarantees.

              Smoothnone0/10

              Smooth's docs describe browser-automation features (live sessions, proxies, persistent auth, structured outputs) but nowhere mention a sandbox/staging mode or any mechanism to isolate test runs from production data or accounts. Community feedback even flags unresolved concerns about data handling and security, but no concrete sandbox capability is described or corroborated.

              • ai-native userRely on versioned APIs with a documented deprecation policy

                weight 2 · round drawn
                Skyvernnone0/10

                No evidence of API versioning scheme or a documented deprecation policy; OpenAPI/spec probes all returned 404 and no changelog or versioning docs appear in the pack. Missing for 10: versioned API endpoints (e.g., /v1/), a published deprecation/support policy, and changelog documentation.

                • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
                • [probe] PROBE docs-md: HTTP 404 at https://skyvern.com/docs.md
                Smoothnone0/10

                No evidence of API versioning scheme or a documented deprecation policy; OpenAPI spec probe returned 404s and docs show no changelog/versioning references. missing for 10: versioned API scheme, deprecation policy documentation, changelog/migration guides.

                • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…

              Auth session persistence — stories about auth session persistence in this arenaAuth session persistence

              Stories about auth session persistence in this arena

              Compat

              1. developerConnect my existing Playwright, Puppeteer, or CDP automation code to the product's browsers instead of rewriting it

                weight 2 · round to Skyvern
                Skyvernpartialclaimed5/10

                Docs show Skyvern's own SDK connects to a cloud Chromium instance over CDP and layers Playwright on top, and describe 'drop-in AI commands on top of Playwright' with fallback to raw selectors, implying some interoperability with existing Playwright code. However there is no explicit guidance or example showing a developer pointing an existing Playwright/Puppeteer/CDP script at Skyvern's managed browser instead of rewriting into Skyvern's task/workflow API, and no independent confirmation of this specific reuse pattern. Missing for 10: explicit BYO-script CDP endpoint docs, Puppeteer-specific support, and hands-on/community verification of dropping in existing automation code unchanged.

                • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                • [claimed-docs] Drop-in AI commands on top of Playwright. Use natural language to act, extract, and validate — or fall back to selectors.
                • [claimed-docs] you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…
                Smoothnone0/10

                Smooth's docs describe its own SDK/task API (structured outputs, sessions, custom tools) but there is no mention of a CDP endpoint, Playwright/Puppeteer connect() compatibility, or any way to point existing automation code at Smooth's browsers; one commenter even notes 'agent-browser helped a lot over playwright but doesn't completely close the gap,' underscoring the absence of such interoperability. Missing for 10: any CDP/WebSocket endpoint, official Playwright/Puppeteer connect examples, or documented browser-endpoint compatibility.

                • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                • [claimed-docs] Set to `"self"` to create a P2P tunnel through your machine, routing traffic via your IP and enabling access to localhost.
                • [community] agent-browser helped a lot over playwright but doesn't completely close the gap.

              Credentials

              1. automation-engineerStore credentials in a vault and have the agent complete logins including TOTP/2FA challenges without exposing secrets to the model

                weight 2 · round to Skyvern
                Skyvernfullcommunity7/10

                Skyvern's docs explicitly describe storing credentials via password-manager vault integrations (Bitwarden, 1Password, Azure Key Vault) and automatically handling TOTP/2FA, email, and SMS verification during login flows, matching the story closely [skyvern-docs-8][skyvern-docs-26]. However, there's no independent/hands-on verification that secrets are never exposed to the LLM, and a community comment raises concern about credentials being handled in plain text, so missing for 10: independent security audit or hands-on confirmation of secret-masking from the model, and clarification addressing the community's plaintext-handling concern.

                • [claimed-docs] Skyvern handles logins with stored credentials, TOTP/authenticator codes, email and SMS verification, magic links, and password manager inte…
                • [claimed-docs] Skyvern handles authentication end-to-end, from simple passwords to multi-factor flows with TOTP codes, email verification, and magic links.
                • [community] you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…
                Smoothnone0/10

                The docs describe persistent sessions (log in once, reuse authentication) but there is no mention of a credential vault, secret injection to avoid model exposure, or TOTP/2FA handling anywhere in the evidence pack. This axis clearly applies to a browser-automation agent product, but no capability matching the story is documented.

                • [claimed-docs] Log in once, then reuse that authentication for future tasks.

              Profiles

              1. developerPersist logged-in browser state in reusable profiles so agents skip the login wall on every subsequent run

                weight 3 · round to Smooth
                Skyvernpartialclaimed6/10

                Skyvern's browser-sessions feature explicitly persists cookies, local storage, and open tabs across operations so 'later operations pick up exactly where the previous one stopped,' and pauses preserve browser state — this directly supports skipping repeated logins. However, the docs don't clearly describe a named 'profile' abstraction, how long sessions persist across truly separate future runs, or how these persisted sessions are managed/reused across different agents or teams. missing for 10: explicit reusable-profile management docs, long-term persistence guarantees across independent runs, independent/hands-on confirmation of skip-login behavior.

                • [claimed-docs] Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.
                • [claimed-docs] Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…
                • [claimed-docs] Skyvern handles logins with stored credentials, TOTP/authenticator codes, email and SMS verification, magic links, and password manager inte…
                Smoothfullclaimed7/10

                Smooth's docs explicitly describe a persistent-sessions feature ('Log in once, then reuse that authentication for future tasks') and a session workflow that maintains a persistent browser session across multi-step tasks, directly matching the story. However, there is no independent/hands-on corroboration of this specific feature working reliably, and no detail on profile management (multiple reusable profiles, storage/export). Missing for 10: independent verification of session persistence in practice, documentation on managing multiple reusable profiles.

                • [claimed-docs] Log in once, then reuse that authentication for future tasks.
                • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…

              Automation depth — how much of the product can run unattendedAutomation depth

              How much of the product can run unattended

              1. ai-native userPerform bulk operations across many items at once

                weight 2 · round drawn
                Skyvernnone0/10

                No evidence describes a bulk-operation feature (e.g., running the same task across a list/CSV of items, batch triggering, or concurrent multi-item processing) — the docs focus on single-task API calls, visual workflows, and MCP integration rather than batch/bulk execution.

                  Smoothnone0/10

                  Smooth's docs describe single-task execution, session workflows, and structured extraction, but nothing about running/orchestrating bulk operations across many items (e.g., batch task queues, parallel task fan-out) is documented or mentioned by users.

                  • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
                  • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                  • [claimed-docs] Extract structured data from the current page by providing a schema.
                • ai-native userSchedule recurring jobs or workflows

                  weight 2 · round drawn
                  Skyvernnone0/10

                  The evidence pack describes Skyvern's workflow builder, API/SDK, MCP integration, and automation features extensively, but contains no mention of scheduling, cron triggers, or recurring job execution anywhere in the docs, GitHub description, or community discussion. Since Skyvern is a workflow/automation platform, scheduling recurring runs is a fair capability to expect, but it's simply absent from the provided evidence.

                    Smoothnone0/10

                    No evidence of scheduling, cron-like triggers, or recurring workflow orchestration; Smooth is documented as a task-execution/browser-automation tool (session workflows, extraction, structured output) with no mention of recurring/scheduled jobs. Missing for 10: any scheduling API, cron/trigger mechanism, or recurring workflow docs.

                    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                    • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.

                  Deployment modes — stories about deployment modes in this arenaDeployment modes

                  Stories about deployment modes in this arena

                  Local

                  1. developerRun the agent against a local browser on my own machine for development, without any cloud account

                    weight 2 · round to Skyvern
                    Skyvernpartialclaimed6/10

                    Skyvern is open-source and its self-hosted docs explicitly state it 'runs entirely on your infrastructure: your servers, your browsers, your LLM API keys' (skyvern-docs-12, skyvern-docs-2), which supports running without a cloud account. However, the core SDK/browser-automation flow described elsewhere connects to a 'cloud Chromium instance over CDP' (skyvern-docs-5), suggesting the default path is cloud-based, and no local-machine dev setup details (docker/local browser config, install steps) are shown. Missing for 10: explicit local-browser dev walkthrough, confirmation that the local-first SDK path bypasses cloud Chromium, and independent hands-on confirmation of local-only operation.

                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                    Smoothnone0/10

                    Smooth is documented as a cloud-hosted browser agent service (task execution via live_url, proxies, zero-data-retention as an 'enterprise' add-on), with no docs describing a local-browser/offline mode; a P2P tunnel feature only lets the cloud agent reach your localhost, not run without an account. Community feedback explicitly asks for self-hosting ('Make it self-hostable, the conversation can change'), confirming no local/no-account mode exists.

                    • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                    • [claimed-docs] Set to `"self"` to create a P2P tunnel through your machine, routing traffic via your IP and enabling access to localhost.
                    • [community] I'm unwilling to send my data to a 3rd party that is so new on the scene... Make it self-hostable, the conversation can change
                    • [community] Way too expensive, I'll wait for a free/open source browser optimized to be used by agents.

                  Framework model support — stories about framework model support in this arenaFramework model support

                  Stories about framework model support in this arena

                  Frameworks

                  1. developerPlug the browser layer into agent frameworks (Claude Agent SDK, Vercel AI SDK, LangChain, CrewAI) through documented adapters

                    weight 2 · round drawn
                    Skyvernnone0/10

                    Skyvern documents a Python/TypeScript/REST SDK and an MCP server that plugs into Claude Desktop, Claude Code, Codex, Cursor, and Windsurf, but there is no evidence of documented adapters for Claude Agent SDK, Vercel AI SDK, LangChain, or CrewAI. missing for 10: any documented integration guide or adapter for LangChain, CrewAI, Vercel AI SDK, or Claude Agent SDK specifically.

                    • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
                    • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                    • [claimed-docs] The Skyvern MCP server lets AI assistants like Claude Desktop, Claude Code, Codex, Cursor, and Windsurf control a browser.
                    Smoothnone0/10

                    Evidence shows Smooth positions itself as a browser tool usable by agents like Claude Code, but there is no documentation of adapters for Claude Agent SDK, Vercel AI SDK, LangChain, or CrewAI specifically. Missing for 10: any documented integration guide or adapter package for these named frameworks.

                    • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
                    • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.

                  Models

                  1. developerBring my own LLM provider — the framework is model-agnostic rather than locked to one vendor's models

                    weight 2 · round to Skyvern
                    Skyvernpartialclaimed6/10

                    Docs state self-hosted Skyvern runs with 'your own LLM API keys' on your own infrastructure, implying model-agnosticism rather than lock-in to a single vendor, but there is no explicit list of supported providers/models or first-party guide on swapping LLM backends, and no independent confirmation of multi-provider support. Missing for 10: an explicit supported-providers list/config docs, and community/hands-on evidence of using non-default LLMs.

                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    Smoothnone0/10

                    Smooth's docs describe it as using its own internal 'small and efficient AI models' for browser-agent tasks, and none of the documentation, CLI, or API references mention any mechanism for developers to configure or swap in their own LLM provider (e.g., OpenAI, Anthropic, custom model endpoints). This suggests the opposite of model-agnosticism — the models are an implementation detail of the service, not a pluggable choice.

                    • [claimed-docs] Smooth uses small and efficient AI models, making it 7x more affordable than browser-use.
                    • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
                    • [probe] PROBE llms.txt: HTTP 200 at https://docs.smooth.sh/llms.txt # Smooth ## Docs - [Introduction](https://docs.smooth.sh/index.md): Welcome to…

                  Nl task execution — stories about nl task execution in this arenaNl task execution

                  Stories about nl task execution in this arena

                  Tasks

                  1. ai agentSubmit a browser task over a hosted HTTP API and receive the result by polling or webhook, without managing any browser myself

                    weight 2 · round drawn
                    Skyvernpartialprobed5/10

                    Docs confirm a hosted REST/SDK API where you submit a prompt+URL (optionally a JSON schema) and Skyvern runs the task on cloud Chromium without the caller managing a browser (skyvern-docs-1, skyvern-docs-5, skyvern-docs-18, skyvern-docs-17). However, there is no explicit documentation of a polling endpoint or webhook callback mechanism, and probes found no discoverable OpenAPI spec, so the exact result-retrieval mechanism described in the story is unconfirmed. Missing for 10: explicit webhook/callback docs, explicit polling endpoint docs, and an accessible API reference confirming these mechanics.

                    • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
                    • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                    • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
                    • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
                    • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
                    Smoothpartialprobed5/10

                    Docs confirm a hosted task-submission model (4-line integration, live_url for tracking, persistent sessions, structured outputs) consistent with an agent submitting tasks without managing a browser, but no evidence pack item explicitly documents a polling endpoint or webhook delivery mechanism, and probes found no public OpenAPI/REST spec. missing for 10: explicit polling endpoint docs, explicit webhook/callback docs, confirmed REST API schema (openapi probe 404s).

                    • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
                    • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
                    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                    • [probe] PROBE openapi: all candidate paths 404 (https://docs.smooth.sh/openapi.json, https://docs.smooth.sh/swagger.json, https://docs.smooth.sh/api…
                    • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
                  2. developerHand the product a natural-language goal and it completes a multi-step web task end to end — navigating, filling forms, and clicking through flows

                    weight 3 · round to Smooth

                    Skyvern's docs and GitHub strongly claim natural-language goal execution across novel, multi-step web flows (forms, logins, CAPTCHAs) via a single prompt/API (skyvern-docs-17, skyvern-docs-18, skyvern-gh-1), but a hands-on community test found it succeeded only on the 'happy path' and concretely failed on a real multi-step flow (costcotravel.com), struggling to hit a tab and failing to click a popup (skyvern-comm-2). This is a specific documented counter-example contradicting the 'completes multi-step task end to end' claim, not just general skepticism. Missing for 10: independent benchmark results, more hands-on trials showing consistent success on complex/unseen sites, and resolution of the reported failure case.

                    • [claimed-docs] It navigates websites it has never seen before, filling forms, extracting data, and completing multi-step tasks via a simple API.
                    • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
                    • [github] Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…
                    • [community] I played with the Geico example, and it seems to do a good job on the happy path. But I tried costcotravel.com... it struggled to hit the 'r…

                    Docs describe exactly this capability: multi-step 'Session Workflow' that navigates URLs, orchestrates sub-tasks, and extracts data, plus a live_url to watch the agent act, and a community commenter confirms 'I just wrote a complex prompt and it did a good job.' This matches the natural-language, end-to-end web task story well. Missing for 10: independently reproducible benchmarks/evals (a commenter explicitly asks for third-party reproducible comparisons and gets no clear answer), and broader hands-on validation beyond a single anecdote.

                    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                    • [claimed-docs] Extract structured data from the current page by providing a schema.
                    • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
                    • [community] I just wrote a complex prompt and it did a good job. How do you do evals or testing of your project?
                    • [community] are your evals / comparisons publicly/3rd party reproducible? If it's 'trust me, I did a fair comparison', that's not going to fly today.

                  Workflows

                  1. automation-engineerCompose repeatable multi-step workflows with loops, conditionals, and parameters instead of one-shot prompts

                    weight 2 · round to Skyvern
                    Skyvernpartialclaimed5/10

                    Skyvern clearly supports multi-step, repeatable workflows via both a code-first SDK and a visual no-code drag-and-drop builder (skyvern-docs-5, skyvern-docs-6, skyvern-docs-22), plus SOP-to-workflow generation and a browser recorder for building reusable automations (skyvern-docs-23, skyvern-docs-24). However, the evidence never explicitly documents loop constructs, conditional branching, or parameterized workflow inputs as first-class workflow-builder features. Missing for 10: explicit documentation of loop/iteration blocks, conditional/branching logic, and named/typed workflow parameters in the workflow builder.

                    • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                    • [claimed-docs] Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.
                    • [claimed-docs] Visual workflow builder for non-developers — drag-and-drop, no code required
                    • [claimed-docs] Browser recorder that converts manual actions into reusable automations
                    • [claimed-docs] SOP upload — describe a process in plain English and Skyvern builds the workflow
                    • [github] a no-code workflow builder to help both technical and non-technical users automate manual workflows on any website, replacing brittle or unr…
                    Smoothpartialclaimed4/10

                    Smooth documents a 'Session Workflow' method for multi-step execution—orchestrating smaller tasks, navigating URLs, and extracting data within a persistent browser session—plus structured outputs and custom tools that let developers build deterministic logic around agent calls. However, there is no explicit documentation of native loop/conditional constructs or parameterized workflow templates; any control flow would rely on the surrounding SDK code rather than a built-in workflow engine. Missing for 10: explicit loop/conditional primitives, parameterization/templating of workflows, and independent evidence of repeatable multi-step automations beyond simple session chaining.

                    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                    • [claimed-docs] Structured outputs allow you to write deterministic code based on the agent's output. To activate structured outputs, set the `response_mode…
                    • [claimed-docs] Custom tools allow you to give Smooth any arbitrary function as a tool.

                  Openness — open source, data portability, and self-hosting storiesOpenness

                  Open source, data portability, and self-hosting stories

                  1. ai-native userDo everything through the API that I can do in the UI

                    weight 2 · round to Skyvern
                    Skyvernpartialprobed6/10

                    Skyvern's docs show a strong code-first path (Python/TS/REST SDKs, page.extract, workflow creation via API) that covers most core automation tasks also available in the dashboard, and MCP/REST access is documented. However, several UI-only tooling features (drag-and-drop visual builder, browser recorder, SOP upload, copilot chat) are described only as dashboard capabilities with no documented API equivalent, and no public OpenAPI/swagger spec was discoverable to confirm full API-UI parity. Missing for 10: documented API equivalents for recorder/SOP-upload/copilot-chat features, and a discoverable OpenAPI reference confirming complete parity.

                    • [claimed-docs] Integrate browser automation into your product with Python, TypeScript, or REST.
                    • [claimed-docs] Use the dashboard to run tasks and build agents visually.
                    • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                    • [claimed-docs] Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.
                    • [claimed-docs] Visual workflow builder for non-developers — drag-and-drop, no code required
                    • [claimed-docs] Browser recorder that converts manual actions into reusable automations
                    • [claimed-docs] SOP upload — describe a process in plain English and Skyvern builds the workflow
                    • [claimed-docs] Copilot chat for building and debugging workflows interactively
                    • [probe] PROBE openapi: all candidate paths 404 (https://skyvern.com/openapi.json, https://skyvern.com/swagger.json, https://skyvern.com/api/openapi.…
                    Smoothnone0/10

                    The evidence pack shows Smooth as an API/CLI/SDK-first browser-automation tool with docs for tasks, sessions, proxies, custom tools, and a live_url for viewing agent actions, but there is no mention of a separate web dashboard/UI or any comparison of UI-only vs API-only capabilities. Without evidence of what a UI offers (or that all UI features are mirrored in the API), the parity claim can't be substantiated. Missing for 10: any documented web UI/dashboard, and an explicit statement or demonstration that all UI actions are also achievable via API.

                    • [claimed-docs] Plug-and-play: Run a task in just 4 lines of code, making it easy to integrate into your workflow.
                    • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                    • [probe] official CLI documented at https://docs.smooth.sh/cli/overview
                  2. ai-native userExport all of my data in open formats and leave

                    weight 3 · round to Skyvern
                    Skyvernpartialclaimed3/10

                    Skyvern is open-source and self-hostable, meaning your data (artifacts, recordings, screenshots, network traffic) stays on your own infrastructure rather than being locked in a vendor's cloud, which implicitly supports data portability. However, there is no explicit documentation of a data export feature, standard open-format export (e.g., JSON/CSV bulk export of run history), or a stated 'leave with your data' workflow. Missing for 10: explicit export functionality/documentation, named open data formats, and any independent confirmation of successful data migration out of the platform.

                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    • [claimed-docs] Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…
                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    Smoothnone0/10

                    Smooth's docs mention Zero Data Retention (deletion of task data) but there is no evidence of a bulk data export feature or open-format export for users to take their data and leave — the closest related item is deletion, not portability. missing for 10: any documented export mechanism, open format specification, or user data portability tooling.

                    • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                  3. ai-native userRead the product's source under an open license

                    weight 2 · round to Skyvern

                    Skyvern is described as open-source with a GitHub repo, and probe evidence confirms it self-identifies as 'open-source' (skyvern-probe-1), but community evidence directly contradicts full open-license access, noting the project is AGPL3 licensed, which is a legally open license but is called out as a practical non-starter/restrictive for many users (skyvern-comm-3). missing for 10: explicit statement of license terms in docs, confirmation of what percentage of the product (cloud vs self-hosted) is actually open-sourced, and independent corroboration that the full source is readable without restriction.

                    • [github] Skyvern can operate on websites it's never seen before, as it's able to map visual elements to actions necessary to complete a workflow, wit…
                    • [github] a no-code workflow builder to help both technical and non-technical users automate manual workflows on any website, replacing brittle or unr…
                    • [probe] PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …
                    • [community] Exciting stuff, my employer would be interested but it's AGPL3 licensed so it's a non-starter for them.
                    Smoothnone0/10

                    Smooth ships a GitHub repo for its SDK/CLI, but there is no evidence of an open-source license for the core product, and community comments explicitly request self-hosting/open-source alternatives ('Make it self-hostable, the conversation can change', 'I'll wait for a free/open source browser'), implying the core service is closed.

                    • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
                    • [community] I'm unwilling to send my data to a 3rd party that is so new on the scene... Make it self-hostable, the conversation can change
                    • [community] Way too expensive, I'll wait for a free/open source browser optimized to be used by agents.
                  4. ai-native userSelf-host the core product

                    weight 3 · round to Skyvern
                    Skyvernfullprobed8/10

                    Skyvern has a dedicated self-hosted docs page stating it 'runs entirely on your infrastructure: your servers, your browsers, your LLM API keys' (skyvern-docs-12, skyvern-docs-2), and community evidence confirms it is genuinely open-source (AGPL3) rather than just marketing language (skyvern-comm-3). Missing for 10: independent hands-on confirmation of a successful self-host deployment and clarity on how AGPL licensing affects commercial self-hosting use.

                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [community] Exciting stuff, my employer would be interested but it's AGPL3 licensed so it's a non-starter for them.
                    • [probe] PROBE llms.txt: HTTP 200 at https://skyvern.com/llms.txt # Skyvern > Skyvern is an open-source, AI-powered browser automation platform. It …
                    Smoothnone0/10

                    Smooth is offered only as a hosted cloud API/SaaS with no documented self-host option, and community feedback explicitly requests self-hosting as a missing capability ('Make it self-hostable, the conversation can change').

                    • [community] I'm unwilling to send my data to a 3rd party that is so new on the scene... Make it self-hostable, the conversation can change
                    • [community] My first question was whether I could use this for sensitive tasks, given that it's not running on our machines. And after poking around for…

                  Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits

                  Free-tier ceilings, usage caps, and rate limits before you have to pay

                  Pricing

                  1. developerSee transparent per-task or per-browser-hour pricing and documented rate/concurrency limits before committing

                    weight 2 · round drawn
                    Skyvernnone0/10

                    The pricing page is referenced only for its target-audience blurb (skyvern-docs-13); no evidence pack item shows actual per-task or per-browser-hour rates, tiers, or documented rate/concurrency limits. Community comments only express general cost concerns ('pretty pricey', wanting cost down 'at scale') without citing concrete published pricing or limits.

                    • [claimed-docs] You're replacing brittle Selenium scripts, integrating browser automation via API, or building workflows into your product.
                    • [community] I tried it out and it's pretty pricey. My OpenAI API bill is $3.20 after using this on a few different pages to test it out... this is alway…
                    • [community] This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…
                    Smoothnone0/10

                    No evidence of documented per-task/per-browser-hour pricing tiers or rate/concurrency limits; only a vague claim of being '7x more affordable' with no actual pricing page or limits documented, and community comments call it 'too expensive' without citing specifics.

                    • [claimed-docs] Smooth uses small and efficient AI models, making it 7x more affordable than browser-use.
                    • [community] Way too expensive, I'll wait for a free/open source browser optimized to be used by agents.
                    • [community] I'm paying a fixed amount on Claude and other agents, so 'more tokens' is 'free' for me. There's a lot of niche tools out there but I think …

                  Privacy posture — data-handling and privacy storiesPrivacy posture

                  Data-handling and privacy stories

                  1. ai-native userChoose where my data is stored (region/residency)

                    weight 2 · round to Skyvern
                    Skyvernpartialclaimed5/10

                    Skyvern offers a self-hosted deployment mode where 'your servers, your browsers, your LLM API keys' run entirely on the user's own infrastructure, which lets an AI-native user control where data resides by choosing their hosting region themselves — but this is achieved only by self-hosting, not via an explicit region/residency selector in the managed cloud product. Missing for 10: documented data residency/region options in the hosted Skyvern Cloud offering, compliance certifications, or explicit multi-region storage controls.

                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    Smoothnone0/10

                    No evidence of data residency/region selection options; only Zero Data Retention (deletion) is mentioned, which is a different capability. Missing for 10: any mention of region choice, data center locations, or residency controls.

                    • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                  2. ai-native userPrevent my data from being used to train AI models

                    weight 3 · round drawn
                    Skyvernpartialclaimed4/10

                    Skyvern offers self-hosted deployment using your own infrastructure and your own LLM API keys, which implicitly lets users avoid sending data to Skyvern-controlled models/training pipelines, but there is no explicit privacy policy, data-retention statement, or 'we do not train on your data' commitment in the evidence for the hosted/cloud offering. Missing for 10: explicit no-training/data-use policy documentation, opt-out mechanism for the cloud product, and independent confirmation of data handling practices.

                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.

                    Smooth documents a 'Zero Data Retention' enterprise feature that lets customers delete all data tied to completed tasks, which is adjacent to preventing data reuse, but there is no explicit statement that data is excluded from model training, and this feature is gated to enterprise tier. Community feedback also notes an absence of any detailed security/privacy documentation despite marketing claims of 'enterprise-grade security', raising trust concerns without disputing the ZDR feature itself. missing for 10: explicit AI-training opt-out policy, default (non-enterprise) privacy guarantees, independent verification of data handling.

                    • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                    • [community] My first question was whether I could use this for sensitive tasks, given that it's not running on our machines. And after poking around for…
                  3. ai-native userControl data retention and deletion

                    weight 2 · round to Smooth

                    Skyvern offers self-hosting (docs-2, docs-12) which gives infrastructure-level control over where data lives, and it captures artifacts (recordings, screenshots, network traffic) per run (docs-11), implying some data exists to manage, but there is no documented retention policy, data deletion API/UI, or export/purge controls for the cloud/hosted product. missing for 10: explicit data retention policy, user-facing deletion/export controls, documentation on how long artifacts/credentials are stored in cloud mode, and independent confirmation that self-hosting actually eliminates vendor-side data retention.

                    • [claimed-docs] Run Skyvern on your own infrastructure with your own LLM keys.
                    • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                    • [claimed-docs] Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…
                    • [community] you are expecting them to pass over their website login credentials and apparently their credit card details too, in plain text. You had bet…

                    Docs confirm a 'Zero Data Retention' feature letting customers delete all data tied to completed tasks, directly addressing retention/deletion control, but it's explicitly gated as an 'enterprise feature' rather than a universal capability, and no detail is given on default retention periods, deletion APIs/CLI commands, or granular controls for non-enterprise users. Community feedback (e.g., concerns about sending data to a third party, no security details found) shows some skepticism but doesn't concretely contradict the ZDR claim itself. missing for 10: default/non-enterprise retention policy, self-serve deletion mechanism (API/CLI), independent verification of ZDR working in practice.

                    • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                    • [community] My first question was whether I could use this for sensitive tasks, given that it's not running on our machines. And after poking around for…
                  4. ai-native userOpt out of telemetry and usage tracking

                    weight 2 · round drawn
                    Skyvernnone0/10

                    No evidence pack item discusses telemetry, usage tracking, or an opt-out mechanism; while self-hosting exists, there is no explicit statement about data collection or opt-out controls for the cloud/hosted product. missing for 10: any mention of telemetry collection, privacy policy on usage data, or an opt-out setting/flag.

                      Smoothnone0/10

                      No evidence describes a telemetry/usage-tracking opt-out control; the only related privacy feature is 'Zero Data Retention' for enterprise customers, which addresses data deletion after tasks rather than disabling telemetry/tracking. Community comments raise general privacy/security concerns but do not confirm or deny an opt-out mechanism.

                      • [claimed-docs] Zero Data Retention is an enterprise feature that provides enhanced data privacy by allowing you to delete all data associated with complete…
                      • [community] My first question was whether I could use this for sensitive tasks, given that it's not running on our machines. And after poking around for…

                    Replay debugging — stories about replay debugging in this arenaReplay debugging

                    Stories about replay debugging in this arena

                    Live

                    1. automation-engineerWatch a session live and take human control mid-run when the agent gets stuck

                      weight 2 · round to Skyvern
                      Skyvernfullclaimed8/10

                      Docs explicitly state a VNC stream lets you watch a live session and take control at any point, plus pause-for-approval human-in-the-loop flows that preserve browser state — directly matching the story. missing for 10: independent/hands-on confirmation of the live takeover UX and details on how control handoff works mid-run beyond the docs description.

                      • [claimed-docs] Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…
                      • [claimed-docs] Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.
                      Smoothpartialclaimed4/10

                      Docs confirm a live_url to watch agent actions in real time (smooth-docs-2), satisfying the 'watch a session live' half of the story, but there is no evidence of any mechanism for a human to intervene or take control mid-run when the agent gets stuck. missing for 10: documented human-takeover/control API or UI, evidence of pausing/resuming agent execution, hands-on confirmation of mid-run intervention.

                      • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.
                      • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…

                    Replay

                    1. automation-engineerDebug a failed agent run from recorded replays — video, screenshots, step-by-step action timelines

                      weight 2 · round to Skyvern
                      Skyvernfullcommunity8/10

                      Skyvern docs explicitly state every run captures session recordings, per-step screenshots, AI reasoning traces, and network traffic for debugging, and community feedback corroborates strong observability into workflow steps. missing for 10: independent hands-on verification of the video/timeline UI itself, and no mention of a true step-by-step interactive timeline scrubber beyond artifact capture.

                      • [claimed-docs] Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…
                      • [community] This is an impressive tool. I especially like the observability around the workflow and the steps it takes to achieve the outcome. We are po…
                      Smoothpartialclaimed3/10

                      Docs mention a `live_url` for viewing agent actions live during a run, but there is no evidence of persisted video recordings, screenshots, or a step-by-step action timeline that can be replayed after a run has finished and failed. Missing for 10: recorded video/screenshot artifacts, post-hoc replay viewer, structured action timeline, and any independent confirmation of replay-based debugging.

                      • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.

                    Scale parallelism — running many jobs at once — concurrency, fleets, queueingScale parallelism

                    Running many jobs at once — concurrency, fleets, queueing

                    Fleets

                    1. automation-engineerRun a fleet of concurrent browser sessions with documented concurrency limits and programmatic session management

                      weight 2 · round drawn
                      Skyvernnone0/10

                      Evidence covers session persistence, VNC control, self-hosting, and SDK/API access, but nowhere documents concurrency limits, fleet-level session orchestration, or programmatic management of multiple simultaneous browser sessions. Missing for 10: documented concurrency limits, APIs for spinning up/managing many parallel sessions, and any scaling/throughput guidance.

                      • [claimed-docs] Human-in-the-loop flows: pause for approval between steps without losing browser state. The VNC stream lets you watch or take control at any…
                      • [claimed-docs] Cookies, local storage, open tabs, and the current page all persist, so later operations pick up exactly where the previous one stopped.
                      • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                      • [claimed-docs] Self-hosted Skyvern runs entirely on your infrastructure: your servers, your browsers, your LLM API keys.
                      Smoothnone0/10

                      The evidence pack covers single-session features (persistent sessions, live URL, proxies, structured output) but contains no documentation of concurrency limits, fleet/pool management, or APIs for running many sessions in parallel. Missing for 10: documented concurrency limits, fleet/pool orchestration APIs, rate-limit or scaling guidance, evidence of parallel session usage.

                      • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                      • [claimed-docs] Log in once, then reuse that authentication for future tasks.
                      • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.

                    Lifecycle

                    1. developerGet webhook notifications when tasks and sessions finish instead of polling for status

                      weight 1 · round drawn
                      Skyvernnone0/10

                      No evidence pack item mentions webhooks, callback URLs, or push notifications for task/session completion; the docs discuss artifacts, VNC streaming, and human-in-the-loop review but nothing about event-driven notification instead of polling. missing for 10: any documentation of webhook/callback support, event subscription API, or notification configuration.

                        Smoothnone0/10

                        No evidence pack item mentions webhooks, callback URLs, or event-driven notifications for task/session completion; the docs describe live_url viewing, structured outputs, and session workflows but nothing about push notifications versus polling.

                        Stealth captcha — stories about stealth captcha in this arenaStealth captcha

                        Stories about stealth captcha in this arena

                        Captcha

                        1. automation-engineerRely on a documented captcha stance — automatic solving, human fallback, or explicit non-support — instead of silent task failures

                          weight 2 · round to Skyvern
                          Skyvernfullclaimed8/10

                          Skyvern's docs give an explicit, detailed captcha stance: automatic detection and solving via its vision model for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstile, FunCaptcha, MTCaptcha, and text/image captchas, avoiding silent failure ambiguity. Missing for 10: independent/hands-on confirmation that captcha solving works reliably in practice (community evidence discusses pricing, mobile UX, and credential handling but not captcha outcomes specifically), and no documented fallback/human-in-the-loop behavior specifically tied to captcha failures.

                          • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…
                          • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically.

                          Smooth explicitly documents an automatic captcha-solving stance ('Auto-CAPTCHA solvers: Bypass CAPTCHA challenges automatically, allowing for uninterrupted task execution'), giving automation engineers a clear documented behavior rather than silent failure. Community reaction (smooth-comm-3) criticizes the ethics/marketing of this feature but does not present a hands-on failure showing the solver doesn't work, so this remains a documented claim rather than a disputed one. Missing for 10: independent/hands-on verification that auto-solving actually succeeds in practice, and no documentation of fallback behavior (e.g., what happens if a captcha can't be auto-solved).

                          • [claimed-docs] Auto-CAPTCHA solvers: Bypass CAPTCHA challenges automatically, allowing for uninterrupted task execution.
                          • [community] So you're shamelessly selling spambots? The marketing here is wild... "proxy rotation"... "auto-CAPTCHA solvers"

                        Posture

                        1. automation-engineerPoint to the vendor's published acceptable-use and anti-abuse posture governing what its stealth and automation features may be used for

                          weight 1 · round drawn
                          Skyvernnone0/10

                          No evidence in the pack of any published acceptable-use policy, terms of service, or anti-abuse statement covering CAPTCHA-solving/stealth automation features; docs describe capabilities (CAPTCHA bypass, bot bypass) but no governance/AUP language is cited. missing for 10: a published acceptable-use policy, anti-abuse terms, or statement on permitted use of stealth/CAPTCHA-bypass features.

                            Smoothnone0/10

                            No evidence anywhere in the pack of a published acceptable-use policy, anti-abuse terms, or guidance on permissible use of the stealth/CAPTCHA-bypass and automation features; docs only describe how to use auto-CAPTCHA and proxy features, not what usage is disallowed. Community commentary even calls out the lack of any such framing (e.g., accusing the marketing of enabling spambots), reinforcing the absence rather than disputing a claim.

                            • [claimed-docs] Auto-CAPTCHA solvers: Bypass CAPTCHA challenges automatically, allowing for uninterrupted task execution.
                            • [community] So you're shamelessly selling spambots? The marketing here is wild... "proxy rotation"... "auto-CAPTCHA solvers"

                          Stealth

                          1. automation-engineerEnable stealth fingerprinting and residential or geo-targeted proxies so legitimate automations aren't blocked as bots

                            weight 2 · round to Smooth
                            Skyvernnone0/10

                            Evidence covers CAPTCHA solving and authentication/2FA handling, but there is no mention anywhere of stealth fingerprinting, browser fingerprint spoofing, or residential/geo-targeted proxy support. Missing for 10: any documentation of proxy configuration, geo-targeting, or anti-fingerprinting/stealth mode features.

                            • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically. This works for reCAPTCHA v2/v3, hCaptcha, Cloudflare Turnstil…
                            • [claimed-docs] Skyvern detects CAPTCHAs using its vision model and solves them automatically.

                            Docs confirm auto-CAPTCHA solving and configurable proxy server parameters plus persistent authenticated sessions, which support anti-bot automation goals, but there is no explicit mention of residential/geo-targeted proxy pools or stealth browser fingerprinting techniques. Missing for 10: explicit residential/geo-targeted proxy options, stealth fingerprinting details, and independent verification that bot-detection evasion actually works in practice.

                            • [claimed-docs] To use a proxy with Smooth, you need to specify the proxy server details in your task parameters.
                            • [claimed-docs] Auto-CAPTCHA solvers: Bypass CAPTCHA challenges automatically, allowing for uninterrupted task execution.
                            • [claimed-docs] Log in once, then reuse that authentication for future tasks.
                            • [community] So you're shamelessly selling spambots? The marketing here is wild... "proxy rotation"... "auto-CAPTCHA solvers"

                          Structured extraction — stories about structured extraction in this arenaStructured extraction

                          Stories about structured extraction in this arena

                          Extraction

                          1. developerExtract typed, schema-validated data (Zod/Pydantic-style) from pages the agent visits, not just raw text

                            weight 3 · round to Smooth
                            Skyvernfullclaimed7/10

                            Skyvern's docs explicitly support structured, schema-based extraction via `page.extract` with a JSON schema or `data_extraction_schema` param, matching the developer's need for typed output rather than raw text (skyvern-docs-4, skyvern-docs-20, skyvern-docs-18). However, evidence only shows JSON-schema validation, not native Zod/Pydantic model binding, and there's no independent/hands-on confirmation of this specific feature. missing for 10: explicit Zod/Pydantic model integration examples, independent verification of extraction accuracy/schema enforcement.

                            • [claimed-docs] you can extract structured data from any page using `page.extract` with a JSON schema, or by passing a `data_extraction_schema` to `page.age…
                            • [claimed-docs] you can extract structured data from any page using page.extract with a JSON schema
                            • [claimed-docs] You provide a natural-language prompt describing the goal, a starting URL, and optionally a JSON schema for structured output.
                            • [claimed-docs] Browser Automation is the code-first way to build multi-step automations. The Skyvern SDK connects to a cloud Chromium instance over CDP, la…
                            Smoothfullclaimed8/10

                            Docs explicitly describe structured outputs via `response_model` for deterministic typed data and a dedicated `session-extract` method to extract structured data from a page by providing a schema, directly matching the story. Missing for 10: explicit Zod/Pydantic code examples and independent/hands-on confirmation that extraction validation works as documented.

                            • [claimed-docs] Structured outputs allow you to write deterministic code based on the agent's output. To activate structured outputs, set the `response_mode…
                            • [claimed-docs] Extract structured data from the current page by providing a schema.
                            • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…

                          Files

                          1. developerMy agent can download files from and upload files to the sites it operates, with the artifacts retrievable afterwards

                            weight 1 · round to Skyvern
                            Skyvernpartialclaimed5/10

                            Docs show Skyvern can log into vendor portals and download PDFs (skyvern-docs-15) and captures per-run artifacts like recordings, screenshots, and network traffic retrievable afterward (skyvern-docs-11), implying file download support, but there is no explicit documentation of file upload capability to sites, nor of a dedicated API/UI for retrieving downloaded artifacts as opposed to just run/debug artifacts. missing for 10: explicit upload-to-site capability documentation, a documented file-download/artifact storage API distinct from debugging screenshots, and independent/hands-on confirmation of file transfer working in practice.

                            • [claimed-docs] Log into vendor portals, find invoices, download PDFs.
                            • [claimed-docs] Every run automatically captures what happened: recordings of the browser session, screenshots at each step, the AI's reasoning, and network…
                            • [claimed-docs] Auto-fill and submit applications on Lever, Greenhouse, and more.
                            Smoothnone0/10

                            The evidence pack describes Smooth's session workflows, structured extraction, live-view, proxies, and persistent auth, but nothing addresses file download/upload during a browser session or persisting artifacts for later retrieval. Since browser automation tools plausibly support file transfer (e.g., downloading a report from a site or uploading a document to a form), this axis applies but is unaddressed.

                            • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                            • [claimed-docs] Extract structured data from the current page by providing a schema.
                            • [claimed-docs] When running a task, you will receive a `live_url`, which can be used to view the agent actions live.

                          Not comparable on these axes

                          1. ai-native userGet AI-generated insights and suggestions from my data inside the product

                            weight 2 · not comparable
                            Skyvernn/a

                            Skyvern is a browser-automation/agent platform for executing web tasks and extracting data per user-specified schemas, not a product that analyzes a user's own data corpus to surface proactive insights or suggestions; this axis is a category mismatch for its purpose.

                              Smoothn/a

                              Smooth is a browser-automation SDK/CLI that lets AI agents navigate the web and extract structured data from pages — it is not a data platform or analytics product with a UI that surfaces AI-generated insights/suggestions from a user's own data. This story targets a different product category (BI/analytics-style in-product insights), so it does not apply to Smooth's browser-agent tooling.

                              • [github] The Smooth CLI is a browser for AI agents, enabling tools like Claude Code to navigate the web quickly, cheaply, and reliably.
                              • [claimed-docs] Session Workflow — Multi-step execution where you can orchestrate smaller tasks, navigate to URLs, and extract data within a persistent brow…
                              • [claimed-docs] Extract structured data from the current page by providing a schema.
                            • ai-native userDefine rules that trigger actions automatically on events

                              weight 3 · not comparable
                              Skyvernpartialclaimed3/10

                              Skyvern documents a Zapier integration, which could allow external events to trigger Skyvern workflows, but there is no evidence of a native rule/trigger engine, webhooks, or scheduled/event-based automation within Skyvern itself. Missing for 10: documented native event triggers or webhook listeners, schedule-based triggers, and any conditional rule engine inside Skyvern's workflow builder.

                              • [claimed-docs] Connect to Zapier
                              • [claimed-docs] Build multi-step automations visually in the Cloud UI with drag-and-drop blocks. No code required. Share templates across your team.
                              Smoothn/a

                              Smooth is a browser-automation/agent-tool product for running tasks on demand, not an event-driven rules/automation-trigger platform; there is no mention of defining rules or triggers that fire actions on events, so this automation-depth axis (workflow/event triggers) is a category mismatch for this product type.

                              • ai-native userVersion, review, and roll back my automations

                                weight 1 · not comparable
                                Skyvernnone0/10

                                No evidence in the pack describes version history, change review, or rollback capabilities for Skyvern workflows/automations — only building, running, sharing templates, and artifact capture (recordings/screenshots) are documented. Missing for 10: workflow version history, diff/review UI, rollback-to-previous-version mechanism, any changelog or audit trail for automation edits.

                                  Smoothn/a

                                  Smooth is a browser-automation/AI-agent-browsing tool (task execution, sessions, structured extraction) rather than an automation-authoring platform with version history or workflow rollback semantics; versioning/review/rollback of 'automations' is not an applicable axis for this product category.