Skip to content

Web Scraping APIs Arena

Crawl4AI vs Riveter

Crawl4AI wins · 3520 (35 drawn)

Agenticness — how well agents can access and operate the productAgenticness

How well agents can access and operate the product

Agent access

  1. ai-native userPoint an agent at llms.txt or agent-oriented docs

    weight 2 · round drawn
    Crawl4AInone0/10

    A direct probe of the docs site found no llms.txt (404), and there is no evidence of any agent-oriented docs format for AI agents to consume; the evidence pack shows only standard human-readable documentation and CLI/MCP references that don't satisfy this story.

    • [probe] PROBE llms.txt: HTTP 404 at https://docs.crawl4ai.com/llms.txt
    Riveternone0/10

    Direct probes show llms.txt returns 404 and no OpenAPI spec is discoverable at any standard path, and no evidence pack item claims an agent-oriented docs format exists; while MCP integration is mentioned, that's a separate capability from machine-readable docs for pointing an agent at.

    • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
    • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
  2. ai-native userRun the product headlessly / in CI for automation

    weight 2 · round to Crawl4AI
    Crawl4AIfullprobed8/10

    Crawl4AI ships a CLI (crwl), a Python async API usable in scripts, and a Dockerized FastAPI server setup explicitly for deployment/automation, all consistent with headless CI use; community evidence confirms production/Docker/n8n integrations. Missing for 10: no explicit CI pipeline example (e.g., GitHub Actions) or headless-mode flag documentation in the pack.

    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
    • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
    • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
    • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
    Riveterpartialclaimed6/10

    Riveter exposes a full API with SDKs (Go example shown), webhooks for async completion, dry_run/max_credits safety controls, and scheduling for recurring automation — all of which support headless, non-interactive use in a pipeline. However, there is no explicit CI/CD example, GitHub Actions integration, or CLI documentation demonstrating a documented headless workflow. Missing for 10: explicit CI/CD or pipeline integration guide, CLI headless invocation docs, independent confirmation of automated/scripted runs.

    • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
    • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
    • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
    • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
    • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
  3. ai-native userConnect an agent via an official MCP server

    weight 3 · round to Riveter
    Crawl4AIpartialprobed7/10

    Official docs explicitly document an MCP (Model Context Protocol) server for self-hosting, confirming Crawl4AI ships a first-party MCP integration point for agents. However, community evidence notes developers commonly struggle with configuring MCP servers for tools like Cursor, indicating real-world friction rather than a seamless plug-and-play experience. Missing for 10: detailed first-party MCP server docs/spec excerpt, independent hands-on confirmation of successful agent connection, and evidence the setup struggles are resolved.

    • [probe] official MCP server documented at https://docs.crawl4ai.com/core/self-hosting/#mcp-model-context-protocol-support
    • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
    • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
    Riveterfullclaimed7/10

    Docs explicitly describe connecting Riveter to Claude, ChatGPT, Cursor, or any MCP-compatible assistant via two connection methods, including a local Node.js-based server option, indicating an official MCP server offering. Missing for 10: no independent/hands-on corroboration of the MCP server working, and no detail on the remote/hosted connection method's implementation.

    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
    • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.
  4. ai-native userUse an official CLI

    weight 2 · round to Crawl4AI
    Crawl4AIfullprobed8/10

    There is a documented official CLI (`crwl`) with deep-crawl and other flags shown in GitHub examples, plus a dedicated docs page confirming it as an official feature. missing for 10: independent/hands-on third-party verification of the CLI's usage and a fuller list of supported CLI commands/flags beyond the single example.

    • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
    • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
    Riveternone0/10

    Evidence shows SDKs (Go), a local MCP server requiring Node.js, and REST API features, but no mention of an official CLI tool for running enrichments or managing the product. The docs and probes (llms.txt, openapi) surface no CLI reference, so this applicable axis is unmet.

    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
    • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.
    • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
  5. ai-native userDrive the product through a documented public API

    weight 3 · round drawn
    Crawl4AIpartialprobed6/10

    Crawl4AI ships a documented Python async API (AsyncWebCrawler.arun), a CLI, and a Dockerized FastAPI server plus an official MCP endpoint, giving AI agents multiple programmatic ways to drive it. However, probes show no discoverable OpenAPI spec or llms.txt for the hosted API, meaning the REST/API surface isn't formally machine-documented in a standard way. Missing for 10: a published OpenAPI/swagger schema, llms.txt, and independent confirmation of API stability/versioning.

    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
    • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
    • [probe] official MCP server documented at https://docs.crawl4ai.com/core/self-hosting/#mcp-model-context-protocol-support
    • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
    • [probe] PROBE llms.txt: HTTP 404 at https://docs.crawl4ai.com/llms.txt
    • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
    Riveterpartialprobed6/10

    Docs describe concrete API mechanics (webhook_url, dry_run, max_credits, SDK auth/retry/pagination handling, Go SDK code sample) showing a real documented public API surface for driving runs programmatically, and MCP/remote-server integration is documented. However, probes for a formal machine-readable spec (openapi.json/swagger.json) and llms.txt all returned 404, so there's no discoverable canonical API reference, undermining full 'documented public API' claims. missing for 10: a public OpenAPI/swagger spec or llms.txt confirming a fully machine-readable API contract, independent third-party confirmation of API usage.

    • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
    • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
    • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
    • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
    • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
    • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
  6. ai-native userIssue scoped/least-privilege API credentials for an agent

    weight 2 · round drawn
    Crawl4AInone0/10

    Crawl4AI is an open-source library/self-hosted tool that explicitly avoids API keys ('No forced API keys'), and there is no evidence of any credential issuance system, scoped tokens, or least-privilege access controls for agents; auth-related evidence only covers browser profile cookies/session state, not API credential scoping.

    • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
    • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
    Riveternone0/10

    Riveter's evidence covers a single API key model, credit caps, and dry-run cost estimation, but there is no mention of scoped or least-privilege credentials, per-agent tokens, or permission scoping for agents. missing for 10: scoped/least-privilege credential issuance, per-agent API key scoping, role/permission-based access control.

    • ai-native userBuild against official SDKs

      weight 2 · round drawn
      Crawl4AIfullprobed7/10

      Crawl4AI ships a first-party Python SDK (AsyncWebCrawler API, extraction strategies, CLI) that is well documented and used directly by developers per docs and GitHub. missing for 10: no official SDKs beyond Python (e.g., JS/TS), no OpenAPI spec (404s found), and no independent benchmarking of SDK stability/versioning.

      • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
      • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
      • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
      • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
      • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
      • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
      Riveterfullclaimed7/10

      Riveter ships an official Go SDK (riveterhq/riveter-go) with documented client code (riveter.EnrichParams), and docs describe SDK-level handling of auth, retries, long-polling, and pagination, indicating a first-party SDK layer built for AI-native workflows. Missing for 10: confirmation of additional language SDKs (e.g., Python/JS) beyond Go, and independent/hands-on corroboration of SDK reliability.

      • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
      • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
    • ai-native userSubscribe to events via webhooks

      weight 2 · round to Riveter
      Crawl4AInone0/10

      No evidence of webhook subscription or event-push capability anywhere in the docs, GitHub features, or community discussion; the product is a crawling library/service with Docker/FastAPI/MCP interfaces but nothing about webhooks.

        Riveterpartialclaimed6/10

        Riveter supports webhooks by passing a webhook_url when starting a run, with Riveter POSTing results back on run.completed, run.stopped, and run.finished events — a real event-notification mechanism for agentic workflows. However this is scoped to a single run's lifecycle rather than a general subscription model (no persistent webhook registration/management endpoint, no broader event catalog, no signature/security details). Missing for 10: a dedicated webhook subscription/management API, documentation of additional event types beyond run lifecycle, and payload signing/verification details.

        • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…

      Agentic features

      1. ai-native userGet AI-generated insights and suggestions from my data inside the product

        weight 2 · round to Riveter
        Crawl4AIpartialclaimed4/10

        Crawl4AI offers LLM-driven structured extraction and adaptive crawling that determines when 'sufficient information' has been gathered, which could generate insight-like structured data from crawled content, but there is no evidence of a dashboard or interface that generates proactive 'insights and suggestions' about a user's own data corpus in the way the story implies. missing for 10: evidence of an insights/suggestions UI or report generation feature, evidence of proactive recommendations rather than raw extraction, independent confirmation of this use case.

        • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
        • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
        • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
        Riveterfullclaimed7/10

        Riveter's core enrichment feature fills columns using AI agents, web search/scrape, and other tools to generate insights directly on user data, and search_agent provides ad hoc AI-researched answers within the product. missing for 10: independent/hands-on corroboration of insight quality, no example of proactive/unprompted suggestions (only prompt-driven enrichment), and no dashboard-level 'insights' UI evidence beyond API/SDK docs.

        • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
        • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
        • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
        • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.
        • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…
      2. ai-native userSet up automations that run autonomously in the background

        weight 2 · round to Riveter
        Crawl4AIpartialcommunity4/10

        Crawl4AI provides Docker/FastAPI deployment, resume checkpoints, and community mentions of bridging to automation tools like n8n and MCP servers, suggesting it can be embedded into autonomous background pipelines, but there is no first-party evidence of a native scheduler, trigger system, or persistent autonomous agent loop within Crawl4AI itself. missing for 10: native scheduling/trigger mechanism, documented autonomous background-run feature, first-party (non-community) evidence of persistent unattended operation, integration guide owned by Crawl4AI rather than third-party community sites.

        • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
        • [github] resume_state parameter to continue from a saved checkpoint
        • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
        • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
        • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
        Riveterpartialclaimed6/10

        Riveter supports scheduling projects to run on a cadence ('every minute' for fast-moving data) and webhook notifications on run completion, which enables autonomous background execution without manual triggering. However, there's no evidence of broader automation orchestration (e.g., conditional triggers, chaining multiple actions, or a dedicated automation/workflow builder) beyond scheduled data refresh. Missing for 10: evidence of multi-step autonomous workflows beyond scheduled enrichment refresh, independent/hands-on confirmation that scheduling works reliably in production, and any automation trigger types beyond time-based schedules.

        • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
        • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.
        • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
      3. ai-native userDelegate tasks to a built-in AI assistant inside the product

        weight 3 · round to Riveter
        Crawl4AInone0/10

        The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

          Riveterpartialclaimed6/10

          Riveter ships an internal 'agent loop' (search_agent, enrichment AI) that autonomously researches, scrapes, and fills data on request, which functions as a built-in AI assistant for delegated research tasks rather than a conversational general-purpose assistant. Missing for 10: evidence of a general chat/task interface for arbitrary delegation, independent hands-on validation, and clarity on how broadly the agent can handle tasks beyond enrichment/search/scrape.

          • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
          • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
          • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.
          • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…
        • ai-native userOperate the product with natural-language commands

          weight 2 · round to Riveter
          Crawl4AIpartialclaimed5/10

          Crawl4AI supports LLM-based extraction where users can specify extraction goals in natural language, and its adaptive crawling engine stops based on a natural-language 'query' describing what information is needed. However, the core interface (CLI, Python API) is still command/flag-based, not a general natural-language command layer for controlling the crawler itself. missing for 10: evidence of a chat-style or NL command interface for the tool's core operations, independent confirmation of how well NL-driven extraction/query works in practice.

          • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
          • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
          • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
          • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
          Riveterfullclaimed7/10

          Riveter explicitly supports building enrichments from natural-language prompts (riveter-docs-2), offers a search_agent that answers questions in natural language without setup (riveter-docs-5), and can be operated via MCP-compatible AI assistants like Claude, ChatGPT, and Cursor (riveter-docs-9), which is the core mechanism for natural-language control. Missing for 10: independent/hands-on confirmation of NL command reliability, and no evidence of a broader NL command surface beyond enrichment/search (e.g., NL-driven scheduling or config changes).

          • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
          • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
          • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
          • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
        • ai-native userApply a preset configuration tuned for research agents that returns structured, citable output

          weight 2 · round to Crawl4AI
          Crawl4AIpartialclaimed5/10

          Crawl4AI supports markdown/structured output and LLM-based structured extraction, and its 'adaptive crawling' feature explicitly determines when 'sufficient information has been gathered to answer your query,' which aligns with a research-agent workflow. However, there is no evidence of an actual named preset/config specifically tuned for research agents nor of output formatted with citations/sources for verifiability. Missing for 10: a documented 'research agent' preset profile, explicit citation/source-tracking in output, and independent confirmation that adaptive crawling output is citable.

          • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
          • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
          • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
          • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
          Riveternone0/10

          Riveter offers enrichment, search_agent, and scrape tools with structured outputs, but there is no evidence of a preset/template configuration specifically tuned for research agents or citable output formatting; missing for 10: a named preset or template targeting research-agent workflows, citation/source-attribution formatting in outputs, and any documentation referencing 'research agent' presets.

          Api quality

          1. ai-native userExplore an interactive API reference with runnable examples

            weight 2 · round drawn
            Crawl4AInone0/10

            Docs show static code snippets (e.g., crawl4ai-docs-1) but there is no evidence of an interactive API reference (like Swagger/OpenAPI UI) or runnable in-browser examples; probes explicitly confirm openapi.json/swagger.json and llms.txt endpoints return 404, indicating no such interactive reference exists.

            • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
            • [probe] PROBE llms.txt: HTTP 404 at https://docs.crawl4ai.com/llms.txt
            • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
            Riveternone0/10

            No evidence of an interactive API reference or runnable examples; probes for llms.txt and OpenAPI/Swagger specs both returned 404s, and docs snippets are static text/code examples only, not interactive/runnable.

            • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
            • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
          2. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

            weight 2 · round drawn
            Crawl4AInone0/10

            Crawl4AI ships a Dockerized FastAPI server (crawl4ai-gh-5), so a machine-readable OpenAPI spec would be a plausible artifact, but direct probes for openapi.json/swagger.json/llms.txt all returned 404 with no alternative spec location documented.

            • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
            • [probe] PROBE llms.txt: HTTP 404 at https://docs.crawl4ai.com/llms.txt
            • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
            Riveternone0/10

            Probes for llms.txt and OpenAPI/swagger spec files all returned 404s, and no documentation mentions a downloadable machine-readable API spec despite having a REST API and SDKs.

            • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
            • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
          3. ai-native userTest against a sandbox environment without touching production data

            weight 1 · round to Riveter
            Crawl4AInone0/10

            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

              Riveterpartialclaimed4/10

              Riveter offers a dry_run mode that validates a request and returns a credit estimate without creating or charging anything, and a max_credits cap that blocks runs before they execute — both function like a lightweight 'test without side effects' capability. However, there's no explicit documentation of a separate sandbox environment or synthetic/test dataset distinct from production data sources (Riveter always operates against live web/data sources when actually run). Missing for 10: a documented sandbox/staging environment, sample or mock datasets, and explicit guidance on testing enrichments without touching real production data sources.

              • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
              • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
            • ai-native userRely on versioned APIs with a documented deprecation policy

              weight 2 · round drawn
              Crawl4AInone0/10

              No evidence of versioned APIs or a documented deprecation policy; probes show no OpenAPI spec, no llms.txt, and no mention of versioning/deprecation practices anywhere in docs or community discussion.

              • [probe] PROBE llms.txt: HTTP 404 at https://docs.crawl4ai.com/llms.txt
              • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
              Riveternone0/10

              No evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI/spec discovery returned 404s, and docs mention SDKs/features but nothing about version numbers or deprecation guarantees. Missing for 10: versioned endpoint scheme (e.g., /v1/), a published deprecation/sunset policy, changelog or migration guides.

              • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
              • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
            • data-engineerThe documented rate limit (requests per second or minute) enforced on my API key before throttling kicks in

              weight 3 · round drawn
              Crawl4AInone0/10

              The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                Riveternone0/10

                There is a mention of SDKs handling retries on 429s, implying rate limiting exists, but no documented numeric rate limit (requests per second/minute) is given anywhere in the evidence pack, and probes for API spec/docs return 404s.

                • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
                • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
                • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…

              Anti bot — getting past bot defenses — CAPTCHAs, fingerprinting, blocksAnti bot

              Getting past bot defenses — CAPTCHAs, fingerprinting, blocks

              Block evasion

              1. ai-native userHave an agent automatically get past a CAPTCHA, login, or form wall without my manual intervention

                weight 2 · round to Crawl4AI
                Crawl4AIpartialcommunity4/10

                Crawl4AI offers persistent browser profiles with saved authentication/cookies and 'undetected browser' support to evade bot detection, plus proxy/retry chains, which partially help with login walls and basic anti-bot evasion. However, there is no evidence of automatic CAPTCHA-solving, and community feedback explicitly calls out login/session handling and bot mitigation as things the user must configure and own themselves rather than fully automatic agent behavior. missing for 10: CAPTCHA-solving capability, evidence of fully hands-off login/session bootstrap, independent confirmation that undetected-browser mode reliably bypasses modern bot walls without manual setup.

                • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                • [github] Automatic retry with proxy chain and fallback fetch function
                • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                Riveternone0/10

                The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                • data-engineerAutomatically retry through a chain of different proxies when anti-bot detection blocks a request

                  weight 2 · round to Crawl4AI
                  Crawl4AIpartialclaimed6/10

                  GitHub feature list explicitly documents 'Automatic retry with proxy chain and fallback fetch function' plus undetected browser support for bot detection bypass, directly matching the story. However, this is a single line-item mention with no detailed docs, configuration examples, or independent/hands-on validation showing it working against real anti-bot systems. Missing for 10: dedicated documentation/tutorial on configuring proxy chains, code examples showing retry-on-block logic, and independent confirmation it succeeds against modern anti-bot defenses.

                  • [github] Automatic retry with proxy chain and fallback fetch function
                  • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                  Riveternone0/10

                  No evidence of proxy rotation, IP chaining, or anti-bot-specific retry logic; only generic SDK retries for 429s/transient failures are mentioned, which is unrelated to proxy chaining against anti-bot blocks.

                  • developerUse an undetected browser mode to bypass sophisticated bot detection systems

                    weight 3 · round to Crawl4AI
                    Crawl4AIpartialcommunity6/10

                    GitHub feature list explicitly claims 'Undetected Browser Support: Bypass sophisticated bot detection systems,' directly matching the story, and this is corroborated by related anti-detection features like persistent browser profiles and proxy chain retries. However, there is no independent/hands-on evidence confirming its effectiveness, and community commentary notes bot mitigation is still something users must handle themselves ('own the policy layer', 'boring production bits: ... bot mitigation'), suggesting real-world limitations. Missing for 10: independent verification of undetected-mode effectiveness, technical documentation on how it works, and resolution of community caveats about needing to handle bot mitigation manually.

                    • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                    • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                    • [github] Automatic retry with proxy chain and fallback fetch function
                    • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                    Riveternone0/10

                    No evidence mentions undetected browser mode, bot-detection bypass, proxies, or stealth automation features; Riveter's evidence only covers enrichment, scraping, and search tooling. Missing for 10: any mention of anti-bot/stealth browser capabilities, CAPTCHA handling, or evasion of bot detection.

                    Proxy rotation

                    1. developerRequest a proxy from a specific country to get geolocation-appropriate content

                      weight 2 · round drawn
                      Crawl4AInone0/10

                      Evidence mentions proxy chain retry/fallback for reliability but nothing about selecting or requesting a proxy from a specific country/geolocation. Missing for 10: documentation of country-specific proxy selection, geolocation targeting API/config, and any example of requesting geo-located content.

                      • [github] Automatic retry with proxy chain and fallback fetch function
                      Riveternone0/10

                      The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                      • developerUse premium residential or datacenter proxies to bypass sites that are hard to scrape

                        weight 3 · round to Crawl4AI
                        Crawl4AIpartialclaimed4/10

                        Crawl4AI supports proxy chains with automatic retry/fallback and undetected browser mode to bypass bot detection, but there is no evidence of built-in support for premium residential/datacenter proxy providers or proxy rotation services—users must bring and configure their own proxies. missing for 10: no documented integration with residential/datacenter proxy providers, no proxy rotation/pool management features, no independent evidence of successful bypass on hard-to-scrape sites using proxies.

                        • [github] Automatic retry with proxy chain and fallback fetch function
                        • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                        Riveternone0/10

                        No evidence mentions proxy support (residential or datacenter) or IP rotation for anti-bot bypass; the docs describe scraping and AI agent interpretation but never address proxy infrastructure.

                        • developerRoute requests through a rotating pool of proxy IPs to avoid blocks

                          weight 3 · round to Crawl4AI
                          Crawl4AIpartialclaimed5/10

                          There is evidence of proxy chain retry/fallback logic (automatic retry with proxy chain and fallback fetch function) and undetected browser support for bot detection bypass, indicating some proxy-rotation and anti-bot capability exists. However, no documentation details how to configure a pool of rotating proxy IPs, proxy list management, or rotation strategy specifics. missing for 10: explicit proxy pool configuration docs, rotation strategy details, independent confirmation of proxy rotation working in practice.

                          • [github] Automatic retry with proxy chain and fallback fetch function
                          • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                          Riveternone0/10

                          The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                          • developerRoute multiple requests through the same proxy IP using a session identifier to maintain a consistent identity

                            weight 2 · round drawn
                            Crawl4AInone0/10

                            Evidence mentions proxy chains for retry/fallback and undetected browser support, but there is no mention of a session identifier mechanism to route multiple requests through the same proxy IP for persistent identity. Missing for 10: sticky-session/proxy-session-ID feature documentation, any example binding a session to a specific proxy IP, and independent confirmation of this capability.

                            • [github] Automatic retry with proxy chain and fallback fetch function
                            • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                            Riveternone0/10

                            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                            Automation depth — how much of the product can run unattendedAutomation depth

                            How much of the product can run unattended

                            1. ai-native userPerform bulk operations across many items at once

                              weight 2 · round drawn
                              Crawl4AIfullclaimed8/10

                              Crawl4AI supports batch/bulk crawling via deep-crawl BFS with max-pages, multi-URL configuration with per-pattern strategies, checkpoint/resume for large jobs, and dockerized/API deployment for scaling bulk crawls. Missing for 10: independent benchmarks of large-scale bulk runs and clearer documentation of concurrency/throughput limits at scale.

                              • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                              • [github] Multi-URL Configuration: Different strategies for different URL patterns in one batch
                              • [github] resume_state parameter to continue from a saved checkpoint
                              • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                              • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                              Riveterfullclaimed8/10

                              Riveter's core enrichment model operates on many rows at once (bulk input data with AI-filled columns), supports batch generation from a prompt/spec, scheduling for ongoing refresh, and examples like pulling every dentist from every practice in a city in one request. Missing for 10: independent/hands-on verification of large-scale bulk runs and no explicit documentation of per-run item limits or throughput benchmarks.

                              • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                              • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                              • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                              • [claimed-docs] It can find every dental practice in a city, then pull every dentist from each one, in a single request.
                              • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.
                              • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…
                            2. ai-native userDefine rules that trigger actions automatically on events

                              weight 3 · round to Riveter
                              Crawl4AInone0/10

                              Crawl4AI is a crawling/extraction library with adaptive crawling, retries, and checkpointing, but there is no evidence of a rules/trigger engine that lets users define conditional event-based automations (e.g., 'if X happens, do Y'). Community notes even highlight that users must build their own automation/policy layer via external tools like n8n rather than Crawl4AI natively supporting this.

                              • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                              • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                              • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                              Riveterpartialclaimed4/10

                              Riveter supports scheduled refresh of projects (time-based automation) and webhook events (run.completed/stopped/finished) that can notify external systems, giving some automation-on-events capability, but there is no evidence of a rules/condition engine that lets users define arbitrary triggers (e.g., 'if data matches X, then do Y') beyond scheduling and run-completion notifications. missing for 10: conditional rule definitions, event-driven branching logic, multi-condition triggers, and any UI/API for building custom automations beyond schedule+webhook.

                              • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                              • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                              • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.
                            3. ai-native userSchedule recurring jobs or workflows

                              weight 2 · round to Riveter
                              Crawl4AInone0/10

                              Crawl4AI provides crawling, extraction, checkpointing, and Docker/API deployment, but no evidence of built-in scheduling or recurring job/workflow orchestration; community notes mention bridging to external tools like n8n for automation, implying no native scheduler exists.

                              • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                              • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                              • [github] resume_state parameter to continue from a saved checkpoint
                              Riveterpartialclaimed6/10

                              Docs state you can 'schedule any project to monitor for changes and keep your data fresh' and refresh as often as every minute, indicating recurring job/workflow scheduling support. However, details are thin — no documentation on schedule configuration (cron-like syntax, timezone, pause/resume), no UI/API endpoint specifics for managing schedules, and no independent or hands-on corroboration. Missing for 10: scheduling API/UI details, configuration options, independent verification of reliability at scale.

                              • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                              • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.

                            Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience

                            Day-to-day developer experience — setup friction, docs, debugging, iteration speed

                            Collaboration

                            1. developerShare scrapers with teammates and manage organizations and role-based permissions

                              weight 2 · round drawn
                              Crawl4AInone0/10

                              The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                Riveternone0/10

                                No evidence pack items mention team sharing, organizations, workspaces, or role-based access control for Riveter; all evidence covers scraping/enrichment functionality and API mechanics only. Missing for 10: any mention of teams, org management, invites, or RBAC/permissions.

                                Deployment flexibility

                                1. developerBuild and deploy custom serverless scraping scripts on the platform without managing my own infrastructure

                                  weight 2 · round to Riveter
                                  Crawl4AInone0/10

                                  Crawl4AI is an open-source library/framework requiring self-hosting via Docker or local Python install; there is no evidence of a managed serverless platform for deploying custom scraping scripts without infrastructure management. Evidence instead shows users must set up Docker containers, browser pools, and monitoring dashboards themselves.

                                  • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                  • [github] Real-time Monitoring Dashboard with live system metrics and browser pool visibility
                                  • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                  • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                                  Riveterpartialclaimed6/10

                                  Riveter's docs show fully managed, serverless-style capabilities (enrichments, scrapes, quick_search, search_agent) that developers configure via natural-language prompts or structured specs and trigger via API/SDK/webhooks with no server management (riveter-docs-1,2,3,4,5,6,11,12). However, this is closer to configuring built-in AI-driven tools than deploying arbitrary custom scraping code/scripts — there's no evidence of a code-upload or custom-script execution environment. Missing for 10: evidence of arbitrary custom code/script deployment (vs. prompt/spec-based enrichment configuration), and independent confirmation of the serverless execution model.

                                  • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                  • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                                  • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                  • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                  • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
                                  • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                  • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
                                  • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
                                2. developerDeploy the scraping service via a Docker container for production use

                                  weight 2 · round to Crawl4AI
                                  Crawl4AIpartialprobed7/10

                                  GitHub docs explicitly advertise a 'Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment' and community mentions of one-click Docker setups for production use corroborate this. However, there's no independent hands-on production deployment report, no details on scaling/orchestration guidance, and no OpenAPI spec confirmed (probe found 404s), leaving some production-readiness details unverified. Missing for 10: independent hands-on verification of the Docker deployment in production, confirmed API schema/OpenAPI docs, and details on scaling/orchestration best practices.

                                  • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                  • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                  • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
                                  Riveternone0/10

                                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                  • developerSelf-host an open-source version of the scraper instead of relying on a hosted cloud service

                                    weight 2 · round to Crawl4AI
                                    Crawl4AIfullprobed9/10

                                    Crawl4AI is explicitly open source with no forced API keys/paywalls, distributed via GitHub, and supports Dockerized self-hosting with a FastAPI server, plus community-documented self-hosting guides (Docker, n8n, MCP for Cursor/Claude) corroborating real-world self-hosted deployments. missing for 10: independent benchmark/uptime evidence of large-scale self-hosted production use.

                                    • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                    • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                    • [probe] official MCP server documented at https://docs.crawl4ai.com/core/self-hosting/#mcp-model-context-protocol-support
                                    Riveternone0/10

                                    Riveter is presented as a hosted API/service (with a local MCP connector for client access to the remote service), but there is no evidence of an open-source, self-hostable version of the scraper itself; docs only describe running a local MCP bridge that still relies on the remote API key.

                                    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
                                    • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.

                                  Integrations

                                  1. developerConnect the scraping API to no-code automation platforms like n8n or Zapier through a prebuilt connector

                                    weight 2 · round drawn
                                    Crawl4AInone0/10

                                    No evidence of a prebuilt n8n/Zapier connector; the only related evidence is community commentary noting developers struggle to bridge Crawl4AI with n8n and a third-party community doc hub with Docker setup guides, not an official connector from Crawl4AI itself.

                                    • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                                    • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                    Riveternone0/10

                                    No evidence of a prebuilt n8n or Zapier connector; only webhook support and MCP/AI-assistant integrations are documented, which are not no-code automation platform connectors.

                                    Library compatibility

                                    1. developerBuild scrapers using popular open-source automation libraries like Playwright, Puppeteer, Selenium, or Scrapy

                                      weight 2 · round drawn
                                      Crawl4AInone0/10

                                      Crawl4AI ships its own AsyncWebCrawler API (built on Playwright internally) rather than exposing compatibility layers for Playwright, Puppeteer, Selenium, or Scrapy code; none of the evidence mentions using these other libraries to build scrapers within Crawl4AI.

                                      • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                      • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                      Riveternone0/10

                                      The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                      Migration lock in

                                      1. developerExport my scraped data and job configurations in a portable format to migrate to another provider without lock-in

                                        weight 3 · round to Crawl4AI
                                        Crawl4AIpartialclaimed4/10

                                        Crawl4AI outputs scraped data in portable formats like Markdown and structured JSON/CSS-XPath extraction, and being open-source with no forced API keys supports a no-lock-in narrative, but there is no documented feature for exporting or migrating job configurations, crawl profiles, or schemas to another provider. missing for 10: explicit config/job export or import tooling, documented migration path to another scraping provider, independent confirmation of format portability.

                                        • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                        • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                        • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                        Riveternone0/10

                                        No evidence of any export/portability feature for scraped data or job configs, nor any mention of migration support or standard/open config formats; API docs focus on running enrichments, webhooks, and SDKs, not data portability.

                                        Quickstart

                                        1. developerPublish my custom scraper to a public marketplace and earn revenue when others use it

                                          weight 1 · round drawn
                                          Crawl4AInone0/10

                                          The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                            Riveternone0/10

                                            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                            • developerRun a ready-made scraper from a marketplace instead of building one from scratch

                                              weight 2 · round drawn
                                              Crawl4AInone0/10

                                              No evidence of a marketplace or gallery of ready-made scrapers/templates that a developer could run out of the box; the evidence only covers building crawlers via code/CLI configuration yourself, not a pre-built scraper marketplace.

                                                Riveternone0/10

                                                The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                • developerStart building immediately using a library of ready-made project templates

                                                  weight 1 · round drawn
                                                  Crawl4AInone0/10

                                                  The evidence shows basic usage snippets, CLI/Docker deployment instructions, and a third-party community docs hub with one-click setups, but no official library of ready-made project templates or starter kits is documented by the vendor.

                                                  • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                  • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                  • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                                  Riveternone0/10

                                                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                  Extraction quality — how faithfully content is extracted — structure, fidelity, edge casesExtraction quality

                                                  How faithfully content is extracted — structure, fidelity, edge cases

                                                  Ai extraction

                                                  1. developerExtract structured data from a page using natural language instructions instead of writing selectors

                                                    weight 3 · round to Riveter
                                                    Crawl4AIpartialclaimed6/10

                                                    Crawl4AI documents LLM-based extraction as an alternative to CSS/XPath selectors, letting developers describe desired structured data rather than write selectors, and this is corroborated by GitHub feature docs (LLM-Driven Extraction, LLMTableExtraction). However, the evidence doesn't show natural-language instruction schemas in detail (e.g., prompt examples), nor independent hands-on validation of extraction quality/accuracy. missing for 10: concrete example of natural-language extraction prompt/schema, independent quality benchmarks or hands-on confirmation of NL-instruction extraction accuracy.

                                                    • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                    • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                    • [github] LLMTableExtraction: Revolutionary table extraction with intelligent chunking for massive tables
                                                    Riveterfullclaimed7/10

                                                    Riveter's core enrichment feature lets developers build extraction jobs from a natural-language prompt with target attributes instead of writing selectors, and AI agents interpret pages semantically so configs survive redesigns, directly matching the story. missing for 10: independent/hands-on verification of extraction accuracy and no live API schema (openapi/llms.txt probes 404) to confirm behavior beyond vendor docs.

                                                    • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                                                    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
                                                    • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.
                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                  2. developerPass a JSON schema so the API returns structured data matching that schema

                                                    weight 2 · round to Crawl4AI
                                                    Crawl4AIpartialclaimed5/10

                                                    Docs mention structured extraction via CSS/XPath/LLM-based extraction and LLM-driven extraction supporting schema-like structured output, implying JSON-schema-guided extraction, but no evidence pack item explicitly shows passing a JSON schema and receiving matching structured JSON output. missing for 10: explicit documented JSON schema parameter/example, sample output matching schema, independent verification of schema conformance.

                                                    • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                    • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                    Riveterpartialclaimed4/10

                                                    Riveter lets you define enrichments via a natural-language prompt or a 'structured spec' with named attributes/columns (riveter-docs-2, riveter-docs-12), which produces structured output, but there is no documented mechanism for passing an arbitrary JSON Schema that the API validates/returns against. missing for 10: explicit JSON Schema input parameter, schema validation of output, and any example showing schema-conformant responses.

                                                    • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                                                    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                  3. ai-native userHave an LLM read a page and decide what structured fields to pull out without pre-written selectors

                                                    weight 2 · round to Riveter
                                                    Crawl4AIpartialclaimed6/10

                                                    Docs and GitHub confirm LLM-based structured extraction supporting arbitrary LLMs (open-source and proprietary), which enables schema-free, LLM-decided field extraction rather than fixed CSS/XPath selectors. However, evidence is thin on how the LLM decides fields (e.g., whether a schema/prompt is still required or if it's fully autonomous field discovery), and there's no hands-on example or independent validation of the LLM extraction path's accuracy. missing for 10: concrete example/walkthrough of LLM freely deciding fields without any schema, independent quality benchmarks on this specific extraction mode.

                                                    • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                    • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                    Riveterfullclaimed8/10

                                                    Docs describe enrichments where AI agents interpret pages and fill arbitrary attribute columns from a natural-language prompt or structured spec (no selectors), with scraping/search tools feeding an AI agent loop that adapts to page structure and redesigns. This directly matches the story of an LLM reading a page and deciding what fields to extract without pre-written selectors. Missing for 10: independent hands-on verification of extraction accuracy and no example showing the LLM's field-selection reasoning in practice.

                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                    • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                                                    • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                    • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.
                                                    • [claimed-docs] It can find every dental practice in a city, then pull every dentist from each one, in a single request.
                                                    • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},
                                                  4. developerPlug in a local or self-hosted LLM as the extraction backend instead of a cloud-only model

                                                    weight 2 · round to Crawl4AI
                                                    Crawl4AIfullclaimed8/10

                                                    Docs and GitHub explicitly state LLM-based extraction supports all LLMs, both open-source and proprietary, and the project is fully open source with no forced API keys, implying local/self-hosted LLM backends can be plugged in for extraction. Missing for 10: explicit step-by-step docs/config example showing pointing extraction at a local model (e.g., Ollama endpoint) and independent hands-on confirmation of this specific workflow.

                                                    • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                    • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                    • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                    Riveternone0/10

                                                    No evidence anywhere in the docs suggests Riveter allows swapping in a local or self-hosted LLM as the extraction engine; the product is presented as a cloud-only enrichment/extraction service with API keys, credits, and hosted agents. Missing for 10: any mention of local model support, self-hosted backend configuration, or BYO-model options.

                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
                                                    • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.

                                                  Basic scraping

                                                  1. developerScrape a web page with a single API call and get its raw HTML back

                                                    weight 3 · round to Crawl4AI
                                                    Crawl4AIpartialprobed6/10

                                                    The docs show a single async call (crawler.arun(url=...)) returning a result object, and result.html/cleaned_html is a documented attribute of Crawl4AI's result, though the sample shown emphasizes result.markdown rather than raw HTML explicitly. This confirms single-call scraping works, but the evidence pack doesn't explicitly show raw HTML retrieval or an OpenAPI-documented single-endpoint HTTP API (openapi probes 404). missing for 10: explicit example of raw HTML field usage, independent confirmation of HTML fidelity/extraction quality.

                                                    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                    • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
                                                    Riveternone0/10

                                                    Riveter's scrape endpoint explicitly returns 'easily parseable text' from a URL, not raw HTML — the opposite of what this story asks for, and no evidence shows an option to retrieve unprocessed HTML.

                                                    • [claimed-docs] A scrape lets you turn a URL into easily parseable text.

                                                  Data safety

                                                  1. data-engineerAutomatically detect and filter personally identifiable information out of scraped content before it reaches storage

                                                    weight 2 · round drawn
                                                    Crawl4AInone0/10

                                                    No evidence anywhere in the pack of built-in PII detection or filtering; extraction features focus on structured/LLM-based data extraction, not privacy compliance. Community commentary explicitly flags PII handling as something the user must own ('not accidentally hoovering up PII' as a 'boring production bit'), reinforcing that this is not a shipped capability.

                                                    • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                    Riveternone0/10

                                                    No evidence anywhere in the pack mentions PII detection, filtering, redaction, or compliance controls for scraped/enriched data; Riveter's documented features cover scraping, enrichment, search, and workflow orchestration but nothing about identifying or removing personal data before storage.

                                                    Document extraction

                                                    1. data-engineerExtract text content from PDFs, Word, Excel, and PowerPoint files without hosting them myself

                                                      weight 2 · round to Riveter
                                                      Crawl4AInone0/10

                                                      No evidence in the pack mentions extraction of PDF, Word, Excel, or PowerPoint file content; all documented capabilities relate to web page crawling, structured/LLM extraction from HTML, and table extraction, not office document formats.

                                                        Riveterpartialclaimed4/10

                                                        Riveter is delivered as a hosted API/SaaS (no self-hosting required) and docs state it 'reads PDFs and images' as part of enrichment workflows, but there is no evidence it extracts text from Word, Excel, or PowerPoint files specifically. missing for 10: explicit support for .docx/.xlsx/.pptx extraction, any extraction-quality benchmarks or examples for Office file formats.

                                                        • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…

                                                      Multimodal extraction

                                                      1. ai-native userGet automatic captions for images on a page so a text-only model can reason about visual content

                                                        weight 2 · round to Riveter
                                                        Crawl4AInone0/10

                                                        No evidence pack item mentions image captioning or alt-text generation for images; the extraction features described (LLM-based structured extraction, table extraction) are unrelated to describing visual content for a text-only model. Missing for 10: any mention of image-to-text captioning, vision-model integration, or alt-text generation feature.

                                                          Riveterpartialclaimed3/10

                                                          Riveter's docs mention it 'reads PDFs and images' and combines results with web data (riveter-docs-18), implying some visual-content ingestion, but there is no explicit description of generating captions or text descriptions of images for downstream reasoning by a text-only model. Missing for 10: explicit captioning/description output format, example enrichment showing image-to-text extraction, and any confirmation this text is usable standalone by a text-only model.

                                                          • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…

                                                        Search integration

                                                        1. developerSearch the web and get full page content from results in a single call instead of just links and snippets

                                                          weight 3 · round drawn
                                                          Crawl4AInone0/10

                                                          Crawl4AI's evidence describes crawling/scraping given URLs, deep-crawl (BFS) from a seed URL, and structured/LLM extraction, but no evidence of a web-search capability that returns full content for search-engine results in one call. Since comparable scraping tools do offer this, the axis applies but no supporting evidence exists here.

                                                          • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                          • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                          • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                                                          Riveternone0/10

                                                          Riveter's quick_search explicitly returns only urls, titles, and snippets (not full page content), and its scrape tool requires a specific URL rather than combining search+content in one call. search_agent returns a single synthesized answer, not full page content per search result, so no evidenced single-call capability matches the story's exact requirement.

                                                          • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                          • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                                          • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…

                                                        Selector extraction

                                                        1. developerExtract specific fields from a page using CSS or XPath selector rules

                                                          weight 3 · round to Crawl4AI
                                                          Crawl4AIpartialclaimed6/10

                                                          Docs explicitly mention structured extraction supporting CSS and XPath selectors alongside LLM-based extraction, confirming the capability exists. However, evidence lacks concrete code examples, schema syntax details, or independent hands-on confirmation of CSS/XPath extraction specifically (most community and GitHub evidence focuses on LLM extraction, crawling, and deployment features instead). Missing for 10: detailed CSS/XPath schema examples, independent verification of selector-based extraction working in practice, documentation depth comparable to LLM extraction features.

                                                          • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                          Riveternone0/10

                                                          Riveter's docs describe AI-driven page interpretation and scraping (turning URLs into parseable text, agents reading pages 'the way a person would') rather than CSS/XPath selector rules; no evidence pack item mentions selector-based extraction at all, and one item explicitly frames the AI approach as an alternative to fragile configuration that would break on redesign, which is the kind of setup selectors typically require.

                                                          • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                          • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.

                                                        Structured data handling

                                                        1. data-engineerExtract data from very large tables using intelligent chunking so it fits within processing limits

                                                          weight 1 · round to Crawl4AI
                                                          Crawl4AIfullclaimed7/10

                                                          GitHub docs explicitly cite 'LLMTableExtraction: Revolutionary table extraction with intelligent chunking for massive tables,' directly matching the story of extracting data from very large tables via chunking. However, missing for 10: independent hands-on validation of chunking behavior on real large tables, and detailed documentation on configuring chunk size/limits or performance benchmarks.

                                                          • [github] LLMTableExtraction: Revolutionary table extraction with intelligent chunking for massive tables
                                                          Riveternone0/10

                                                          Riveter's evidence covers enrichment, scraping, search, and workflow automation, but there is no mention of chunking large tables, row batching, pagination for extraction limits, or handling of very large datasets to fit processing constraints. missing for 10: any mention of chunking strategy, table size limits, batching large extractions, or row-splitting logic.

                                                          Js rendering — handling JavaScript-heavy pages — rendering, waiting, dynamic contentJs rendering

                                                          Handling JavaScript-heavy pages — rendering, waiting, dynamic content

                                                          Headless rendering

                                                          1. developerRender JavaScript-heavy single-page applications and get the fully rendered HTML

                                                            weight 3 · round to Crawl4AI
                                                            Crawl4AIpartialclaimed5/10

                                                            Crawl4AI is built on a real browser (AsyncWebCrawler with undetected browser support, browser profiles, etc.), which implies it can render JS-heavy SPAs and return rendered HTML/markdown, but the evidence pack never explicitly documents JS execution/wait-for-selector behavior or confirms fully-rendered HTML output for SPAs. Missing for 10: explicit documentation of JS rendering/execution settings (e.g., wait_for, js_code, page load strategies), and independent/hands-on confirmation that dynamic SPA content is captured correctly.

                                                            • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                            • [github] Undetected Browser Support: Bypass sophisticated bot detection systems
                                                            • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                                                            • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                            Riveternone0/10

                                                            Riveter is a data-enrichment/scraping/AI-agent tool focused on turning URLs into text and filling data columns; there is no evidence it renders JS-heavy SPAs into fully rendered HTML (e.g., headless browser rendering, DOM snapshot output). The 'scrape' feature converts URLs to 'easily parseable text', not full rendered HTML, so this capability is unevidenced.

                                                            • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                          2. developerHave the API wait for a specific selector to appear before returning the rendered page

                                                            weight 2 · round drawn
                                                            Crawl4AInone0/10

                                                            No evidence in the pack mentions a wait_for/selector-based config option or any mechanism to delay page return until a specific CSS/XPath selector appears; the docs snippets shown only cover basic arun usage, extraction, and CLI/MCP features. missing for 10: documentation or example of a wait_for_selector or similar parameter, confirmation it blocks return until element renders, any community/hands-on validation of this feature.

                                                            • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                            • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                            Riveternone0/10

                                                            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                            Interactive automation

                                                            1. developerAccess a managed remote browser sandbox for interactive, manual browsing workflows

                                                              weight 2 · round drawn
                                                              Crawl4AInone0/10

                                                              The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                Riveternone0/10

                                                                The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                • developerKeep interacting with an already-scraped page, clicking and filling forms to reach content behind a login wall

                                                                  weight 2 · round to Crawl4AI
                                                                  Crawl4AIpartialcommunity4/10

                                                                  The Browser Profiler feature (crawl4ai-gh-3) supports persistent authentication states and cookies, which can help reach content behind a login wall, but there is no direct evidence of interactive session APIs for clicking or filling forms mid-crawl. Community commentary (crawl4ai-comm-3) even flags login/session handling as one of the 'boring production bits' users must handle themselves, suggesting it's not a polished, first-class capability. missing for 10: explicit documentation of click/fill/form-interaction APIs, session-persistence across multiple interactive steps, and independent confirmation of successful login-wall traversal.

                                                                  • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                                                                  • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                  Riveternone0/10

                                                                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                  • developerScript page interactions like clicking, filling inputs, and scrolling before content is returned

                                                                    weight 3 · round drawn
                                                                    Crawl4AInone0/10

                                                                    The evidence pack describes many Crawl4AI features (extraction, deep-crawl, browser profiles, proxy retry, docker/MCP/CLI) but never mentions scripting page interactions such as clicking, filling inputs, or scrolling before extraction. This is a fair capability to expect from a browser-based crawler, but no evidence in the pack documents it.

                                                                      Riveternone0/10

                                                                      The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                      Render configuration

                                                                      1. developerControl the browser viewport width and height when rendering a page

                                                                        weight 1 · round drawn
                                                                        Crawl4AInone0/10

                                                                        Crawl4AI is a browser-based crawler and viewport control (headless=True/False, viewport_width/height config) is a real feature category, but no evidence pack item mentions viewport, browser window size, or rendering dimensions configuration.

                                                                          Riveternone0/10

                                                                          The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                          Session persistence

                                                                          1. developerPass my own session cookies so the API fetches pages requiring authentication

                                                                            weight 2 · round to Crawl4AI
                                                                            Crawl4AIpartialclaimed6/10

                                                                            GitHub docs mention a Browser Profiler that creates and manages persistent profiles with saved authentication states and cookies, indicating support for passing session/auth state into crawls. However, there's no explicit first-party documentation snippet showing how to directly inject custom session cookies into the arun() API call, and no independent hands-on confirmation of this specific workflow. Missing for 10: direct API-level example of passing cookies, independent verification of authenticated-page fetching working reliably.

                                                                            • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                                                                            Riveternone0/10

                                                                            No evidence in the pack mentions passing custom session cookies, auth headers, or cookie jars for scraping/enrichment requests; the docs describe scrapes, quick_search, and enrichments but not authenticated-page access via user-supplied cookies.

                                                                            • developerReuse a persistent browser profile with saved cookies and login state across multiple requests

                                                                              weight 2 · round to Crawl4AI
                                                                              Crawl4AIfullclaimed7/10

                                                                              GitHub docs explicitly describe a 'Browser Profiler' feature for creating and managing persistent profiles with saved authentication states, cookies, and settings, directly matching the story. Missing for 10: no independent/hands-on corroboration of profile reuse across multiple requests, and no first-party code sample demonstrating loading a saved profile in arun/AsyncWebCrawler calls.

                                                                              • [github] Browser Profiler: Create and manage persistent profiles with saved authentication states, cookies, and settings.
                                                                              Riveternone0/10

                                                                              No evidence Riveter offers persistent browser profiles, saved cookies, or login-state reuse across requests; its scraping is described as AI-agent page interpretation, not a session/profile management feature.

                                                                              Openness — open source, data portability, and self-hosting storiesOpenness

                                                                              Open source, data portability, and self-hosting stories

                                                                              1. ai-native userDo everything through the API that I can do in the UI

                                                                                weight 2 · round drawn
                                                                                Crawl4AIpartialprobed5/10

                                                                                Crawl4AI is API/library-first (Python API, CLI, Docker/FastAPI server) and the only 'UI' surface mentioned is a monitoring dashboard for the Docker deployment, so most functionality is inherently API-native; however probes found no OpenAPI spec (404s) to confirm full parity/documentation of the API surface, and there's no explicit claim that dashboard-only features (e.g., live monitoring) are also exposed via API. missing for 10: explicit API/OpenAPI documentation confirming parity, evidence that dashboard-specific features (metrics, browser pool visibility) are also API-accessible.

                                                                                • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                                                • [github] Real-time Monitoring Dashboard with live system metrics and browser pool visibility
                                                                                • [probe] PROBE openapi: all candidate paths 404 (https://docs.crawl4ai.com/openapi.json, https://docs.crawl4ai.com/swagger.json, https://docs.crawl4a…
                                                                                • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
                                                                                Riveterpartialprobed5/10

                                                                                Docs show many core capabilities (building enrichments via prompt/spec, scraping, quick_search, search_agent, webhooks, dry_run) are all API-accessible, suggesting broad parity, but there is no explicit statement of full UI/API parity and some UI-highlighted features like scheduling refresh (riveter-docs-13, riveter-docs-17) aren't confirmed as API-exposed. Additionally, probes show no discoverable OpenAPI spec (riveter-probe-2) or llms.txt (riveter-probe-1), undermining confidence that the API surface is fully documented/openly specified. missing for 10: explicit parity statement, API access to scheduling/monitoring feature, published OpenAPI spec for verification.

                                                                                • [claimed-docs] You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.
                                                                                • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                                                                • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
                                                                                • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                                                                                • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.
                                                                                • [probe] PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt
                                                                                • [probe] PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…
                                                                              2. ai-native userExport all of my data in open formats and leave

                                                                                weight 3 · round to Crawl4AI
                                                                                Crawl4AIfullclaimed7/10

                                                                                Crawl4AI is fully open-source and self-hosted, and its core output is markdown/JSON (open, non-proprietary formats) with no forced API keys or paywalls, meaning there is no vendor silo to 'leave' in the first place. Structured extraction (CSS/XPath/LLM) further lets users get data out in standard formats. Missing for 10: no explicit bulk 'export all my data' feature, no documented data-portability/migration tooling, and no independent hands-on confirmation of full data portability beyond architecture inference.

                                                                                • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                                                • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                                                Riveternone0/10

                                                                                No evidence of any data export feature or open-format export capability; the evidence only covers enrichment, scraping, search, and API integration features, with no mention of exporting data or portability guarantees. missing for 10: export functionality documentation, supported open formats (CSV/JSON/etc), any data-portability or account-closure workflow.

                                                                                • ai-native userRead the product's source under an open license

                                                                                  weight 2 · round to Crawl4AI
                                                                                  Crawl4AIpartialcommunity6/10

                                                                                  The project is explicitly described as open source (GitHub repo, docs stating 'Open Source: No forced API keys, no paywalls'), and community posts confirm it as an 'amazing open-source library', supporting readable source code. However, no specific license name (e.g., Apache-2.0, MIT) is cited in the evidence pack, so the exact open-license terms are unconfirmed. Missing for 10: explicit license identification/text, independent confirmation of license permissiveness.

                                                                                  • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                  • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                                                  • [community] Crawl4AI is an amazing open-source library that solves many LLM-scraping headaches.
                                                                                  Riveternone0/10

                                                                                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                  • ai-native userSelf-host the core product

                                                                                    weight 3 · round to Crawl4AI
                                                                                    Crawl4AIfullprobed8/10

                                                                                    Crawl4AI is open-source with a Dockerized FastAPI setup for deployment, explicit self-hosting docs (including MCP support), and community confirmation of running it themselves via Docker/n8n setups. Missing for 10: independent hands-on verification of a full self-hosted production deployment at scale, and more detail on resource/infra requirements.

                                                                                    • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                                                    • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                    • [probe] official MCP server documented at https://docs.crawl4ai.com/core/self-hosting/#mcp-model-context-protocol-support
                                                                                    • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                                                                    Riveternone0/10

                                                                                    Riveter is presented as a hosted API/SaaS product (with local MCP server option only for connecting AI clients, not for self-hosting the core enrichment engine); no evidence of open-source code, self-hosting instructions, or a downloadable core product exists in the pack.

                                                                                    • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
                                                                                    • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.

                                                                                  Output formats — stories about output formats in this arenaOutput formats

                                                                                  Stories about output formats in this arena

                                                                                  Content formats

                                                                                  1. developerReceive scraped content as clean markdown instead of raw HTML

                                                                                    weight 3 · round to Crawl4AI
                                                                                    Crawl4AIfullcommunity8/10

                                                                                    First-party docs show result.markdown as the direct output from crawler.arun(), and community sentiment corroborates it as a core value proposition for LLM-scraping. Missing for 10: independent hands-on verification of markdown quality/cleanliness and details on markdown customization options (e.g., filters).

                                                                                    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                    • [community] Crawl4AI is an amazing open-source library that solves many LLM-scraping headaches.
                                                                                    Riveterpartialclaimed5/10

                                                                                    Docs state a scrape 'turns a URL into easily parseable text,' implying cleaned output rather than raw HTML, but there's no explicit mention of markdown formatting or output schema. Missing for 10: explicit confirmation that scrape output is markdown-formatted, example output showing markdown structure, independent verification of output cleanliness.

                                                                                    • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                  2. developerChoose exactly which output format is returned, such as markdown, HTML, text, or frontmatter

                                                                                    weight 2 · round to Crawl4AI
                                                                                    Crawl4AIpartialclaimed4/10

                                                                                    Evidence confirms markdown output (result.markdown) and structured/CSS/XPath/LLM extraction, but the pack contains no explicit mention of selectable HTML, text, or frontmatter output formats. Missing for 10: documented options for raw/cleaned HTML output, plain text output, and frontmatter format selection.

                                                                                    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                    • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                                                    Riveternone0/10

                                                                                    Riveter's evidence pack covers enrichment, scraping, search, webhooks, and credit controls but never mentions selectable output formats like markdown, HTML, text, or frontmatter for returned data.

                                                                                    • developerReceive scraped content as structured JSON

                                                                                      weight 3 · round to Riveter
                                                                                      Crawl4AIpartialclaimed6/10

                                                                                      Docs confirm structured extraction via CSS/XPath/LLM strategies producing structured data (JSON-like) and LLM-driven structured data extraction, plus table extraction into structured form, supporting the core capability. However, the evidence never explicitly shows a JSON output example or schema, and there's no first-party confirmation of a dedicated JSON output mode/field beyond the markdown example shown. missing for 10: an explicit documented JSON output example/schema, independent hands-on confirmation of JSON structure quality.

                                                                                      • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                                                      • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                                                      • [github] LLMTableExtraction: Revolutionary table extraction with intelligent chunking for massive tables
                                                                                      • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                      Riveterfullclaimed7/10

                                                                                      Riveter's enrichments and scrapes explicitly return structured, parseable data (columns, urls/titles/snippets, webhook payloads of 'full results'), and SDK examples show structured attribute objects returned from calls, indicating outputs are consumable as structured JSON rather than raw text. missing for 10: an explicit statement of JSON schema/response format in docs, and independent/hands-on confirmation of the JSON structure (API docs endpoints 404 in probes).

                                                                                      • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                                                      • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                      • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                                                                      • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                      • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
                                                                                      • [claimed-docs] run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},

                                                                                    Llm ready output

                                                                                    1. ai-native userGet clean LLM-ready text directly instead of dealing with blocking, rendering, and messy HTML myself

                                                                                      weight 3 · round to Crawl4AI
                                                                                      Crawl4AIfullcommunity8/10

                                                                                      Core value proposition is documented directly: result.markdown provides clean LLM-ready markdown output from arun(), avoiding manual HTML parsing, plus structured/LLM-based extraction options and community confirmation it 'solves many LLM-scraping headaches.' Missing for 10: independent benchmarking of markdown output quality across diverse sites, and more detail on how blocking/anti-bot handling integrates seamlessly with the output pipeline.

                                                                                      • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                      • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                                                      • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                                                      • [community] Crawl4AI is an amazing open-source library that solves many LLM-scraping headaches.
                                                                                      Riveterpartialclaimed5/10

                                                                                      Docs claim a scrape converts any URL into 'easily parseable text' and that AI agents interpret pages 'the way a person would', directly addressing the ask for clean, LLM-ready text instead of raw HTML. However, all evidence is vendor documentation with no independent hands-on verification of output cleanliness, no example output shown, and no explicit mention of handling JS rendering/blocking obstacles beyond the general claim. Missing for 10: independent corroboration of scrape text quality, concrete example output, and explicit handling of anti-bot/rendering blockers.

                                                                                      • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                      • [claimed-docs] Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.
                                                                                      • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…
                                                                                      • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                                                                    2. ai-native userRequest semantically chunked output instead of one large content blob, so it feeds cleanly into a retrieval pipeline

                                                                                      weight 2 · round to Crawl4AI
                                                                                      Crawl4AIpartialclaimed4/10

                                                                                      Evidence only shows 'intelligent chunking' applied specifically to massive table extraction (LLMTableExtraction), not a general semantic chunking mode for arbitrary page content feeding a RAG pipeline. Structured/LLM extraction exists but nothing documents configurable chunk sizes, overlap, or semantic-boundary chunking of markdown output. Missing for 10: documented general-purpose content chunking strategy (e.g. semantic/topic-based chunking of markdown), configurable chunk size/overlap, and independent confirmation it integrates cleanly into retrieval pipelines.

                                                                                      • [github] LLMTableExtraction: Revolutionary table extraction with intelligent chunking for massive tables
                                                                                      • [claimed-docs] Structured Extraction: Parse repeated patterns with CSS, XPath, or LLM-based extraction.
                                                                                      • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                      Riveternone0/10

                                                                                      Riveter's evidence describes enrichments, scrapes, searches, and structured row outputs, but nothing indicates a semantic-chunking output mode designed for retrieval pipelines (e.g., configurable chunk size/overlap, chunk metadata). Structured rows/columns are not the same as semantic chunking for RAG ingestion, and no such feature is documented.

                                                                                      • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                                                      • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                      • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…

                                                                                    Visual capture

                                                                                    1. developerCapture a screenshot of a full page or a specific selected area

                                                                                      weight 2 · round drawn
                                                                                      Crawl4AInone0/10

                                                                                      No evidence in the pack mentions screenshot capture, full-page or selector-based screenshots, or any image output capability of Crawl4AI.

                                                                                        Riveternone0/10

                                                                                        The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                        Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits

                                                                                        Free-tier ceilings, usage caps, and rate limits before you have to pay

                                                                                        Cost optimization

                                                                                        1. developerLet the API automatically pick the cheapest configuration that still succeeds

                                                                                          weight 2 · round drawn
                                                                                          Crawl4AInone0/10

                                                                                          No evidence of any auto-selection of cheapest model/config that still meets quality requirements; there's no cost-based routing, budget optimizer, or fallback-on-price logic described anywhere in the docs or community reports. Adaptive crawling stops when enough info is gathered, but that's about crawl coverage, not cost-based configuration selection.

                                                                                            Riveternone0/10

                                                                                            Riveter offers cost controls like dry_run estimates and max_credits caps that refuse overpriced requests, but there is no evidence the API automatically searches for or selects the cheapest configuration that still succeeds — it only estimates/caps, it doesn't auto-optimize. Missing for 10: any documentation of automatic configuration search/optimization for cost, fallback logic that retries cheaper options, or an API parameter that lets Riveter choose the minimal successful config itself.

                                                                                            • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                            • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
                                                                                          • developerBlock ads on the target page to speed up scraping requests

                                                                                            weight 1 · round drawn
                                                                                            Crawl4AInone0/10

                                                                                            No evidence in the pack mentions ad-blocking or resource-blocking features to speed up crawling; while Crawl4AI has various performance and crawling features, none reference blocking ads specifically.

                                                                                              Riveternone0/10

                                                                                              The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                              • developerBlock images and CSS resources by default to reduce bandwidth and speed up requests

                                                                                                weight 1 · round drawn
                                                                                                Crawl4AInone0/10

                                                                                                No evidence in the pack mentions blocking images/CSS resources or any bandwidth-saving resource-filtering feature; none of the docs, GitHub, or community citations reference this capability.

                                                                                                  Riveternone0/10

                                                                                                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                                  • ai-native userSet how much reasoning effort an autonomous agent spends on a data-gathering task (low, medium, high)

                                                                                                    weight 2 · round drawn
                                                                                                    Crawl4AInone0/10

                                                                                                    Crawl4AI has adaptive crawling that stops when 'enough' info is gathered, but there is no evidence of a configurable reasoning-effort dial (low/medium/high) for agent tasks.

                                                                                                      Riveternone0/10

                                                                                                      No evidence anywhere in the pack of a control that lets users set reasoning effort (low/medium/high) for an agent's data-gathering task; only credit caps and dry-run cost estimation are documented, which are cost controls, not reasoning-effort controls.

                                                                                                      Cost transparency

                                                                                                      1. developerWhether failed, blocked, or empty-result requests still consume my billing quota

                                                                                                        weight 2 · round drawn
                                                                                                        Crawl4AInone0/10

                                                                                                        The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                                          Riveternone0/10

                                                                                                          The docs describe dry_run cost estimation and max_credits caps that prevent overage, but nothing states whether a failed, blocked, or empty-result run still consumes credits. Missing for 10: explicit policy on billing for failed/empty/blocked runs, any refund or non-charge guarantee for zero-result enrichments.

                                                                                                          • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                                          • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
                                                                                                        • developerSet a spending cap or usage alert so proxy/credit consumption doesn't silently blow past my budget

                                                                                                          weight 3 · round to Riveter
                                                                                                          Crawl4AInone0/10

                                                                                                          Crawl4AI is an open-source self-hosted crawler with no billing/credit system mentioned anywhere in the evidence; there is no spending cap, usage alert, or budget-tracking feature documented for proxy/LLM credit consumption.

                                                                                                            Riveterpartialclaimed6/10

                                                                                                            Riveter offers per-request cost control via dry_run (credit estimate before charging) and max_credits (hard ceiling that returns 422 credit_cap_exceeded with nothing charged), which directly prevents a single run from blowing past a set budget. However, there's no evidence of an account-wide spending cap, recurring usage alerts, or a dashboard/notification system for cumulative consumption across runs. Missing for 10: account/org-level budget cap, proactive usage alerts/notifications, historical spend tracking dashboard.

                                                                                                            • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                                            • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.

                                                                                                          Performance tuning

                                                                                                          1. developerTrade off latency against completeness by controlling exactly when content is returned

                                                                                                            weight 1 · round to Riveter
                                                                                                            Crawl4AIpartialclaimed6/10

                                                                                                            Crawl4AI offers explicit levers to trade latency for completeness: adaptive crawling that stops once 'sufficient information' is gathered, deep-crawl with max-pages limits, and resume_state checkpointing to control scope of a crawl before returning results. However, evidence is first-party docs/GitHub only, with no independent benchmarks or hands-on confirmation of how well the adaptive stopping heuristic tunes latency-vs-completeness in practice. Missing for 10: independent verification of adaptive-crawl accuracy/latency tradeoffs, and explicit developer-facing controls (e.g., a 'depth' or 'confidence threshold' parameter) documented with examples.

                                                                                                            • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                                                                                                            • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                            • [github] resume_state parameter to continue from a saved checkpoint
                                                                                                            Riveterfullclaimed7/10

                                                                                                            Riveter explicitly exposes multiple latency/completeness tradeoffs: quick_search returns fast synchronous structured snippets, search_agent runs a fuller AI research loop for one question, and full enrichments can be tracked via wait_for_result long-polling or async webhook callbacks — giving a developer direct control over when and how complete the returned content is. missing for 10: no independent/hands-on benchmarks or third-party confirmation of actual latency differences between these modes.

                                                                                                            • [claimed-docs] A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…
                                                                                                            • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
                                                                                                            • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                                            • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…

                                                                                                          Privacy posture — data-handling and privacy storiesPrivacy posture

                                                                                                          Data-handling and privacy stories

                                                                                                          1. ai-native userChoose where my data is stored (region/residency)

                                                                                                            weight 2 · round to Crawl4AI
                                                                                                            Crawl4AIpartialclaimed3/10

                                                                                                            Crawl4AI is open-source and self-hosted (Dockerized Setup), which implicitly lets users control where data is processed/stored by choosing their own deployment infrastructure, but there is no explicit documentation, configuration option, or claim about region/data-residency selection. missing for 10: explicit data residency/region configuration options, documentation addressing compliance/residency requirements, any mention of storage location control beyond generic self-hosting.

                                                                                                            • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                                                                            • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                                            Riveternone0/10

                                                                                                            No evidence in the pack mentions data residency, regional storage options, or compliance controls for where data is stored; the docs focus entirely on enrichment features and API mechanics. Missing for 10: any mention of region selection, data residency options, or storage location controls.

                                                                                                            • ai-native userControl data retention and deletion

                                                                                                              weight 2 · round drawn
                                                                                                              Crawl4AInone0/10

                                                                                                              No evidence describes explicit data retention/deletion controls (e.g., cache TTLs, purge commands, GDPR-style export/delete APIs); the only related item is a vague self-hosted/open-source claim about accessing your own data, which does not address retention or deletion policy.

                                                                                                              • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                                              Riveternone0/10

                                                                                                              No evidence in the pack addresses data retention policies, deletion controls, or data lifecycle management for Riveter's stored enrichment data, run results, or scraped content.

                                                                                                              • ai-native userOpt out of telemetry and usage tracking

                                                                                                                weight 2 · round drawn
                                                                                                                Crawl4AInone0/10

                                                                                                                Crawl4AI is an open-source, self-hosted library (no forced API keys/paywalls), which suggests limited built-in telemetry, but no evidence pack item mentions a telemetry system, opt-out flag, or privacy/usage-tracking policy at all.

                                                                                                                  Riveternone0/10

                                                                                                                  No evidence pack item mentions telemetry, usage tracking, analytics collection, or any opt-out mechanism for Riveter; the docs focus entirely on enrichment, scraping, and API features. Missing for 10: any mention of telemetry practices, privacy policy, or opt-out settings.

                                                                                                                  Scale reliability — behavior under load — scaling limits, uptime, failure handlingScale reliability

                                                                                                                  Behavior under load — scaling limits, uptime, failure handling

                                                                                                                  Ai driven crawling

                                                                                                                  1. ai-native userRely on adaptive crawling that automatically stops once enough information has been gathered to answer my query

                                                                                                                    weight 2 · round to Crawl4AI
                                                                                                                    Crawl4AIpartialclaimed6/10

                                                                                                                    First-party docs explicitly describe an adaptive crawling feature using 'information foraging algorithms' that stops once sufficient information is gathered to answer a query, directly matching the story. However, there is no independent/hands-on corroboration of this specific feature's effectiveness, and no benchmark or user report validating its stopping accuracy. missing for 10: independent verification of adaptive-stop behavior, quantitative accuracy/efficiency data, community confirmation of real-world use.

                                                                                                                    • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                                                                                                                    Riveterpartialclaimed3/10

                                                                                                                    Riveter's search_agent and enrichment agent loop imply some autonomous research process that fills a cell with an AI-researched answer, suggesting the agent decides when it has enough data, but there is no explicit documentation of stopping criteria or adaptive crawling behavior tied to query sufficiency. missing for 10: explicit description of adaptive stopping/crawling logic, evidence of how the agent determines 'enough information', independent confirmation of this behavior in practice.

                                                                                                                    • [claimed-docs] A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…
                                                                                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.

                                                                                                                  Batch processing

                                                                                                                  1. data-engineerBatch scrape thousands of URLs asynchronously

                                                                                                                    weight 3 · round drawn
                                                                                                                    Crawl4AIpartialcommunity6/10

                                                                                                                    Crawl4AI supports async crawling (AsyncWebCrawler/arun), multi-URL batch configuration with per-pattern strategies, deep-crawl CLI options, retry/proxy fallback, and resume-from-checkpoint for long jobs, all pointing toward large-scale async scraping. However, there's no explicit documentation of a dedicated 'arun_many' or thousands-of-URLs batch API, concurrency/throughput benchmarks, or first-party evidence of tested scale at 'thousands of URLs'; community comments note buyers must build their own policy/quality/rate-limiting layer for production scale. Missing for 10: documented high-concurrency batch API (e.g., arun_many) with concurrency controls, published benchmarks/case studies at thousands-of-URL scale, and independent confirmation of reliability at that scale.

                                                                                                                    • [claimed-docs] async with AsyncWebCrawler() as crawler: result = await crawler.arun(url="https://crawl4ai.com") print(result.markdown)
                                                                                                                    • [github] Multi-URL Configuration: Different strategies for different URL patterns in one batch
                                                                                                                    • [github] Automatic retry with proxy chain and fallback fetch function
                                                                                                                    • [github] resume_state parameter to continue from a saved checkpoint
                                                                                                                    • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                                    • [community] Promising foundation if you're willing to own the policy layer + quality gates.
                                                                                                                    • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                                                                    Riveterpartialclaimed6/10

                                                                                                                    Riveter's enrichment engine explicitly processes rows of URLs with scraping, runs asynchronously (webhook_url on completion), and SDKs handle retries, long-polling, and pagination — all core pieces for async batch scraping. However, there's no explicit documentation of scale limits, concurrency handling, or a tested example at thousands-of-URLs volume. Missing for 10: explicit large-scale (thousands of URLs) benchmarks or case studies, concurrency/rate-limit guidance for very large batches.

                                                                                                                    • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                                                                                    • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                                                    • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                                                    • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
                                                                                                                    • [claimed-docs] It can find every dental practice in a city, then pull every dentist from each one, in a single request.
                                                                                                                  2. developerApply different crawl configurations to different URL patterns within a single batch job

                                                                                                                    weight 1 · round to Crawl4AI
                                                                                                                    Crawl4AIfullclaimed7/10

                                                                                                                    GitHub docs explicitly advertise 'Multi-URL Configuration: Different strategies for different URL patterns in one batch,' directly matching the story. However, this is only a single-line feature mention with no first-party documentation example, API detail, or independent hands-on confirmation. missing for 10: detailed docs/tutorial showing per-pattern config syntax, independent/community validation of this specific feature in practice.

                                                                                                                    • [github] Multi-URL Configuration: Different strategies for different URL patterns in one batch
                                                                                                                    Riveternone0/10

                                                                                                                    No evidence describes applying different crawl configurations per URL pattern within one batch/enrichment job; docs mention scraping, searching, and enrichment generally but not per-pattern configuration rules. missing for 10: any mention of per-URL-pattern rules or configuration scoping within a single job, examples or docs showing mixed crawl settings in one batch.

                                                                                                                    Concurrency

                                                                                                                    1. data-engineerSpin up many concurrent scraping sessions to gather data at scale

                                                                                                                      weight 3 · round to Crawl4AI
                                                                                                                      Crawl4AIpartialcommunity6/10

                                                                                                                      Crawl4AI supports batch/multi-URL crawling, deep crawl with max-pages, checkpoint resume, Docker/FastAPI deployment with a monitoring dashboard showing browser pool visibility, and retry/proxy chains—together implying support for concurrent, at-scale scraping. However, there's no explicit documentation of concurrency limits, session pooling configuration, or benchmarks proving many-simultaneous-session throughput, and community commentary notes users must build their own rate-limiting/production policy layer. Missing for 10: explicit concurrency/session-pool configuration docs, load/scale benchmarks, and independent verification of large-scale concurrent runs.

                                                                                                                      • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                                      • [github] Automatic retry with proxy chain and fallback fetch function
                                                                                                                      • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                                                                                      • [github] Real-time Monitoring Dashboard with live system metrics and browser pool visibility
                                                                                                                      • [github] resume_state parameter to continue from a saved checkpoint
                                                                                                                      • [github] Multi-URL Configuration: Different strategies for different URL patterns in one batch
                                                                                                                      • [community] Promising foundation if you're willing to own the policy layer + quality gates.
                                                                                                                      • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                                                                      Riveterpartialclaimed5/10

                                                                                                                      Riveter's enrichment engine processes many rows in a single run and can chain scrapes/searches (e.g., finding every dental practice then every dentist in one request), implying built-in batch/bulk scraping at scale, and SDKs handle retries/pagination for large jobs. However, there is no explicit documentation of concurrency limits, parallel session management, or throughput guarantees for scraping specifically. Missing for 10: explicit concurrency/session limits, performance benchmarks, and independent evidence of scaling to many simultaneous scrape sessions.

                                                                                                                      • [claimed-docs] An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.
                                                                                                                      • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                                                      • [claimed-docs] It can find every dental practice in a city, then pull every dentist from each one, in a single request.
                                                                                                                      • [claimed-docs] the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…
                                                                                                                      • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.

                                                                                                                    Crawl compliance

                                                                                                                    1. data-engineerConfigure the crawler to respect robots.txt rules and target-site rate limits automatically

                                                                                                                      weight 2 · round drawn
                                                                                                                      Crawl4AInone0/10

                                                                                                                      No documentation or feature evidence shows Crawl4AI automatically respects robots.txt or enforces target-site rate limits; the only relevant community evidence explicitly notes that 'robots/ToS, rate limiting' are things the operator must own themselves, i.e., not built-in automation.

                                                                                                                      • [community] Promising foundation if you're willing to own the policy layer + quality gates.
                                                                                                                      • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                                                                      Riveternone0/10

                                                                                                                      No evidence anywhere in the docs mentions robots.txt compliance or rate-limit configuration; the pack only covers scraping features, retries, credits, and MCP integration. This is a fair axis for a web-scraping/crawling product, but absence of evidence means it cannot be credited as delivered.

                                                                                                                      Fault tolerance

                                                                                                                      1. data-engineerResume a crashed deep crawl from a saved checkpoint instead of restarting from scratch

                                                                                                                        weight 2 · round to Crawl4AI
                                                                                                                        Crawl4AIpartialclaimed6/10

                                                                                                                        GitHub evidence confirms a resume_state parameter to continue a deep crawl from a saved checkpoint, directly matching the story. However, there's no documentation detail on how checkpoints are saved automatically during a crash, how frequently state is persisted, or independent hands-on confirmation of this working in practice. missing for 10: first-party docs walkthrough of checkpoint save/resume workflow, independent/community verification of crash-recovery behavior.

                                                                                                                        • [github] resume_state parameter to continue from a saved checkpoint
                                                                                                                        Riveternone0/10

                                                                                                                        No evidence describes checkpointing or resuming a crashed deep crawl; the docs mention webhooks, dry runs, and credit caps but nothing about saving/resuming crawl state after a crash.

                                                                                                                        Scheduling monitoring

                                                                                                                        1. data-engineerMonitor target pages for content changes, such as price or listing updates, and get notified as they happen

                                                                                                                          weight 2 · round to Riveter
                                                                                                                          Crawl4AInone0/10

                                                                                                                          Crawl4AI is a crawling/extraction library with deep-crawl, retry, and dashboard monitoring features, but nothing in the evidence describes scheduled re-crawling, diff/change-detection, or alerting/notification mechanisms for tracking content changes like price or listing updates over time.

                                                                                                                            Riveterpartialclaimed7/10

                                                                                                                            Riveter explicitly supports scheduling projects to monitor for changes, refreshing as often as every minute, and can POST results to a webhook_url when a run finishes, which together deliver change-monitoring plus notification. However, the webhook fires on run completion rather than a dedicated 'content changed' diff event, and there's no independent/hands-on evidence of this workflow in production. Missing for 10: independent corroboration of the schedule+webhook pipeline in practice, and explicit diff/change-detection logic distinguishing 'changed' vs 'unchanged' pages.

                                                                                                                            • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                                                                                                                            • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.
                                                                                                                            • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                                                          • data-engineerMonitor job performance, validate data quality, and receive alerts when something fails

                                                                                                                            weight 2 · round to Riveter
                                                                                                                            Crawl4AIpartialcommunity3/10

                                                                                                                            There is a documented real-time monitoring dashboard with live system metrics and browser pool visibility, which covers basic job performance monitoring, and automatic retry with proxy/fallback chains aids reliability. However, there is no evidence of data quality validation features or an alerting/notification system for failures, and community feedback explicitly notes users must 'own the policy layer + quality gates' themselves. Missing for 10: data quality validation tooling, failure alerting/notification integration, and independent confirmation of the monitoring dashboard's depth.

                                                                                                                            • [github] Real-time Monitoring Dashboard with live system metrics and browser pool visibility
                                                                                                                            • [github] Automatic retry with proxy chain and fallback fetch function
                                                                                                                            • [community] Promising foundation if you're willing to own the policy layer + quality gates.
                                                                                                                            • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                                                                            Riveterpartialclaimed5/10

                                                                                                                            Riveter supports webhook alerts on run completion/stop/finish events and scheduled monitoring for data freshness, giving some job-status alerting and monitoring capability, but there is no explicit data-quality validation feature (e.g., schema/anomaly checks) or job performance dashboards described. missing for 10: explicit data quality validation tooling, job performance metrics/dashboard, and independent confirmation of alerting reliability.

                                                                                                                            • [claimed-docs] Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…
                                                                                                                            • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                                                                                                                            • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                                                            • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
                                                                                                                          • developerMonitor live system metrics and worker/browser pool status through a real-time dashboard

                                                                                                                            weight 1 · round to Crawl4AI
                                                                                                                            Crawl4AIpartialclaimed6/10

                                                                                                                            GitHub evidence explicitly claims a 'Real-time Monitoring Dashboard with live system metrics and browser pool visibility,' directly matching the story, but this is a single first-party mention with no independent hands-on corroboration, screenshots, or docs detail on what metrics/UI it exposes. missing for 10: independent/community confirmation of the dashboard working, detailed docs on metrics tracked, screenshots or setup instructions.

                                                                                                                            • [github] Real-time Monitoring Dashboard with live system metrics and browser pool visibility
                                                                                                                            Riveternone0/10

                                                                                                                            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                                                            • developerSchedule scraping jobs to run automatically at specific times

                                                                                                                              weight 2 · round to Riveter
                                                                                                                              Crawl4AInone0/10

                                                                                                                              No evidence of built-in scheduling functionality (cron-like triggers or job scheduler) — Crawl4AI is a crawling/extraction library and CLI/Docker deployment, with community notes suggesting users must bridge to external automation tools like n8n for production workflows including scheduling. Missing for 10: any native scheduler, cron integration, or documented recurring-job API.

                                                                                                                              • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                                                                                                                              • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                                                                                                              • [github] Dockerized Setup: Optimized Docker image with FastAPI server for easy deployment.
                                                                                                                              Riveterpartialclaimed5/10

                                                                                                                              Riveter supports scheduling projects to monitor for changes and refresh data as often as every minute, which implies automatic recurring scraping jobs, but there's no detail on specifying exact times/cron-like scheduling, timezone control, or a documented scheduling API/UI. missing for 10: explicit scheduling configuration details (time-of-day, cron syntax, timezone), independent/hands-on confirmation of scheduling reliability, and API endpoint documentation for creating/managing schedules.

                                                                                                                              • [claimed-docs] Schedule any project to monitor for changes and keep your data fresh.
                                                                                                                              • [claimed-docs] For fast moving data like scores or election results, you can refresh as often as every minute.

                                                                                                                            Site crawling

                                                                                                                            1. data-engineerRun a deep crawl using a breadth-first strategy with a configurable maximum page limit

                                                                                                                              weight 2 · round to Crawl4AI
                                                                                                                              Crawl4AIfullprobed8/10

                                                                                                                              CLI evidence explicitly shows `--deep-crawl bfs --max-pages 10`, directly matching the requested breadth-first strategy with configurable page limit, and the official CLI docs corroborate this exists as a documented feature. Missing for 10: no independent hands-on report validating large-scale BFS crawl behavior/performance at scale, and no Python API example (only CLI) confirming programmatic configurability.

                                                                                                                              • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                                              • [probe] official CLI documented at https://docs.crawl4ai.com/core/cli/
                                                                                                                              Riveternone0/10

                                                                                                                              The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                                                              • developerCrawl an entire website and get content from all its pages with one request

                                                                                                                                weight 3 · round to Crawl4AI
                                                                                                                                Crawl4AIfullcommunity8/10

                                                                                                                                Crawl4AI supports deep/BFS crawling with a max-pages parameter via CLI (--deep-crawl bfs --max-pages 10), plus adaptive crawling that decides when enough pages have been gathered, and resume_state for continuing large crawls — directly enabling whole-site crawling in one request/command. Community feedback confirms it's used for scraping at scale, though notes production concerns like rate limiting and bot mitigation as caveats. Missing for 10: independent benchmark of full-site crawl completeness/performance and clearer documentation of concurrency limits at scale.

                                                                                                                                • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                                                • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                                                                                                                                • [github] resume_state parameter to continue from a saved checkpoint
                                                                                                                                • [community] Promising foundation if you're willing to own the policy layer + quality gates.
                                                                                                                                • [community] Worth calling out the boring production bits: robots/ToS, rate limiting, bot mitigation, login/session handling, and not accidentally hoover…
                                                                                                                                Riveterpartialclaimed5/10

                                                                                                                                Riveter's docs describe single-URL 'scrape' and 'quick_search' calls, but the marketing example of finding every dental practice in a city and pulling data from each one in a single request shows it can aggregate content across multiple pages/sources in one enrichment run, which approximates whole-site crawling. There is no explicit sitemap-style 'crawl entire website' feature or evidence of full-domain page enumeration. missing for 10: explicit full-site/sitemap crawl feature, evidence of automatically discovering and traversing all pages of a single domain, independent confirmation of multi-page crawl behavior.

                                                                                                                                • [claimed-docs] A scrape lets you turn a URL into easily parseable text.
                                                                                                                                • [claimed-docs] It can find every dental practice in a city, then pull every dentist from each one, in a single request.
                                                                                                                                • [claimed-docs] It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…
                                                                                                                              • developerInstantly discover all URLs on a website without fully crawling it

                                                                                                                                weight 2 · round drawn
                                                                                                                                Crawl4AInone0/10

                                                                                                                                Evidence shows deep-crawl (BFS) and adaptive crawling features that limit or stop crawling, but these still involve fetching and parsing pages rather than instantly enumerating a site's URL list (e.g., via sitemap parsing) without crawling. No probe or doc confirms a dedicated 'discover URLs only' mode. Missing for 10: sitemap.xml/URL-discovery feature, evidence of URL enumeration without page fetches, independent confirmation of instant discovery.

                                                                                                                                • [github] crwl https://docs.crawl4ai.com --deep-crawl bfs --max-pages 10
                                                                                                                                • [claimed-docs] Crawl4AI now features intelligent adaptive crawling that knows when to stop! Using advanced information foraging algorithms, it determines w…
                                                                                                                                Riveternone0/10

                                                                                                                                The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                                                                                                                                Not comparable on these axes

                                                                                                                                1. ai-native userPlug MCP servers into this product so it can use their tools

                                                                                                                                  weight 3 · not comparable
                                                                                                                                  Crawl4AIn/a

                                                                                                                                  Crawl4AI is a web-crawling library/service, not an agent that consumes external tools; the evidence shows it exposes an official MCP *server* (crawl4ai-probe-3) so that agents like Cursor/Claude can plug into it, which is the reverse relationship from the story's 'plug MCP servers into this product' framing. There is no evidence of Crawl4AI acting as an MCP client consuming other servers' tools, and this role doesn't fit its product category.

                                                                                                                                  • [probe] official MCP server documented at https://docs.crawl4ai.com/core/self-hosting/#mcp-model-context-protocol-support
                                                                                                                                  • [community] New developers often struggle with production configurations—specifically how to use Crawl4AI with MCP servers for Cursor, or how to bridge …
                                                                                                                                  • [community] Built crawl4ai.dev as a community-driven documentation hub with one-click Docker setups for n8n/FastAPI and production-ready MCP server guid…
                                                                                                                                  Riveternone0/10

                                                                                                                                  All MCP-related evidence describes Riveter acting as an MCP *server* that other AI assistants (Claude, ChatGPT, Cursor) connect to, not Riveter acting as a client that can plug in and consume external MCP servers' tools. No evidence shows Riveter itself importing/connecting to third-party MCP servers to extend its own tool use.

                                                                                                                                  • [claimed-docs] Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.
                                                                                                                                  • [claimed-docs] Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.
                                                                                                                                2. ai-native userVersion, review, and roll back my automations

                                                                                                                                  weight 1 · not comparable
                                                                                                                                  Crawl4AIn/a

                                                                                                                                  Crawl4AI is a web-crawling/extraction library, not an automation/workflow builder with versioned pipelines; version control, review, and rollback of automations is a category error for this product type. No evidence pack item addresses versioning or rollback of crawl configurations/workflows.

                                                                                                                                    Riveternone0/10

                                                                                                                                    No evidence of version history, review workflows, or rollback capability for automations/enrichments; the pack only covers run execution, credit control, and data enrichment features. Missing for 10: versioning of automation configs, review/approval workflow, rollback/undo mechanism.

                                                                                                                                    • developerWhether exceeding my plan's monthly credit or request quota triggers overage charges or a hard cutoff

                                                                                                                                      weight 3 · not comparable
                                                                                                                                      Crawl4AIn/a

                                                                                                                                      Crawl4AI is an open-source, self-hosted crawling library/tool with no paywalls or paid plans (crawl4ai-docs-4), so there is no metered plan with credits, quotas, or overage billing for this question to apply to.

                                                                                                                                      • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                                                                      Riveternone0/10

                                                                                                                                      The evidence describes credit estimation, dry_run, and max_credits cap that refuses requests at 422 before charging, but there is no mention of plan-level monthly credit/request quotas, nor whether exceeding them triggers overage billing or a hard cutoff. missing for 10: any documentation of monthly plan quotas, overage billing policy, or hard-cutoff behavior when a subscription limit is exceeded.

                                                                                                                                      • [claimed-docs] dry_run: true — validate the request and return a credit estimate without creating or charging anything.
                                                                                                                                      • [claimed-docs] max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.
                                                                                                                                    • data-engineerThe maximum concurrent sessions or requests allowed on my pricing tier and the cost to raise that cap

                                                                                                                                      weight 2 · not comparable
                                                                                                                                      Crawl4AIn/a

                                                                                                                                      Crawl4AI is an open-source, self-hosted library/Docker deployment with no vendor-imposed pricing tiers or session caps (crawl4ai-docs-4 emphasizes no forced API keys or paywalls); concurrency is governed by the user's own infrastructure, not a purchasable tier, so the pricing-limits axis doesn't apply to this product category.

                                                                                                                                      • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                                                                      Riveternone0/10

                                                                                                                                      No evidence pack items mention pricing tiers, concurrency/rate limits, or the cost to raise such caps; only credit-cap controls (max_credits) and dry-run estimates are documented, which address per-request cost, not concurrent session limits or tier-based upgrade pricing.

                                                                                                                                      • ai-native userPrevent my data from being used to train AI models

                                                                                                                                        weight 3 · not comparable
                                                                                                                                        Crawl4AIn/a

                                                                                                                                        Crawl4AI is a self-hosted, open-source crawling library that runs on the user's own infrastructure; there is no vendor-hosted service that ingests user data for model training, so a 'prevent training on my data' opt-out is a category mismatch rather than a missing feature. Any LLM training concerns would pertain to whichever third-party LLM the user chooses to plug in, not to Crawl4AI itself.

                                                                                                                                        • [claimed-docs] Open Source: No forced API keys, no paywalls—everyone can access their data.
                                                                                                                                        • [github] LLM-Driven Extraction: Supports all LLMs (open-source and proprietary) for structured data extraction.
                                                                                                                                        Rivetern/a

                                                                                                                                        Riveter is a data enrichment/web-scraping tool, not an AI model provider or chat assistant whose data usage/training policies would be relevant here; the evidence contains no mention of AI training data policy or opt-out controls, and this axis is a category error for the product type.

                                                                                                                                        • data-engineerCheck a public status page showing uptime history and past incident postmortems before committing to the service

                                                                                                                                          weight 2 · not comparable
                                                                                                                                          Crawl4AIn/a

                                                                                                                                          Crawl4AI is an open-source self-hosted crawling library/tool, not a hosted SaaS with an uptime/SLA obligation; a public status page with incident postmortems is not a fair expectation for this product category.

                                                                                                                                            Riveternone0/10

                                                                                                                                            No evidence of a public status page, uptime history, or incident postmortems anywhere in the evidence pack; only product feature docs and API references are present.