Firecrawl vs Riveter
Firecrawl
Mendable AI, Inc.
Firecrawl wins · 30–18 (46 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to FirecrawlA direct probe confirms Firecrawl publishes a working llms.txt at docs.firecrawl.dev/llms.txt with structured agent-readable documentation links, letting an AI agent be pointed directly at it to navigate Firecrawl's docs. Missing for 10: explicit first-party announcement/documentation describing llms.txt support as a deliberate feature, and independent community confirmation of agents successfully using it.
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.firecrawl.dev/llms.txt # Firecrawl Docs ## English ### v2 #### Documentation ##### Get Started …”
Riveternone0/10Direct probes show llms.txt returns 404 and no OpenAPI spec is discoverable at any standard path, and no evidence pack item claims an agent-oriented docs format exists; while MCP integration is mentioned, that's a separate capability from machine-readable docs for pointing an agent at.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round drawnFirecrawl offers an API-first product (async scraping, webhooks, CLI, SDKs) that is well-suited to headless/CI use, and docs confirm a CLI and webhook-based async event delivery for automation pipelines. However, there's no explicit CI-specific documentation (e.g., GitHub Actions examples, Docker image for CI), and community comments note some daemon/CLI limitations rather than confirming robust CI usage. missing for 10: explicit CI/headless deployment docs or examples, independent confirmation of stable CLI/daemon behavior in automated pipelines, containerization guidance for CI environments.
- [claimed-docs] “Webhooks Async event delivery”
- [probe] “official CLI documented at https://docs.firecrawl.dev/sdks/cli”
- [github] “Scrape thousands of URLs asynchronously”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Riveter exposes a full API with SDKs (Go example shown), webhooks for async completion, dry_run/max_credits safety controls, and scheduling for recurring automation — all of which support headless, non-interactive use in a pipeline. However, there is no explicit CI/CD example, GitHub Actions integration, or CLI documentation demonstrating a documented headless workflow. Missing for 10: explicit CI/CD or pipeline integration guide, CLI headless invocation docs, independent confirmation of automated/scripted runs.
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
ai-native userConnect an agent via an official MCP server
weight 3 · round to FirecrawlFirecrawl is a scraping/data-extraction service (not itself an agent), and it documents an official MCP server for connecting AI tools/agents to Firecrawl, corroborated by a dedicated GitHub repo (firecrawl-mcp-server). Missing for 10: independent hands-on testing of the MCP server itself and details on tool/resource coverage exposed via MCP.
- [claimed-docs] “MCP Server: Connect Firecrawl to any AI tool via the Model Context Protocol”
- [probe] “official MCP server documented at https://github.com/mendableai/firecrawl-mcp-server”
Docs explicitly describe connecting Riveter to Claude, ChatGPT, Cursor, or any MCP-compatible assistant via two connection methods, including a local Node.js-based server option, indicating an official MCP server offering. Missing for 10: no independent/hands-on corroboration of the MCP server working, and no detail on the remote/hosted connection method's implementation.
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
ai-native userUse an official CLI
weight 2 · round to FirecrawlFirecrawl ships an official CLI (docs.firecrawl.dev/sdks/cli) that installs, authenticates, and adds skills to coding agents, directly matching an AI-native CLI story. However, community feedback notes real limitations in CLI/daemon mode (e.g., inability to return HTML), suggesting it's not fully mature. Missing for 10: independent hands-on verification of full CLI feature parity, and no comparison data beyond one critical community comment.
- [claimed-docs] “One command installs the Firecrawl CLI, authenticates in your browser, and adds skills to every detected coding agent.”
- [probe] “official CLI documented at https://docs.firecrawl.dev/sdks/cli”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Riveternone0/10Evidence shows SDKs (Go), a local MCP server requiring Node.js, and REST API features, but no mention of an official CLI tool for running enrichments or managing the product. The docs and probes (llms.txt, openapi) surface no CLI reference, so this applicable axis is unmet.
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
ai-native userDrive the product through a documented public API
weight 3 · round to FirecrawlFirecrawl is API-first: docs cover scrape/crawl/search/extract endpoints, schema-based structured output, webhooks, and SDKs/CLI, all confirmed by an extensive llms.txt-indexed documentation site and GitHub feature list. Missing for 10: a discoverable machine-readable OpenAPI/Swagger spec (probe returned 404s) and independent third-party confirmation of API completeness.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Search the web and get full page content from results in one call.”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [github] “Use a schema to get structured data:”
- [claimed-docs] “Webhooks Async event delivery”
- [probe] “PROBE llms.txt: HTTP 200 at https://docs.firecrawl.dev/llms.txt # Firecrawl Docs ## English ### v2 #### Documentation ##### Get Started …”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.firecrawl.dev/openapi.json, https://docs.firecrawl.dev/swagger.json, https://docs.firec…”
- [probe] “official CLI documented at https://docs.firecrawl.dev/sdks/cli”
Docs describe concrete API mechanics (webhook_url, dry_run, max_credits, SDK auth/retry/pagination handling, Go SDK code sample) showing a real documented public API surface for driving runs programmatically, and MCP/remote-server integration is documented. However, probes for a formal machine-readable spec (openapi.json/swagger.json) and llms.txt all returned 404, so there's no discoverable canonical API reference, undermining full 'documented public API' claims. missing for 10: a public OpenAPI/swagger spec or llms.txt confirming a fully machine-readable API contract, independent third-party confirmation of API usage.
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round drawnFirecrawlnone0/10No evidence of scoped or least-privilege API key/credential management for agents; documentation covers scraping, crawling, MCP, CLI, and webhooks but nothing about API key scopes, permissions, or credential issuance controls.
Riveternone0/10Riveter's evidence covers a single API key model, credit caps, and dry-run cost estimation, but there is no mention of scoped or least-privilege credentials, per-agent tokens, or permission scoping for agents. missing for 10: scoped/least-privilege credential issuance, per-agent API key scoping, role/permission-based access control.
ai-native userBuild against official SDKs
weight 2 · round to RiveterThe only concrete artifact tied to 'SDKs' in the evidence is the CLI documented at docs.firecrawl.dev/sdks/cli, implying an SDKs section exists, but no evidence pack item names or links a Python/Node/other language SDK, shows install/usage snippets, or corroborates community usage. Missing for 10: explicit language SDK docs/links, code examples, independent/community confirmation of SDK usage.
- [probe] “official CLI documented at https://docs.firecrawl.dev/sdks/cli”
- [claimed-docs] “One command installs the Firecrawl CLI, authenticates in your browser, and adds skills to every detected coding agent.”
Riveter ships an official Go SDK (riveterhq/riveter-go) with documented client code (riveter.EnrichParams), and docs describe SDK-level handling of auth, retries, long-polling, and pagination, indicating a first-party SDK layer built for AI-native workflows. Missing for 10: confirmation of additional language SDKs (e.g., Python/JS) beyond Go, and independent/hands-on corroboration of SDK reliability.
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
ai-native userSubscribe to events via webhooks
weight 2 · round to FirecrawlFirecrawl's docs explicitly document a Webhooks feature for async event delivery, directly matching the story. Missing for 10: details on event types, payload schema, retry/security guarantees, and independent/hands-on confirmation of webhook usage.
- [claimed-docs] “Webhooks Async event delivery”
Riveter supports webhooks by passing a webhook_url when starting a run, with Riveter POSTing results back on run.completed, run.stopped, and run.finished events — a real event-notification mechanism for agentic workflows. However this is scoped to a single run's lifecycle rather than a general subscription model (no persistent webhook registration/management endpoint, no broader event catalog, no signature/security details). Missing for 10: a dedicated webhook subscription/management API, documentation of additional event types beyond run lifecycle, and payload signing/verification details.
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to RiveterFirecrawlnone0/10The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)
Riveter's core enrichment feature fills columns using AI agents, web search/scrape, and other tools to generate insights directly on user data, and search_agent provides ad hoc AI-researched answers within the product. missing for 10: independent/hands-on corroboration of insight quality, no example of proactive/unprompted suggestions (only prompt-driven enrichment), and no dashboard-level 'insights' UI evidence beyond API/SDK docs.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to RiveterFirecrawl supports webhooks for async event delivery and crawling jobs that run without blocking, which enables background/autonomous data-retrieval workflows, and its MCP server/CLI let agents trigger these jobs programmatically. However there's no evidence of a scheduling/trigger system (e.g., cron-like recurring jobs) or persistent autonomous 'automation' orchestration beyond one-off crawl/extract jobs with webhook callbacks. Missing for 10: scheduled/recurring job support, autonomous multi-step automation orchestration, independent confirmation of long-running background automation reliability.
- [claimed-docs] “Webhooks Async event delivery”
- [github] “Scrape thousands of URLs asynchronously”
- [claimed-docs] “MCP Server: Connect Firecrawl to any AI tool via the Model Context Protocol”
- [probe] “official MCP server documented at https://github.com/mendableai/firecrawl-mcp-server”
Riveter supports scheduling projects to run on a cadence ('every minute' for fast-moving data) and webhook notifications on run completion, which enables autonomous background execution without manual triggering. However, there's no evidence of broader automation orchestration (e.g., conditional triggers, chaining multiple actions, or a dedicated automation/workflow builder) beyond scheduled data refresh. Missing for 10: evidence of multi-step autonomous workflows beyond scheduled enrichment refresh, independent/hands-on confirmation that scheduling works reliably in production, and any automation trigger types beyond time-based schedules.
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to RiveterFirecrawl exposes an 'AI agent' mode where a user describes what they need and the agent searches/navigates/retrieves without URLs, with configurable reasoning effort (firecrawl-gh-1, firecrawl-gh-2) — a limited form of task delegation to an embedded AI. However, this is a narrow scraping/search agent, not a general-purpose in-product assistant, and there's no evidence of a broader conversational assistant UI for delegating arbitrary tasks. Missing for 10: evidence of a general-purpose conversational assistant interface, examples of delegated multi-step tasks beyond search/navigate, and independent confirmation of this agent's real-world reliability.
Riveter ships an internal 'agent loop' (search_agent, enrichment AI) that autonomously researches, scrapes, and fills data on request, which functions as a built-in AI assistant for delegated research tasks rather than a conversational general-purpose assistant. Missing for 10: evidence of a general chat/task interface for arbitrary delegation, independent hands-on validation, and clarity on how broadly the agent can handle tasks beyond enrichment/search/scrape.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
ai-native userOperate the product with natural-language commands
weight 2 · round to RiveterFirecrawl offers a natural-language 'search agent' mode ('Describe what you need... No URLs required') and lets users tune agent reasoning effort, which supports NL-driven operation, and its MCP/CLI integrations let AI agents invoke it conversationally through coding assistants. However, most of the product's core surface (scrape, crawl, extract, map) is still driven by structured API calls/schemas rather than free-form natural language commands. Missing for 10: evidence of full NL command coverage across all core endpoints (not just the search agent), and independent hands-on confirmation that NL commands reliably work end-to-end.
- [github] “Describe what you need, and our AI agent searches, navigates, and retrieves it. No URLs required.”
- [github] “Set how much reasoning the agent spends on the task”
- [claimed-docs] “One command installs the Firecrawl CLI, authenticates in your browser, and adds skills to every detected coding agent.”
- [claimed-docs] “MCP Server: Connect Firecrawl to any AI tool via the Model Context Protocol”
- [probe] “official MCP server documented at https://github.com/mendableai/firecrawl-mcp-server”
Riveter explicitly supports building enrichments from natural-language prompts (riveter-docs-2), offers a search_agent that answers questions in natural language without setup (riveter-docs-5), and can be operated via MCP-compatible AI assistants like Claude, ChatGPT, and Cursor (riveter-docs-9), which is the core mechanism for natural-language control. Missing for 10: independent/hands-on confirmation of NL command reliability, and no evidence of a broader NL command surface beyond enrichment/search (e.g., NL-driven scheduling or config changes).
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
ai-native userApply a preset configuration tuned for research agents that returns structured, citable output
weight 2 · round drawnFirecrawlnone0/10Firecrawl offers general scraping, structured JSON extraction, and search, but the evidence pack shows no dedicated preset/mode tuned specifically for research agents that returns citable, source-attributed output — no citation formatting, source-tracking, or research-agent-specific configuration is documented.
Riveternone0/10Riveter offers enrichment, search_agent, and scrape tools with structured outputs, but there is no evidence of a preset/template configuration specifically tuned for research agents or citable output formatting; missing for 10: a named preset or template targeting research-agent workflows, citation/source-attribution formatting in outputs, and any documentation referencing 'research agent' presets.
Api quality
ai-native userExplore an interactive API reference with runnable examples
weight 2 · round drawnFirecrawlnone0/10No evidence of an interactive API reference or runnable-example playground; the OpenAPI/swagger probe explicitly returned 404s at all candidate paths, and docs items only describe features, not an interactive reference experience.
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.firecrawl.dev/openapi.json, https://docs.firecrawl.dev/swagger.json, https://docs.firec…”
Riveternone0/10No evidence of an interactive API reference or runnable examples; probes for llms.txt and OpenAPI/Swagger specs both returned 404s, and docs snippets are static text/code examples only, not interactive/runnable.
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnFirecrawlnone0/10A direct probe for OpenAPI/Swagger spec files at all standard locations (openapi.json, swagger.json, etc.) returned 404s, and no other evidence pack item mentions a downloadable machine-readable API spec; only an llms.txt documentation index was found, which is not an OpenAPI-equivalent spec.
Riveternone0/10Probes for llms.txt and OpenAPI/swagger spec files all returned 404s, and no documentation mentions a downloadable machine-readable API spec despite having a REST API and SDKs.
ai-native userTest against a sandbox environment without touching production data
weight 1 · round to RiveterFirecrawlnone0/10The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)
Riveter offers a dry_run mode that validates a request and returns a credit estimate without creating or charging anything, and a max_credits cap that blocks runs before they execute — both function like a lightweight 'test without side effects' capability. However, there's no explicit documentation of a separate sandbox environment or synthetic/test dataset distinct from production data sources (Riveter always operates against live web/data sources when actually run). Missing for 10: a documented sandbox/staging environment, sample or mock datasets, and explicit guidance on testing enrichments without touching real production data sources.
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnFirecrawlnone0/10The docs reference a 'v2' API version (firecrawl-probe-1), showing some versioning exists, but there is no evidence of a documented deprecation policy, version support timelines, or migration guides, and an OpenAPI spec could not even be located (firecrawl-probe-2). Missing for 10: explicit deprecation policy documentation, versioning/support lifecycle statements, migration guidance for older API versions.
Riveternone0/10No evidence of API versioning scheme or a documented deprecation policy; probes for OpenAPI/spec discovery returned 404s, and docs mention SDKs/features but nothing about version numbers or deprecation guarantees. Missing for 10: versioned endpoint scheme (e.g., /v1/), a published deprecation/sunset policy, changelog or migration guides.
data-engineerThe documented rate limit (requests per second or minute) enforced on my API key before throttling kicks in
weight 3 · round drawnFirecrawlnone0/10No evidence pack item documents specific rate limits (requests per second/minute) per API key or plan tier; only general product features and community commentary are present.
Riveternone0/10There is a mention of SDKs handling retries on 429s, implying rate limiting exists, but no documented numeric rate limit (requests per second/minute) is given anywhere in the evidence pack, and probes for API spec/docs return 404s.
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…”
Anti bot — getting past bot defenses — CAPTCHAs, fingerprinting, blocksAnti bot
Getting past bot defenses — CAPTCHAs, fingerprinting, blocks
Block evasion
ai-native userHave an agent automatically get past a CAPTCHA, login, or form wall without my manual intervention
weight 2 · round to FirecrawlFirecrawl's docs support form-filling, clicking, and navigating via a 'Browser Sandbox' for interactive workflows (firecrawl-docs-3, firecrawl-docs-8), and community comments reference actual CAPTCHA 'solves' being consumed at cost (firecrawl-comm-6), suggesting some automated CAPTCHA handling exists in practice. However, there is no first-party documentation explicitly claiming automatic CAPTCHA bypass or login-wall traversal, and community sentiment flags cost/reliability friction rather than seamless unattended operation. Missing for 10: explicit vendor documentation of CAPTCHA-solving/login automation, and independent hands-on confirmation that it reliably completes login flows without manual steps.
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [community] “same setup here for news pages. tier 3 is where my money went, 320 solves a day and 10gb of proxy gone in two days.”
- [community] “Quite useful. Currently we do overpay for the services [referring to Firecrawl-like scraping services].”
data-engineerAutomatically retry through a chain of different proxies when anti-bot detection blocks a request
weight 2 · round drawnFirecrawlnone0/10No evidence describes proxy rotation or anti-bot retry chains; the only relevant community comment explicitly states Firecrawl lacks a proxy service, which is core to bypassing anti-bot blocks.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
developerUse an undetected browser mode to bypass sophisticated bot detection systems
weight 3 · round drawnFirecrawlnone0/10Evidence mentions a 'Browser Sandbox' for managed browser sessions and general scraping/crawling features, but there is no documentation or claim of a stealth/undetected browser mode specifically designed to bypass sophisticated bot detection. Community comments (e.g., proxy tiers, captcha solves) hint indirectly at anti-bot infrastructure but do not confirm an official 'undetected mode' feature. missing for 10: explicit stealth/undetected browser mode docs, technical details on bypassing bot detection (fingerprint spoofing, TLS/JA3 randomization, etc.), independent verification of bypass success.
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [community] “same setup here for news pages. tier 3 is where my money went, 320 solves a day and 10gb of proxy gone in two days.”
Riveternone0/10No evidence mentions undetected browser mode, bot-detection bypass, proxies, or stealth automation features; Riveter's evidence only covers enrichment, scraping, and search tooling. Missing for 10: any mention of anti-bot/stealth browser capabilities, CAPTCHA handling, or evasion of bot detection.
Proxy rotation
developerRequest a proxy from a specific country to get geolocation-appropriate content
weight 2 · round drawnFirecrawlnone0/10No evidence in the pack shows Firecrawl offering country-specific or geolocation proxy selection; in fact a community comment explicitly states Firecrawl lacks a proxy service entirely, and no docs or GitHub references mention proxy/geolocation features.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
developerUse premium residential or datacenter proxies to bypass sites that are hard to scrape
weight 3 · round drawnFirecrawlnone0/10The evidence pack contains no vendor documentation mentioning residential or datacenter proxy support; in fact a community source explicitly states 'Firecrawl... don't have proxy service which is the heart of any crawler and scraper' (firecrawl-comm-3). No official docs or GitHub features reference proxy rotation, IP pools, or anti-bot proxy tiers.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
developerRoute requests through a rotating pool of proxy IPs to avoid blocks
weight 3 · round drawnFirecrawlnone0/10No first-party documentation or GitHub evidence claims a rotating proxy pool feature; in fact community commentary explicitly states Firecrawl 'don't have proxy service which is the heart of any crawler and scraper.' Without vendor claims to dispute, this is simply unevidenced.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
developerRoute multiple requests through the same proxy IP using a session identifier to maintain a consistent identity
weight 2 · round drawnFirecrawlnone0/10No evidence that Firecrawl exposes a session-identifier parameter to pin requests to the same proxy IP; the closest evidence is a community comment stating Firecrawl lacks its own proxy service entirely, which undercuts rather than supports this specific anti-bot capability.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round drawnFirecrawl explicitly supports bulk operations at scale: crawling entire websites, scraping thousands of URLs asynchronously, batch discovery of URLs, and async webhook delivery for large jobs. This directly matches an AI-native user's need to operate across many items at once. Missing for 10: independent hands-on benchmarks validating throughput/reliability at scale and more detail on rate limits/error handling for bulk jobs.
- [github] “Crawl an entire website and get content from all pages.”
- [github] “Discover all URLs on a website instantly.”
- [github] “Scrape thousands of URLs asynchronously”
- [claimed-docs] “Webhooks Async event delivery”
Riveter's core enrichment model operates on many rows at once (bulk input data with AI-filled columns), supports batch generation from a prompt/spec, scheduling for ongoing refresh, and examples like pulling every dentist from every practice in a city in one request. Missing for 10: independent/hands-on verification of large-scale bulk runs and no explicit documentation of per-run item limits or throughput benchmarks.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “It can find every dental practice in a city, then pull every dentist from each one, in a single request.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to RiveterFirecrawl offers webhooks for async event delivery (e.g., notifying when a crawl job completes), which is the only automation-adjacent capability in the evidence; there's no documented rule-definition engine or conditional trigger system for defining custom actions on events. Missing for 10: a rules/trigger engine, conditional logic, or action-chaining beyond simple webhook notifications, and any independent confirmation of automation depth.
- [claimed-docs] “Webhooks Async event delivery”
Riveter supports scheduled refresh of projects (time-based automation) and webhook events (run.completed/stopped/finished) that can notify external systems, giving some automation-on-events capability, but there is no evidence of a rules/condition engine that lets users define arbitrary triggers (e.g., 'if data matches X, then do Y') beyond scheduling and run-completion notifications. missing for 10: conditional rule definitions, event-driven branching logic, multi-condition triggers, and any UI/API for building custom automations beyond schedule+webhook.
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to RiveterFirecrawlnone0/10Firecrawl offers webhooks for async event delivery and async crawling/scraping, but there is no evidence of a scheduler or recurring-job/workflow feature (e.g., cron-based crawls or scheduled scrape jobs).
Docs state you can 'schedule any project to monitor for changes and keep your data fresh' and refresh as often as every minute, indicating recurring job/workflow scheduling support. However, details are thin — no documentation on schedule configuration (cron-like syntax, timezone, pause/resume), no UI/API endpoint specifics for managing schedules, and no independent or hands-on corroboration. Missing for 10: scheduling API/UI details, configuration options, independent verification of reliability at scale.
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnFirecrawlnone0/10Firecrawl is a web scraping/extraction API and toolset; there is no evidence of automation versioning, review workflows, or rollback capabilities for crawl/scrape configurations or workflows in any of the docs, GitHub, or community sources.
Riveternone0/10No evidence of version history, review workflows, or rollback capability for automations/enrichments; the pack only covers run execution, credit control, and data enrichment features. Missing for 10: versioning of automation configs, review/approval workflow, rollback/undo mechanism.
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Collaboration
developerShare scrapers with teammates and manage organizations and role-based permissions
weight 2 · round drawnFirecrawlnone0/10No evidence pack items mention team collaboration, organizations, workspaces, or role-based access control for sharing scrapers; documentation focuses on scraping, extraction, CLI, and MCP features only.
Deployment flexibility
developerBuild and deploy custom serverless scraping scripts on the platform without managing my own infrastructure
weight 2 · round to RiveterFirecrawlnone0/10Firecrawl's evidence shows a fixed API/SDK/CLI for scraping, crawling, extracting, and search, plus webhooks and an MCP server — but nothing about writing and deploying custom serverless scripts or actor-style code that runs on Firecrawl's own infrastructure (unlike platforms such as Apify Actors). No docs, GitHub, or community evidence mentions custom script deployment or a functions/actors runtime.
Riveter's docs show fully managed, serverless-style capabilities (enrichments, scrapes, quick_search, search_agent) that developers configure via natural-language prompts or structured specs and trigger via API/SDK/webhooks with no server management (riveter-docs-1,2,3,4,5,6,11,12). However, this is closer to configuring built-in AI-driven tools than deploying arbitrary custom scraping code/scripts — there's no evidence of a code-upload or custom-script execution environment. Missing for 10: evidence of arbitrary custom code/script deployment (vs. prompt/spec-based enrichment configuration), and independent confirmation of the serverless execution model.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
developerDeploy the scraping service via a Docker container for production use
weight 2 · round drawnFirecrawlnone0/10The evidence confirms Firecrawl is open source (AGPL-3.0) and self-hostable, but no citation mentions Docker, docker-compose, or containerized deployment instructions for production use.
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
developerSelf-host an open-source version of the scraper instead of relying on a hosted cloud service
weight 2 · round to FirecrawlFirecrawl is explicitly confirmed open source under AGPL-3.0 with the cloud version noted as having 'additional features', confirming self-hosting is possible but with reduced functionality (firecrawl-gh-6). Community commentary corroborates this, noting the self-hosted version lacks the proxy service considered 'the heart' of a scraper and other missing capabilities like screenshots (firecrawl-comm-3, firecrawl-comm-4). Missing for 10: first-party self-hosting setup/docker docs, explicit feature-parity comparison, and independent hands-on confirmation of a smooth self-host deployment experience.
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
- [community] “Interesting... Looks like it would be good for RAG. Maybe add Ollama support for local hosting?”
Riveternone0/10Riveter is presented as a hosted API/service (with a local MCP connector for client access to the remote service), but there is no evidence of an open-source, self-hostable version of the scraper itself; docs only describe running a local MCP bridge that still relies on the remote API key.
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
Integrations
developerConnect the scraping API to no-code automation platforms like n8n or Zapier through a prebuilt connector
weight 2 · round drawnFirecrawlnone0/10No evidence of a prebuilt n8n or Zapier connector; docs mention MCP server, CLI, SDKs, and webhooks but nothing about no-code automation platform integrations.
Library compatibility
developerBuild scrapers using popular open-source automation libraries like Playwright, Puppeteer, Selenium, or Scrapy
weight 2 · round drawnFirecrawlnone0/10Firecrawl is a hosted scraping/crawling API with its own primitives (scrape, crawl, extract, browser sandbox) rather than a framework for developers to write Playwright/Puppeteer/Selenium/Scrapy scripts; there is no documented support for plugging in or building on these open-source libraries. A community comment even notes Firecrawl internally uses Puppeteer (not user-selectable) and lacks the openness these libraries provide, contradicting any claim of multi-library dev flexibility.
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
Migration lock in
developerExport my scraped data and job configurations in a portable format to migrate to another provider without lock-in
weight 3 · round to FirecrawlFirecrawl's outputs (markdown/HTML/structured JSON) are inherently portable formats, and its open-source AGPL-3.0 license means self-hosting/forking is possible, reducing lock-in — but there is no documented feature for exporting job configurations, crawl settings, or webhooks setups for migration to another provider. missing for 10: explicit job-configuration export/import tooling, migration guides, or documented data-portability features beyond raw scrape output formats.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
- [community] “Finally, people starting to realize that AGPL means you can just fork and remove everything you don't like (including branding).”
Quickstart
developerPublish my custom scraper to a public marketplace and earn revenue when others use it
weight 1 · round drawnFirecrawlnone0/10The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)
developerRun a ready-made scraper from a marketplace instead of building one from scratch
weight 2 · round drawnFirecrawlnone0/10No evidence of a marketplace of ready-made scrapers/templates that developers can pick up and run; Firecrawl's evidence covers building scraping/crawling calls via API, CLI, MCP, and SDKs, not a curated marketplace of pre-built scrapers.
developerStart building immediately using a library of ready-made project templates
weight 1 · round drawnFirecrawlnone0/10Evidence shows CLI, SDKs, MCP server, and API docs, but nothing about a library of ready-made project templates or starter projects to jumpstart development.
Extraction quality — how faithfully content is extracted — structure, fidelity, edge casesExtraction quality
How faithfully content is extracted — structure, fidelity, edge cases
Ai extraction
developerExtract structured data from a page using natural language instructions instead of writing selectors
weight 3 · round to RiveterFirecrawl's Extract feature lets developers get structured JSON via schemas and its agent can be described in natural language to find/retrieve content without URLs, but the evidence pack shows schema-based extraction more than fully free-form natural-language field extraction replacing selectors. Missing for 10: explicit documentation of prompt-only (no schema) extraction, and independent hands-on confirmation of extraction accuracy.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Use a schema to get structured data:”
- [github] “Describe what you need, and our AI agent searches, navigates, and retrieves it. No URLs required.”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
Riveter's core enrichment feature lets developers build extraction jobs from a natural-language prompt with target attributes instead of writing selectors, and AI agents interpret pages semantically so configs survive redesigns, directly matching the story. missing for 10: independent/hands-on verification of extraction accuracy and no live API schema (openapi/llms.txt probes 404) to confirm behavior beyond vendor docs.
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
developerPass a JSON schema so the API returns structured data matching that schema
weight 2 · round to FirecrawlFirecrawl's docs and GitHub explicitly advertise passing a JSON schema to extract structured data ("Use a schema to get structured data") and general structured JSON extraction from URLs, PDFs, and other formats. Missing for 10: independent/hands-on confirmation of schema-conformance accuracy and edge-case handling beyond vendor docs.
- [github] “Use a schema to get structured data:”
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
Riveter lets you define enrichments via a natural-language prompt or a 'structured spec' with named attributes/columns (riveter-docs-2, riveter-docs-12), which produces structured output, but there is no documented mechanism for passing an arbitrary JSON Schema that the API validates/returns against. missing for 10: explicit JSON Schema input parameter, schema validation of output, and any example showing schema-conformant responses.
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
ai-native userHave an LLM read a page and decide what structured fields to pull out without pre-written selectors
weight 2 · round to RiveterFirecrawl's docs and GitHub note schema-based structured extraction ("Use a schema to get structured data") and general LLM-driven content extraction to JSON, which aligns with selector-free, LLM-decided field extraction. However, evidence doesn't show prompt-only (schema-less) extraction quality, nor independent verification of how well the LLM infers fields without any schema hints. missing for 10: evidence of extraction working from a pure natural-language prompt without any schema, and independent/hands-on validation of extraction accuracy.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Use a schema to get structured data:”
- [github] “Describe what you need, and our AI agent searches, navigates, and retrieves it. No URLs required.”
Docs describe enrichments where AI agents interpret pages and fill arbitrary attribute columns from a natural-language prompt or structured spec (no selectors), with scraping/search tools feeding an AI agent loop that adapts to page structure and redesigns. This directly matches the story of an LLM reading a page and deciding what fields to extract without pre-written selectors. Missing for 10: independent hands-on verification of extraction accuracy and no example showing the LLM's field-selection reasoning in practice.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
- [claimed-docs] “It can find every dental practice in a city, then pull every dentist from each one, in a single request.”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
developerPlug in a local or self-hosted LLM as the extraction backend instead of a cloud-only model
weight 2 · round drawnFirecrawlnone0/10No evidence that Firecrawl allows swapping in a local or self-hosted LLM as the extraction backend; a community comment even suggests adding Ollama support as a future wish, implying it isn't currently offered.
- [community] “Interesting... Looks like it would be good for RAG. Maybe add Ollama support for local hosting?”
Riveternone0/10No evidence anywhere in the docs suggests Riveter allows swapping in a local or self-hosted LLM as the extraction engine; the product is presented as a cloud-only enrichment/extraction service with API keys, credits, and hosted agents. Missing for 10: any mention of local model support, self-hosted backend configuration, or BYO-model options.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
Basic scraping
developerScrape a web page with a single API call and get its raw HTML back
weight 3 · round to FirecrawlFirst-party docs explicitly state that Firecrawl's scrape endpoint extracts content from any URL as markdown, HTML, or structured JSON in a single call, directly matching the story. A community comment raises a narrow caveat about HTML not being returned in a separate 'daemon mode', but this does not contradict the main scrape API. Missing for 10: independent hands-on confirmation of raw HTML output quality/fidelity for the primary scrape endpoint.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Riveternone0/10Riveter's scrape endpoint explicitly returns 'easily parseable text' from a URL, not raw HTML — the opposite of what this story asks for, and no evidence shows an option to retrieve unprocessed HTML.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
Data safety
data-engineerAutomatically detect and filter personally identifiable information out of scraped content before it reaches storage
weight 2 · round drawnFirecrawlnone0/10No evidence in the pack mentions PII detection, redaction, or filtering capabilities; Firecrawl's documented features cover scraping, extraction, crawling, and structured output but nothing about privacy/PII compliance controls.
Riveternone0/10No evidence anywhere in the pack mentions PII detection, filtering, redaction, or compliance controls for scraped/enriched data; Riveter's documented features cover scraping, enrichment, search, and workflow orchestration but nothing about identifying or removing personal data before storage.
Document extraction
data-engineerExtract text content from PDFs, Word, Excel, and PowerPoint files without hosting them myself
weight 2 · round to FirecrawlFirecrawl explicitly documents converting local PDFs, DOCX, XLSX, HTML and more into Markdown or structured JSON as a hosted (cloud) service, directly matching the story of extracting text from PDFs/Word/Excel/PowerPoint without self-hosting. Missing for 10: explicit mention of PowerPoint (.pptx) support and independent hands-on confirmation of file-parsing quality/accuracy.
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
Riveter is delivered as a hosted API/SaaS (no self-hosting required) and docs state it 'reads PDFs and images' as part of enrichment workflows, but there is no evidence it extracts text from Word, Excel, or PowerPoint files specifically. missing for 10: explicit support for .docx/.xlsx/.pptx extraction, any extraction-quality benchmarks or examples for Office file formats.
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
Multimodal extraction
ai-native userGet automatic captions for images on a page so a text-only model can reason about visual content
weight 2 · round to RiveterFirecrawlnone0/10No evidence Firecrawl generates automatic image captions or alt-text descriptions for visual content; evidence only covers text/HTML/markdown extraction, crawling, and structured data extraction.
Riveter's docs mention it 'reads PDFs and images' and combines results with web data (riveter-docs-18), implying some visual-content ingestion, but there is no explicit description of generating captions or text descriptions of images for downstream reasoning by a text-only model. Missing for 10: explicit captioning/description output format, example enrichment showing image-to-text extraction, and any confirmation this text is usable standalone by a text-only model.
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
Search integration
developerSearch the web and get full page content from results in a single call instead of just links and snippets
weight 3 · round to FirecrawlFirecrawl's docs explicitly advertise a search endpoint that returns full page content from results in one call, matching the story exactly, and this is backed by broader scrape/extract capabilities showing it can fetch full markdown/HTML/structured content rather than just snippets. Missing for 10: independent hands-on verification of the search+content endpoint specifically (community evidence discusses scraping/crawling generally but not this exact combined search feature).
- [claimed-docs] “Search the web and get full page content from results in one call.”
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Describe what you need, and our AI agent searches, navigates, and retrieves it. No URLs required.”
Riveternone0/10Riveter's quick_search explicitly returns only urls, titles, and snippets (not full page content), and its scrape tool requires a specific URL rather than combining search+content in one call. search_agent returns a single synthesized answer, not full page content per search result, so no evidenced single-call capability matches the story's exact requirement.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
Selector extraction
developerExtract specific fields from a page using CSS or XPath selector rules
weight 3 · round drawnFirecrawlnone0/10Evidence shows Firecrawl's extraction relies on schema-based/LLM extraction (firecrawl-gh-7) and general markdown/HTML/JSON output (firecrawl-docs-1), but nothing in the pack documents CSS or XPath selector-based field extraction rules. Missing for 10: any mention of CSS selector or XPath rule support in scrape/extract config, docs page confirming selector-based extraction, or independent confirmation of this capability.
- [github] “Use a schema to get structured data:”
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
Riveternone0/10Riveter's docs describe AI-driven page interpretation and scraping (turning URLs into parseable text, agents reading pages 'the way a person would') rather than CSS/XPath selector rules; no evidence pack item mentions selector-based extraction at all, and one item explicitly frames the AI approach as an alternative to fragile configuration that would break on redesign, which is the kind of setup selectors typically require.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
Structured data handling
data-engineerExtract data from very large tables using intelligent chunking so it fits within processing limits
weight 1 · round drawnFirecrawlnone0/10No evidence pack items mention table extraction, large-table handling, or intelligent chunking strategies for oversized data; the evidence only covers general scraping, crawling, and structured extraction features. missing for 10: any mention of table-specific extraction, chunking mechanisms, or handling of oversized documents/tables to fit token/processing limits.
Riveternone0/10Riveter's evidence covers enrichment, scraping, search, and workflow automation, but there is no mention of chunking large tables, row batching, pagination for extraction limits, or handling of very large datasets to fit processing constraints. missing for 10: any mention of chunking strategy, table size limits, batching large extractions, or row-splitting logic.
Js rendering — handling JavaScript-heavy pages — rendering, waiting, dynamic contentJs rendering
Handling JavaScript-heavy pages — rendering, waiting, dynamic content
Headless rendering
developerRender JavaScript-heavy single-page applications and get the fully rendered HTML
weight 3 · round to FirecrawlDocs confirm Firecrawl scrapes pages with an actual browser session ('Browser Sandbox... managed browser sessions for interactive workflows', 'click, fill forms, extract dynamic content'), and community evidence confirms it uses a real headless browser (Puppeteer) to render pages rather than static HTTP fetch, which supports JS-heavy SPA rendering. Output can be returned as HTML per docs-1. Missing for 10: independent benchmark/proof of correctly rendering complex SPAs, and community notes it uses Puppeteer not Playwright with some limitations in certain modes.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
Riveternone0/10Riveter is a data-enrichment/scraping/AI-agent tool focused on turning URLs into text and filling data columns; there is no evidence it renders JS-heavy SPAs into fully rendered HTML (e.g., headless browser rendering, DOM snapshot output). The 'scrape' feature converts URLs to 'easily parseable text', not full rendered HTML, so this capability is unevidenced.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
developerHave the API wait for a specific selector to appear before returning the rendered page
weight 2 · round drawnFirecrawlnone0/10No evidence pack item mentions waiting for a specific CSS selector before returning rendered content; only general mentions of scraping, interactive actions, and browser sandboxing are present without detail on selector-based wait conditions.
Interactive automation
developerAccess a managed remote browser sandbox for interactive, manual browsing workflows
weight 2 · round to FirecrawlFirecrawl docs explicitly mention a 'Browser Sandbox' offering managed browser sessions for interactive workflows, plus 'scrape, then keep working with it: click, fill forms, extract dynamic content' — directly matching the story. However, this is only a single doc snippet with no detail on session persistence, remote access UI, or manual/human-driven browsing versus API-driven automation, and no independent/community corroboration of this specific feature. Missing for 10: detailed documentation on session duration/access model, evidence of true manual/interactive human use (vs agent-driven), and third-party confirmation.
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
developerKeep interacting with an already-scraped page, clicking and filling forms to reach content behind a login wall
weight 2 · round to FirecrawlFirecrawl's docs explicitly describe an interactive workflow — 'Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper' — plus a 'Browser Sandbox' for managed interactive browser sessions, directly matching the story. Missing for 10: independent/hands-on corroboration that clicking/filling forms actually reaches login-walled content, and more detail on session persistence across interactions.
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
developerScript page interactions like clicking, filling inputs, and scrolling before content is returned
weight 3 · round to FirecrawlFirecrawl's docs explicitly describe scripting page interactions—click, fill forms, extract dynamic content, navigate deeper—after an initial scrape, and mention a managed Browser Sandbox for interactive workflows, directly matching the story of clicking/filling/scrolling before content is returned. Missing for 10: detailed API reference for the specific 'actions' parameter (e.g. scroll behavior), and independent/hands-on confirmation from community sources that these interaction primitives work reliably in practice.
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
Render configuration
developerControl the browser viewport width and height when rendering a page
weight 1 · round drawnFirecrawlnone0/10No evidence in the pack mentions viewport width/height, mobile emulation, or screen size configuration for rendering pages; the docs mention scraping, actions, and a browser sandbox but nothing about viewport control.
Session persistence
developerPass my own session cookies so the API fetches pages requiring authentication
weight 2 · round drawnFirecrawlnone0/10No evidence pack item mentions passing custom cookies, headers, or session/auth tokens to Firecrawl's scrape API; only generic scraping, crawling, and browser-sandbox features are documented.
developerReuse a persistent browser profile with saved cookies and login state across multiple requests
weight 2 · round drawnFirecrawlnone0/10The evidence mentions a 'Browser Sandbox' for managed sessions and interactive workflows, but nothing describes persisting cookies/login state or reusing a browser profile across multiple separate requests. No docs, SDK, or community evidence confirms this capability.
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to FirecrawlFirecrawl is fundamentally API-first — scrape, crawl, extract, search, and structured data features are all exposed via API/SDKs and docs, and there is no evidence of a rich standalone UI with capabilities withheld from the API. However, the evidence pack lacks a discoverable OpenAPI spec (probe found 404s) and does not explicitly confirm dashboard-only features (e.g., billing, team management, job monitoring) are also API-accessible. missing for 10: a published OpenAPI/swagger spec, explicit confirmation that all dashboard/UI-only functions (usage analytics, team/billing management, job history) are API-reachable, and independent verification of full UI/API parity.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Search the web and get full page content from results in one call.”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [github] “Crawl an entire website and get content from all pages.”
- [github] “Scrape thousands of URLs asynchronously”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.firecrawl.dev/openapi.json, https://docs.firecrawl.dev/swagger.json, https://docs.firec…”
- [probe] “official CLI documented at https://docs.firecrawl.dev/sdks/cli”
Docs show many core capabilities (building enrichments via prompt/spec, scraping, quick_search, search_agent, webhooks, dry_run) are all API-accessible, suggesting broad parity, but there is no explicit statement of full UI/API parity and some UI-highlighted features like scheduling refresh (riveter-docs-13, riveter-docs-17) aren't confirmed as API-exposed. Additionally, probes show no discoverable OpenAPI spec (riveter-probe-2) or llms.txt (riveter-probe-1), undermining confidence that the API surface is fully documented/openly specified. missing for 10: explicit parity statement, API access to scheduling/monitoring feature, published OpenAPI spec for verification.
- [claimed-docs] “You can build one from a natural-language prompt or a structured spec, and Riveter will generate the rows for you.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
- [probe] “PROBE llms.txt: HTTP 404 at https://docs.riveterhq.com/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://docs.riveterhq.com/openapi.json, https://docs.riveterhq.com/swagger.json, https://docs.rivet…”
ai-native userExport all of my data in open formats and leave
weight 3 · round to FirecrawlFirecrawl outputs are natively in open formats (markdown, HTML, structured JSON) and the core engine is open source (AGPL-3.0), letting a user self-host and avoid lock-in to the hosted service. However there's no explicit 'export all your account/config data' feature documented, and community notes only touch on forking rights, not a formal data-export path. Missing for 10: a documented account-data export/migration flow, evidence of exporting crawl history/settings, and independent confirmation users have actually migrated off the hosted service.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
- [community] “Finally, people starting to realize that AGPL means you can just fork and remove everything you don't like (including branding).”
Riveternone0/10No evidence of any data export feature or open-format export capability; the evidence only covers enrichment, scraping, search, and API integration features, with no mention of exporting data or portability guarantees. missing for 10: export functionality documentation, supported open formats (CSV/JSON/etc), any data-portability or account-closure workflow.
ai-native userRead the product's source under an open license
weight 2 · round to FirecrawlFirecrawl's GitHub repo confirms it is open source under the AGPL-3.0 license, with community discussion also confirming this (including implications of AGPL forking rights). Source is publicly readable on GitHub with an OSI-approved-family open license. Missing for 10: no evidence of clarity on which parts of the cloud-only features are excluded from the open license, and no independent audit of full repo completeness.
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
- [community] “Finally, people starting to realize that AGPL means you can just fork and remove everything you don't like (including branding).”
ai-native userSelf-host the core product
weight 3 · round to FirecrawlFirecrawl's GitHub repo confirms the core product is open source under AGPL-3.0 and can be self-hosted, with the hosted cloud version offering extra features (firecrawl-gh-6). However, community reports note self-hosted/simple versions lack key production features like proxy support and have functional limitations (e.g., daemon mode restrictions, no HTML return) compared to the cloud offering (firecrawl-comm-3, firecrawl-comm-4). Missing for 10: official self-hosting setup docs/guide in the evidence pack, and confirmation that self-hosted deployment achieves full feature parity with the hosted service.
- [github] “Firecrawl is open source under the AGPL-3.0 license. The cloud version at firecrawl.dev includes additional features”
- [community] “As I see, you use Puppeteer, not Playwright. Also, both Firecrawl and Firecrawl Simple are really simple, and most importantly don't have pr…”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Riveternone0/10Riveter is presented as a hosted API/SaaS product (with local MCP server option only for connecting AI clients, not for self-hosting the core enrichment engine); no evidence of open-source code, self-hosting instructions, or a downloadable core product exists in the pack.
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
Output formats — stories about output formats in this arenaOutput formats
Stories about output formats in this arena
Content formats
developerReceive scraped content as clean markdown instead of raw HTML
weight 3 · round to FirecrawlFirst-party docs explicitly state extraction as markdown (alongside HTML/JSON) and support converting local files to markdown, confirming clean markdown output is a core, well-documented feature. Missing for 10: independent hands-on confirmation specifically praising markdown output quality/cleanliness.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
Docs state a scrape 'turns a URL into easily parseable text,' implying cleaned output rather than raw HTML, but there's no explicit mention of markdown formatting or output schema. Missing for 10: explicit confirmation that scrape output is markdown-formatted, example output showing markdown structure, independent verification of output cleanliness.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
developerChoose exactly which output format is returned, such as markdown, HTML, text, or frontmatter
weight 2 · round to FirecrawlDocs confirm output as markdown, HTML, or structured JSON (firecrawl-docs-1, firecrawl-docs-6), and a community comment notes a daemon-mode limitation where HTML return is unsupported in some contexts, suggesting partial reliability. No explicit mention of 'frontmatter' or 'text' formats, and no documentation snippet showing a format-selection parameter/API example. Missing for 10: explicit mention of frontmatter/text format options, and a documented parameter/example showing developers selecting formats.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
developerReceive scraped content as structured JSON
weight 3 · round to FirecrawlFirecrawl docs explicitly support extracting content as structured JSON, including with a defined schema, alongside markdown/HTML options, and this extends to document formats like PDFs/DOCX as well. Missing for 10: independent hands-on confirmation of JSON output quality/schema fidelity beyond vendor docs and GitHub README.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [github] “Use a schema to get structured data:”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
Riveter's enrichments and scrapes explicitly return structured, parseable data (columns, urls/titles/snippets, webhook payloads of 'full results'), and SDK examples show structured attribute objects returned from calls, indicating outputs are consumable as structured JSON rather than raw text. missing for 10: an explicit statement of JSON schema/response format in docs, and independent/hands-on confirmation of the JSON structure (API docs endpoints 404 in probes).
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “run, err := client.Enrich(ctx, riveter.EnrichParams{ Prompt: "Research each company", Attributes: []string{"CEO", "Employee Count"},”
Llm ready output
ai-native userGet clean LLM-ready text directly instead of dealing with blocking, rendering, and messy HTML myself
weight 3 · round to FirecrawlFirecrawl's core value proposition is turning any URL into clean markdown/structured JSON, handling rendering, JS-heavy pages, and blocking via a managed browser sandbox, explicitly for LLM/RAG use cases. Docs and GitHub confirm markdown/HTML/JSON extraction, PDF/DOCX conversion, and managed browser sessions abstracting away rendering complexity, though community comments note some limitations (e.g., proxy/anti-bot gaps, missing HTML in some modes). Missing for 10: independent benchmark of output cleanliness vs raw HTML scraping, and resolution of community-reported edge-case limitations (daemon mode HTML issue).
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Scrape a page, then keep working with it: click, fill forms, extract dynamic content, or navigate deeper.”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
- [claimed-docs] “Browser Sandbox Managed browser sessions for interactive workflows”
- [github] “Crawl an entire website and get content from all pages.”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Docs claim a scrape converts any URL into 'easily parseable text' and that AI agents interpret pages 'the way a person would', directly addressing the ask for clean, LLM-ready text instead of raw HTML. However, all evidence is vendor documentation with no independent hands-on verification of output cleanliness, no example output shown, and no explicit mention of handling JS rendering/blocking obstacles beyond the general claim. Missing for 10: independent corroboration of scrape text quality, concrete example output, and explicit handling of anti-bot/rendering blockers.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “Riveter uses AI agents that interpret pages the way a person would, so the same configuration keeps working after a redesign.”
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
ai-native userRequest semantically chunked output instead of one large content blob, so it feeds cleanly into a retrieval pipeline
weight 2 · round drawnFirecrawlnone0/10Firecrawl's evidence covers markdown/HTML/structured JSON extraction, crawling, and PDF/DOCX conversion, but nothing describes a semantic chunking feature or chunked output mode for retrieval pipelines. The axis applies (chunked output is a plausible feature for a scraping/RAG-prep tool) but no evidence shows it exists.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [claimed-docs] “Turn local PDFs, DOCX, XLSX, HTML, and more into Markdown or structured JSON”
- [github] “Use a schema to get structured data:”
Riveternone0/10Riveter's evidence describes enrichments, scrapes, searches, and structured row outputs, but nothing indicates a semantic-chunking output mode designed for retrieval pipelines (e.g., configurable chunk size/overlap, chunk metadata). Structured rows/columns are not the same as semantic chunking for RAG ingestion, and no such feature is documented.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
Visual capture
developerCapture a screenshot of a full page or a specific selected area
weight 2 · round drawnFirecrawlnone0/10The evidence pack lists output formats as markdown/HTML/JSON but never mentions screenshot capture, full-page or selector-based, as a capability. A community comment even raises it as an open question ('does it support screenshots?') without confirmation, so there's no evidence the capability exists.
- [claimed-docs] “Extract content from any URL as markdown, HTML, or structured JSON”
- [community] “It being vibe coded aside, does it support screenshots? I noticed the daemon mode has a lot of weird limitations too like not being able to …”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Cost optimization
developerLet the API automatically pick the cheapest configuration that still succeeds
weight 2 · round drawnFirecrawlnone0/10No evidence of automatic cost-optimal configuration selection; docs mention manual controls like reasoning effort but nothing about the API choosing cheapest successful config automatically.
Riveternone0/10Riveter offers cost controls like dry_run estimates and max_credits caps that refuse overpriced requests, but there is no evidence the API automatically searches for or selects the cheapest configuration that still succeeds — it only estimates/caps, it doesn't auto-optimize. Missing for 10: any documentation of automatic configuration search/optimization for cost, fallback logic that retries cheaper options, or an API parameter that lets Riveter choose the minimal successful config itself.
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
developerBlock ads on the target page to speed up scraping requests
weight 1 · round drawnFirecrawlnone0/10No evidence pack item mentions ad-blocking or any option to strip ads/trackers on target pages to speed up scraping; only general scraping, crawling, and extraction features are documented.
developerBlock images and CSS resources by default to reduce bandwidth and speed up requests
weight 1 · round drawnFirecrawlnone0/10No evidence pack mentions blocking images or CSS resources, resource-type filtering, or bandwidth-saving scrape options; only general scraping/crawling features are documented.
ai-native userSet how much reasoning effort an autonomous agent spends on a data-gathering task (low, medium, high)
weight 2 · round to FirecrawlGitHub README explicitly states the agent lets users 'set how much reasoning the agent spends on the task,' directly matching the story, but there's no detailed documentation confirming discrete low/medium/high levels or pricing-tied reasoning-effort controls. Missing for 10: first-party docs specifying the exact reasoning-effort parameter/levels, independent confirmation of how this affects cost/limits.
Cost transparency
developerWhether exceeding my plan's monthly credit or request quota triggers overage charges or a hard cutoff
weight 3 · round drawnFirecrawlnone0/10No evidence in the pack addresses billing behavior when a plan's credit/request quota is exceeded — nothing on overage charges vs. hard cutoffs. This is a fair pricing question for a paid API product, so absence of evidence yields none. Missing for 10: any pricing/billing docs describing quota overage policy, hard-stop vs auto-billing behavior, or community reports confirming either.
Riveternone0/10The evidence describes credit estimation, dry_run, and max_credits cap that refuses requests at 422 before charging, but there is no mention of plan-level monthly credit/request quotas, nor whether exceeding them triggers overage billing or a hard cutoff. missing for 10: any documentation of monthly plan quotas, overage billing policy, or hard-cutoff behavior when a subscription limit is exceeded.
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
developerWhether failed, blocked, or empty-result requests still consume my billing quota
weight 2 · round drawnFirecrawlnone0/10No evidence pack items discuss billing/credit treatment for failed, blocked, or empty-result requests; documentation snippets cover features (scrape, crawl, MCP, webhooks) but not quota/credit consumption policy.
Riveternone0/10The docs describe dry_run cost estimation and max_credits caps that prevent overage, but nothing states whether a failed, blocked, or empty-result run still consumes credits. Missing for 10: explicit policy on billing for failed/empty/blocked runs, any refund or non-charge guarantee for zero-result enrichments.
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
developerSet a spending cap or usage alert so proxy/credit consumption doesn't silently blow past my budget
weight 3 · round to RiveterFirecrawlnone0/10No evidence of spending caps, budget alerts, or usage-limit notifications; community comments even describe unexpectedly high consumption ('10gb of proxy gone in two days') with no mention of a cap/alert mechanism to prevent overage.
- [community] “same setup here for news pages. tier 3 is where my money went, 320 solves a day and 10gb of proxy gone in two days.”
- [community] “Quite useful. Currently we do overpay for the services [referring to Firecrawl-like scraping services].”
Riveter offers per-request cost control via dry_run (credit estimate before charging) and max_credits (hard ceiling that returns 422 credit_cap_exceeded with nothing charged), which directly prevents a single run from blowing past a set budget. However, there's no evidence of an account-wide spending cap, recurring usage alerts, or a dashboard/notification system for cumulative consumption across runs. Missing for 10: account/org-level budget cap, proactive usage alerts/notifications, historical spend tracking dashboard.
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
Performance tuning
developerTrade off latency against completeness by controlling exactly when content is returned
weight 1 · round to RiveterFirecrawl offers async webhooks for event delivery and an agent 'reasoning effort' setting that trades speed for thoroughness, plus async bulk scraping — all of which let a developer influence when/how much content comes back, but there's no explicit documented parameter (e.g., wait-time or completeness threshold) framed as a direct latency-vs-completeness control on the standard scrape/crawl endpoints. missing for 10: explicit sync-return timeout/partial-completeness parameter, independent benchmarking of latency vs completeness tradeoffs, and hands-on confirmation of the reasoning-effort knob's effect.
- [github] “Set how much reasoning the agent spends on the task”
- [claimed-docs] “Webhooks Async event delivery”
- [github] “Scrape thousands of URLs asynchronously”
Riveter explicitly exposes multiple latency/completeness tradeoffs: quick_search returns fast synchronous structured snippets, search_agent runs a fuller AI research loop for one question, and full enrichments can be tracked via wait_for_result long-polling or async webhook callbacks — giving a developer direct control over when and how complete the returned content is. missing for 10: no independent/hands-on benchmarks or third-party confirmation of actual latency differences between these modes.
- [claimed-docs] “A quick_search lets you quickly web search a query, and pull structured results with urls, titles, and snippets — synchronously, in one requ…”
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
Plan scale limits
data-engineerThe maximum concurrent sessions or requests allowed on my pricing tier and the cost to raise that cap
weight 2 · round drawnFirecrawlnone0/10No evidence pack item documents rate limits, concurrency caps per pricing tier, or the cost to raise them; only unrelated product feature docs and community anecdotes about usage cost are present. missing for 10: documented per-tier concurrency/request limits, documented pricing to upgrade limits, any rate-limit or quota API reference.
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnFirecrawlnone0/10No evidence in the pack mentions data residency, regional storage options, or compliance controls for where scraped data is processed/stored; the open-source AGPL version could theoretically be self-hosted for residency control, but this is not documented anywhere in the evidence.
Riveternone0/10No evidence in the pack mentions data residency, regional storage options, or compliance controls for where data is stored; the docs focus entirely on enrichment features and API mechanics. Missing for 10: any mention of region selection, data residency options, or storage location controls.
ai-native userControl data retention and deletion
weight 2 · round drawnFirecrawlnone0/10No evidence pack item mentions data retention policies, deletion controls, or privacy settings for stored crawl/scrape data; this is a fair question for a cloud scraping/data API but no documentation addresses it.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnFirecrawlnone0/10No evidence in the pack mentions telemetry, usage tracking, analytics collection, or an opt-out setting/flag for Firecrawl's CLI, SDK, or self-hosted deployment; while the open-source AGPL nature suggests self-hosting is possible, nothing documents a telemetry toggle or privacy control.
Scale reliability — behavior under load — scaling limits, uptime, failure handlingScale reliability
Behavior under load — scaling limits, uptime, failure handling
Ai driven crawling
ai-native userRely on adaptive crawling that automatically stops once enough information has been gathered to answer my query
weight 2 · round to RiveterFirecrawlnone0/10The evidence describes crawling, scraping, and AI agent search/reasoning controls (e.g., firecrawl-gh-1, firecrawl-gh-2), but nothing documents adaptive crawling that automatically halts once sufficient information has been gathered to answer a specific query — crawls appear to run to full site discovery or fixed limits rather than stopping based on information sufficiency.
Riveter's search_agent and enrichment agent loop imply some autonomous research process that fills a cell with an AI-researched answer, suggesting the agent decides when it has enough data, but there is no explicit documentation of stopping criteria or adaptive crawling behavior tied to query sufficiency. missing for 10: explicit description of adaptive stopping/crawling logic, evidence of how the agent determines 'enough information', independent confirmation of this behavior in practice.
- [claimed-docs] “A search_agent call asks one question and gets one AI-researched answer back — the same agent loop that fills a single enrichment cell, with…”
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
Batch processing
data-engineerBatch scrape thousands of URLs asynchronously
weight 3 · round to FirecrawlFirecrawl's GitHub docs explicitly advertise batch/async scraping of thousands of URLs, plus webhook-based async event delivery for pipeline integration, and crawl/map endpoints for URL discovery at scale, aligning well with the data-engineer scale story. Missing for 10: independent hands-on benchmarks proving reliability at thousands-of-URL scale and details on rate limits/retry/error handling under batch load.
- [github] “Scrape thousands of URLs asynchronously”
- [claimed-docs] “Webhooks Async event delivery”
- [github] “Crawl an entire website and get content from all pages.”
- [github] “Discover all URLs on a website instantly.”
Riveter's enrichment engine explicitly processes rows of URLs with scraping, runs asynchronously (webhook_url on completion), and SDKs handle retries, long-polling, and pagination — all core pieces for async batch scraping. However, there's no explicit documentation of scale limits, concurrency handling, or a tested example at thousands-of-URLs volume. Missing for 10: explicit large-scale (thousands of URLs) benchmarks or case studies, concurrency/rate-limit guidance for very large batches.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “It can find every dental practice in a city, then pull every dentist from each one, in a single request.”
developerApply different crawl configurations to different URL patterns within a single batch job
weight 1 · round drawnFirecrawlnone0/10The evidence pack covers crawling, scraping, extraction, webhooks, and CLI/MCP features, but nothing describes per-URL-pattern configuration overrides within a single crawl/batch job (e.g., different scrape options for different path patterns). No docs or community evidence mention such rule-based configuration.
Riveternone0/10No evidence describes applying different crawl configurations per URL pattern within one batch/enrichment job; docs mention scraping, searching, and enrichment generally but not per-pattern configuration rules. missing for 10: any mention of per-URL-pattern rules or configuration scoping within a single job, examples or docs showing mixed crawl settings in one batch.
Concurrency
data-engineerSpin up many concurrent scraping sessions to gather data at scale
weight 3 · round to FirecrawlFirecrawl explicitly supports scraping 'thousands of URLs asynchronously' and full-site crawling with async webhooks for event delivery, which supports scaling to many concurrent scrape jobs. However, there is no documentation of concurrency limits, session management, or dedicated infrastructure for spinning up many parallel sessions, and community feedback raises cost/efficiency concerns at scale (proxy usage, cost overpay) without directly disputing the concurrency capability itself. Missing for 10: explicit concurrency/rate-limit documentation, first-party benchmarks or case studies of large-scale concurrent scraping, and independent verification of scale claims.
- [github] “Scrape thousands of URLs asynchronously”
- [github] “Crawl an entire website and get content from all pages.”
- [claimed-docs] “Webhooks Async event delivery”
- [community] “same setup here for news pages. tier 3 is where my money went, 320 solves a day and 10gb of proxy gone in two days.”
- [community] “I made newsagents.app and I ended up using the extract API from kagi and falling back to cloudflare's browser API for problem pages. That lo…”
Riveter's enrichment engine processes many rows in a single run and can chain scrapes/searches (e.g., finding every dental practice then every dentist in one request), implying built-in batch/bulk scraping at scale, and SDKs handle retries/pagination for large jobs. However, there is no explicit documentation of concurrency limits, parallel session management, or throughput guarantees for scraping specifically. Missing for 10: explicit concurrency/session limits, performance benchmarks, and independent evidence of scaling to many simultaneous scrape sessions.
- [claimed-docs] “An enrichment takes rows of input data and fills in new columns using AI, web searches, web scrapes, and other tools.”
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “It can find every dental practice in a city, then pull every dentist from each one, in a single request.”
- [claimed-docs] “the SDKs handle auth, retries (429s and transient GET failures), the wait long-poll, polling until a run finishes (wait_for_result), and pag…”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
Crawl compliance
data-engineerConfigure the crawler to respect robots.txt rules and target-site rate limits automatically
weight 2 · round drawnFirecrawlnone0/10No documentation or evidence describes robots.txt compliance settings or automatic rate-limit throttling; the only related community comment (firecrawl-comm-8) suggests sites must proactively disallow the crawler, which doesn't confirm built-in respect for robots.txt as a configurable, automatic behavior.
- [community] “Excellent, another kind of copyright theft as a service that assumes your site is ripe for scraping unless you disallow yet another agent (F…”
Riveternone0/10No evidence anywhere in the docs mentions robots.txt compliance or rate-limit configuration; the pack only covers scraping features, retries, credits, and MCP integration. This is a fair axis for a web-scraping/crawling product, but absence of evidence means it cannot be credited as delivered.
Fault tolerance
data-engineerResume a crashed deep crawl from a saved checkpoint instead of restarting from scratch
weight 2 · round drawnFirecrawlnone0/10No evidence of checkpointing or resuming crawls from saved state; docs mention crawling, webhooks, and async scraping but nothing about crash recovery or resumable checkpoints.
Operational transparency
data-engineerCheck a public status page showing uptime history and past incident postmortems before committing to the service
weight 2 · round drawnFirecrawlnone0/10No evidence pack item mentions a public status page, uptime history, or incident postmortems for Firecrawl; the docs and community threads cover product features and complaints but nothing about SLA/uptime transparency.
Scheduling monitoring
data-engineerMonitor target pages for content changes, such as price or listing updates, and get notified as they happen
weight 2 · round to RiveterFirecrawlnone0/10The evidence pack shows scraping, crawling, extraction, and webhook-based async event delivery, but no dedicated change-tracking/monitoring feature (e.g., diffing pages over time, price/listing change alerts) is documented anywhere in the pack.
- [claimed-docs] “Webhooks Async event delivery”
- [github] “Crawl an entire website and get content from all pages.”
- [github] “Scrape thousands of URLs asynchronously”
Riveter explicitly supports scheduling projects to monitor for changes, refreshing as often as every minute, and can POST results to a webhook_url when a run finishes, which together deliver change-monitoring plus notification. However, the webhook fires on run completion rather than a dedicated 'content changed' diff event, and there's no independent/hands-on evidence of this workflow in production. Missing for 10: independent corroboration of the schedule+webhook pipeline in practice, and explicit diff/change-detection logic distinguishing 'changed' vs 'unchanged' pages.
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
data-engineerMonitor job performance, validate data quality, and receive alerts when something fails
weight 2 · round to RiveterFirecrawlnone0/10Evidence shows webhooks for async event delivery but nothing about job performance dashboards, data quality validation, or failure alerting mechanisms for a data-engineering monitoring workflow.
Riveter supports webhook alerts on run completion/stop/finish events and scheduled monitoring for data freshness, giving some job-status alerting and monitoring capability, but there is no explicit data-quality validation feature (e.g., schema/anomaly checks) or job performance dashboards described. missing for 10: explicit data quality validation tooling, job performance metrics/dashboard, and independent confirmation of alerting reliability.
- [claimed-docs] “Pass webhook_url in the JSON body when starting a run and Riveter POSTs the full results to your URL when it finishes (events: run.completed…”
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “dry_run: true — validate the request and return a credit estimate without creating or charging anything.”
- [claimed-docs] “max_credits — a ceiling. When the estimate's maximum exceeds it, the request is refused with 422 credit_cap_exceeded and nothing is charged.”
developerMonitor live system metrics and worker/browser pool status through a real-time dashboard
weight 1 · round drawnFirecrawlnone0/10No evidence of a real-time dashboard for monitoring system metrics, worker pool, or browser pool status; evidence only covers scraping/crawling features, CLI, MCP server, and community discussion unrelated to monitoring dashboards.
developerSchedule scraping jobs to run automatically at specific times
weight 2 · round to RiveterFirecrawlnone0/10No evidence of scheduled/cron-based scraping jobs; Firecrawl's evidence covers crawling, scraping, webhooks, and async batch scraping, but nothing about scheduling jobs to run at specific times.
Riveter supports scheduling projects to monitor for changes and refresh data as often as every minute, which implies automatic recurring scraping jobs, but there's no detail on specifying exact times/cron-like scheduling, timezone control, or a documented scheduling API/UI. missing for 10: explicit scheduling configuration details (time-of-day, cron syntax, timezone), independent/hands-on confirmation of scheduling reliability, and API endpoint documentation for creating/managing schedules.
- [claimed-docs] “Schedule any project to monitor for changes and keep your data fresh.”
- [claimed-docs] “For fast moving data like scores or election results, you can refresh as often as every minute.”
Site crawling
data-engineerRun a deep crawl using a breadth-first strategy with a configurable maximum page limit
weight 2 · round to FirecrawlEvidence confirms Firecrawl can crawl an entire website and discover all URLs (firecrawl-gh-3, firecrawl-gh-4), which implies a crawl feature suitable for a data-engineer's bulk scraping needs, but nothing in the pack explicitly documents a breadth-first crawl strategy or a configurable maximum page limit parameter. Missing for 10: explicit mention of BFS traversal mode, documented maxPages/limit parameter, and independent confirmation that these controls work at scale.
developerCrawl an entire website and get content from all its pages with one request
weight 3 · round to FirecrawlGitHub docs explicitly state 'Crawl an entire website and get content from all pages' with supporting features like URL discovery and async scraping of thousands of URLs, directly matching the story. Missing for 10: independent hands-on validation specifically of full-site crawl completeness/reliability at scale (community comments discuss cost/proxy issues but not crawl-completeness failures).
Riveter's docs describe single-URL 'scrape' and 'quick_search' calls, but the marketing example of finding every dental practice in a city and pulling data from each one in a single request shows it can aggregate content across multiple pages/sources in one enrichment run, which approximates whole-site crawling. There is no explicit sitemap-style 'crawl entire website' feature or evidence of full-domain page enumeration. missing for 10: explicit full-site/sitemap crawl feature, evidence of automatically discovering and traversing all pages of a single domain, independent confirmation of multi-page crawl behavior.
- [claimed-docs] “A scrape lets you turn a URL into easily parseable text.”
- [claimed-docs] “It can find every dental practice in a city, then pull every dentist from each one, in a single request.”
- [claimed-docs] “It reads PDFs and images, calls third party APIs as part of a workflow, and combines those results with data pulled from the web in a single…”
developerInstantly discover all URLs on a website without fully crawling it
weight 2 · round to FirecrawlFirecrawl explicitly offers a 'Map' capability described as 'Discover all URLs on a website instantly,' distinct from full crawling, directly matching the story. This is a first-party GitHub claim but lacks independent hands-on corroboration or detail on accuracy/limits at scale. missing for 10: independent/hands-on verification of speed and completeness, documentation of limits on very large sites.
Not comparable on these axes
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · not comparableFirecrawln/aFirecrawl is a web-scraping/crawling API/service that itself exposes an MCP server for other AI tools to consume (firecrawl-docs-5, firecrawl-probe-3); it is not an agent or orchestration platform designed to consume external MCP servers as a client. The evidence pack shows no capability for Firecrawl to import/plug in third-party MCP servers to extend its own tool use, and this is not a fair capability to expect of a scraping API — it's the reverse role.
- [claimed-docs] “MCP Server: Connect Firecrawl to any AI tool via the Model Context Protocol”
- [probe] “official MCP server documented at https://github.com/mendableai/firecrawl-mcp-server”
Riveternone0/10All MCP-related evidence describes Riveter acting as an MCP *server* that other AI assistants (Claude, ChatGPT, Cursor) connect to, not Riveter acting as a client that can plug in and consume external MCP servers' tools. No evidence shows Riveter itself importing/connecting to third-party MCP servers to extend its own tool use.
- [claimed-docs] “Use Riveter from Claude, ChatGPT, Cursor, or any MCP-compatible AI assistant. Pick one of the two ways to connect.”
- [claimed-docs] “Runs on your machine and needs Node.js and an API key. Use it when your client cannot reach remote servers.”
ai-native userPrevent my data from being used to train AI models
weight 3 · not comparableFirecrawlnone0/10No evidence pack item mentions any data-training opt-out, privacy policy on model training use, or data retention/DPA controls for Firecrawl's scraped or user data; the only related community comment raises concerns about scraping others' data, not about protecting the user's own data from AI training use.
Rivetern/aRiveter is a data enrichment/web-scraping tool, not an AI model provider or chat assistant whose data usage/training policies would be relevant here; the evidence contains no mention of AI training data policy or opt-out controls, and this axis is a category error for the product type.