Rank #7 of 8 in Agent Sandboxes & Code Execution
Install
npm i @cloudflare/sandboxShowcase


Products
Cloudflare, product by product →Cloudflare ships more than one product — each judged line competes in its own arena on the same stories as everyone else.
| Line | Arena | Rank | PA Score | Agent-ready |
|---|---|---|---|---|
| Cloudflare Workers | Edge & App Platforms | #2/6 | 34/100 | 63/100 |
| Cloudflare Sandboxesthis page | Agent Sandboxes & Code Execution | #7/8 | 21/100 | 26/100 |
| Cloudflare AI Gateway | Model Gateways & Routers | #5/7 | 19/100 | 22/100 |
Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardwareevidence →
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode executionevidence →
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experienceevidence →
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation securityevidence →
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scaleevidence →
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycleevidence →
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistenceevidence →
Stories about snapshot persistence in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Cloudflare Sandboxes stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Webhooks · Scoped API keys · MCP server · Versioning policy · Official CLI
Subscribe to events via webhooks
—–
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
—–
Download a machine-readable API spec (OpenAPI or equivalent)
~4/10
unlocks → Interactive API docs · MCP server
Rely on versioned APIs with a documented deprecation policy
—–
Test against a sandbox environment without touching production data
✓7/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
~6/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
n/an/a
Operate the product with natural-language commands
—0/10
Plug MCP servers into this product so it can use their tools
—–
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode execution
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation security
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scale
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence
Stories about snapshot persistence in this arena
Sorted by importance (agentic first) (high → low) · 51/51 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Xcommunity | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Xcommunity | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | full | 7/10 | Xcommunity | |
Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds C Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 3 | full | 8/10 | Xcommunity | |
Execute untrusted, AI-generated code without risking my own infrastructure C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 3 | partial | 7/10 | Xcommunity | |
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human C Agent lifecycle | ai-native user | Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience | 3 | partial | 7/10 | Xcommunity | |
Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems C Isolation | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | full | 7/10 | Xcommunity | |
Snapshot a sandbox and later restore or fork new sandboxes from that snapshot C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 3 | partial | 5/10 | Xcommunity | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 4/10 | Cclaimed | |
Restrict or allow the sandbox's network egress with explicit policy C Network policy | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | disputed | 3/10 | Dcontradicted | |
Pay per second only for the compute a sandbox actually uses G Pricing | platform-engineer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 3 | none | 0/10 | ||
Run large concurrent fleets of sandboxes with documented concurrency limits C Scale | platform-engineer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Start sandboxes with documented sub-second-to-few-second cold starts C Latency | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Expose a port from the sandbox on a public preview URL to reach services running inside C Preview access | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | full | 9/10 | Cclaimed | |
Read, write, upload, and download files in the sandbox filesystem via the SDK C Files | developer | Code execution — stories about code execution in this arenaCode execution | 2 | full | 9/10 | Cclaimed | |
Run arbitrary shell commands and install packages inside the sandbox C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 2 | full | 9/10 | Cclaimed | |
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results back C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | full | 8/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 7/10 | Xcommunity | |
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs C Isolation | ai-native user | Isolation security — stories about isolation security in this arenaIsolation security | 2 | partial | 6/10 | Xcommunity | |
Keep a sandbox session running for hours or days for long agent tasks C Scale | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 2 | disputed | 5/10 | Dcontradicted | |
Pause a running sandbox and resume it later with filesystem and memory state intact C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 2 | partial | 5/10 | Xcommunity | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 5/10 | Xcommunity | |
Set timeouts so sandboxes shut down automatically and stop billing when idle or done G Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | disputed | 3/10 | Dcontradicted | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Define custom sandbox templates or bring my own container image C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | none | 0/10 | ||
Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipe C Agent workloads | ai-native user | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Attach GPUs to sandboxed workloads C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | untested | none yet | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Run a headless browser or full desktop environment inside the sandbox C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 4/10 | Cclaimed | |
Start building with a free tier or included credits without talking to sales G Pricing | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | none | 0/10 |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 35 stories with headroom
What would move Cloudflare Sandboxes’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
Cloudflare Sandboxes is an execution/infrastructure SDK (not itself an agent), so an official MCP server for agent connectivity is a fair axis to ask about—but the evidence pack contains no mention of MCP, an MCP server, or any protocol for connecting AI agents beyond raw SDK APIs (exec, files, sessions, tunnels, etc.).
Performance scale — stories about performance scale in this arenaRun large concurrent fleets of sandboxes with documented concurrency limits
nonemoves PA Scoreimpact 30
Docs describe per-sandbox features (sessions, exec, files) and one architectural note that multiplexing avoids subrequest limits, but there is no documentation of concurrency limits, fleet-level scaling guidance, or how many sandboxes/instances can run concurrently.
Performance scale — stories about performance scale in this arenaStart sandboxes with documented sub-second-to-few-second cold starts
nonemoves PA Scoreimpact 30
Missing: official docs quantifying cold-start latency, benchmark data, or a first-party performance page addressing startup time.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Cloudflare Sandboxes is built entirely on Cloudflare's proprietary Workers/Durable Objects/container infrastructure, and no evidence in the pack mentions any open-source release, self-hosted deployment option, or ability to run the core product outside Cloudflare's platform.
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPay per second only for the compute a sandbox actually uses
nonemoves PA Scoreimpact 30
No documentation in the evidence pack describes per-second or usage-based billing for sandbox compute; the only pricing-related evidence is community commentary describing flat, expensive per-vCPU pricing (comm-3, comm-6) and the absence of automatic idle shutdown, meaning engineers must build their own cleanup to avoid being billed for idle sandboxes (comm-7).
Agenticness — how well agents can access and operate the productOperate the product with natural-language commands
nonemoves Built-in AIimpact 30
Missing: any NL command parsing/interface, evidence of a conversational control layer, examples of natural-language-driven sandbox operations.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
The evidence pack shows the Sandbox SDK is an API/library used from Workers code (exec, files, sessions, etc.) but nowhere mentions an official CLI tool for AI-native workflows; interaction is entirely via SDK calls or the general Wrangler CLI, not a dedicated Sandbox CLI.
Showing the top 8 of 35 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map4 surfaces · 26 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Sandbox docs26 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Set up automations that run autonomously in the background
- Download a machine-readable API spec (OpenAPI or equivalent)
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Read, write, upload, and download files in the sandbox filesystem via the SDK
- Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results back
- Execute untrusted, AI-generated code without risking my own infrastructure
- Run arbitrary shell commands and install packages inside the sandbox
- My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human
- Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems
- Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs
- Restrict or allow the sandbox's network egress with explicit policy
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Keep a sandbox session running for hours or days for long agent tasks
- Set timeouts so sandboxes shut down automatically and stop billing when idle or done
- Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds
- Expose a port from the sandbox on a public preview URL to reach services running inside
- Pause a running sandbox and resume it later with filesystem and memory state intact
- Snapshot a sandbox and later restore or fork new sandboxes from that snapshot
Hacker News15 stories
- Run the product headlessly / in CI for automation
- Set up automations that run autonomously in the background
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Execute untrusted, AI-generated code without risking my own infrastructure
- My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human
- Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems
- Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs
- Restrict or allow the sandbox's network egress with explicit policy
- Do everything through the API that I can do in the UI
- Keep a sandbox session running for hours or days for long agent tasks
- Set timeouts so sandboxes shut down automatically and stop billing when idle or done
- Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds
- Pause a running sandbox and resume it later with filesystem and memory state intact
- Snapshot a sandbox and later restore or fork new sandboxes from that snapshot
llms.txt3 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
4 of 13 testable claims verified · 5 contradicted → integrity 0/100
35 distinct capability claims found in Cloudflare Sandboxes’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
4
Verified
4
Unverified
5
Contradicted
15
Undersold
Verified (4)
“Create point-in-time snapshots of directories and restore them from R2”
Snapshot a sandbox and later restore or fork new sandboxes from that snapshotpartialproof ↗
“VM-based isolation with each sandbox running in its own VM”
Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systemsfullproof ↗
“Get or create a sandbox instance by stable ID to reconnect to the same sandbox”
Spin up an isolated sandbox with one API/SDK call and get a live environment in secondsfullproof ↗
“SDK enables running untrusted code securely in isolated environments”
Execute untrusted, AI-generated code without risking my own infrastructurepartialproof ↗
Unverified (15)
“Run shell commands in the sandbox and capture stdout, stderr, and exit codes”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“Write and read files inside the sandbox filesystem via SDK calls”
Read, write, upload, and download files in the sandbox filesystem via the SDKfullproof ↗
“Execute Python and JavaScript code and get back rich outputs like charts and tables”
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results backfullproof ↗
“Expose a sandbox port as a public preview URL to reach services running inside”
Expose a port from the sandbox on a public preview URL to reach services running insidefullproof ↗
“Get zero-config *.trycloudflare.com tunnel URLs for exposed ports”
Expose a port from the sandbox on a public preview URL to reach services running insidefullproof ↗
“Mount S3-compatible buckets (R2, S3, GCS) as local filesystems for persistent storage across sandbox lifecycles”
Read, write, upload, and download files in the sandbox filesystem via the SDKfullproof ↗
“Pass input via stdin to commands, avoiding shell injection risks”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“Execute a command and stream results in real time via Server-Sent Events”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“Write binary data and files larger than 32 MiB to the sandbox via streaming, replacing base64 encoding”
Read, write, upload, and download files in the sandbox filesystem via the SDKfullproof ↗
“Create directories recursively in the sandbox filesystem”
Read, write, upload, and download files in the sandbox filesystem via the SDKfullproof ↗
“Run pandas-based Python data analysis with automatic return of last expression result”
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results backfullproof ↗
“Pass stdin data to commands enabling arbitrary input without shell injection risk”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“Receive real-time output callbacks distinguishing stdout and stderr streams”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“writeFile supports ReadableStream input for binary/large file writes”
Read, write, upload, and download files in the sandbox filesystem via the SDKfullproof ↗
“Expose a port and get a preview URL for accessing services running in the sandbox”
Expose a port from the sandbox on a public preview URL to reach services running insidefullproof ↗
Contradicted (8)
“Create isolated execution contexts/sessions within a sandbox, each with its own shell state, env vars, and working directory”
Keep a sandbox session running for hours or days for long agent tasksdisputedproof ↗
“Code execution contexts maintain state (variables, imports, functions) across multiple executions”
Keep a sandbox session running for hours or days for long agent tasksdisputedproof ↗
“Multiplex all SDK calls over a single persistent connection to avoid subrequest limits on concurrent operations”
Run large concurrent fleets of sandboxes with documented concurrency limitsnoneproof ↗
“Create additional sessions for separate workflows within the same sandbox”
Keep a sandbox session running for hours or days for long agent tasksdisputedproof ↗
“Set a default command timeout for all commands in a session”
Set timeouts so sandboxes shut down automatically and stop billing when idle or donedisputedproof ↗
“Run Docker commands inside a sandbox container”
Define custom sandbox templates or bring my own container imagenoneproof ↗
“Intercept and handle outbound HTTP traffic from sandboxes using Workers”
Restrict or allow the sandbox's network egress with explicit policydisputedproof ↗
“Create a named session with custom environment variables and working directory”
Keep a sandbox session running for hours or days for long agent tasksdisputedproof ↗
Undersold (15)
Point an agent at llms.txt or agent-oriented docspartialproof ↗
Run the product headlessly / in CI for automationpartialproof ↗
Drive the product through a documented public APIfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)partialproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a humanpartialproof ↗
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runspartialproof ↗
Do everything through the API that I can do in the UIfullproof ↗
Export all of my data in open formats and leavepartialproof ↗
Pause a running sandbox and resume it later with filesystem and memory state intactpartialproof ↗
Claims outside our story set (8)
Real capability claims found in Cloudflare Sandboxes’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Watch the sandbox filesystem in real time for changes using native inotify”
source ↗“Connect browser-based terminal UIs to sandbox shells via WebSocket with xterm.js addon supporting reconnection and resize”
source ↗“Route incoming HTTP and WebSocket requests to the correct sandbox container”
source ↗“SDK supports three transport protocols for communication between Durable Object and container”
source ↗“Access KV, R2, Durable Objects, and other Cloudflare bindings from within a sandbox”
source ↗“Clone repositories, manage branches, and automate Git operations from the sandbox”
source ↗“Sandboxes are positioned for building AI agents that execute code, IDEs, data analysis platforms, and CI/CD systems”
source ↗“Watch specific file patterns in a directory with an event callback”
source ↗
Business model
Runs on Cloudflare Containers: requires the $5/mo Workers Paid plan, then usage-based billing for container vCPU-seconds, memory, disk, and egress beyond included allowances.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% · openapi.json 100% (30d, checked every 6h since Sep 8 '26)
