Rank #3 of 8 in Agent Sandboxes & Code Execution
Showcase


Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardwareevidence →
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode executionevidence →
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experienceevidence →
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation securityevidence →
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scaleevidence →
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycleevidence →
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistenceevidence →
Stories about snapshot persistence in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Runloop stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Webhooks · Machine-readable spec · Versioning policy
Subscribe to events via webhooks
—–
Build against official SDKs
~6/10
Issue scoped/least-privilege API credentials for an agent
~6/10
Connect an agent via an official MCP server
✓7/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
✓8/10
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
n/an/a
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
—0/10
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
~7/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode execution
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation security
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scale
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence
Stories about snapshot persistence in this arena
Sorted by importance (agentic first) (high → low) · 51/51 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 7/10 | Tprobed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 7/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | full | 8/10 | Cclaimed | |
Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds C Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 3 | full | 9/10 | Cclaimed | |
Execute untrusted, AI-generated code without risking my own infrastructure C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 3 | full | 8/10 | Tprobed | |
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human C Agent lifecycle | ai-native user | Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience | 3 | full | 8/10 | Cclaimed | |
Snapshot a sandbox and later restore or fork new sandboxes from that snapshot C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 3 | full | 8/10 | Cclaimed | |
Start sandboxes with documented sub-second-to-few-second cold starts C Latency | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | partial | 6/10 | Cclaimed | |
Run large concurrent fleets of sandboxes with documented concurrency limits C Scale | platform-engineer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 4/10 | Cclaimed | |
Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems C Isolation | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | partial | 3/10 | Tprobed | |
Pay per second only for the compute a sandbox actually uses G Pricing | platform-engineer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Restrict or allow the sandbox's network egress with explicit policy C Network policy | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Define custom sandbox templates or bring my own container image C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | full | 8/10 | Cclaimed | |
Expose a port from the sandbox on a public preview URL to reach services running inside C Preview access | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | full | 8/10 | Cclaimed | |
Keep a sandbox session running for hours or days for long agent tasks C Scale | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 2 | full | 8/10 | Cclaimed | |
Pause a running sandbox and resume it later with filesystem and memory state intact C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 2 | full | 8/10 | Cclaimed | |
Run arbitrary shell commands and install packages inside the sandbox C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 2 | full | 7/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Tprobed | |
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs C Isolation | ai-native user | Isolation security — stories about isolation security in this arenaIsolation security | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results back C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | partial | 5/10 | Cclaimed | |
Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipe C Agent workloads | ai-native user | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | partial | 5/10 | Cclaimed | |
Set timeouts so sandboxes shut down automatically and stop billing when idle or done G Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | partial | 5/10 | Cclaimed | |
Read, write, upload, and download files in the sandbox filesystem via the SDK C Files | developer | Code execution — stories about code execution in this arenaCode execution | 2 | partial | 4/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Run a headless browser or full desktop environment inside the sandbox C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | 0/10 | ||
Attach GPUs to sandboxed workloads C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | untested | none yet | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Start building with a free tier or included credits without talking to sales G Pricing | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | full | 8/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 4/10 | Cclaimed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 33 stories with headroom
What would move Runloop’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Missing: any documentation of an MCP-client integration point, config for adding external MCP servers, or examples of Devboxes/agents consuming external MCP tool servers.
Isolation security — stories about isolation security in this arenaRestrict or allow the sandbox's network egress with explicit policy
nonemoves PA Scoreimpact 30
No evidence pack item describes network egress controls, firewall rules, or allow/deny policies for Devbox sandboxes; tunnels and secrets features are documented but network egress policy configuration is never mentioned.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
nonemoves PA Scoreimpact 30
Runloop is presented entirely as a hosted cloud platform (Devboxes, Agents API, benchmarks) with no evidence of an on-premises or self-hostable deployment option, open-source core, or self-hosting documentation.
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPay per second only for the compute a sandbox actually uses
nonemoves PA Scoreimpact 30
No evidence pack item states an explicit per-second or usage-based billing model; the only tangential mention is that suspend/resume can 'save costs' (runloop-docs-27), which doesn't establish a per-second pricing mechanism.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence anywhere in the pack addresses AI training data usage, opt-out policies, or data-retention commitments for model training; the docs focus on sandbox infrastructure (Devboxes, secrets, gateways) but never mention training-data policy.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
Missing: any documentation of a webhooks API, subscription/callback URLs, or event delivery mechanism to external endpoints.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
No evidence of an interactive API reference with runnable examples; the explicit probe for OpenAPI/swagger specs returned 404s across all candidate paths, and docs mention SDKs/CLI but not a runnable API explorer.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Runloop offers SDKs and a CLI but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes to common OpenAPI paths (openapi.json, swagger.json, etc.) all returned 404, confirming absence rather than just lack of mention.
Showing the top 8 of 33 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map4 surfaces · 33 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs32 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Connect an agent via an official MCP server
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipe
- Read, write, upload, and download files in the sandbox filesystem via the SDK
- Define custom sandbox templates or bring my own container image
- Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results back
- Execute untrusted, AI-generated code without risking my own infrastructure
- Run arbitrary shell commands and install packages inside the sandbox
- My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human
- Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems
- Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Start sandboxes with documented sub-second-to-few-second cold starts
- Keep a sandbox session running for hours or days for long agent tasks
- Start building with a free tier or included credits without talking to sales
- Set timeouts so sandboxes shut down automatically and stop billing when idle or done
- Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds
- Expose a port from the sandbox on a public preview URL to reach services running inside
- Pause a running sandbox and resume it later with filesystem and memory state intact
- Snapshot a sandbox and later restore or fork new sandboxes from that snapshot
runloop.ai7 stories
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Execute untrusted, AI-generated code without risking my own infrastructure
- My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human
- Start sandboxes with documented sub-second-to-few-second cold starts
- Run large concurrent fleets of sandboxes with documented concurrency limits
- Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds
llms.txt3 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
2 of 9 testable claims verified · 0 contradicted → integrity 22/100
13 distinct capability claims found in Runloop’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
2
Verified
7
Unverified
0
Contradicted
24
Undersold
Verified (2)
“Provides secure sandboxed execution environments (Devboxes) for running code safely”
Execute untrusted, AI-generated code without risking my own infrastructurefullproof ↗
“Official CLI (rli) offers both an interactive terminal UI and traditional CLI commands for managing resources”
Unverified (8)
“CLI can build a custom sandbox image/blueprint from a Dockerfile”
Define custom sandbox templates or bring my own container imagefullproof ↗
“Offers optimized public prebuilt blueprints for common environments”
Define custom sandbox templates or bring my own container imagefullproof ↗
“Snapshots save a Devbox's disk state and can be used to create new Devboxes from that saved state”
Snapshot a sandbox and later restore or fork new sandboxes from that snapshotfullproof ↗
“Devbox tunnels expose ports securely over a simple public URL”
Expose a port from the sandbox on a public preview URL to reach services running insidefullproof ↗
“CLI can open a secure SSH shell session directly into a devbox, handling key setup automatically”
Run arbitrary shell commands and install packages inside the sandboxfullproof ↗
“Agent Gateways let agents call LLM APIs (Anthropic, OpenAI) without ever exposing the underlying API keys”
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runspartialproof ↗
“Supports running 10,000+ parallel sandboxes concurrently”
Run large concurrent fleets of sandboxes with documented concurrency limitspartialproof ↗
“10GB sandbox images start up in under 2 seconds”
Start sandboxes with documented sub-second-to-few-second cold startspartialproof ↗
Undersold (24)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIfullproof ↗
Issue scoped/least-privilege API credentials for an agentpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipepartialproof ↗
Read, write, upload, and download files in the sandbox filesystem via the SDKpartialproof ↗
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results backpartialproof ↗
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a humanfullproof ↗
Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systemspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Keep a sandbox session running for hours or days for long agent tasksfullproof ↗
Start building with a free tier or included credits without talking to salesfullproof ↗
Set timeouts so sandboxes shut down automatically and stop billing when idle or donepartialproof ↗
Spin up an isolated sandbox with one API/SDK call and get a live environment in secondsfullproof ↗
Pause a running sandbox and resume it later with filesystem and memory state intactfullproof ↗
Claims outside our story set (3)
Real capability claims found in Runloop’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“CLI can submit a benchmark job that auto-provisions devboxes, runs agents, scores results, and compares agents side-by-side”
source ↗“Can run agents against well-known open-source benchmarks like Terminal Bench 2 and AIME”
source ↗“Supports creating custom scorers to evaluate agents on dimensions like security, cost, performance, and compliance”
source ↗
Business model
New accounts get $50 in credits; the Basic plan is a free subscription billed purely on metered devbox CPU/RAM/storage usage, Pro adds a $250/mo base, and Enterprise/VPC is custom.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
