Rank #4 of 8 in Agent Sandboxes & Code Execution
Install
npm install -g maritime-cliTry itExperimental
See what an agent can do with Maritime before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -si https://maritime.sh/api/agentsrecorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardwareevidence →
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode executionevidence →
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experienceevidence →
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation securityevidence →
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scaleevidence →
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycleevidence →
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistenceevidence →
Stories about snapshot persistence in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 1 paid · 0 enterprise · 18 not stated in evidence
Follow the green: where the map greys out is where Maritime stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → MCP server · Machine-readable spec · Versioning policy · API sandbox · Full data export
Subscribe to events via webhooks
✓7/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
✓8/10
Connect an agent via an official MCP server
—0/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
—–
Get AI-generated insights and suggestions from my data inside the product
—–
Set up automations that run autonomously in the background
✓8/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware
Stories about capabilities hardware in this arena
Code execution — stories about code execution in this arenaCode execution
Stories about code execution in this arena
Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience
Day-to-day developer experience — setup friction, docs, debugging, iteration speed
Isolation security — stories about isolation security in this arenaIsolation security
Stories about isolation security in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Performance scale — stories about performance scale in this arenaPerformance scale
Stories about performance scale in this arena
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle
Creating, updating, and tearing down resources across their lifecycle
Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence
Stories about snapshot persistence in this arena
Sorted by importance (agentic first) (high → low) · 51/51 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds C Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 3 | full | 7/10 | Tprobed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 6/10 | Cclaimed | |
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human C Agent lifecycle | ai-native user | Dev experience — day-to-day developer experience — setup friction, docs, debugging, iteration speedDev experience | 3 | partial | 6/10 | Cclaimed | |
Execute untrusted, AI-generated code without risking my own infrastructure C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 3 | partial | 5/10 | Tprobed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Pay per second only for the compute a sandbox actually uses G Pricing | platform-engineer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Start sandboxes with documented sub-second-to-few-second cold starts C Latency | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Rely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems C Isolation | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | none | untested | none yet | |
Restrict or allow the sandbox's network egress with explicit policy C Network policy | platform-engineer | Isolation security — stories about isolation security in this arenaIsolation security | 3 | none | untested | none yet | |
Run large concurrent fleets of sandboxes with documented concurrency limits C Scale | platform-engineer | Performance scale — stories about performance scale in this arenaPerformance scale | 3 | none | untested | none yet | |
Snapshot a sandbox and later restore or fork new sandboxes from that snapshot C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 3 | none | untested | none yet | |
Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipe C Agent workloads | ai-native user | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | full | 8/10 | Cclaimed | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | full | 8/10 | Cclaimed | |
Run a headless browser or full desktop environment inside the sandbox C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | full | 7/10 | Cclaimed | |
Define custom sandbox templates or bring my own container image C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | partial | 6/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Tprobed | |
Expose a port from the sandbox on a public preview URL to reach services running inside C Preview access | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | partial | 6/10 | Cclaimed | |
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs C Isolation | ai-native user | Isolation security — stories about isolation security in this arenaIsolation security | 2 | partial | 6/10 | Cclaimed | |
Keep a sandbox session running for hours or days for long agent tasks C Scale | developer | Performance scale — stories about performance scale in this arenaPerformance scale | 2 | partialpaid | 5/10 | Tprobed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 4/10 | Cclaimed | |
Attach GPUs to sandboxed workloads C Workloads | developer | Capabilities hardware — stories about capabilities hardware in this arenaCapabilities hardware | 2 | none | 0/10 | ||
Execute code in multiple language runtimes (Python, JavaScript, and more) and get rich results back C Runtimes | developer | Code execution — stories about code execution in this arenaCode execution | 2 | none | 0/10 | ||
Set timeouts so sandboxes shut down automatically and stop billing when idle or done G Lifecycle | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Pause a running sandbox and resume it later with filesystem and memory state intact C Snapshots | developer | Snapshot persistence — stories about snapshot persistence in this arenaSnapshot persistence | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Read, write, upload, and download files in the sandbox filesystem via the SDK C Files | developer | Code execution — stories about code execution in this arenaCode execution | 2 | none | untested | none yet | |
Run arbitrary shell commands and install packages inside the sandbox C Untrusted code | developer | Code execution — stories about code execution in this arenaCode execution | 2 | none | untested | none yet | |
Start building with a free tier or included credits without talking to sales G Pricing | developer | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | full | 8/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 38 stories with headroom
What would move Maritime’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Maritime is presented as infrastructure for hosting/deploying user-built AI agents, not as a product with its own built-in assistant for task delegation.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Maritime is a platform for deploying and hosting agents built with frameworks (CrewAI, LangGraph, etc.), but the evidence pack shows no mention of MCP server integration or a mechanism for agents to consume external MCP tool servers.
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
Maritime is a platform for deploying/hosting agents (not itself an agent), so an official MCP server is a fair axis to expect, but no evidence pack item mentions MCP, an MCP server, or MCP-compatible endpoints — only REST API, CLI, SDKs, and webhooks are documented.
Performance scale — stories about performance scale in this arenaRun large concurrent fleets of sandboxes with documented concurrency limits
nonemoves PA Scoreimpact 30
No evidence of documented concurrency limits or guidance for running large fleets of sandboxes at scale; docs cover provisioning, idempotency, and billing per-agent but never state concurrency ceilings, throughput benchmarks, or fleet-scale operational guidance.
Performance scale — stories about performance scale in this arenaStart sandboxes with documented sub-second-to-few-second cold starts
nonemoves PA Scoreimpact 30
Maritime's docs describe sleep/wake behavior (e.g., 'Sleeping agents wake automatically', 'puts it to sleep when idle, and wakes it on the next visit') and an 'always-on' add-on for millisecond reaction, but no evidence anywhere documents actual cold-start latency numbers (sub-second to few-second) for waking a sandboxed agent from sleep.
Isolation security — stories about isolation security in this arenaRely on a documented hard isolation boundary (microVM or equivalent) between sandboxes and my systems
nonemoves PA Scoreimpact 30
Missing: any explicit statement of sandbox isolation technology, security model documentation, or third-party audit/confirmation of hard isolation.
Isolation security — stories about isolation security in this arenaRestrict or allow the sandbox's network egress with explicit policy
nonemoves PA Scoreimpact 30
No evidence pack item mentions network egress controls, firewall rules, or any explicit allow/deny policy for outbound traffic from the sandbox/agent VM; the closest material covers API key scoping and secret management, not network policy.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
Missing: an export command or documented data-portability feature, open-format export of agent state/configs, any migration-out guide.
Showing the top 8 of 38 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map5 surfaces · 23 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
docs23 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Schedule recurring jobs or workflows
- Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipe
- Run a headless browser or full desktop environment inside the sandbox
- Define custom sandbox templates or bring my own container image
- Execute untrusted, AI-generated code without risking my own infrastructure
- My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a human
- Give an agent a sandbox where host secrets and credentials are unreachable by the code it runs
- Do everything through the API that I can do in the UI
- Keep a sandbox session running for hours or days for long agent tasks
- Start building with a free tier or included credits without talking to sales
- Spin up an isolated sandbox with one API/SDK call and get a live environment in seconds
- Expose a port from the sandbox on a public preview URL to reach services running inside
llms.txt5 stories
- Point an agent at llms.txt or agent-oriented docs
- Define custom sandbox templates or bring my own container image
- Execute untrusted, AI-generated code without risking my own infrastructure
- Do everything through the API that I can do in the UI
- Keep a sandbox session running for hours or days for long agent tasks
OpenAPI spec4 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -si https://maritime.sh/api/agentsreproduced$ curl -si https://maritime.sh/api/agents
HTTP/2 403
content-type: application/json
date: Tue, 15 Sep 2026 17:37:05 GMT
server: railway-hikari
vary: rsc, next-router-state-tree, next-router-prefetch, next-router-segment-prefetch
x-railway-request-id: DFk8XTq8QLSYK_ubBhdwDg
x-railway-edge: lax1
x-hikari-trace: lax1.ez9k
{"error":"Not authenticated"}
$npx -y maritime-cli --versionreproduced$ npx -y maritime-cli --version \|/-\|/-\|1.7.0 \
$curl -s https://maritime.sh/llms.txt | head -4reproduced$ curl -s https://maritime.sh/llms.txt | head -4 # Maritime > Maritime is a cloud platform for deploying, hosting, and managing AI agents. Agents run on serverless infrastructure with a sleep/wake architecture (they sleep when idle and wake in about a second), so a typical agent costs $1 per agent per month on a flat plan with no usage metering. Deploy from the CLI, the dashboard, or clone a Template: a complete, working agent (prompts, tools, integrations, memory, triggers, and hosting) shared as a single link.
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
4 of 11 testable claims verified · 0 contradicted → integrity 36/100
16 distinct capability claims found in Maritime’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
4
Verified
7
Unverified
0
Contradicted
12
Undersold
Verified (5)
“Agent provisioning is idempotent via an externalId tag, returning existing agent or creating a new one”
Drive the product through a documented public APIfullproof ↗
“Generate a personal API key via CLI and use it as a Bearer token, same as the CLI itself uses”
“Every CLI command supports a --json flag returning a single JSON value for scripting/automation”
“Every CLI command supports a --json flag returning a single JSON value for scripting/automation”
Run the product headlessly / in CI for automationfullproof ↗
“Point Claude Code or a Maritime agent at a provider's runbook page to inventory and recreate old workloads as Maritime agents using read-only credentials”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Unverified (9)
“Any Docker image implementing a three-endpoint contract can run as an agent, regardless of framework (CrewAI, LangGraph, AutoGen, custom scripts)”
Define custom sandbox templates or bring my own container imagepartialproof ↗
“Subscribe a URL to receive signed JSON webhook notifications on agent lifecycle state changes”
“Issue narrowly scoped API keys so a leaked key only grants limited permissions”
Issue scoped/least-privilege API credentials for an agentfullproof ↗
“Provision a persistent Linux desktop (XFCE, Chromium, LibreOffice) per end user, controllable by a model”
Run a headless browser or full desktop environment inside the sandboxfullproof ↗
“Templates are cloneable agent packages that publishers can share via link or a public library”
Define custom sandbox templates or bring my own container imagepartialproof ↗
“Cron triggers using a standard five-field expression wake an agent on schedule and deliver a prompt as a chat message”
“Bulk-import environment variables from a .env file or stdin, then hot-reload them”
Perform bulk operations across many items at oncepartialproof ↗
“Ready-made templates are pre-configured real agent frameworks deployable with one image pull”
Define custom sandbox templates or bring my own container imagepartialproof ↗
“Point Claude Code or a Maritime agent at a provider's runbook page to inventory and recreate old workloads as Maritime agents using read-only credentials”
Run coding agents like Claude Code or Codex inside the sandbox following the vendor's own recipefullproof ↗
Undersold (12)
Set up automations that run autonomously in the backgroundfullproof ↗
Operate the product with natural-language commandspartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Execute untrusted, AI-generated code without risking my own infrastructurepartialproof ↗
My agent can provision its own sandbox, execute code, read the results, and tear it down — end to end without a humanpartialproof ↗
Give an agent a sandbox where host secrets and credentials are unreachable by the code it runspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Keep a sandbox session running for hours or days for long agent taskspartialproof ↗
Start building with a free tier or included credits without talking to salesfullproof ↗
Spin up an isolated sandbox with one API/SDK call and get a live environment in secondsfullproof ↗
Expose a port from the sandbox on a public preview URL to reach services running insidepartialproof ↗
Claims outside our story set (4)
Real capability claims found in Maritime’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Live-stream an agent's build logs on its Overview page”
source ↗“An always-on add-on keeps an agent running continuously for low-latency, round-the-clock reaction”
source ↗“Deploy and start talking to your first agent in under 5 minutes”
source ↗“Supports migrating existing workloads from AWS, Heroku, or a VPS”
source ↗
Business model
Flat subscription plans priced per machine (about $1/machine/month, free plan with 3 machines); no metering of tokens, sessions, or awake time.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
