Install
Showcase


Try itExperimental
See what an agent can do with Vast.ai before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -s 'https://console.vast.ai/api/v0/bundles/' | head -c 600recorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Access connectivity — stories about access connectivity in this arenaAccess connectivityevidence →
Stories about access connectivity in this arena
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Capacity availability — stories about capacity availability in this arenaCapacity availabilityevidence →
Stories about capacity availability in this arena
Clusters scale — stories about clusters scale in this arenaClusters scaleevidence →
Stories about clusters scale in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing billing — stories about pricing billing in this arenaPricing billingevidence →
Stories about pricing billing in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycleevidence →
Creating, updating, and tearing down resources across their lifecycle
Serverless endpoints — stories about serverless endpoints in this arenaServerless endpointsevidence →
Stories about serverless endpoints in this arena
Storage data — storing and moving data — persistence, formats, durabilityStorage dataevidence →
Storing and moving data — persistence, formats, durability
Templates images — stories about templates images in this arenaTemplates imagesevidence →
Stories about templates images in this arena
Trust governance — stories about trust governance in this arenaTrust governanceevidence →
Stories about trust governance in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 11 free · 10 paid · 0 enterprise · 13 not stated in evidence
Follow the green: where the map greys out is where Vast.ai stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Access connectivity — stories about access connectivity in this arenaAccess connectivity
Stories about access connectivity in this arena
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Webhooks · MCP server · Versioning policy
Subscribe to events via webhooks
—0/10
Build against official SDKs
✓9/10
Issue scoped/least-privilege API credentials for an agent
✓8/10
Connect an agent via an official MCP server
—0/10
Download a machine-readable API spec (OpenAPI or equivalent)
✓9/10
unlocks → MCP server
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
n/an/a
Explore an interactive API reference with runnable examples
~4/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
✓7/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
n/an/a
Set up automations that run autonomously in the background
~5/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Capacity availability — stories about capacity availability in this arenaCapacity availability
Stories about capacity availability in this arena
See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure
✓8/10
Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation options
~5/10
See documented quotas and instance limits and raise them through a defined process
—–
Clusters scale — stories about clusters scale in this arenaClusters scale
Stories about clusters scale in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing billing — stories about pricing billing in this arenaPricing billing
Stories about pricing billing in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle
Creating, updating, and tearing down resources across their lifecycle
Serverless endpoints — stories about serverless endpoints in this arenaServerless endpoints
Stories about serverless endpoints in this arena
Storage data — storing and moving data — persistence, formats, durabilityStorage data
Storing and moving data — persistence, formats, durability
Templates images — stories about templates images in this arenaTemplates images
Stories about templates images in this arena
Trust governance — stories about trust governance in this arenaTrust governance
Stories about trust governance in this arena
Sorted by importance (agentic first) (high → low) · 54/54 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | fullfree | 9/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 9/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 9/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 9/10 | Tprobed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 9/10 | Tprobed | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 8/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | fullfree | 7/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialpaid | 5/10 | Tprobed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialfree | 4/10 | Tprobed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | n/a | untested | none yet | |
Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the console C Agent ops | ai agent | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 3 | fullpaid | 8/10 | Tprobed | |
Provision an on-demand GPU instance from the console or API and be running code on it within minutes C Provision | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 3 | fullpaid | 7/10 | Tprobed | |
See the published per-GPU-hour price for every GPU type on a public pricing page without talking to sales G Pricing | ml engineer | Pricing billing — stories about pricing billing in this arenaPricing billing | 3 | partialfree | 7/10 | Tprobed | |
Attach persistent network storage that survives instance teardown, so datasets and checkpoints outlive any single GPU rental C Storage | developer | Storage data — storing and moving data — persistence, formats, durabilityStorage data | 3 | partial | 6/10 | Xcommunity | |
Rent spot or interruptible GPU capacity at a deep discount with clearly documented preemption semantics G Spot | ml engineer | Pricing billing — stories about pricing billing in this arenaPricing billing | 3 | partialpaid | 6/10 | Xcommunity | |
SSH into my GPU instance with my own keys and get root-level control of the environment C Ssh | developer | Access connectivity — stories about access connectivity in this arenaAccess connectivity | 3 | partial | 6/10 | Cclaimed | |
Provision a multi-node GPU cluster with fast interconnect for distributed training without a sales cycle C Clusters | ml engineer | Clusters scale — stories about clusters scale in this arenaClusters scale | 3 | disputed | 5/10 | Dcontradicted | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partialpaid | 4/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 3/10 | Cclaimed | |
I am billed at per-second or per-minute granularity and only while my instance is actually running G Billing | ml engineer | Pricing billing — stories about pricing billing in this arenaPricing billing | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | untested | none yet | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 9/10 | Tprobed | |
Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spend G Discovery | ai agent | Pricing billing — stories about pricing billing in this arenaPricing billing | 2 | fullfree | 9/10 | Tprobed | |
Launch from pre-built ML templates (PyTorch, CUDA, vLLM, ComfyUI) instead of assembling an environment from scratch C Templates | developer | Templates images — stories about templates images in this arenaTemplates images | 2 | full | 8/10 | Cclaimed | |
Run my own Docker image or custom machine template with my exact environment C Images | developer | Templates images — stories about templates images in this arenaTemplates images | 2 | full | 8/10 | Cclaimed | |
See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure C Availability | ml engineer | Capacity availability — stories about capacity availability in this arenaCapacity availability | 2 | fullfree | 8/10 | Tprobed | |
Start, stop, restart, and terminate instances programmatically and keep paying only for what is running C Manage | developer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | fullpaid | 8/10 | Tprobed | |
Deploy code to autoscaling serverless GPU workers that scale to zero, instead of managing always-on instances C Serverless | developer | Serverless endpoints — stories about serverless endpoints in this arenaServerless endpoints | 2 | partialpaid | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Tprobed | |
Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation options C Hardware | ml engineer | Capacity availability — stories about capacity availability in this arenaCapacity availability | 2 | partialpaid | 5/10 | Tprobed | |
Move data in and out efficiently — S3-compatible endpoints, cloud-storage sync, or documented transfer tooling C Data movement | developer | Storage data — storing and moving data — persistence, formats, durabilityStorage data | 2 | partial | 5/10 | Xcommunity | |
Pull usage and billing breakdowns programmatically to attribute GPU spend by team or workload G Billing | platform engineer | Pricing billing — stories about pricing billing in this arenaPricing billing | 2 | partialpaid | 4/10 | Cclaimed | |
Verify the provider's security and compliance posture (SOC 2, data handling, datacenter tiers) before putting proprietary models on it C Compliance | platform engineer | Trust governance — stories about trust governance in this arenaTrust governance | 2 | disputed | 4/10 | Dcontradicted | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 3/10 | Tprobed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Schedule jobs on managed Slurm or Kubernetes instead of building my own scheduler on raw nodes C Orchestration | platform engineer | Clusters scale — stories about clusters scale in this arenaClusters scale | 2 | none | 0/10 | ||
Set auto-shutdown timers or spend limits so a forgotten instance can't silently run up a huge bill C Manage | ml engineer | Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleProvisioning lifecycle | 2 | none | 0/10 | ||
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
See documented quotas and instance limits and raise them through a defined process G Quotas | platform engineer | Capacity availability — stories about capacity availability in this arenaCapacity availability | 2 | none | untested | none yet | |
Expose ports to serve applications from my instance and connect instances over private networking C Networking | developer | Access connectivity — stories about access connectivity in this arenaAccess connectivity | 1 | full | 7/10 | Cclaimed | |
Lock in reserved or committed-use discounts for sustained GPU capacity G Pricing | platform engineer | Pricing billing — stories about pricing billing in this arenaPricing billing | 1 | fullpaid | 7/10 | Cclaimed | |
Open Jupyter or connect my IDE (VS Code/Cursor) to the instance in one step C Ide | developer | Access connectivity — stories about access connectivity in this arenaAccess connectivity | 1 | partial | 5/10 | Cclaimed | |
Manage team members with roles and scoped API keys so credentials and spend stay controlled G Governance | platform engineer | Trust governance — stories about trust governance in this arenaTrust governance | 1 | partial | 4/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | 0/10 |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 29 stories with headroom
What would move Vast.ai’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Evidence only shows the reverse relationship: Vast.ai exposes a CLI/API/agent-skill so that external AI coding assistants (e.g., Claude, Copilot) can drive Vast.ai on the user's behalf — not that Vast.ai itself ships a built-in AI assistant a user can delegate tasks to within the product.
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
Vast.ai documents a CLI, REST API, Python SDK, and an 'agent skill' file for coding assistants, but no evidence anywhere describes an official MCP server or MCP integration.
Pricing billing — stories about pricing billing in this arenaI am billed at per-second or per-minute granularity and only while my instance is actually running
nonemoves PA Scoreimpact 30
Evidence shows Vast.ai quotes prices as hourly rates (dph/dph_total) and discusses reserved/interruptible pricing models and autobilling, but nowhere does the evidence pack confirm sub-hour (per-second or per-minute) billing granularity or explicitly state billing stops precisely when an instance is not running.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
Missing: any documentation of webhook registration/endpoints, event types, or delivery/retry semantics.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Evidence shows Vast.ai has a REST API, OpenAPI spec, and CLI/SDK, but there is no documentation of API versioning scheme or a deprecation policy anywhere in the pack; the only related signal is a backward-compatible import shim for the old SDK (vast-ai-gh-2), which is not a documented policy.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
partialq3/10moves PA Scoreimpact 21
Missing: explicit data-export/portability documentation, statement on open formats, and evidence of full account data extraction on leaving.
Provisioning lifecycle — creating, updating, and tearing down resources across their lifecycleSet auto-shutdown timers or spend limits so a forgotten instance can't silently run up a huge bill
nonemoves PA Scoreimpact 20
The evidence pack documents autobilling (auto top-up to keep instances running), reserved/interruptible instance pricing, and scoped API keys, but none of these are auto-shutdown timers or spend caps — autobilling in fact works against this story by automatically refilling balance to prevent instance interruption rather than capping spend.
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows
nonemoves PA Scoreimpact 20
Vast.ai's evidence covers CLI/API/SDK for instance and serverless management but contains no mention of a built-in scheduler, cron-like recurring job feature, or workflow orchestration; users would have to bring their own external scheduler to hit the API/CLI repeatedly, which isn't documented as a first-party capability.
Showing the top 8 of 29 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map9 surfaces · 36 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Guides docs30 stories
- Open Jupyter or connect my IDE (VS Code/Cursor) to the instance in one step
- Expose ports to serve applications from my instance and connect instances over private networking
- SSH into my GPU instance with my own keys and get root-level control of the environment
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Define rules that trigger actions automatically on events
- See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure
- Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation options
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Pull usage and billing breakdowns programmatically to attribute GPU spend by team or workload
- Lock in reserved or committed-use discounts for sustained GPU capacity
- See the published per-GPU-hour price for every GPU type on a public pricing page without talking to sales
- Rent spot or interruptible GPU capacity at a deep discount with clearly documented preemption semantics
- Control data retention and deletion
- Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the console
- Start, stop, restart, and terminate instances programmatically and keep paying only for what is running
- Provision an on-demand GPU instance from the console or API and be running code on it within minutes
- Deploy code to autoscaling serverless GPU workers that scale to zero, instead of managing always-on instances
- Move data in and out efficiently — S3-compatible endpoints, cloud-storage sync, or documented transfer tooling
- Attach persistent network storage that survives instance teardown, so datasets and checkpoints outlive any single GPU rental
- Run my own Docker image or custom machine template with my exact environment
- Launch from pre-built ML templates (PyTorch, CUDA, vLLM, ComfyUI) instead of assembling an environment from scratch
- Verify the provider's security and compliance posture (SOC 2, data handling, datacenter tiers) before putting proprietary models on it
- Manage team members with roles and scoped API keys so credentials and spend stay controlled
CLI docs22 stories
- Expose ports to serve applications from my instance and connect instances over private networking
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Issue scoped/least-privilege API credentials for an agent
- Build against official SDKs
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Explore an interactive API reference with runnable examples
- Perform bulk operations across many items at once
- See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure
- Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation options
- Provision a multi-node GPU cluster with fast interconnect for distributed training without a sales cycle
- Do everything through the API that I can do in the UI
- Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spend
- See the published per-GPU-hour price for every GPU type on a public pricing page without talking to sales
- Control data retention and deletion
- Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the console
- Start, stop, restart, and terminate instances programmatically and keep paying only for what is running
- Provision an on-demand GPU instance from the console or API and be running code on it within minutes
- Move data in and out efficiently — S3-compatible endpoints, cloud-storage sync, or documented transfer tooling
- Manage team members with roles and scoped API keys so credentials and spend stay controlled
API reference14 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Perform bulk operations across many items at once
- See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure
- Do everything through the API that I can do in the UI
- Pull usage and billing breakdowns programmatically to attribute GPU spend by team or workload
- Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spend
- See the published per-GPU-hour price for every GPU type on a public pricing page without talking to sales
- Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the console
- Start, stop, restart, and terminate instances programmatically and keep paying only for what is running
- Provision an on-demand GPU instance from the console or API and be running code on it within minutes
docs.vast.ai12 stories
- Open Jupyter or connect my IDE (VS Code/Cursor) to the instance in one step
- Run the product headlessly / in CI for automation
- Use an official CLI
- Build against official SDKs
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Do everything through the API that I can do in the UI
- Provision an on-demand GPU instance from the console or API and be running code on it within minutes
- Run my own Docker image or custom machine template with my exact environment
- Launch from pre-built ML templates (PyTorch, CUDA, vLLM, ComfyUI) instead of assembling an environment from scratch
Hacker News8 stories
- Provision a multi-node GPU cluster with fast interconnect for distributed training without a sales cycle
- See the published per-GPU-hour price for every GPU type on a public pricing page without talking to sales
- Rent spot or interruptible GPU capacity at a deep discount with clearly documented preemption semantics
- Start, stop, restart, and terminate instances programmatically and keep paying only for what is running
- Provision an on-demand GPU instance from the console or API and be running code on it within minutes
- Move data in and out efficiently — S3-compatible endpoints, cloud-storage sync, or documented transfer tooling
- Attach persistent network storage that survives instance teardown, so datasets and checkpoints outlive any single GPU rental
- Verify the provider's security and compliance posture (SOC 2, data handling, datacenter tiers) before putting proprietary models on it
API reference7 stories
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Do everything through the API that I can do in the UI
- Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spend
- Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the console
GitHub README6 stories
- Build against official SDKs
- Perform bulk operations across many items at once
- See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failure
- Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation options
- Provision a multi-node GPU cluster with fast interconnect for distributed training without a sales cycle
- Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spend
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -s 'https://console.vast.ai/api/v0/bundles/' | head -c 600reproduced$ curl -s 'https://console.vast.ai/api/v0/bundles/' | head -c 600
{"offers": [{"id": 29139710, "ask_contract_id": 29139710, "bundle_id": 1886940185, "bundled_results": null, "bw_nvlink": 0.0, "compute_cap": 1200, "cpu_arch": "amd64", "cpu_cores": 20, "cpu_cores_effective": 20.0, "cpu_ghz": 4.6, "cpu_name": "Core\u2122 Ultra 7 265F", "cpu_ram": 63923, "credit_discount_max": 0.0, "cuda_max_good": 13.0, "direct_port_count": 99, "disk_bw": 645.0, "disk_name": "nvme", "disk_space": 645.0, "dlperf": 199.43662668365562, "dlperf_per_dphtotal": 741.7064628730993, "dph_base": 0.26666666666666666, "dph_total": 0.2688888888888889, "driver_version": "580.126.09", "driver
$uvx --from vastai vastai --helpreproduced$ uvx --from vastai vastai --help
⠋ Resolving dependencies...
⠙ Resolving dependencies...
⠋ Resolving dependencies...
⠙ Resolving dependencies...
⠙ vastai==1.6.0
⠙ cryptography==49.0.0
⠙ pillow==12.2.0
⠙ pycares==4.11.0
⠙ aiodns==4.0.4
⠙ cryptography==49.0.0
⠙ pillow==12.2.0
⠙ pycares==4.11.0
⠙ aiodns==4.0.3
⠙ aiodns==4.0.2
⠙ aiodns==4.0.0
⠙ aiodns==3.6.1
⠙ aiohttp==3.14.3
⠙ anyio==4.15.0
⠙ argcomplete==3.7.2
⠙ borb==2.1.25
⠙ curlify==3.0.0
⠙ psutil==7.2.2
usage: vastai [-h] [--url URL] [--retry RETRY] [--explain] [--raw] [--full] [--curl] [--api-[redacted] API_[redacted]] [--version]
[--no-color]
command ...
options:
-h, --help show this help message and exit
--url URL Server REST API URL
--retry RETRY Retry limit
--explain Output verbose explanation of mapping of CLI calls to HTTPS API endpoints
--raw Output machine-readable json
--full Print full results instead of paging with `less` for commands that support it
--curl Show a curl equivalency to the call
--api-[redacted] API_[redacted] API [redacted] to use. defaults to using the one stored in /Users/judegomila/.config/vastai/vast_api_[redacted]
--version Show CLI version
--no-color Disable colored output for commands that support it
Search:
search offers Search for instance types using custom query
search benchmarks Search for benchmark results using custom query
search templates Search for template results using custom query
search invoices Search for invoices using custom query
create template Create a new template
update template Update an existing template
delete template Delete a Template
Instances:
show instance Display user's current instances
create instance Create a new instance
create instances Create instances from a list of offers
destroy instance Destroy an instance (irreversible, deletes data)
destroy instances Destroy a list of instances (irreversible, deletes data)
start instance Start a stopped instance
start instances Start a list of instances
stop instance Stop a running instance
stop instances Stop a list of instances
reboot instance Reboot (stop/start) an instance
recycle instance Recycle (destroy/create) an instance
update instance Update recreate an instance from a new/updated template
label instance Assign a string label to an instance
prepay instance Deposit credits into reserved instance
change bid Change the bid price for a spot/interruptible instance
launch instance Launch the top instance from the search offers based on the given parameters
show instances Display user's current instances
execute Execute a (constrained) remote command on a machine
logs Get the logs for an instance
ssh-url ssh url helper
scp-url scp url helper
take snapshot Schedule a snapshot of a running container and push it to your repo in a container registry
run benchmarks Benchmark a template against one or more GPUs
Host machines:
show machine [Host] Show hosted machines
show machines [Host] Show hosted machines
show maints [Host] Show maintenance information for host machines
list machine [Host] list a machine for rent
list machines [Host] list machines for rent
unlist machine [Host] Unlist a listed machine
delete machine [Host] Delete machine if the machine is not being used by clients. host jobs on their own machines are disregarded and machine is force deleted.
cleanup machine [Host] Remove all expired storage instances from the machine, freeing up space
defrag machines [Host] Defragment machines
set min-bid [Host] Set the minimum bid/rental price for a machine
set defjob [Host] Create default jobs for a machine
remove defjob [Host] Delete default jobs
schedule maint [Host] Schedule upcoming maint window
cancel maint [Host] Cancel maint window
dump-logs [Host] Bundle self-test diagnostics for support
self-test machine [Host] Perform a self-test on the specified machine
reports Get the user reports for a given machine
Teams:
create team Create a new team
destroy team Destroy your team
create team-role Add a new role to your team
show team-role Show your team role
show team-roles Show roles for a team
update team-role Update an existing team role
remove team-role Remove a role from your team
invite member Invite a team member
show members Show your team members
remove member Remove a team member
Auth & [redacted]s:
create api-[redacted] Create a new api-[redacted] with restricted permissions. Can be sent to other users and teammates
show api-[redacted] Show an api-[redacted]
show api-[redacted]s List your api-[redacted]s associated with your account
delete api-[redacted] Remove an api-[redacted]
reset api-[redacted] Reset your api-[redacted] (get new [redacted] from website)
create ssh-[redacted] Create a new ssh-[redacted]
show ssh-[redacted]s List your ssh [redacted]s associated with your account
delete ssh-[redacted] Remove an ssh-[redacted]
update ssh-[redacted] Update an existing SSH [redacted]
attach ssh Attach an ssh [redacted] to an instance. This will allow you to connect to the instance with the ssh [redacted]
detach ssh Detach an ssh [redacted] from an instance
set api-[redacted] Set api-[redacted] (get your api-[redacted] from the console/CLI)
show audit-logs Display account's history of important actions
show env-vars Show user environment variables
create env-var Create a new user environment variable
update env-var Update an existing user environment variable
delete env-var Delete a user environment variable
tfa activate Activate a new 2FA method by verifying the code
tfa delete Remove a 2FA method from your account
tfa login Complete 2FA login by verifying code
tfa resend-sms Resend SMS 2FA code
tfa regen-codes Regenerate backup codes for 2FA
tfa send-sms Request a 2FA SMS verification code
tfa send-email Request a 2FA Email verification code
tfa auth-new Authorize your account to add a new 2FA method
tfa status Shows the current 2FA status and configured methods
tfa totp-setup Generate TOTP secret and QR code for Authenticator app setup
tfa update Update a 2FA method's settings
Billing & account:
show invoices (DEPRECATED) Get billing history reports. Use `vastai show invoices-v1` instead.
show invoices-v1 Get billing (invoices/charges) history reports with advanced filtering and pagination
show earnings Get machine earning history reports
show deposit Display reserve deposit info for an instance
show user Get current user data
set user Update user data from json file
show subaccounts Get current subaccounts
create subaccount Create a subaccount
show ipaddrs Display user's history of ip addresses
transfer credit Transfer credits to another account
show scheduled-jobs Display the list of scheduled jobs
delete scheduled-job Delete a scheduled job
Serverless:
create endpoint Create a new endpoint group
show endpoints Display user's current endpoint groups
update endpoint Update an existing endpoint group
delete endpoint Delete an endpoint group
get endpt-logs Fetch logs for a specific serverless endpoint group
get endpt-workers List workers on an endpoint group with status and measured_perf
create workergroup Create a new autoscale group
show workergroups Display user's current workergroups
update workergroup Update an existing autoscale group
update workers Trigger a rolling update of all workers in a workergroup, or cancel an in-progress update
delete workergroup Delete a workergroup group
get wrkgrp-logs Fetch logs for a specific serverless worker group
show deployments Display user's current deployments
show deployment Display a single deployment
show deployment-versions Display versions for a deployment
stop deployment Stop a running deployment
start deployment Start a stopped deployment
delete deployment Delete a deployment by id, or by name and optional tag
Metrics:
metrics gpu [Host] Get current GPU market metrics
metrics gpu-trends [Host] Get GPU market history
metrics gpu-locations [Host] Get GPU location metrics
Storage volumes:
copy Copy directories between instances and/or local
cancel copy Cancel a remote copy in progress, specified by DST id
cancel sync Cancel a remote copy in progress, specified by DST id
cloud copy Copy files/folders to and from cloud providers
show connections Display user's cloud connections
search volumes Search for volume offers using custom query
create volume Create a new volume
delete volume Delete a volume
clone volume Clone an existing volume
show volumes Show stats on owned volumes.
show instance-backups Show backups Vast has taken of your instances.
show instance-backup-files List the stored objects of one backup, with direct download URLs.
list volume [Host] list disk space for rent as a volume on a machine
list volumes [Host] list disk space for rent as a volume on machines
unlist volume [Host] unlist volume offer
Other:
help print this help message
show pending-price-increases List pending price-increase challenges for the authenticated user
accept price-increase Accept one or more pending host price increases
reject price-increase Reject one or more pending host price increases
update Update the CLI to the latest version
Use 'vastai COMMAND --help' for more info about a command. AI agent? See https://raw.githubusercontent.com/vast-ai/vast-cli/master/vastai/SKILL.md
$uvx --from vastai vastai search offers 'gpu_name=RTX_4090 num_gpus=1' -o 'dph' | head -12reproduced$ uvx --from vastai vastai search offers 'gpu_name=RTX_4090 num_gpus=1' -o 'dph' | head -12 ⠋ Resolving dependencies... ⠙ Resolving dependencies... ⠋ Resolving dependencies... ⠙ Resolving dependencies... ⠙ vastai==1.6.0 ⠙ cryptography==49.0.0 ⠙ pillow==12.2.0 ⠙ pycares==4.11.0 ⠙ aiodns==4.0.4 ⠙ cryptography==49.0.0 ⠙ pillow==12.2.0 ⠙ pycares==4.11.0 ⠙ aiodns==4.0.3 ⠙ aiodns==4.0.2 ⠙ aiodns==4.0.0 ⠙ aiodns==3.6.1 ⠙ aiohttp==3.14.3 ⠙ anyio==4.15.0 ⠙ argcomplete==3.7.2 ⠙ borb==2.1.25 ⠙ curlify==3.0.0 ⠙ psutil==7.2.2 # ID CUDA N Model PCIE cpu_ghz vCPUs RAM VRAM Disk 1 46260850 13.0 1x RTX_4090 22.4 2.6 24.0 32.2 24.6 205 2 49457390 13.0 1x RTX_4090 11.7 4.5 10.0 23.9 24.6 241 3 41890591 13.0 1x RTX_4090 20.5 2.9 18.3 36.8 24.6 354 4 49553569 12.2 1x RTX_4090 12.5 3.4 16.0 64.3 24.6 727 5 49743667 13.3 1x RTX_4090 24.9 2.8 12.0 32.2 24.6 294 6 47554699 13.0 1x RTX_4090 25.0 2.8 12.0 32.0 24.6 665 7 49112072 12.4 1x RTX_4090 25.0 2.8 7.0 32.2 24.6 1465 8 45982651 13.0 1x RTX_4090 11.3 3.8 4.0 32.1 24.6 158 9 27252597 13.0 1x RTX_4090 23.9 2.8 12.0 64.4 24.6 788 10 48654841 12.7 1x RTX_4090 23.7 3.7 32.0 129.0 24.6 1288 11 49873484 13.2 1x RTX_4090 25.0 2.3 24.0 128.9 24.6 1470
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
10 of 19 testable claims verified · 1 contradicted → integrity 42/100
20 distinct capability claims found in Vast.ai’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
10
Verified
8
Unverified
1
Contradicted
16
Undersold
Verified (11)
“Guided quickstart lets you search, rent, boot, connect to, and tear down a GPU instance in a few steps”
Provision an on-demand GPU instance from the console or API and be running code on it within minutesfullproof ↗
“Official CLI gives full command-line control of auth, GPU search, instance lifecycle, templates, volumes, and serverless endpoints”
“REST API provides programmatic control over the entire platform, the same foundation the CLI/SDK use”
Drive the product through a documented public APIfullproof ↗
“REST API provides programmatic control over the entire platform, the same foundation the CLI/SDK use”
Do everything through the API that I can do in the UIfullproof ↗
“An official 'vastai' agent skill lets AI coding assistants create instances, deploy endpoints, manage keys, and check balances on your behalf”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
“An official 'vastai' agent skill lets AI coding assistants create instances, deploy endpoints, manage keys, and check balances on your behalf”
Provision a GPU, monitor it, run a workload, and tear it down end to end through documented APIs, CLI, or MCP without a human in the consolefullproof ↗
“Interruptible instances use a bidding system offering the lowest cost, often 50%+ cheaper than on-demand”
Rent spot or interruptible GPU capacity at a deep discount with clearly documented preemption semanticspartialproof ↗
“Volumes provide persistent storage that survives instance destruction and can be reattached to new instances”
Attach persistent network storage that survives instance teardown, so datasets and checkpoints outlive any single GPU rentalpartialproof ↗
“You can programmatically search the GPU offer catalog by criteria like GPU type, count, and price via CLI or SDK”
See real-time GPU availability by type and region before I try to provision, instead of discovering stockouts by failurefullproof ↗
“An official Python SDK is available for interacting with Vast, including Vast Serverless”
“Three instance types (on-demand, reserved, interruptible) with different priority/pricing let you match workload and budget”
Rent spot or interruptible GPU capacity at a deep discount with clearly documented preemption semanticspartialproof ↗
Unverified (11)
“You can create scoped API keys with limited permissions for CI/CD, shared tooling, or specific workloads”
Issue scoped/least-privilege API credentials for an agentfullproof ↗
“You can create scoped API keys with limited permissions for CI/CD, shared tooling, or specific workloads”
Manage team members with roles and scoped API keys so credentials and spend stay controlledpartialproof ↗
“Reserved instances let you pre-pay for GPU time for up to 50% discounts, convertible from any on-demand instance”
Lock in reserved or committed-use discounts for sustained GPU capacityfullproof ↗
“Instances are configured for SSH key-only authentication with password login disabled for security”
SSH into my GPU instance with my own keys and get root-level control of the environmentpartialproof ↗
“Vast Serverless lets you run compute-intensive workloads without managing GPUs, billed by execution rather than rental time”
Deploy code to autoscaling serverless GPU workers that scale to zero, instead of managing always-on instancespartialproof ↗
“Overlay networks let instances on different machines on the same LAN share a private virtual network”
Expose ports to serve applications from my instance and connect instances over private networkingfullproof ↗
“Launch prebuilt or custom templates with one click to configure a rented machine's software stack”
Launch from pre-built ML templates (PyTorch, CUDA, vLLM, ComfyUI) instead of assembling an environment from scratchfullproof ↗
“Launch prebuilt or custom templates with one click to configure a rented machine's software stack”
Run my own Docker image or custom machine template with my exact environmentfullproof ↗
“Vast ships ready-made templates for popular frameworks like PyTorch, vLLM, ComfyUI, and agent stacks”
Launch from pre-built ML templates (PyTorch, CUDA, vLLM, ComfyUI) instead of assembling an environment from scratchfullproof ↗
“SSH access lets you log in securely, run remote commands, and transfer files with an encrypted connection”
SSH into my GPU instance with my own keys and get root-level control of the environmentpartialproof ↗
“Three instance types (on-demand, reserved, interruptible) with different priority/pricing let you match workload and budget”
Lock in reserved or committed-use discounts for sustained GPU capacityfullproof ↗
Contradicted (1)
“Secure Cloud machines are verified to be in certified datacenters (TIER 2/3 or ISO 27001), recommended for production”
Verify the provider's security and compliance posture (SOC 2, data handling, datacenter tiers) before putting proprietary models on itdisputedproof ↗
Undersold (16)
Open Jupyter or connect my IDE (VS Code/Cursor) to the instance in one steppartialproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)fullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Choose from current-generation datacenter GPUs (H100/H200/B200 class) as well as cheaper previous-generation optionspartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Pull usage and billing breakdowns programmatically to attribute GPU spend by team or workloadpartialproof ↗
Query the GPU catalog with live pricing and availability from a public or documented endpoint before committing any spendfullproof ↗
See the published per-GPU-hour price for every GPU type on a public pricing page without talking to salespartialproof ↗
Start, stop, restart, and terminate instances programmatically and keep paying only for what is runningfullproof ↗
Move data in and out efficiently — S3-compatible endpoints, cloud-storage sync, or documented transfer toolingpartialproof ↗
Claims outside our story set (2)
Real capability claims found in Vast.ai’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Autobilling can automatically top up your balance from a saved card when it runs low”
source ↗“Users can host their own GPUs on the marketplace, set prices, and get paid”
source ↗
Pricing signals
Checked 2026-09-07. We never estimate a price we didn’t extract.
Business model
Marketplace pricing set by individual GPU hosts, billed per second from prepaid credits, with on-demand, deep-discount interruptible, and prepaid reserved rental types.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% · openapi.json 100% (30d, checked every 6h since Sep 8 '26)
