Local LLM Runtimes arenaLocal LLM Runtimes
Tools for running large language models on your own hardware, judged on model support breadth, performance across CPU/GPU/Apple Silicon, API compatibility, quantization flexibility, and fit for agentic and power-user workflows.
92 user stories · 644 judged cells · updated 2026-09-04 · Evidence as of 2026-09-04
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | 53/100 | 47/100 | 0/100 | 43/100 | untested | ★ 49.1k▲ 14.1k/yr | 4/42 verified | 0/100 integrity | |||
| 2 | 25/100 | untested | 0/100 | 62/100 | 30/100 | ★ 91.7k▲ 25.5k/yr | 21/40 verified | 43/100 integrity | |||
| 3 | free-tier vs LocalAI ↗ | 56/100 | 28/100 | 0/100 | 31/100 | 0/100 | 27/39 verified · 1 disputed | 72/100 integrity | |||
| 4 | free-tier vs LocalAI ↗ | 42/100 | 12/100 | 0/100 | 39/100 | untested | ★ 180.8k▲ 56.2k/yrnpm 518.6k/wk | 27/39 verified · 1 disputed | 40/100 integrity | ||
| 5 | 30/100 | 10/100 | 0/100 | 57/100 | 0/100 | ★ 26k▲ 8.6k/yr | 26/35 verified · 2 disputed | 30/100 integrity | |||
| 6 | 17/100 | 0/100 | 0/100 | 57/100 | untested | ★ 128.1k▲ 36.5k/yr | 22/38 verified | 69/100 integrity | |||
| 7 | 15/100 | 19/100 | 8/100 | 35/100 | untested | ★ 44.5k▲ 14.4k/yr | 8/26 verified | 14/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 8/10 | none 0/10 | none 0/10 | fullT 8/10 | none 0/10 | none 0/10 | partialT 4/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | partialX 6/10 | partialC 6/10 | partialC 6/10 | fullX 8/10 | none 0/10 | partialC 7/10 | partialX 6/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | none 0/10 | none 0/10 | n/a | partialX 6/10 | partialC 6/10 | fullC 8/10 | n/a |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | none 0/10 | none 0/10 | none 0/10 | n/a | n/a | fullC 7/10 | n/a |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | fullX 8/10 | fullX 8/10 | none 0/10 | fullX 8/10 | none 0/10 | fullC 8/10 | fullT 7/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | fullT 8/10 | partialC 4/10 | fullX 8/10 | fullX 8/10 | partialT 5/10 | fullT 8/10 | partialT 5/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | n/a | none 0/10 | n/a | none 0/10 | partialC 3/10 | n/a |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | fullC 7/10 | none 0/10 | partialX 5/10 | none 0/10 | none 0/10 | partialT 6/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | none 0/10 | none 0/10 | n/a | n/a | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productConnect a coding agent to this product as a working backend | ai-native | fullC 8/10 | partialC 4/10 | partialX 6/10 | partialX 7/10 | partialT 6/10 | fullC 8/10 | partialT 4/10 |
Agentic features
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | n/a | none 0/10 | n/a | partialX 6/10 | none 0/10 | partialC 5/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | n/a | none 0/10 | n/a | none 0/10 | none 0/10 | partialC 4/10 | n/a |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | none 0/10 | none 0/10 | n/a | partialX 6/10 | partialC 6/10 | fullC 8/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | partialC 5/10 | none 0/10 | n/a | partialX 6/10 | partialC 5/10 | partialC 6/10 | partialX 6/10 |
Api quality
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialT 4/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | n/a | n/a | n/a | none 0/10 | n/a | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | none 0/10 | none 0/10 | partialX 5/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | n/a | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | n/a |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | n/a | n/a | n/a | none 0/10 | none 0/10 | none 0/10 | n/a |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | n/a | n/a | n/a | n/a | none 0/10 | none 0/10 | n/a |
Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem
Build and install
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesBuild the runtime from source with minimal external dependencies | developer | none 0/10 | fullX 8/10 | partialC 3/10 | none 0/10 | partialC 4/10 | none 0/10 | none 0/10 |
| Ecosystem — integrations, plugins, and third-party ecosystem storiesRun the runtime inside a container for reproducible deployment | developer | partialX 3/10 | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Ecosystem — integrations, plugins, and third-party ecosystem storiesInstall the runtime quickly using a standard package manager | developer | none 0/10 | partialX 4/10 | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Ecosystem — integrations, plugins, and third-party ecosystem storiesInstall using prebuilt binaries or packages instead of compiling from source | developer | partialT 5/10 | partialX 6/10 | fullC 7/10 | fullX 7/10 | none 0/10 | none 0/10 | fullX 8/10 |
Community contribution
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesContribute code and become a recognized collaborator through the project's open-source process | developer | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Language bindings
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesCall the runtime from official client libraries in languages like Python or JavaScript | developer | fullC 8/10 | none 0/10 | partialX 5/10 | none 0/10 | partialT 4/10 | partialT 6/10 | none 0/10 |
Licensing and cost
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesWhether commercial or enterprise use requires a paid license or subscription beyond the free community edition | power-user | none 0/10 | none 0/10 | n/a | none 0/10 | none 0/10 | n/a | n/a |
Maintenance health
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesHow quickly the project ships patches for critical bugs and security vulnerabilities based on its public release history | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Model portability
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesWhether downloaded model files and caches can be reused by other runtimes without re-downloading or re-converting them | developer | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Privacy control
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ecosystem — integrations, plugins, and third-party ecosystem storiesRun inference entirely on my own machine so my data and prompts never leave my device | power-user | fullT 8/10 | fullX 9/10 | fullX 8/10 | fullX 9/10 | fullX 8/10 | fullC 9/10 | fullX 9/10 |
Model support — which models run and how well — coverage, formats, update cadenceModel support
Architecture coverage
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Model support — which models run and how well — coverage, formats, update cadenceRun hundreds of different model architectures including LLMs, MoE, multi-modal, and embedding models | developer | partialT 5/10 | partialX 6/10 | fullX 9/10 | partialX 5/10 | partialC 5/10 | partialC 6/10 | partialX 6/10 |
| Model support — which models run and how well — coverage, formats, update cadenceServe embedding models for retrieval and search applications | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Custom assistants
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Model support — which models run and how well — coverage, formats, update cadenceCreate specialized custom assistants configured for specific tasks | power-user | none 0/10 | partialC 3/10 | partialC 5/10 | partialC 4/10 | partialC 5/10 | fullC 8/10 | partialX 5/10 |
Hybrid cloud local
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Model support — which models run and how well — coverage, formats, update cadenceConnect to cloud AI providers alongside local models within the same interface | power-user | partialT 6/10 | n/a | n/a | none 0/10 | fullC 8/10 | none 0/10 | none 0/10 |
| Model support — which models run and how well — coverage, formats, update cadenceOffload very large models to a hosted cloud tier without downloading them when my local hardware is insufficient | power-user | fullC 7/10 | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Model support — which models run and how well — coverage, formats, update cadenceThe pricing tiers, rate limits, and usage caps that apply when offloading inference to the vendor's hosted cloud tier | power-user | none 0/10 | n/a | n/a | n/a | none 0/10 | n/a | n/a |
Model hub download
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Model support — which models run and how well — coverage, formats, update cadenceDownload and run open models directly from Hugging Face | power-user | partialX 5/10 | fullX 8/10 | fullX 8/10 | fullX 8/10 | fullC 8/10 | fullC 9/10 | none 0/10 |
Multi modal support
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Model support — which models run and how well — coverage, formats, update cadenceRun vision-language models that understand images alongside text | power-user | partialX 5/10 | fullX 8/10 | none 0/10 | none 0/10 | none 0/10 | partialC 3/10 | fullC 8/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 5/10 | partialC 6/10 | n/a | partialX 6/10 | partialT 4/10 | partialT 6/10 | partialT 6/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | disputedD 3/10 | partialX 5/10 | n/a | none 0/10 | none 0/10 | partialC 5/10 | n/a |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | partialX 5/10 | fullC 7/10 | partialC 6/10 | none 0/10 | partialC 5/10 | none 0/10 | partialC 5/10 |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | fullX 8/10 | fullX 9/10 | fullX 8/10 | fullX 8/10 | fullC 8/10 | fullC 9/10 | fullX 9/10 |
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource usePerformance hardware
Distributed serving
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useDisaggregate prefill and decode phases for optimized large-scale serving | developer | n/a | none 0/10 | partialC 6/10 | n/a | n/a | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useDistribute inference across multiple GPUs using tensor, pipeline, or data parallelism | developer | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 |
Gpu acceleration
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun inference on specialized accelerators like TPUs or Gaudi through plugin support | developer | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun models larger than my available VRAM using combined CPU+GPU offload | power-user | none 0/10 | fullX 8/10 | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | partialX 5/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useWhy GPU acceleration failed and silently fell back to CPU through clear diagnostic output | power-user | partialX 4/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | disputedD 3/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun models on NVIDIA, AMD, or other GPU vendors using vendor-specific acceleration kernels | power-user | partialX 5/10 | fullX 8/10 | fullC 8/10 | partialX 4/10 | none 0/10 | partialC 6/10 | disputedD 5/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAccelerate inference on AMD GPUs via a Vulkan backend without needing a full ROCm install | power-user | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 |
Memory management
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useControl how context memory is allocated when running multiple model instances concurrently | power-user | none 0/10 | none 0/10 | partialX 6/10 | partialC 6/10 | none 0/10 | none 0/10 | partialX 2/10 |
Platform acceleration
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useGet accelerated inference on Apple Silicon via native ARM and Metal optimizations | power-user | partialX 4/10 | fullX 9/10 | partialC 3/10 | partialX 5/10 | none 0/10 | none 0/10 | partialX 5/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun inference on diverse CPU architectures beyond x86 and ARM, such as PowerPC | developer | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useLeverage advanced x86 CPU instruction sets like AVX, AVX2, AVX512, and AMX for faster inference | power-user | none 0/10 | fullX 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Startup footprint
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useGet a fast cold start from a lightweight runtime binary instead of waiting seconds before inference begins | power-user | none 0/10 | fullX 7/10 | none 0/10 | disputedD 4/10 | none 0/10 | none 0/10 | partialX 5/10 |
Throughput optimization
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAchieve high serving throughput via continuous batching and chunked prefill | power-user | none 0/10 | none 0/10 | fullX 9/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRely on paged memory management for attention key/value cache to maximize concurrent request capacity without memory fragmentation | developer | none 0/10 | none 0/10 | fullX 9/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useThe runtime reserves dedicated capacity so throughput holds steady when multiple agents or sessions issue requests concurrently | power-user | partialC 4/10 | none 0/10 | partialX 5/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useSpeed up repeated-prompt workloads using prefix caching | power-user | none 0/10 | none 0/10 | partialX 6/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAccelerate generation speed using speculative decoding techniques | power-user | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | partialC 4/10 | fullX 6/10 | n/a | n/a | partialC 4/10 | fullC 7/10 | partialX 6/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | fullC 7/10 | fullX 7/10 | none 0/10 | partialC 5/10 | partialC 6/10 | fullC 8/10 | fullX 8/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | partialC 4/10 | partialC 3/10 | n/a | partialC 4/10 | partialC 4/10 | partialC 5/10 | fullX 7/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | partialC 5/10 | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | fullX 8/10 |
Quantization formats — stories about quantization formats in this arenaQuantization formats
Adapters
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Quantization formats — stories about quantization formats in this arenaEfficiently serve multiple LoRA adapters on top of a base model | developer | none 0/10 | none 0/10 | fullX 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
File formats
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Quantization formats — stories about quantization formats in this arenaWhether upgrading the runtime can break compatibility with previously downloaded quantized model files | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Quantization formats — stories about quantization formats in this arenaLoad and run models packaged in the GGUF format | power-user | none 0/10 | fullX 8/10 | fullX 8/10 | partialX 5/10 | partialC 6/10 | fullC 9/10 | fullX 7/10 |
Quantization levels
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Quantization formats — stories about quantization formats in this arenaReduce memory footprint using integer quantization ranging from very low-bit to 8-bit precision | power-user | partialX 4/10 | fullX 9/10 | fullX 7/10 | none 0/10 | none 0/10 | partialC 3/10 | none 0/10 |
| Quantization formats — stories about quantization formats in this arenaLoad models quantized in formats like FP8, INT4, GPTQ, or AWQ | developer | none 0/10 | none 0/10 | fullX 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Serving api — serving models over an API — endpoints, compatibility, reliabilityServing api
Api compatibility
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityCall the server through an Anthropic-compatible messages endpoint | developer | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 |
| Serving api — serving models over an API — endpoints, compatibility, reliabilityLaunch a local OpenAI-compatible API server for any loaded model | developer | partialT 6/10 | partialC 6/10 | fullX 8/10 | fullX 9/10 | fullC 7/10 | fullC 9/10 | partialC 5/10 |
Deployment modes
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityRun the runtime headlessly with no GUI for use in servers or CI pipelines | developer | partialX 6/10 | fullC 8/10 | fullX 7/10 | fullX 8/10 | none 0/10 | fullC 7/10 | fullX 8/10 |
Generation controls
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityStream generated tokens back to my application as they are produced | developer | none 0/10 | partialX 4/10 | fullC 7/10 | partialX 4/10 | partialT 5/10 | partialC 4/10 | partialC 4/10 |
| Serving api — serving models over an API — endpoints, compatibility, reliabilityConstrain model output to structured formats like JSON using grammars | developer | none 0/10 | fullC 7/10 | partialC 6/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Serving api — serving models over an API — endpoints, compatibility, reliabilityUse native tool-calling and reasoning-parser support in my requests | developer | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 |
Model lifecycle
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityAssign a custom identifier to a loaded model for consistent reference in API calls | developer | none 0/10 | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | fullC 7/10 | none 0/10 |
| Serving api — serving models over an API — endpoints, compatibility, reliabilityLoad and switch between multiple models without restarting the server | power-user | fullX 8/10 | none 0/10 | partialC 4/10 | partialC 6/10 | partialC 4/10 | partialC 5/10 | none 0/10 |
Remote serving
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityServe models over my local network for access from other devices | power-user | partialX 4/10 | partialC 6/10 | partialX 5/10 | partialX 6/10 | partialC 5/10 | partialC 6/10 | partialC 5/10 |
Scale limits
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityThe documented maximum concurrent requests or connections the local server can handle before throughput degrades | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Server configuration
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Serving api — serving models over an API — endpoints, compatibility, reliabilityOverride low-level engine settings like memory locking or mmap behavior instead of being limited to opinionated defaults | power-user | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingUx tooling
Ai assisted setup
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingRely on an AI assistant to recommend which local model best fits my hardware and task before I download it | ai-native | none 0/10 | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Chat interface
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingChat with local models using a built-in graphical chat interface | power-user | partialX 6/10 | fullC 7/10 | n/a | fullX 9/10 | fullX 7/10 | fullC 8/10 | fullX 7/10 |
Cli tooling
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingStart an interactive chat session with a model directly from the terminal | developer | fullX 8/10 | fullX 8/10 | none 0/10 | fullC 8/10 | none 0/10 | fullC 7/10 | fullX 8/10 |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingSearch, download, and manage models from a command-line interface | developer | fullX 8/10 | partialC 6/10 | none 0/10 | fullC 8/10 | none 0/10 | fullC 7/10 | none 0/10 |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingLoad a model with custom GPU offload and context length settings from the command line | developer | none 0/10 | partialX 6/10 | none 0/10 | fullX 8/10 | none 0/10 | none 0/10 | partialT 6/10 |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingStart and stop the local model server from the command line | developer | partialX 3/10 | partialC 6/10 | none 0/10 | fullC 9/10 | none 0/10 | partialC 5/10 | partialT 5/10 |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingLaunch popular third-party coding agent CLIs pre-configured to use my local models with a single command | developer | fullC 7/10 | none 0/10 | n/a | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Document intelligence
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingChat with my own documents entirely offline using automatic retrieval-augmented generation | ai-native | none 0/10 | none 0/10 | n/a | fullX 7/10 | none 0/10 | none 0/10 | none 0/10 |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingHave an AI agent draft and edit documents in an integrated workspace with changes saved automatically | ai-native | n/a | n/a | n/a | fullX 7/10 | none 0/10 | none 0/10 | n/a |
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingDictate speech that gets transcribed in real time by an on-device model | ai-native | n/a | n/a | n/a | fullC 5/10 | none 0/10 | partialC 5/10 | partialC 4/10 |
Local model management
| Story | Persona | |||||||
|---|---|---|---|---|---|---|---|---|
| Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingManage my downloaded models, saved prompts, and per-model configurations in one place | power-user | partialX 3/10 | none 0/10 | n/a | partialX 7/10 | partialC 4/10 | partialC 5/10 | none 0/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.