Skip to content

Arena

Local LLM Runtimes arenaLocal LLM Runtimes

Tools for running large language models on your own hardware, judged on model support breadth, performance across CPU/GPU/Apple Silicon, API compatibility, quantization flexibility, and fit for agentic and power-user workflows.

92 user stories · 644 judged cells · updated 2026-09-04 · Evidence as of 2026-09-04

Buyer checklist →Procurement report →

Leaderboard — every product ranked by evidenceLeaderboard

Best by user type — persona-weighted winnersBest by user type

Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.

Best for developer

vLLM logo

vLLM

41/100

Runner-up: llama.cpp logo llama.cpp (24/100)

31 developer stories scored

Best for ai-native

LocalAI logo

LocalAI

36/100

Runner-up: llamafile logo llamafile (31/100)

34 ai-native stories scored

Best for power-user

llama.cpp logo

llama.cpp

45/100

Runner-up: vLLM logo vLLM (40/100)

27 power-user stories scored

Story matrix — every product × every judged storyStory matrix

92/92 stories shown · legend

Agenticness — how well agents can access and operate the productAgenticness

Agent access

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docsai-native
fullT
8/10
none
0/10
none
0/10
fullT
8/10
none
0/10
none
0/10
partialT
4/10
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automationai-native
partialX
6/10
partialC
6/10
partialC
6/10
fullX
8/10
none
0/10
partialC
7/10
partialX
6/10
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their toolsai-native
none
0/10
none
0/10
n/a
partialX
6/10
partialC
6/10
fullC
8/10
n/a
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP serverai-native
none
0/10
none
0/10
none
0/10
n/a
n/a
fullC
7/10
n/a
Agenticness — how well agents can access and operate the productUse an official CLIai-native
fullX
8/10
fullX
8/10
none
0/10
fullX
8/10
none
0/10
fullC
8/10
fullT
7/10
Agenticness — how well agents can access and operate the productDrive the product through a documented public APIai-native
fullT
8/10
partialC
4/10
fullX
8/10
fullX
8/10
partialT
5/10
fullT
8/10
partialT
5/10
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agentai-native
none
0/10
n/a
none
0/10
n/a
none
0/10
partialC
3/10
n/a
Agenticness — how well agents can access and operate the productBuild against official SDKsai-native
fullC
7/10
none
0/10
partialX
5/10
none
0/10
none
0/10
partialT
6/10
none
0/10
Agenticness — how well agents can access and operate the productSubscribe to events via webhooksai-native
none
0/10
none
0/10
n/a
n/a
none
0/10
none
0/10
n/a
Agenticness — how well agents can access and operate the productConnect a coding agent to this product as a working backendai-native
fullC
8/10
partialC
4/10
partialX
6/10
partialX
7/10
partialT
6/10
fullC
8/10
partialT
4/10

Agentic features

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the productai-native
n/a
none
0/10
n/a
partialX
6/10
none
0/10
partialC
5/10
none
0/10
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the backgroundai-native
n/a
none
0/10
n/a
none
0/10
none
0/10
partialC
4/10
n/a
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the productai-native
none
0/10
none
0/10
n/a
partialX
6/10
partialC
6/10
fullC
8/10
none
0/10
Agenticness — how well agents can access and operate the productOperate the product with natural-language commandsai-native
partialC
5/10
none
0/10
n/a
partialX
6/10
partialC
5/10
partialC
6/10
partialX
6/10

Api quality

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examplesai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)ai-native
none
0/10
none
0/10
none
0/10
none
0/10
partialT
4/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production dataai-native
n/a
n/a
n/a
none
0/10
n/a
none
0/10
n/a
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policyai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Automation depth — how much of the product can run unattendedAutomation depth

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Automation depth — how much of the product can run unattendedPerform bulk operations across many items at onceai-native
none
0/10
none
0/10
partialX
5/10
none
0/10
none
0/10
none
0/10
none
0/10
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on eventsai-native
n/a
none
0/10
n/a
none
0/10
none
0/10
none
0/10
n/a
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflowsai-native
n/a
n/a
n/a
none
0/10
none
0/10
none
0/10
n/a
Automation depth — how much of the product can run unattendedVersion, review, and roll back my automationsai-native
n/a
n/a
n/a
n/a
none
0/10
none
0/10
n/a

Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem

Build and install

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesBuild the runtime from source with minimal external dependenciesdeveloper
none
0/10
fullX
8/10
partialC
3/10
none
0/10
partialC
4/10
none
0/10
none
0/10
Ecosystem — integrations, plugins, and third-party ecosystem storiesRun the runtime inside a container for reproducible deploymentdeveloper
partialX
3/10
fullC
8/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Ecosystem — integrations, plugins, and third-party ecosystem storiesInstall the runtime quickly using a standard package managerdeveloper
none
0/10
partialX
4/10
fullC
8/10
none
0/10
none
0/10
none
0/10
none
0/10
Ecosystem — integrations, plugins, and third-party ecosystem storiesInstall using prebuilt binaries or packages instead of compiling from sourcedeveloper
partialT
5/10
partialX
6/10
fullC
7/10
fullX
7/10
none
0/10
none
0/10
fullX
8/10

Community contribution

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesContribute code and become a recognized collaborator through the project's open-source processdeveloper
none
0/10
partialC
5/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Language bindings

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesCall the runtime from official client libraries in languages like Python or JavaScriptdeveloper
fullC
8/10
none
0/10
partialX
5/10
none
0/10
partialT
4/10
partialT
6/10
none
0/10

Licensing and cost

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesWhether commercial or enterprise use requires a paid license or subscription beyond the free community editionpower-user
none
0/10
none
0/10
n/a
none
0/10
none
0/10
n/a
n/a

Maintenance health

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesHow quickly the project ships patches for critical bugs and security vulnerabilities based on its public release historydeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Model portability

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesWhether downloaded model files and caches can be reused by other runtimes without re-downloading or re-converting themdeveloper
none
0/10
none
0/10
partialC
3/10
none
0/10
none
0/10
none
0/10
none
0/10

Privacy control

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ecosystem — integrations, plugins, and third-party ecosystem storiesRun inference entirely on my own machine so my data and prompts never leave my devicepower-user
fullT
8/10
fullX
9/10
fullX
8/10
fullX
9/10
fullX
8/10
fullC
9/10
fullX
9/10

Model support — which models run and how well — coverage, formats, update cadenceModel support

Architecture coverage

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Model support — which models run and how well — coverage, formats, update cadenceRun hundreds of different model architectures including LLMs, MoE, multi-modal, and embedding modelsdeveloper
partialT
5/10
partialX
6/10
fullX
9/10
partialX
5/10
partialC
5/10
partialC
6/10
partialX
6/10
Model support — which models run and how well — coverage, formats, update cadenceServe embedding models for retrieval and search applicationsdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Custom assistants

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Model support — which models run and how well — coverage, formats, update cadenceCreate specialized custom assistants configured for specific taskspower-user
none
0/10
partialC
3/10
partialC
5/10
partialC
4/10
partialC
5/10
fullC
8/10
partialX
5/10

Hybrid cloud local

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Model support — which models run and how well — coverage, formats, update cadenceConnect to cloud AI providers alongside local models within the same interfacepower-user
partialT
6/10
n/a
n/a
none
0/10
fullC
8/10
none
0/10
none
0/10
Model support — which models run and how well — coverage, formats, update cadenceOffload very large models to a hosted cloud tier without downloading them when my local hardware is insufficientpower-user
fullC
7/10
none
0/10
n/a
none
0/10
none
0/10
none
0/10
none
0/10
Model support — which models run and how well — coverage, formats, update cadenceThe pricing tiers, rate limits, and usage caps that apply when offloading inference to the vendor's hosted cloud tierpower-user
none
0/10
n/a
n/a
n/a
none
0/10
n/a
n/a

Model hub download

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Model support — which models run and how well — coverage, formats, update cadenceDownload and run open models directly from Hugging Facepower-user
partialX
5/10
fullX
8/10
fullX
8/10
fullX
8/10
fullC
8/10
fullC
9/10
none
0/10

Multi modal support

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Model support — which models run and how well — coverage, formats, update cadenceRun vision-language models that understand images alongside textpower-user
partialX
5/10
fullX
8/10
none
0/10
none
0/10
none
0/10
partialC
3/10
fullC
8/10

Openness — open source, data portability, and self-hosting storiesOpenness

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UIai-native
partialT
5/10
partialC
6/10
n/a
partialX
6/10
partialT
4/10
partialT
6/10
partialT
6/10
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leaveai-native
disputedD
3/10
partialX
5/10
n/a
none
0/10
none
0/10
partialC
5/10
n/a
Openness — open source, data portability, and self-hosting storiesRead the product's source under an open licenseai-native
partialX
5/10
fullC
7/10
partialC
6/10
none
0/10
partialC
5/10
none
0/10
partialC
5/10
Openness — open source, data portability, and self-hosting storiesSelf-host the core productai-native
fullX
8/10
fullX
9/10
fullX
8/10
fullX
8/10
fullC
8/10
fullC
9/10
fullX
9/10

Performance hardware — raw speed and hardware efficiency — throughput, latency, resource usePerformance hardware

Distributed serving

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useDisaggregate prefill and decode phases for optimized large-scale servingdeveloper
n/a
none
0/10
partialC
6/10
n/a
n/a
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useDistribute inference across multiple GPUs using tensor, pipeline, or data parallelismdeveloper
none
0/10
none
0/10
fullC
8/10
none
0/10
none
0/10
partialC
3/10
none
0/10

Gpu acceleration

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun inference on specialized accelerators like TPUs or Gaudi through plugin supportdeveloper
none
0/10
none
0/10
fullC
8/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun models larger than my available VRAM using combined CPU+GPU offloadpower-user
none
0/10
fullX
8/10
none
0/10
partialC
3/10
none
0/10
none
0/10
partialX
5/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useWhy GPU acceleration failed and silently fell back to CPU through clear diagnostic outputpower-user
partialX
4/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
disputedD
3/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun models on NVIDIA, AMD, or other GPU vendors using vendor-specific acceleration kernelspower-user
partialX
5/10
fullX
8/10
fullC
8/10
partialX
4/10
none
0/10
partialC
6/10
disputedD
5/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAccelerate inference on AMD GPUs via a Vulkan backend without needing a full ROCm installpower-user
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
partialC
4/10

Memory management

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useControl how context memory is allocated when running multiple model instances concurrentlypower-user
none
0/10
none
0/10
partialX
6/10
partialC
6/10
none
0/10
none
0/10
partialX
2/10

Platform acceleration

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useGet accelerated inference on Apple Silicon via native ARM and Metal optimizationspower-user
partialX
4/10
fullX
9/10
partialC
3/10
partialX
5/10
none
0/10
none
0/10
partialX
5/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRun inference on diverse CPU architectures beyond x86 and ARM, such as PowerPCdeveloper
none
0/10
none
0/10
fullC
8/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useLeverage advanced x86 CPU instruction sets like AVX, AVX2, AVX512, and AMX for faster inferencepower-user
none
0/10
fullX
8/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Startup footprint

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useGet a fast cold start from a lightweight runtime binary instead of waiting seconds before inference beginspower-user
none
0/10
fullX
7/10
none
0/10
disputedD
4/10
none
0/10
none
0/10
partialX
5/10

Throughput optimization

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAchieve high serving throughput via continuous batching and chunked prefillpower-user
none
0/10
none
0/10
fullX
9/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useRely on paged memory management for attention key/value cache to maximize concurrent request capacity without memory fragmentationdeveloper
none
0/10
none
0/10
fullX
9/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useThe runtime reserves dedicated capacity so throughput holds steady when multiple agents or sessions issue requests concurrentlypower-user
partialC
4/10
none
0/10
partialX
5/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useSpeed up repeated-prompt workloads using prefix cachingpower-user
none
0/10
none
0/10
partialX
6/10
none
0/10
none
0/10
none
0/10
none
0/10
Performance hardware — raw speed and hardware efficiency — throughput, latency, resource useAccelerate generation speed using speculative decoding techniquespower-user
none
0/10
none
0/10
partialC
6/10
none
0/10
none
0/10
none
0/10
none
0/10

Privacy posture — data-handling and privacy storiesPrivacy posture

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)ai-native
partialC
4/10
fullX
6/10
n/a
n/a
partialC
4/10
fullC
7/10
partialX
6/10
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI modelsai-native
fullC
7/10
fullX
7/10
none
0/10
partialC
5/10
partialC
6/10
fullC
8/10
fullX
8/10
Privacy posture — data-handling and privacy storiesControl data retention and deletionai-native
partialC
4/10
partialC
3/10
n/a
partialC
4/10
partialC
4/10
partialC
5/10
fullX
7/10
Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage trackingai-native
partialC
5/10
none
0/10
n/a
none
0/10
none
0/10
none
0/10
fullX
8/10

Quantization formats — stories about quantization formats in this arenaQuantization formats

Adapters

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Quantization formats — stories about quantization formats in this arenaEfficiently serve multiple LoRA adapters on top of a base modeldeveloper
none
0/10
none
0/10
fullX
8/10
none
0/10
none
0/10
none
0/10
none
0/10

File formats

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Quantization formats — stories about quantization formats in this arenaWhether upgrading the runtime can break compatibility with previously downloaded quantized model filesdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Quantization formats — stories about quantization formats in this arenaLoad and run models packaged in the GGUF formatpower-user
none
0/10
fullX
8/10
fullX
8/10
partialX
5/10
partialC
6/10
fullC
9/10
fullX
7/10

Quantization levels

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Quantization formats — stories about quantization formats in this arenaReduce memory footprint using integer quantization ranging from very low-bit to 8-bit precisionpower-user
partialX
4/10
fullX
9/10
fullX
7/10
none
0/10
none
0/10
partialC
3/10
none
0/10
Quantization formats — stories about quantization formats in this arenaLoad models quantized in formats like FP8, INT4, GPTQ, or AWQdeveloper
none
0/10
none
0/10
fullX
8/10
none
0/10
none
0/10
none
0/10
none
0/10

Serving api — serving models over an API — endpoints, compatibility, reliabilityServing api

Api compatibility

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityCall the server through an Anthropic-compatible messages endpointdeveloper
none
0/10
none
0/10
partialC
6/10
none
0/10
none
0/10
fullC
8/10
none
0/10
Serving api — serving models over an API — endpoints, compatibility, reliabilityLaunch a local OpenAI-compatible API server for any loaded modeldeveloper
partialT
6/10
partialC
6/10
fullX
8/10
fullX
9/10
fullC
7/10
fullC
9/10
partialC
5/10

Deployment modes

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityRun the runtime headlessly with no GUI for use in servers or CI pipelinesdeveloper
partialX
6/10
fullC
8/10
fullX
7/10
fullX
8/10
none
0/10
fullC
7/10
fullX
8/10

Generation controls

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityStream generated tokens back to my application as they are produceddeveloper
none
0/10
partialX
4/10
fullC
7/10
partialX
4/10
partialT
5/10
partialC
4/10
partialC
4/10
Serving api — serving models over an API — endpoints, compatibility, reliabilityConstrain model output to structured formats like JSON using grammarsdeveloper
none
0/10
fullC
7/10
partialC
6/10
none
0/10
none
0/10
none
0/10
none
0/10
Serving api — serving models over an API — endpoints, compatibility, reliabilityUse native tool-calling and reasoning-parser support in my requestsdeveloper
none
0/10
none
0/10
partialC
6/10
none
0/10
none
0/10
partialC
6/10
none
0/10

Model lifecycle

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityAssign a custom identifier to a loaded model for consistent reference in API callsdeveloper
none
0/10
none
0/10
none
0/10
fullC
8/10
none
0/10
fullC
7/10
none
0/10
Serving api — serving models over an API — endpoints, compatibility, reliabilityLoad and switch between multiple models without restarting the serverpower-user
fullX
8/10
none
0/10
partialC
4/10
partialC
6/10
partialC
4/10
partialC
5/10
none
0/10

Remote serving

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityServe models over my local network for access from other devicespower-user
partialX
4/10
partialC
6/10
partialX
5/10
partialX
6/10
partialC
5/10
partialC
6/10
partialC
5/10

Scale limits

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityThe documented maximum concurrent requests or connections the local server can handle before throughput degradesdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Server configuration

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Serving api — serving models over an API — endpoints, compatibility, reliabilityOverride low-level engine settings like memory locking or mmap behavior instead of being limited to opinionated defaultspower-user
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingUx tooling

Ai assisted setup

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingRely on an AI assistant to recommend which local model best fits my hardware and task before I download itai-native
none
0/10
none
0/10
n/a
none
0/10
none
0/10
none
0/10
none
0/10

Chat interface

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingChat with local models using a built-in graphical chat interfacepower-user
partialX
6/10
fullC
7/10
n/a
fullX
9/10
fullX
7/10
fullC
8/10
fullX
7/10

Cli tooling

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingStart an interactive chat session with a model directly from the terminaldeveloper
fullX
8/10
fullX
8/10
none
0/10
fullC
8/10
none
0/10
fullC
7/10
fullX
8/10
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingSearch, download, and manage models from a command-line interfacedeveloper
fullX
8/10
partialC
6/10
none
0/10
fullC
8/10
none
0/10
fullC
7/10
none
0/10
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingLoad a model with custom GPU offload and context length settings from the command linedeveloper
none
0/10
partialX
6/10
none
0/10
fullX
8/10
none
0/10
none
0/10
partialT
6/10
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingStart and stop the local model server from the command linedeveloper
partialX
3/10
partialC
6/10
none
0/10
fullC
9/10
none
0/10
partialC
5/10
partialT
5/10
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingLaunch popular third-party coding agent CLIs pre-configured to use my local models with a single commanddeveloper
fullC
7/10
none
0/10
n/a
none
0/10
none
0/10
none
0/10
none
0/10

Document intelligence

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingChat with my own documents entirely offline using automatic retrieval-augmented generationai-native
none
0/10
none
0/10
n/a
fullX
7/10
none
0/10
none
0/10
none
0/10
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingHave an AI agent draft and edit documents in an integrated workspace with changes saved automaticallyai-native
n/a
n/a
n/a
fullX
7/10
none
0/10
none
0/10
n/a
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingDictate speech that gets transcribed in real time by an on-device modelai-native
n/a
n/a
n/a
fullC
5/10
none
0/10
partialC
5/10
partialC
4/10

Local model management

StoryPersona
Ollama logoOllama
llama.cpp logollama.cpp
vLLM logovLLM
LM Studio logoLM Studio
Jan logoJan
LocalAI logoLocalAI
llamafile logollamafile
Ux tooling — the working surface itself — layout, ergonomics, quality-of-life toolingManage my downloaded models, saved prompts, and per-model configurations in one placepower-user
partialX
3/10
none
0/10
n/a
partialX
7/10
partialC
4/10
partialC
5/10
none
0/10
Verdict✓ fullclear evidence~ partialwith caveats! disputedevidence conflicts— noneno evidence foundn/aquestion doesn't apply to this kind of product
ProofT probedtested by usX communityusers back itC claimedvendor claim onlyD contradictedevidence disagrees⚿ auth-gatedprobe hit a live sign-in wall — verified reachable, untestable keylessly
quality 0–10 · PA Score /100 · A–D = evidence confidence · full guide

Adjacent arenas — categories often shopped togetherAdjacent arenas

Shopping this category often means shopping these too.