AI Assistants — procurement report
ProductArena · rankings as of 2026-09-16 · evidence as of 2026-09-16 · 9 products · 52 judged requirements · 468 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | ChatGPT | 38.3 | 50.5 | 50/52 | C |
| 2 | Claude | 28.6 | 52.7 | 50/52 | C |
| 3 | Grok | 26.1 | 23.1 | 51/52 | C |
| 4 | Perplexity | 25.0 | 18.1 | 52/52 | C |
| 5 | Poke | 23.6 | 18.6 | 49/52 | C |
| 6 | Martin | 19.3 | 16.3 | 43/52 | C |
| 7 | Gemini | 16.5 | 36.6 | 45/52 | C |
| 8 | Muse | 8.3 | 11.1 | 45/52 | D |
| 9 | Microsoft Copilot | 4.3 | 18.6 | 48/52 | C |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
The current #1/#2 gap in this arena is not close enough to qualify for the multi-judge uncertainty pass (or the pass has not covered it yet) — no extra caveat applies beyond the per-product confidence grades above.
Buyer checklist (RFP)
The arena's 52 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/ai-assistants/checklist
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Agents tasks — stories about agents tasks in this arenaAgents tasks
Stories about agents tasks in this arena
- power-userDelegate a multi-step task that the assistant works on autonomously in the background and returns for my reviewmust-have
- power-userHave the assistant operate a web browser on my behalf to research and complete tasks on websitesmust-have
- power-userLet the assistant see and operate applications on my computer to complete workshould-have
- power-userSchedule recurring or one-off tasks that run automatically and come back to me with resultsshould-have
Apps devices — stories about apps devices in this arenaApps devices
Stories about apps devices in this arena
- power-userUse an official desktop app with OS-level shortcuts and access to what is on my screenshould-have
- knowledge-workerUse full-featured official mobile apps for iOS and Androidshould-have
- power-userBuild and share custom assistants with their own instructions and knowledgeshould-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Connectors apps — stories about connectors apps in this arenaConnectors apps
Stories about connectors apps in this arena
- knowledge-workerConnect my cloud drive, email, and calendar so the assistant can search and use them in answersmust-have
- power-userBrowse a directory of third-party apps and connectors and add them to the assistantshould-have
Files analysis — stories about files analysis in this arenaFiles analysis
Stories about files analysis in this arena
- power-userHave the assistant write and run code on my data to produce charts, computed answers, and downloadable filesmust-have
- knowledge-workerUpload documents, spreadsheets, and PDFs and get accurate analysis of their contentsmust-have
- knowledge-workerHave the assistant create and iteratively edit documents, presentations, and other files I can exportshould-have
Memory context — stories about memory context in this arenaMemory context
Stories about memory context in this arena
- power-userHave the assistant remember relevant context from previous chats and apply it in new conversationsmust-have
- knowledge-workerOrganize related chats and files into a project or space that shares context and instructionsshould-have
- power-userSet persistent custom instructions and preferences that shape every responsenice-to-have
Multimodal — stories about multimodal in this arenaMultimodal
Stories about multimodal in this arena
- knowledge-workerGenerate and edit images from natural-language promptsshould-have
- knowledge-workerShare screenshots and photos and have the assistant accurately interpret what is in themshould-have
- knowledge-workerHave a natural, real-time voice conversation with the assistantshould-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Research answers — stories about research answers in this arenaResearch answers
Stories about research answers in this arena
- knowledge-workerLaunch a deep research run that autonomously searches many sources and returns a cited reportmust-have
- knowledge-workerGet answers grounded in current web results with citations back to the sourcesshould-have
Trust controls — stories about trust controls in this arenaTrust controls
Stories about trust controls in this arena
- knowledge-workerControl whether my conversations are used to train modelsmust-have
- team-adminManage members, permissions, and data policies for my organization's workspaceshould-have
- knowledge-workerExport my complete chat history and account datanice-to-have
Appendix: recorded probes
Hands-on probe recordings — transcripts/videos a human can replay, the strongest evidence tier. Watch them at https://ultrametric.ai/productarena/proofs
- Grok
curl -si https://api.x.ai/v1/models | head -3; curl -s https://api.x.ai/v1/modelsterminal · recorded 2026-09-14 · exit 0 - Grok
curl -s https://docs.x.ai/.well-known/mcp/server-card.json | head -12terminal · recorded 2026-09-14 · exit 0 - Grok
curl -s https://docs.x.ai/llms.txt | grep -iE 'documentation|grok/overview|grok/user-guide'terminal · recorded 2026-09-14 · exit 0 - Martin
curl -s https://docs.trymartin.com/introduction.md | head -8terminal · recorded 2026-09-14 · exit 0 - Martin
curl -s https://docs.trymartin.com/llms.txt | grep -E 'introduction|calendar|email-triage'terminal · recorded 2026-09-14 · exit 0 - Poke
curl -si -X POST https://poke.com/api/v1/inbound/api-message -H 'Content-Type: application/json' -d '{"message":"ping"}'terminal · recorded 2026-09-14 · exit 0 - Poke
curl -s https://poke.com/docs/api | grep -o 'api/v1/inbound/api-message'; curl -s https://poke.com/docs/mcp-servers | grep -io 'mcp server'terminal · recorded 2026-09-14 · exit 0
Cite as: ProductArena by Ultrametric Inc, AI Assistants arena, rankings as of 2026-09-16 — https://ultrametric.ai/productarena/arena/ai-assistants
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).