Skip to content

Rank #3 of 9 in AI Assistants

Grok logo

Grok

xAI (SpaceXAI) · commercial

no public signals

xAI ships more than one product — each judged line competes in its own arena on the same stories as everyone else.

LineArenaRankPA Score
Grokthis pageAI Assistants#3/926/100

Not yet judged (6 — no arena where they compete): Grok in X · Grok Build · Grok Bot · Grok Imagine · Grokipedia · xAI API

Try itExperimental

See what an agent can do with Grok before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$curl -si https://api.x.ai/v1/models | head -3; curl -s https://api.x.ai/v1/modelsrecorded session — replayed, not live
recorded 2026-09-14 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

37.6/100

Agents tasks — stories about agents tasks in this arenaAgents tasksevidence →

Stories about agents tasks in this arena

10.8/100

Apps devices — stories about apps devices in this arenaApps devicesevidence →

Stories about apps devices in this arena

8.0/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

6.0/100

Connectors apps — stories about connectors apps in this arenaConnectors appsevidence →

Stories about connectors apps in this arena

62.4/100

Files analysis — stories about files analysis in this arenaFiles analysisevidence →

Stories about files analysis in this arena

11.3/100

Memory context — stories about memory context in this arenaMemory contextevidence →

Stories about memory context in this arena

0.0/100

Multimodal — stories about multimodal in this arenaMultimodalevidence →

Stories about multimodal in this arena

73.3/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

6.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Research answers — stories about research answers in this arenaResearch answersevidence →

Stories about research answers in this arena

31.2/100

Trust controls — stories about trust controls in this arenaTrust controlsevidence →

Stories about trust controls in this arena

12.0/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 0 free · 1 paid · 0 enterprise · 18 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10T

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full7/10C

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3disputed6/10D

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/a0/10

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10X

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10X

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10T

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial4/10C

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial3/10T

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial3/10X

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1noneuntestednone yet

Connect my cloud drive, email, and calendar so the assistant can search and use them in answers C

Connectors

knowledge-workerConnectors apps — stories about connectors apps in this arenaConnectors apps3full8/10C

Delegate a multi-step task that the assistant works on autonomously in the background and returns for my review C

Agent mode

power-userAgents tasks — stories about agents tasks in this arenaAgents tasks3partial6/10X

Launch a deep research run that autonomously searches many sources and returns a cited report C

Research

knowledge-workerResearch answers — stories about research answers in this arenaResearch answers3partial6/10C

Upload documents, spreadsheets, and PDFs and get accurate analysis of their contents C

Files

knowledge-workerFiles analysis — stories about files analysis in this arenaFiles analysis3partial5/10C

Control whether my conversations are used to train models C

Data controls

knowledge-workerTrust controls — stories about trust controls in this arenaTrust controls3none0/10

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3none0/10

Have the assistant operate a web browser on my behalf to research and complete tasks on websites C

Agent mode

power-userAgents tasks — stories about agents tasks in this arenaAgents tasks3none0/10

Have the assistant write and run code on my data to produce charts, computed answers, and downloadable files C

Analysis

power-userFiles analysis — stories about files analysis in this arenaFiles analysis3none0/10

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3noneuntestednone yet

Have the assistant remember relevant context from previous chats and apply it in new conversations C

Memory

power-userMemory context — stories about memory context in this arenaMemory context3noneuntestednone yet

Generate and edit images from natural-language prompts C

Images

knowledge-workerMultimodal — stories about multimodal in this arenaMultimodal2full9/10C

Share screenshots and photos and have the assistant accurately interpret what is in them C

Images

knowledge-workerMultimodal — stories about multimodal in this arenaMultimodal2full7/10C

Browse a directory of third-party apps and connectors and add them to the assistant C

Connectors

power-userConnectors apps — stories about connectors apps in this arenaConnectors apps2partial6/10C

Have a natural, real-time voice conversation with the assistant C

Voice

knowledge-workerMultimodal — stories about multimodal in this arenaMultimodal2full6/10X

Manage members, permissions, and data policies for my organization's workspace C

Admin

team-adminTrust controls — stories about trust controls in this arenaTrust controls2partialpaid6/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial5/10T

Get answers grounded in current web results with citations back to the sources C

Research

knowledge-workerResearch answers — stories about research answers in this arenaResearch answers2partial4/10X

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial4/10C

Use full-featured official mobile apps for iOS and Android C

Apps

knowledge-workerApps devices — stories about apps devices in this arenaApps devices2partial4/10C

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2none0/10

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2none0/10

Have the assistant create and iteratively edit documents, presentations, and other files I can export G

Artifacts

knowledge-workerFiles analysis — stories about files analysis in this arenaFiles analysis2none0/10

Let the assistant see and operate applications on my computer to complete work C

Agent mode

power-userAgents tasks — stories about agents tasks in this arenaAgents tasks2none0/10

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2none0/10

Organize related chats and files into a project or space that shares context and instructions C

Projects

knowledge-workerMemory context — stories about memory context in this arenaMemory context2none0/10

Use an official desktop app with OS-level shortcuts and access to what is on my screen C

Apps

power-userApps devices — stories about apps devices in this arenaApps devices2none0/10

Build and share custom assistants with their own instructions and knowledge C

Custom bots

power-userApps devices — stories about apps devices in this arenaApps devices2noneuntestednone yet

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2noneuntestednone yet

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2noneuntestednone yet

Schedule recurring or one-off tasks that run automatically and come back to me with results C

Tasks

power-userAgents tasks — stories about agents tasks in this arenaAgents tasks2noneuntestednone yet

Export my complete chat history and account data G

Data controls

knowledge-workerTrust controls — stories about trust controls in this arenaTrust controls1noneuntestednone yet

Set persistent custom instructions and preferences that shape every response C

Memory

power-userMemory context — stories about memory context in this arenaMemory context1noneuntestednone yet

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1noneuntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 42 stories with headroom

What would move Grok’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events

    nonemoves PA Scoreimpact 30

    Grok's docs describe connectors, tool integrations, and a persistent 'Grok Bot' cloud agent, but nothing describes user-defined rules that trigger actions automatically on external events (e.g., webhooks, triggers, if-this-then-that automations, or scheduled event listeners).

  2. Agents tasks — stories about agents tasks in this arenaHave the assistant operate a web browser on my behalf to research and complete tasks on websites

    nonemoves PA Scoreimpact 30

    The evidence shows connectors for APIs, email, calendar, Salesforce, and a vague 'Grok Bot' cloud computer teammate, but nothing documents Grok actually driving a web browser to navigate sites or complete tasks (no browsing/computer-use agent feature is described).

  3. Files analysis — stories about files analysis in this arenaHave the assistant write and run code on my data to produce charts, computed answers, and downloadable files

    nonemoves PA Scoreimpact 30

    The evidence describes Grok's coding assistant (Grok Build), connectors for email/calendar/files, image/video generation, and multi-agent research, but nothing documents a code-interpreter style capability where Grok writes and executes code against user-supplied data to produce charts, computed answers, or downloadable output files.

  4. Memory context — stories about memory context in this arenaHave the assistant remember relevant context from previous chats and apply it in new conversations

    nonemoves PA Scoreimpact 30

    The evidence pack shows conversation storage (grok-docs-21) and cross-device sync of conversations/settings (grok-docs-41), but nothing describing a memory feature that recalls relevant context from past chats and applies it to new, unrelated conversations.

  5. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    No evidence pack item mentions any data export feature, open format (e.g., JSON/Markdown conversation export) or account data portability tool for Grok; documentation covers connectors, models, and billing but never data export/leave-with-your-data capability.

  6. Openness — open source, data portability, and self-hosting storiesSelf-host the core product

    nonemoves PA Scoreimpact 30

    Grok is a closed, proprietary API/cloud service; no evidence anywhere in the pack of open-sourced weights or a self-hostable core model/product, only hosted APIs, connectors, and SaaS features.

  7. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    The evidence pack documents data retention (prompts and responses stored on xAI's servers, grok-docs-21) and vague 'enterprise-grade privacy protections' for team workspaces (grok-docs-13), but nowhere mentions an opt-out toggle, training-data exclusion setting, or policy statement about excluding user conversations from model training.

  8. Trust controls — stories about trust controls in this arenaControl whether my conversations are used to train models

    nonemoves PA Scoreimpact 30

    Docs mention 'enterprise-grade privacy protections' for team workspaces and that prompts are stored on xAI servers, but there is no documented setting or toggle letting a knowledge-worker control whether their conversations are used to train models.

Showing the top 8 of 42 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 25 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$curl -si https://api.x.ai/v1/models | head -3; curl -s https://api.x.ai/v1/modelsreproduced
$ curl -si https://api.x.ai/v1/models | head -3; curl -s https://api.x.ai/v1/models
HTTP/2 401

date: Mon, 14 Sep 2026 19:12:51 GMT

content-type: application/json

{"code":"unauthenticated:no-credentials","error":"No credentials presented."}
$curl -s https://docs.x.ai/.well-known/mcp/server-card.json | head -12reproduced
$ curl -s https://docs.x.ai/.well-known/mcp/server-card.json | head -12
{
  "serverInfo": {
    "name": "xai-docs-mcp",
    "version": "1.0.0"
  },
  "transport": {
    "type": "http",
    "url": "https://docs.x.ai/api/mcp"
  },
  "capabilities": {
    "tools": true,
    "resources": false,
$curl -s https://docs.x.ai/llms.txt | grep -iE 'documentation|grok/overview|grok/user-guide'reproduced
$ curl -s https://docs.x.ai/llms.txt | grep -iE 'documentation|grok/overview|grok/user-guide'
# SpaceXAI API Documentation
> Official SpaceXAI (xAI) developer documentation at https://docs.x.ai — Grok models, REST API reference, SDKs, guides, pricing, and app docs. The API is served at https://api.x.ai.
Every documentation page is available as markdown: append `.md` to its URL, or request the page URL with an `Accept: text/markdown` header.
- [Get Started](https://docs.x.ai/grok/overview.md)

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

4 of 9 testable claims verified · 0 contradictedintegrity 44/100

19 distinct capability claims found in Grok’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

4

Verified

5

Unverified

0

Contradicted

15

Undersold

Verified (4)
Unverified (12)
Undersold (15)
Claims outside our story set (3)

Real capability claims found in Grok’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Requests can get higher scheduling priority via a 'priority' service tier setting

    source ↗
  • Video model supports text-to-video, image-to-video, and reference-to-video generation at native 1080p

    source ↗
  • Paid SuperGrok plans raise usage limits and unlock added features via a shared weekly allowance

    source ↗
Suggest a story for these →

Business model

free-tiersubscription-flatenterprise-custom

Free tier with usage limits; SuperGrok, Plus, and Heavy plans share one weekly usage pool across Chat, Imagine, Voice, and Build; Business/Enterprise add org management.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score25 (Sep 14 '26)26 (Sep 16 '26)
Agent-ready40 (Sep 14 '26)45 (Sep 16 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data