ChatGPT vs Muse
free-tier · subscription-flat · subscription-per-seat · enterprise-custom
·free-tier · subscription-flat
ChatGPT wins · 33–1 (9 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · round to ChatGPTChatGPT has first-party documentation for MCP support, explicitly describing configuring MCP servers (e.g., Context7) so ChatGPT/Codex can access third-party tools like documentation, browser, or Figma, with config stored in config.toml and scoping options for trusted projects, plus cross-client portability once configured. missing for 10: independent/hands-on community corroboration of MCP server usage specifically (community evidence covers plugins/browsing but not MCP directly).
- [claimed-docs] “let it interact with developer tools like your browser or Figma”
- [claimed-docs] “Once you configure your MCP servers, you can switch among those clients without redoing setup.”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
- [claimed-docs] “to add Context7 (a free MCP server for developer documentation)”
- [claimed-docs] “you can also scope MCP servers to a project with `.codex/config.toml` (trusted projects only)”
- [claimed-docs] “Codex stores MCP configuration in `config.toml`”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation”
Musenone0/10No evidence Muse supports plugging in external MCP servers; it only describes its agent building its own tools internally, which is a different mechanism.
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [claimed-docs] “Muse has its own computer, with a file system and a terminal, so it can write its own code and build the tools a task needs.”
ai-native userUse an official CLI
weight 2 · round to ChatGPTChatGPT ships an official Codex CLI with documented commands (--search, --cd, /memories, /fast, MCP config via config.toml, codex-security bulk-scan), and the llms.txt probe confirms 'Codex CLI' as a first-party interface for building and automating dev workflows, satisfying the AI-native official CLI story. Missing for 10: independent/hands-on community verification of the CLI itself, since community evidence only covers the general ChatGPT web product, not Codex CLI usage.
- [claimed-docs] “In Codex CLI, use /memories in an interactive session to control whether the current chat can use existing local memories or become an input…”
- [claimed-docs] “Run codex from the directory you want Codex to work in, or pass --cd (-C) to set it explicitly.”
- [claimed-docs] “In the CLI, pass `--search` to fetch live results for one run”
- [claimed-docs] “pass `--search` to fetch live results for one run”
- [claimed-docs] “Use the Codex SDK to automate coding tasks, including jobs in CI.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Codex reads AGENTS.md files before doing any work.”
- [claimed-docs] “Use rules to control which commands Codex can run outside the sandbox.”
- [claimed-docs] “Use `/fast on`, `/fast off`, or `/fast status` in the CLI to change or inspect the current setting.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [claimed-docs] “Use `npx @openai/codex-security bulk-scan` to review repositories in one campaign.”
- [probe] “PROBE llms.txt: HTTP 200 at https://learn.chatgpt.com/llms.txt # Codex > Build with the Codex CLI, IDE extension, and cloud automation to s…”
ai-native userDrive the product through a documented public API
weight 3 · round to ChatGPTThere is documented evidence of programmatic/scriptable access via the Codex CLI and Codex SDK ('Use the Codex SDK to automate coding tasks, including jobs in CI', CLI flags like --search, --cd, /fast, config.toml), which lets an AI-native user drive parts of the product outside the chat UI. However, a probe for a standard documented public API (openapi.json/swagger) at the docs site returned 404 on all candidate paths, and none of the evidence describes a general, versioned public API for driving ChatGPT itself (as opposed to Codex-specific tooling or MCP client configuration). Missing for 10: an explicit REST/GraphQL API reference for ChatGPT product actions, OpenAPI/swagger spec, and independent corroboration that non-Codex ChatGPT features are API-drivable.
- [claimed-docs] “Use the Codex SDK to automate coding tasks, including jobs in CI.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [claimed-docs] “Use `npx @openai/codex-security bulk-scan` to review repositories in one campaign.”
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · round to ChatGPTChatGPT/Codex docs show some least-privilege mechanisms for agent tool access — MCP servers can be scoped to a trusted project via `.codex/config.toml` (chatgpt-docs-55), sandbox command rules can restrict what Codex can run outside the sandbox (chatgpt-docs-95), and workspace admins can control access to browser use, plugins, and network access (chatgpt-docs-88), plus GPT builders can configure scoped third-party API actions (chatgpt-docs-34, chatgpt-docs-44). However there is no evidence of a user-facing mechanism to actually issue/generate a scoped or least-privilege API credential/token specifically for an agent to use. missing for 10: explicit scoped API key/token issuance workflow, granularity/expiry controls on such credentials, and any documentation or independent confirmation that a user can mint a restricted-permission credential for agent use.
- [claimed-docs] “you can also scope MCP servers to a project with `.codex/config.toml` (trusted projects only)”
- [claimed-docs] “Use rules to control which commands Codex can run outside the sandbox.”
- [claimed-docs] “Your workspace administrator can control access to ChatGPT Work, plugins, browser use, and network access.”
- [claimed-docs] “Allow GPT builders to use approved workspace apps or configure actions that interact with permitted third-party APIs.”
- [claimed-docs] “configure actions that interact with permitted third-party APIs”
Musenone0/10Evidence shows Muse has approval cards, activity logs, and permission tracking for connected accounts, but nothing about issuing scoped or least-privilege API credentials/tokens for an agent to use with external services — this is a consumer personal-agent product, not a developer credential-management tool.
ai-native userBuild against official SDKs
weight 2 · round to ChatGPTThe pack contains a single line noting 'Use the Codex SDK to automate coding tasks, including jobs in CI' (chatgpt-docs-91), indicating an official SDK exists for Codex/automation, but there is no linked documentation, API reference, code samples, or independent corroboration of this SDK's use. The openapi probe found no public API spec, reinforcing that developer-facing SDK documentation is thin in this evidence set. missing for 10: dedicated SDK documentation/reference pages, code examples, language support details, independent developer reports of building with the SDK.
- [claimed-docs] “Use the Codex SDK to automate coding tasks, including jobs in CI.”
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
ai-native userSubscribe to events via webhooks
weight 2 · round to ChatGPTChatGPT's Automations/Scheduled Tasks support event-driven triggers ('act on pull request feedback without polling on a fixed cadence') tied to Gmail, Slack, and GitHub events, which functions like a webhook-subscription mechanism, but this is not exposed as a general-purpose webhook subscription API — it's limited to a few pre-integrated services with no documented endpoint/config for arbitrary webhook URLs. Missing for 10: a generic webhook subscription API/config for third-party or custom events, documentation of payload/security handling, and any independent confirmation of reliability.
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Use an event trigger to triage new email, summarize channel activity, or act on pull request feedback without polling on a fixed cadence.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
Musenone0/10No evidence in the pack mentions webhooks or any event-subscription API for developers; Muse's documented features focus on agent tasks, permissions, and integrations, not outbound webhook subscriptions. missing for 10: any webhook/event API documentation, developer subscription mechanism, or third-party confirmation of webhook support.
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to ChatGPTChatGPT can analyze uploaded/connected data (files, projects, local folders, connected apps like Google Drive/SharePoint/Salesforce) and generate insights, summaries, and suggestions (e.g. 'Search the web, browse websites, compare sources, read files, analyze data, and summarize findings' and 'Create a deck, analyze files, draft a report, build a project plan'), with projects giving persistent context for tailored suggestions. Missing for 10: independent hands-on evidence specifically validating data-insight quality/accuracy (community evidence largely covers unrelated features like web search accuracy issues, not data-analysis insight quality).
- [claimed-docs] “Search the web, browse websites, compare sources, read files, analyze data, and summarize findings.”
- [claimed-docs] “Create a deck, analyze files, draft a report, build a project plan”
- [claimed-docs] “Use a project to organize related chats and give ChatGPT the context it needs.”
- [claimed-docs] “A local project gives chats access to one or more folders on your computer, such as a collection of source files or a codebase.”
- [claimed-docs] “Plugins can connect ChatGPT to the tools and information you use for work, such as Google Drive, SharePoint, Salesforce, or Gong.”
- [claimed-docs] “Turn research and analysis into documents, presentations, spreadsheets, and other finished work.”
Muse links personal data sources (email, calendar, social, health, finances) and proactively tracks tasks/goals, implying some data-driven suggestions, but the evidence never explicitly describes AI-generated insights or recommendations derived from that data—most documentation focuses on agentic task execution and tool-building rather than analytic insight generation. missing for 10: explicit description of insight/recommendation generation from user data, independent hands-on validation of this specific capability.
- [claimed-docs] “Link your email, calendar, Instagram, Facebook, health and fitness data, and finances.”
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to ChatGPTChatGPT's Scheduled Tasks/Automations feature explicitly lets users schedule recurring or event-triggered tasks (Gmail, Slack, GitHub events) to run autonomously in the background, with a dashboard to review active, paused, and completed runs. This is well-documented first-party functionality directly matching the story. Missing for 10: independent hands-on verification of background automation reliability and no community corroboration specifically about the automations/scheduled-tasks feature.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
- [claimed-docs] “Use an event trigger to triage new email, summarize channel activity, or act on pull request feedback without polling on a fixed cadence.”
Musedisputedcontradicted5/10Muse's Goals tab, activity log, and approval-card workflow (muse-docs-5, muse-docs-8, muse-docs-13) show first-party support for autonomous background task tracking and execution, and muse-docs-7 explicitly describes agents building tools for unattended tasks. However, an independent hands-on stress test found the agent's control plane starts timing out under load, exposing concrete failure points in its background-automation architecture (muse-comm-1), directly undercutting reliability claims for autonomous operation. missing for 10: independent confirmation that background automations run reliably at scale, more detail on scheduling/triggering mechanisms for autonomous runs.
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
- [community] “Blog post title/finding: the author stress-tested Meta Muse until its agent control plane started timing out, exposing failure points in the…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to ChatGPTChatGPT ships extensive built-in agentic delegation: Work/Codex modes carry tasks through to reviewable results, scheduled/background tasks and event-triggered automations, computer use and browser control, subagent delegation, and voice-initiated task delegation — all first-party features, not third-party add-ons. Missing for 10: independent hands-on verification of these newer agentic features (Work, subagents, Computer Use) beyond vendor docs.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Turn on Work in the switcher when you want ChatGPT to carry a larger task through to a reviewable result.”
- [claimed-docs] “Run projects in parallel, work with files, use your computer, and keep long-running work moving from one desktop workspace.”
- [claimed-docs] “Unlike Work on the Web, local Work can operate on resources that remain on your computer without requiring you to upload files to a cloud co…”
- [claimed-docs] “In ChatGPT Work, ask ChatGPT to delegate independent work to subagents.”
- [claimed-docs] “ChatGPT Voice can start separate tasks for longer work, check existing tasks, and send follow-up instructions.”
- [claimed-docs] “Browser lets ChatGPT open websites, gather current information, and take action while you stay in control.”
- [claimed-docs] “With Computer Use, ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “enter `/goal` to start Goal mode”
Muse is explicitly a built-in AI agent with its own compute, terminal, and file system to build tools and complete delegated tasks (goals, side chats, autonomous tool creation), with approval flows and activity logs for oversight. Missing for 10: independent hands-on verification of delegation quality beyond stress-test edge cases and more detail on task breadth/reliability at scale.
- [claimed-docs] “Muse has its own computer, with a file system and a terminal, so it can write its own code and build the tools a task needs.”
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
ai-native userOperate the product with natural-language commands
weight 2 · round to ChatGPTChatGPT is fundamentally a natural-language interface: users type or speak requests and it carries out tasks (browsing, coding, file edits, image generation, scheduling, computer use, voice) via plain conversational commands, per extensive first-party docs and corroborating community reports of natural-language driven capability (e.g. running Docker/filesystem commands via prompts). Missing for 10: independent hands-on verification specifically of newer agentic features (Work, Goal mode, subagents) beyond vendor docs.
- [claimed-docs] “Browser lets ChatGPT open websites, gather current information, and take action while you stay in control.”
- [claimed-docs] “With Computer Use, ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “ChatGPT Voice lets you talk through ideas and coordinate tasks in Chat, Work, and Codex in the ChatGPT desktop app. Start work, check progre…”
- [claimed-docs] “In the ChatGPT desktop app, enter `/goal` to start Goal mode.”
- [claimed-docs] “In ChatGPT Work, ask ChatGPT to delegate independent work to subagents.”
- [community] “I am beyond astounded. I was able to run a Docker image, utilize the fs inside of the container, and exit the container. Docker system comma…”
- [community] “This is huge, essentially adding what people have been building with LangChain Tools into the core product. The browser and file-upload/inte…”
Muse is designed around a chat interface where users talk to it in natural language to manage goals, tasks, and tool-building ('talk to it about these tasks in chat', side chats for topic-specific NL interaction). This is corroborated by first-party docs describing conversational control as the primary interface. missing for 10: independent hands-on confirmation of broad natural-language command coverage beyond goals/chat, and detail on command reliability at scale (comm-1 notes control-plane timeouts under stress).
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “some people still wanted separate context for certain topics — which made sense as projects grew more complex. So we built side chats.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
- [community] “Blog post title/finding: the author stress-tested Meta Muse until its agent control plane started timing out, exposing failure points in the…”
Api quality
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnChatGPTnone0/10The probe explicitly checked for an OpenAPI/machine-readable spec at all standard locations and found only 404s, with no other evidence pack item showing a downloadable API spec for ChatGPT itself.
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
Agents tasks — stories about agents tasks in this arenaAgents tasks
Stories about agents tasks in this arena
Agent mode
power-userDelegate a multi-step task that the assistant works on autonomously in the background and returns for my review
weight 3 · round to ChatGPTChatGPT's Tasks/automations feature explicitly supports scheduling and running multi-step work in the background (including event-triggered runs from Gmail/Slack/GitHub), with a review interface for active/completed runs, and 'Work' mode is described as carrying larger tasks through to a reviewable result; Codex further supports background coding tasks and subagent delegation. Missing for 10: independent hands-on verification of background task reliability/quality and no third-party review of the review/approval workflow.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “Turn on Work in the switcher when you want ChatGPT to carry a larger task through to a reviewable result.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
- [claimed-docs] “In ChatGPT Work, ask ChatGPT to delegate independent work to subagents.”
- [claimed-docs] “Start a Codex task to run the tests and investigate anything that doesn't pass.”
Musedisputedcontradicted6/10Muse's docs describe autonomous background task handling via the Goals tab, self-built tools, activity logs and approval cards for review — matching the delegate-and-review story [muse-docs-5][muse-docs-7][muse-docs-8][muse-docs-13][muse-docs-10]. However, an independent hands-on stress test found the agent's control plane timing out under load, exposing failure points in the very autonomous-task architecture being claimed [muse-comm-1], and another commenter notes the assistant is less capable than power-user expectations for building/using tools [muse-comm-2]. Missing for 10: reproducible evidence of reliable long-running multi-step task completion, and resolution/acknowledgment of the reported control-plane timeout issue.
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
- [claimed-docs] “structured approval cards with clear accept/reject actions, and secure storage for your credentials.”
- [community] “Blog post title/finding: the author stress-tested Meta Muse until its agent control plane started timing out, exposing failure points in the…”
- [community] “"I, too, have a personal agent but my friends using Muse or Instinct approach the functionality I use mine for quite easily... a sophisticat…”
power-userHave the assistant operate a web browser on my behalf to research and complete tasks on websites
weight 3 · round to ChatGPTChatGPT's Browser feature explicitly lets it open websites, gather information, and take action on the user's behalf while staying in control, and can complete multi-step tasks like comparing options or completing actions on a website; this is reinforced by Computer Use (operating GUIs) and Work/Codex browser tab integration for signed-in site tasks. missing for 10: independent hands-on verification of complex multi-step website task completion, and detail on reliability/success rate limits.
- [claimed-docs] “Browser lets ChatGPT open websites, gather current information, and take action while you stay in control.”
- [claimed-docs] “Use it to compare options, complete a multi-step task on a website, or review a page you're building.”
- [claimed-docs] “With Computer Use, ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “ChatGPT can see and operate graphical user interfaces on macOS or Windows”
- [claimed-docs] “Bring an open tab into a ChatGPT Work or Codex chat and work with the website where you're already signed in.”
- [claimed-docs] “Follow the sign-in request and enter your details in the sign-in flow, not in the chat. This doesn't connect your local browser profile.”
- [claimed-docs] “With Site tools (WebMCP), ChatGPT Work and Codex can use actions offered by a website in the desktop app's built-in browser.”
Musenone0/10Evidence describes Muse having its own computer/terminal to build tools, task tracking (Goals tab), and permissions/approval cards, but nothing in the pack mentions browser automation or web navigation on the user's behalf. Missing for 10: any mention of browser control, web research execution, or site-task completion capability.
power-userLet the assistant see and operate applications on my computer to complete work
weight 2 · round to ChatGPTChatGPT's Computer Use feature explicitly lets it see and operate GUIs on macOS/Windows, including testing desktop/mobile app flows, backed by Browser and Work agent capabilities that carry out multi-step tasks on the user's behalf. Documentation is extensive and detailed across multiple first-party sources describing screen operation, browser control, and file/app interaction. Missing for 10: independent hands-on verification of Computer Use reliability in practice.
- [claimed-docs] “With Computer Use, ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “ChatGPT can see and operate graphical user interfaces on macOS or Windows”
- [claimed-docs] “ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “Testing a macOS app, Windows app, iOS simulator flow, or another desktop app that ChatGPT is building.”
- [claimed-docs] “Browser lets ChatGPT open websites, gather current information, and take action while you stay in control.”
- [claimed-docs] “Use it to compare options, complete a multi-step task on a website, or review a page you're building.”
- [claimed-docs] “Run projects in parallel, work with files, use your computer, and keep long-running work moving from one desktop workspace.”
- [claimed-docs] “Unlike Work on the Web, local Work can operate on resources that remain on your computer without requiring you to upload files to a cloud co…”
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”
Muse describes having its own computer/terminal to build tools and links external accounts (email, calendar, social, health, finance) via structured approvals, which lets it act on the user's behalf across services, but the evidence never shows it seeing or directly operating applications on the user's own device (e.g., screen/GUI control) — it operates its own sandboxed environment, not the user's computer. Missing for 10: evidence of direct screen/GUI control of the user's local applications, and independent confirmation that this cross-app automation reliably works beyond linked-account integrations.
- [claimed-docs] “Muse has its own computer, with a file system and a terminal, so it can write its own code and build the tools a task needs.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [claimed-docs] “structured approval cards with clear accept/reject actions, and secure storage for your credentials.”
- [claimed-docs] “Link your email, calendar, Instagram, Facebook, health and fitness data, and finances.”
Tasks
power-userSchedule recurring or one-off tasks that run automatically and come back to me with results
weight 2 · round to ChatGPTChatGPT's Automations/Scheduled Tasks feature explicitly supports scheduling recurring or one-off tasks that run in the background and report back, plus event-triggered tasks from Gmail/Slack/GitHub, with a UI to review active, paused, and completed runs. missing for 10: independent/hands-on community corroboration of scheduled task reliability and no detail on notification/result-delivery mechanics beyond docs.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
- [claimed-docs] “Use an event trigger to triage new email, summarize channel activity, or act on pull request feedback without polling on a fixed cadence.”
Muse's Goals tab and activity log suggest it can track and manage ongoing tasks with an audit trail of completed/planned actions, implying some form of persistent task tracking, but there is no explicit documentation of recurring/scheduled triggers or a scheduling UI. Community evidence also shows the agent's control plane can time out under stress, raising doubts about reliability for automated recurring runs. missing for 10: explicit scheduling/recurrence configuration docs, evidence of one-off vs recurring task setup, and independent confirmation that scheduled tasks reliably complete and report back.
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [community] “Blog post title/finding: the author stress-tested Meta Muse until its agent control plane started timing out, exposing failure points in the…”
Apps devices — stories about apps devices in this arenaApps devices
Stories about apps devices in this arena
Apps
power-userUse an official desktop app with OS-level shortcuts and access to what is on my screen
weight 2 · round to ChatGPTThe evidence confirms an official ChatGPT desktop app (macOS/Windows) that can 'use your computer' and run 'Computer Use' to see and operate GUIs, including screenshot-based screen access, plus desktop-only features like floating Pets controls and multi-browser support. However, there is no documentation of OS-level global keyboard shortcuts (e.g., a system-wide hotkey to invoke the app) or of general 'what's on my screen' querying outside the specific Computer Use/testing use case. Missing for 10: explicit OS-level shortcut/hotkey documentation, general screen-content awareness beyond Computer Use testing scenarios, independent hands-on confirmation of these desktop-specific features.
- [claimed-docs] “Run projects in parallel, work with files, use your computer, and keep long-running work moving from one desktop workspace.”
- [claimed-docs] “Work in Edge, Brave, Opera, or Vivaldi as well as Chrome from the ChatGPT desktop app.”
- [claimed-docs] “With Computer Use, ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “ChatGPT can see and operate graphical user interfaces on macOS or Windows”
- [claimed-docs] “ChatGPT can see and operate graphical user interfaces on macOS or Windows.”
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”
- [claimed-docs] “Type a request or start a voice conversation from the floating Pets controls in the ChatGPT desktop app on macOS and Windows.”
- [claimed-docs] “Testing a macOS app, Windows app, iOS simulator flow, or another desktop app that ChatGPT is building.”
Musenone0/10Evidence describes Muse as an AI agent with its own computer/file system, chat, and permission tracking, but there is no mention of an official desktop app, OS-level keyboard shortcuts, or screen access/observation capability. This appears to be a mobile/app-based assistant rather than a desktop tool with screen access.
knowledge-workerUse full-featured official mobile apps for iOS and Android
weight 2 · round to ChatGPTThe iOS App Store listing (chatgpt-docs-20/21/41/52/58/59) documents rich mobile features—voice mode, photo upload, image generation, stickers—confirming a full-featured iOS app, and doc references (chatgpt-docs-67) explicitly mention continuing tasks 'in ChatGPT Work on the web, iOS, or Android,' indicating Android parity. Missing for 10: no dedicated Android app store evidence pack, no independent hands-on reviews of the mobile apps' feature completeness or performance.
- [claimed-docs] “Photo upload—Snap or upload a picture to transcribe a handwritten recipe or get info about a landmark.”
- [claimed-docs] “Advanced Voice Mode–Tap the soundwave icon to have a real-time convo on the go.”
- [claimed-docs] “Advanced Voice Mode–Tap the soundwave icon to have a real-time convo on the go. Settle a dinner table debate, or practice a new language.”
- [claimed-docs] “Snap or upload a picture to transcribe a handwritten recipe or get info about a landmark.”
- [claimed-docs] “Image generation–Generate original images from a description, or transform existing ones with a few simple words.”
- [claimed-docs] “Choose an image and style, remix with emojis, and make personalized stickers that are ready to share in your group chat.”
- [claimed-docs] “continue a task that needs a website account in ChatGPT Work on the web, iOS, or Android”
There is clear evidence of an iOS app (App Store reviews cited via apps.apple.com), but no evidence at all of an Android app, and no detailed description of feature parity between platforms. missing for 10: Android app existence, feature-completeness claims for either platform, independent verification of app quality.
- [claimed-docs] “I find it more pleasant than Gemini or ChatGPT which is my default for this type of work.”
- [claimed-docs] “I’d been curious to try OpenClaw, but didn’t want to deal with the security concerns or headache of setting it up.”
Custom bots
power-userBuild and share custom assistants with their own instructions and knowledge
weight 2 · round to ChatGPTDocs confirm a 'GPT builder' concept for creating custom GPTs with configurable actions/connected apps and a dedicated 'gpts-and-sharing' doc implying sharing, but no evidence details custom instructions or uploading a knowledge base, and no independent/hands-on corroboration of the sharing flow. Missing for 10: explicit mention of setting custom instructions, uploading knowledge files, and community/hands-on validation of building and sharing a custom GPT.
- [claimed-docs] “Allow GPT builders to use approved workspace apps or configure actions that interact with permitted third-party APIs.”
- [claimed-docs] “configure actions that interact with permitted third-party APIs”
- [claimed-docs] “GPT builders can use either connected apps or custom actions, but not both in the same GPT.”
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to ChatGPTChatGPT/Codex docs show concrete bulk-style automation: 'bulk-scan' scanning many repositories in one campaign (chatgpt-docs-98), a CLI for scanning multiple repos over time (chatgpt-docs-97), running multiple projects in parallel (chatgpt-docs-32), and delegating work to subagents (chatgpt-docs-103), plus background scheduled tasks with multiple runs (chatgpt-docs-37). This covers meaningful bulk/parallel automation but is concentrated in coding/security contexts rather than general-purpose bulk operations across arbitrary item sets, and there's no independent/hands-on corroboration of these bulk claims. Missing for 10: evidence of bulk operations on non-code items (e.g., bulk document/data processing), and independent verification that bulk-scan/subagent delegation works reliably at scale.
- [claimed-docs] “Use `npx @openai/codex-security bulk-scan` to review repositories in one campaign.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [claimed-docs] “Run projects in parallel, work with files, use your computer, and keep long-running work moving from one desktop workspace.”
- [claimed-docs] “In ChatGPT Work, ask ChatGPT to delegate independent work to subagents.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
Musenone0/10No evidence describes Muse performing bulk operations across many items; docs mention goals tracking, avatars, side chats, and tool-building, but nothing about batch/bulk processing of items. The community stress-test post even highlights the agent timing out under load rather than handling bulk tasks smoothly.
ai-native userDefine rules that trigger actions automatically on events
weight 3 · round to ChatGPTChatGPT's automations/scheduled tasks feature explicitly supports event-driven triggers (Gmail, Slack, GitHub) instead of just fixed schedules, letting users define rules that fire actions on external events, and lets them review active/paused/completed runs. Missing for 10: independent or hands-on verification that these event triggers work reliably in practice, and detail on how flexible/general the rule definitions can be beyond the three named integrations.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “Use an event trigger to triage new email, summarize channel activity, or act on pull request feedback without polling on a fixed cadence.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
Musenone0/10The evidence describes Muse's goal tracking, tool-building, and approval workflows, but there is no mention of user-defined conditional rules or event-triggered automations (e.g., 'when X happens, do Y'). Missing for 10: any documentation of a rules/trigger engine, event-based automation configuration, or examples of users setting conditional actions.
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to ChatGPTChatGPT's Scheduled Tasks/Automations feature explicitly supports recurring background jobs ('Schedule recurring tasks to run in the background,' with review of active/paused/completed runs) and even event-based triggers (Gmail, Slack, GitHub) as an alternative to fixed cadence. This directly matches the story of scheduling recurring jobs/workflows with first-party docs. Missing for 10: independent/hands-on community corroboration of the scheduling feature specifically (evidence pack community items don't address automations/scheduling) and more detail on reliability/limits in practice.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
- [claimed-docs] “Use an event trigger to triage new email, summarize channel activity, or act on pull request feedback without polling on a fixed cadence.”
Musenone0/10While Muse has a Goals tab that tracks ongoing tasks and an activity log, there is no evidence describing recurring/scheduled job execution, cron-like automation, or repeat-workflow triggers. missing for 10: explicit scheduling/recurrence feature, evidence of periodic execution, workflow automation configuration.
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnChatGPT supports scheduled/recurring automations with a review of active/paused/completed task runs (chatgpt-docs-11, chatgpt-docs-37, chatgpt-docs-47), which gives some review and status visibility, but there is no evidence of version history, diffing between automation versions, or a rollback mechanism to revert an automation to a prior state. missing for 10: explicit versioning of automations, diff/comparison between versions, and a documented rollback/revert feature.
- [claimed-docs] “Schedule recurring tasks to run in the background.”
- [claimed-docs] “Schedule recurring tasks to run in the background... Review active, paused, and completed tasks and recent runs in Scheduled.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
Muse provides an activity log/audit trail and approval cards that support reviewing what the agent has done and plans to do (muse-docs-8, muse-docs-13, muse-docs-10), which covers the 'review' portion of the story. However, there is no evidence of explicit versioning of automations/skills or a rollback mechanism to revert an automation to a prior state — the 'Forget' skill only removes stored information, not automation history. missing for 10: explicit versioning of automations, a rollback/undo feature for agent actions, independent confirmation these review tools work reliably in practice.
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
- [claimed-docs] “structured approval cards with clear accept/reject actions, and secure storage for your credentials.”
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
Connectors apps — stories about connectors apps in this arenaConnectors apps
Stories about connectors apps in this arena
Connectors
power-userBrowse a directory of third-party apps and connectors and add them to the assistant
weight 2 · round to ChatGPTChatGPT documents a Plugins/connectors ecosystem — installing named plugins (Slack, Codex Security), connecting to third-party tools like Google Drive, SharePoint, Salesforce, and Gong, and GPT builders choosing 'connected apps' or custom actions from an approved list — which together describe browsing and adding third-party integrations to the assistant. Missing for 10: independent/hands-on confirmation of the actual browse-and-add directory UI/UX and its breadth beyond the named examples.
- [claimed-docs] “Install the Slack plugin to summarize channels or draft replies.”
- [claimed-docs] “Install the Codex Security plugin to scan authorized code and confirm plausible vulnerability findings.”
- [claimed-docs] “Plugins can connect ChatGPT to the tools and information you use for work, such as Google Drive, SharePoint, Salesforce, or Gong.”
- [claimed-docs] “Allow GPT builders to use approved workspace apps or configure actions that interact with permitted third-party APIs.”
- [claimed-docs] “GPT builders can use either connected apps or custom actions, but not both in the same GPT.”
- [claimed-docs] “configure actions that interact with permitted third-party APIs”
Muse lets users link specific third-party data sources like email, calendar, Instagram, Facebook, health and finance data (muse-docs-11), which shows some connector-adding capability, but there is no evidence of a browsable directory or marketplace of third-party apps/connectors to explore and add. Missing for 10: an actual app/connector directory UI, discovery/browsing mechanism, and independent confirmation of a marketplace beyond ad-hoc account linking.
- [claimed-docs] “Link your email, calendar, Instagram, Facebook, health and fitness data, and finances.”
knowledge-workerConnect my cloud drive, email, and calendar so the assistant can search and use them in answers
weight 3 · round drawnDocs show ChatGPT can connect to Google Drive/SharePoint via plugins and can trigger on Gmail events, giving cloud-drive and email connectivity, and MCP/plugins provide a general connector framework for third-party tools. However, there is no explicit mention of calendar integration, and all evidence is vendor documentation with no independent/hands-on confirmation that these connectors actually surface content in answers. Missing for 10: explicit calendar connector support, independent corroboration of connector functionality, and clearer 'search across all three data sources in one answer' evidence.
- [claimed-docs] “Plugins can connect ChatGPT to the tools and information you use for work, such as Google Drive, SharePoint, Salesforce, or Gong.”
- [claimed-docs] “Install the Slack plugin to summarize channels or draft replies.”
- [claimed-docs] “Scheduled tasks can now start when a supported event occurs in Gmail, Slack, or GitHub.”
- [claimed-docs] “scheduled tasks can run when a supported Gmail, Slack, or GitHub event occurs”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
Muse docs explicitly support linking email and calendar (and other data sources) with approval/permission controls, but no evidence mentions a cloud drive connector (e.g., Google Drive/Dropbox), and there's no hands-on confirmation that these connected sources are actually searched/used in answers beyond permission approval cards. missing for 10: explicit cloud drive integration, independent evidence of search/use in answers.
- [claimed-docs] “Link your email, calendar, Instagram, Facebook, health and fitness data, and finances.”
- [claimed-docs] “structured approval cards with clear accept/reject actions, and secure storage for your credentials.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
Files analysis — stories about files analysis in this arenaFiles analysis
Stories about files analysis in this arena
Analysis
power-userHave the assistant write and run code on my data to produce charts, computed answers, and downloadable files
weight 3 · round to ChatGPTChatGPT's Code Interpreter/Data Analysis capability (documented via file drafting, spreadsheet/PDF handling, interactive visualizations, and downloadable finished files) lets it write and execute code against uploaded data to produce charts and computed answers, then export results as documents/spreadsheets/PDFs. Docs explicitly describe analyzing files, building interactive visualizations, and downloading completed files, and community evidence corroborates real code/file execution capability (e.g., running code/containers, file interpretation plugins). Missing for 10: no dedicated first-party doc page specifically titled 'Code Interpreter/Data Analysis' in this pack, and no hands-on example showing a specific chart-from-CSV walkthrough.
- [claimed-docs] “Draft and refine documents, presentations, spreadsheets, and PDF files. Review the result, ask for specific changes, and download the comple…”
- [claimed-docs] “make interactive [visualizations], and build or share websites and apps with [Sites]”
- [claimed-docs] “Create a deck, analyze files, draft a report, build a project plan”
- [claimed-docs] “Search the web, browse websites, compare sources, read files, analyze data, and summarize findings.”
- [claimed-docs] “Draft and refine [documents, presentations, spreadsheets, and PDF files]. Review the result, ask for specific changes, and download the comp…”
- [community] “This is huge, essentially adding what people have been building with LangChain Tools into the core product. The browser and file-upload/inte…”
Muse's docs claim it 'has its own computer, with a file system and a terminal' and can 'write its own code and build the tools a task needs,' which supports general code-execution capability, but there's no explicit mention of producing charts, computed answers, or downloadable files for user data analysis. A community stress-test report also notes the agent's control plane can time out under load, raising doubts about reliability for heavier computational tasks. Missing for 10: explicit documentation or examples of chart generation, data analysis outputs, downloadable file creation, and independent verification of successful code-execution results.
- [claimed-docs] “Muse has its own computer, with a file system and a terminal, so it can write its own code and build the tools a task needs.”
- [claimed-docs] “If a task needs a tool that doesn't exist, your agent builds it for you.”
- [community] “Blog post title/finding: the author stress-tested Meta Muse until its agent control plane started timing out, exposing failure points in the…”
Artifacts
knowledge-workerHave the assistant create and iteratively edit documents, presentations, and other files I can export
weight 2 · round to ChatGPTDocs strongly support creating and iteratively editing documents, presentations, spreadsheets, and PDFs, with review/annotation and export/download workflows (chatgpt-docs-8, chatgpt-docs-30, chatgpt-docs-60, chatgpt-docs-66, chatgpt-docs-68, chatgpt-docs-70), plus a dedicated artifacts/canvas viewer supporting annotations and comments for iterative refinement (chatgpt-docs-7, chatgpt-docs-18, chatgpt-docs-24, chatgpt-docs-101). Missing for 10: independent hands-on corroboration specifically of document/presentation export workflows (community evidence pack focuses on other capabilities like coding/search, not file export).
- [claimed-docs] “Draft and refine documents, presentations, spreadsheets, and PDF files.”
- [claimed-docs] “Draft and refine documents, presentations, spreadsheets, and PDF files. Review the result, ask for specific changes, and download the comple…”
- [claimed-docs] “Draft and refine [documents, presentations, spreadsheets, and PDF files]. Review the result, ask for specific changes, and download the comp…”
- [claimed-docs] “Create a deck, analyze files, draft a report, build a project plan”
- [claimed-docs] “Turn research and analysis into documents, presentations, spreadsheets, and other finished work.”
- [claimed-docs] “Use ChatGPT Work to turn notes, docs, research, or meeting materials into a structured deck.”
- [claimed-docs] “Annotations let you point to a specific part of a file and tell ChatGPT what to change.”
- [claimed-docs] “a document editor can provide tools to find a section or add a comment”
- [claimed-docs] “Use annotations to point at a specific part of a supported preview and request a focused revision.”
Musenone0/10Muse's evidence focuses on its file system/terminal, tool-building, goals, avatars, and privacy controls, but nothing describes creating or iteratively editing documents, presentations, or exportable files for knowledge-work tasks. missing for 10: any mention of document/presentation creation, editing UI, or export functionality.
Files
knowledge-workerUpload documents, spreadsheets, and PDFs and get accurate analysis of their contents
weight 3 · round to ChatGPTDocs confirm document/spreadsheet/PDF upload and analysis (draft/refine documents, presentations, spreadsheets, PDFs; 'analyze files'; 'read files, analyze data, and summarize findings'), which directly matches the story. However, evidence is entirely first-party documentation with no independent hands-on validation of accuracy specifically for uploaded file analysis (community evidence covers unrelated web-search accuracy issues, not file analysis). Missing for 10: independent/hands-on verification of analysis accuracy on real uploaded documents/spreadsheets/PDFs, and detail on limits (file size, complex spreadsheet formulas, OCR quality).
- [claimed-docs] “Draft and refine documents, presentations, spreadsheets, and PDF files.”
- [claimed-docs] “Draft and refine documents, presentations, spreadsheets, and PDF files. Review the result, ask for specific changes, and download the comple…”
- [claimed-docs] “Draft and refine [documents, presentations, spreadsheets, and PDF files]. Review the result, ask for specific changes, and download the comp…”
- [claimed-docs] “Create a deck, analyze files, draft a report, build a project plan”
- [claimed-docs] “Search the web, browse websites, compare sources, read files, analyze data, and summarize findings.”
Memory context — stories about memory context in this arenaMemory context
Stories about memory context in this arena
Memory
power-userHave the assistant remember relevant context from previous chats and apply it in new conversations
weight 3 · round to ChatGPTChatGPT's Memories feature explicitly carries useful context from earlier chats into future conversations, with dedicated docs on memory controls (personalize.md, customization/memories.md) and even CLI-level controls (/memories) for whether a chat can use or contribute to memories. Missing for 10: independent/hands-on community corroboration of cross-chat memory recall working reliably, and detail on limits/scope of what memories retain.
- [claimed-docs] “Memories let ChatGPT and Codex carry useful context from earlier work into future work.”
- [claimed-docs] “Memories let ChatGPT carry useful context from earlier chats into future work.”
- [claimed-docs] “In Codex CLI, use /memories in an interactive session to control whether the current chat can use existing local memories or become an input…”
- [claimed-docs] “use `/memories` to choose whether a chat can use local memories or contribute to future memories”
Muse's docs imply persistent memory (a 'Forget skill' to remove remembered info about topics/people, plus 'side chats' for separate context threads and a Goals tab tracking ongoing tasks) suggesting it retains and reuses context across conversations, but there is no first-party or hands-on description of how memory is surfaced or applied in new chats. Missing for 10: explicit documentation of cross-chat memory recall/application, independent verification that remembered context actually surfaces in new conversations, and detail on memory scope/limits.
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
- [claimed-docs] “some people still wanted separate context for certain topics — which made sense as projects grew more complex. So we built side chats.”
- [claimed-docs] “So we built side chats.”
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
power-userSet persistent custom instructions and preferences that shape every response
weight 1 · round to ChatGPTChatGPT supports persistent custom instructions/personalization and Memories that carry context across chats, plus Codex's global AGENTS.md for persistent personal instructions in coding workflows, and fine-grained control via /memories. Missing for 10: independent hands-on verification that custom instructions reliably shape *every* response over long-term use, and no detail on limits/scope conflicts between memories and per-chat overrides.
- [claimed-docs] “Memories let ChatGPT and Codex carry useful context from earlier work into future work.”
- [claimed-docs] “Memories let ChatGPT carry useful context from earlier chats into future work.”
- [claimed-docs] “In Codex, these personal instructions are stored in your global `AGENTS.md` file.”
- [claimed-docs] “use `/memories` to choose whether a chat can use local memories or contribute to future memories”
- [claimed-docs] “In Codex CLI, use /memories in an interactive session to control whether the current chat can use existing local memories or become an input…”
Projects
knowledge-workerOrganize related chats and files into a project or space that shares context and instructions
weight 2 · round to ChatGPTChatGPT's Projects feature explicitly lets users organize related chats, share context/instructions, and even attach local folders/files for shared context, directly matching the story. Memories also reinforce carrying context across chats. Missing for 10: independent/hands-on corroboration of the Projects feature working as described, and detail on instruction-sharing UI beyond docs claims.
- [claimed-docs] “Use a project to organize related chats and give ChatGPT the context it needs.”
- [claimed-docs] “A local project gives chats access to one or more folders on your computer, such as a collection of source files or a codebase.”
- [claimed-docs] “Memories let ChatGPT and Codex carry useful context from earlier work into future work.”
- [claimed-docs] “Memories let ChatGPT carry useful context from earlier chats into future work.”
Muse offers 'side chats' for separate context on topics as work grows complex and a Goals tab for tracking tasks, suggesting some contextual organization, but there is no explicit 'project/space' construct that bundles chats+files+shared instructions together. Missing for 10: explicit project/workspace container, file-attachment grouping, shared custom instructions across chats, and independent verification of how side chats share or isolate context.
- [claimed-docs] “The Goals tab is where Muse keeps track of all this for you, and you can interact directly there or talk to it about these tasks in chat.”
- [claimed-docs] “some people still wanted separate context for certain topics — which made sense as projects grew more complex. So we built side chats.”
- [claimed-docs] “So we built side chats.”
Multimodal — stories about multimodal in this arenaMultimodal
Stories about multimodal in this arena
Images
knowledge-workerGenerate and edit images from natural-language prompts
weight 2 · round to ChatGPTChatGPT's docs explicitly cover image generation and editing from natural-language prompts, including editing via Comment/annotations, reference images, and Canvas view for reviewing multiple images, plus mobile app support for generating/transforming images and creating stickers. This directly matches the knowledge-worker story of generating and editing images conversationally. Missing for 10: no independent/hands-on community corroboration specifically for image generation quality or reliability.
- [claimed-docs] “Ask ChatGPT to generate or edit images.”
- [claimed-docs] “use Comment to add precise feedback to one or more images”
- [claimed-docs] “Generate or edit images, make interactive visualizations, and build or share websites and apps with Sites.”
- [claimed-docs] “Ask ChatGPT to generate or edit images. Use image generation for UI assets, banners, backgrounds, illustrations”
- [claimed-docs] “Switch between Focused view to inspect one image and Canvas view to see the images generated in the same chat.”
- [claimed-docs] “In Canvas view, use Comment to add precise feedback to one or more images.”
- [claimed-docs] “Add a reference image when you want ChatGPT to transform an existing asset or use it as visual guidance.”
- [claimed-docs] “Image generation–Generate original images from a description, or transform existing ones with a few simple words.”
- [claimed-docs] “Choose an image and style, remix with emojis, and make personalized stickers that are ready to share in your group chat.”
Musenone0/10The evidence pack contains no mention of image generation, editing, or any visual/multimodal content creation capability for Muse; it focuses on agent/task automation, avatars, chat organization, and privacy controls. missing for 10: any documentation of image generation from prompts, image editing tools, or multimodal output examples.
knowledge-workerShare screenshots and photos and have the assistant accurately interpret what is in them
weight 2 · round to ChatGPTFirst-party docs confirm ChatGPT accepts photo/screenshot uploads and interprets their content (e.g., transcribing a handwritten recipe or identifying a landmark from a photo), which directly matches the knowledge-worker use case of sharing images for analysis. Missing for 10: independent/hands-on evidence corroborating accuracy of image interpretation, and explicit documentation of screenshot-specific analysis (e.g., UI screenshots) rather than just general photo uploads.
- [claimed-docs] “Photo upload—Snap or upload a picture to transcribe a handwritten recipe or get info about a landmark.”
- [claimed-docs] “Snap or upload a picture to transcribe a handwritten recipe or get info about a landmark.”
- [claimed-docs] “Image generation–Generate original images from a description, or transform existing ones with a few simple words.”
Musenone0/10No evidence in the pack discusses image, screenshot, or photo interpretation capabilities of Muse; all citations focus on agent architecture, avatars, goals, side chats, and privacy controls. Missing for 10: any mention of image/vision input support, screenshot sharing, or accuracy of visual interpretation.
Voice
knowledge-workerHave a natural, real-time voice conversation with the assistant
weight 2 · round to ChatGPTChatGPT ships Advanced Voice Mode for real-time spoken conversation on mobile ('Tap the soundwave icon to have a real-time convo on the go') and ChatGPT Voice on desktop that lets users talk through ideas, start work, check progress, or change direction without switching to typing. This directly matches the knowledge-worker story of natural, real-time voice conversation. Missing for 10: independent hands-on corroboration of voice quality/latency and any community verification beyond vendor docs.
- [claimed-docs] “Advanced Voice Mode–Tap the soundwave icon to have a real-time convo on the go.”
- [claimed-docs] “Advanced Voice Mode–Tap the soundwave icon to have a real-time convo on the go. Settle a dinner table debate, or practice a new language.”
- [claimed-docs] “ChatGPT Voice lets you talk through ideas and coordinate tasks in Chat, Work, and Codex in the ChatGPT desktop app. Start work, check progre…”
- [claimed-docs] “Start work, check progress, or change direction without switching back to typing.”
- [claimed-docs] “ChatGPT Voice lets you talk through ideas and coordinate tasks in Chat, Work, and Codex in the ChatGPT desktop app.”
- [claimed-docs] “ChatGPT Voice can start separate tasks for longer work, check existing tasks, and send follow-up instructions.”
Musenone0/10No evidence describes real-time voice conversation capability for Muse; the evidence pack focuses on agent tool-building, goals, side chats, avatar, and privacy features, with no mention of voice interaction or speech interface. Missing for 10: any documentation of voice input/output, real-time conversational voice mode, or hands-on account of speaking with Muse.
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to ChatGPTThere is evidence of programmatic access via the Codex CLI/SDK for coding automation and CI jobs, and a security-scan CLI, but nothing shows that the broad set of UI-only capabilities (Canvas/image editing, Computer Use, Voice, Work desktop automations, Sites, MCP-in-chat, memories, scheduled tasks) are exposed through a general ChatGPT API with parity to the UI. A probe for an OpenAPI/swagger spec on the docs site also returned 404s, reinforcing the lack of a documented full API surface. missing for 10: documented API endpoints mirroring Canvas/image tools, Computer Use, Voice, Work automations, Sites, and MCP-in-chat; independent confirmation of API-UI feature parity.
- [claimed-docs] “Use the Codex SDK to automate coding tasks, including jobs in CI.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnChatGPTnone0/10No evidence in the pack of a data export feature producing open/portable formats, nor any mention of account data export or deletion workflow. Evidence only covers general product features (Codex, Work, MCP, image gen, etc.) with nothing about exporting personal chat/data history or leaving the platform with your data intact.
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnChatGPTnone0/10No evidence in the pack addresses data residency or regional storage choice for ChatGPT; the evidence only touches general data controls (e.g., Computer Use screenshots) without any region-selection capability.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnChatGPTnone0/10The evidence pack contains no documentation or mention of data-training opt-out controls, privacy settings, or 'improve the model for everyone' toggles; only tangential data-controls mention (chatgpt-docs-89) refers to Computer Use screenshots, not training opt-out generally. Missing for 10: any reference to training-data opt-out settings, business/API data-usage policies, or privacy dashboard controls.
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”
Musenone0/10The evidence pack covers Muse's agent features, permissions, forget skill, and audit trail, but contains no mention of any opt-out or control over whether user data is used to train AI models. missing for 10: explicit opt-out/consent setting for model training, any privacy policy language addressing training-data usage.
ai-native userControl data retention and deletion
weight 2 · round to MuseThe only relevant evidence is a passing reference that 'ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use,' implying some data-control/retention settings exist, but no documentation details how to view, export, or delete data, set retention periods, or manage memory deletion. Missing for 10: explicit data retention/export/delete documentation, memory deletion controls, and independent confirmation that these controls work as described.
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”
Muse offers a 'Forget' skill to remove information about specific topics/people and provides an activity log/audit trail of actions, giving users some visibility and control, but there's no documentation of full account/data deletion, export, or granular retention policy settings. Missing for 10: explicit data export/deletion controls, retention period settings, and independent verification that 'Forget' fully purges underlying data rather than just suppressing recall.
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
- [claimed-docs] “Tapping on your Muse avatar will show you its full activity log as well as permissions you’ve approved.”
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnChatGPTnone0/10None of the evidence addresses telemetry/usage-tracking opt-out controls for ChatGPT; the docs cover data controls tangentially (e.g., data usage for Computer Use screenshots) but nothing about disabling telemetry or usage analytics.
Musenone0/10No evidence pack items mention a telemetry/usage-tracking opt-out setting; there is a 'Forget skill' for topic memory and privacy page references, but nothing about disabling telemetry or usage analytics collection. Missing for 10: any documentation of a telemetry/analytics opt-out toggle, privacy policy language on usage data collection, or user reports confirming such a control exists.
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
Research answers — stories about research answers in this arenaResearch answers
Stories about research answers in this arena
Research
knowledge-workerLaunch a deep research run that autonomously searches many sources and returns a cited report
weight 3 · round to ChatGPTDocs show ChatGPT can search the web, browse multiple sources, compare them, and produce cited results in-chat (docs-86, docs-90), and can turn research into finished documents/reports (docs-66, docs-68, docs-60). However, there is no explicit mention of a dedicated 'Deep Research' autonomous multi-source research mode/run with a structured long-form cited report as its own distinct feature — the evidence only shows general web-search-with-citations and document drafting capabilities. missing for 10: an explicitly named deep-research mode, evidence of autonomous long-running multi-source research runs, and a structured final cited-report output format.
- [claimed-docs] “Search results and citations appear in the chat when ChatGPT uses web search.”
- [claimed-docs] “Search the web, browse websites, compare sources, read files, analyze data, and summarize findings.”
- [claimed-docs] “Create a deck, analyze files, draft a report, build a project plan”
- [claimed-docs] “Turn research and analysis into documents, presentations, spreadsheets, and other finished work.”
- [claimed-docs] “Draft and refine [documents, presentations, spreadsheets, and PDF files]. Review the result, ask for specific changes, and download the comp…”
- [claimed-docs] “In the CLI, pass `--search` to fetch live results for one run”
Musenone0/10No evidence describes Muse performing a deep research run that autonomously searches multiple sources and returns a cited report; the pack focuses on avatar personalization, goals/side chats, permissions, and tool-building, none of which document a research/citation workflow. Missing for 10: any mention of multi-source search, report generation, or citation output.
knowledge-workerGet answers grounded in current web results with citations back to the sources
weight 2 · round to ChatGPTChatGPT's web search docs explicitly state that search results and citations appear in chat when web search is used, and describe searching, browsing, comparing sources, and summarizing findings — directly matching the story. Community feedback shows mixed satisfaction with search accuracy (e.g., comparisons to Perplexity, a weather inaccuracy) but doesn't concretely show citations failing to appear, so this doesn't rise to a dispute. Missing for 10: independent verification of citation accuracy/consistency across many queries and no first-party detail on citation formatting/source diversity.
- [claimed-docs] “Search results and citations appear in the chat when ChatGPT uses web search.”
- [claimed-docs] “Search the web, browse websites, compare sources, read files, analyze data, and summarize findings.”
- [claimed-docs] “In the CLI, pass `--search` to fetch live results for one run”
- [community] “AKA Bing Search in ChatGPT. So it is not it's own search engine and is still using Bing for its results just like the rest of them.”
- [community] “I gave it a quick spin and my initial impression is much worse than perplexity.”
- [community] “I asked it the current weather in my area and the temperature was off by 23 degrees F.”
Trust controls — stories about trust controls in this arenaTrust controls
Stories about trust controls in this arena
Data controls
knowledge-workerExport my complete chat history and account data
weight 1 · round drawnChatGPTnone0/10The evidence pack contains no mention of data export, account data download, or chat history export features anywhere in the docs, community, or probe results; there is only a passing reference to 'data controls' applying to Computer Use content, which does not address exporting complete chat history or account data. Missing for 10: any documentation of an export data feature, its scope (chats, files, settings), format, or process, and any independent corroboration it works.
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”
Musenone0/10Evidence shows a 'Forget' skill and audit trail/activity log, but there is no mention of a data export or account data download feature for chat history or account data. missing for 10: export/download tool for chat history, account data export functionality, documentation of data portability process.
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
knowledge-workerControl whether my conversations are used to train models
weight 3 · round drawnChatGPTnone0/10The evidence pack covers ChatGPT's agentic/feature capabilities (Codex, Work, MCP, Computer Use, etc.) but contains no documentation or mention of data controls, training opt-out settings, or 'Improve the model for everyone' toggles that let a user control whether their conversations are used for model training.
Musenone0/10The evidence pack shows privacy-adjacent features like a 'Forget skill' for memory and audit trails/permissions, but nothing addresses whether users can opt their conversations out of model training specifically. This is a fair axis for a personal AI agent product, but no evidence documents a training-data control. Missing for 10: any documentation of a training opt-out setting, data-usage policy toggle, or explicit statement about whether conversations are used for model training.
- [claimed-docs] “You can also use the Forget skill to have Muse forget information about specific topics or people.”
- [claimed-docs] “Muse shows people a complete audit trail of everything it has done and plans to do.”
Not comparable on these axes
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · not comparableChatGPT's Browser and MCP features let a user direct the agent to fetch and use third-party documentation or websites (e.g. Context7 for developer docs, browsing arbitrary URLs), which would technically allow pointing it at an llms.txt or agent-oriented doc page. The evidence pack also shows ChatGPT's own docs site publishes an llms.txt (probe-1), indicating familiarity with the convention, but there is no explicit documented feature or example of a user instructing ChatGPT to consume an llms.txt file specifically. Missing for 10: explicit product support/example for llms.txt ingestion, confirmation that browsing normalizes/parses such files as agent-context rather than generic web content, independent hands-on verification.
- [probe] “PROBE llms.txt: HTTP 200 at https://learn.chatgpt.com/llms.txt # Codex > Build with the Codex CLI, IDE extension, and cloud automation to s…”
- [claimed-docs] “Browser lets ChatGPT open websites, gather current information, and take action while you stay in control.”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
- [claimed-docs] “Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation”
- [claimed-docs] “to add Context7 (a free MCP server for developer documentation)”
ai-native userRun the product headlessly / in CI for automation
weight 2 · not comparableChatGPT's Codex offers a CLI and SDK explicitly documented for headless/CI use ('Use the Codex SDK to automate coding tasks, including jobs in CI'), plus CLI flags (--search, /fast) and a bulk-scan command (npx @openai/codex-security bulk-scan) that support non-interactive automation workflows. missing for 10: independent/hands-on confirmation of actual CI pipeline runs, and more detail on authentication/headless setup specifics for CI environments.
- [claimed-docs] “Use the Codex SDK to automate coding tasks, including jobs in CI.”
- [claimed-docs] “Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [claimed-docs] “Use `npx @openai/codex-security bulk-scan` to review repositories in one campaign.”
- [claimed-docs] “Use `/fast on`, `/fast off`, or `/fast status` in the CLI to change or inspect the current setting.”
- [claimed-docs] “In the CLI, pass `--search` to fetch live results for one run”
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableChatGPT/Codex has official first-party documentation for MCP support: users can add official MCP servers (e.g., Context7 for docs, Figma, browser tools) to extend the agent, configuration is stored in config.toml, servers can be scoped per-project, and setup carries across clients without redoing it. This directly evidences 'connecting an agent via an official MCP server.' Missing for 10: independent/hands-on community corroboration of MCP server usage and broader detail on the range of officially supported/verified servers beyond the Context7 example.
- [claimed-docs] “let it interact with developer tools like your browser or Figma”
- [claimed-docs] “Once you configure your MCP servers, you can switch among those clients without redoing setup.”
- [claimed-docs] “Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.”
- [claimed-docs] “to add Context7 (a free MCP server for developer documentation)”
- [claimed-docs] “you can also scope MCP servers to a project with `.codex/config.toml` (trusted projects only)”
- [claimed-docs] “Codex stores MCP configuration in `config.toml`”
ai-native userExplore an interactive API reference with runnable examples
weight 2 · not comparableChatGPTnone0/10The evidence pack contains no documentation of an interactive API reference or runnable code examples; a direct probe for OpenAPI/swagger specs on the docs site returned 404 across all candidate paths, and no other citation mentions an API reference sandbox or runnable snippets.
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
ai-native userTest against a sandbox environment without touching production data
weight 1 · not comparableCodex documentation indicates commands run inside a sandbox by default, with rules governing what can run 'outside the sandbox' (chatgpt-docs-95), and Codex Security's scanning is scoped to repos you own/have permission to assess (chatgpt-docs-97/98), implying isolated execution rather than direct production access. However, there is no explicit documentation describing a dedicated 'test/staging' environment distinct from production data, no detail on how production systems/data are excluded, and no independent/hands-on confirmation of this isolation. Missing for 10: explicit sandbox-vs-production data separation docs, details on network/data isolation guarantees, and independent verification of the sandbox boundary holding in practice.
- [claimed-docs] “Use rules to control which commands Codex can run outside the sandbox.”
- [claimed-docs] “Use its command-line interface (CLI) to scan repositories you own or have permission to assess, review findings over time”
- [claimed-docs] “Use `npx @openai/codex-security bulk-scan` to review repositories in one campaign.”
Musen/aMuse is a consumer personal AI agent (linking real email, calendar, social accounts) rather than a developer/testing product with sandbox vs production environments; the story's axis of testing against a sandbox without touching production data is a category error for this kind of product.
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · not comparableChatGPTnone0/10No evidence of a versioned API or a documented deprecation policy is present; the OpenAPI probe explicitly returned 404s and none of the docs mention API versioning or deprecation practices.
- [probe] “PROBE openapi: all candidate paths 404 (https://learn.chatgpt.com/openapi.json, https://learn.chatgpt.com/swagger.json, https://learn.chatgp…”
ai-native userRead the product's source under an open license
weight 2 · not comparableChatGPTn/aChatGPT is a closed-source, proprietary SaaS product; open-sourcing its source code is not a plausible axis for this kind of product (unlike an open-source framework or tool), so this is a category mismatch rather than a missing capability.
ai-native userSelf-host the core product
weight 3 · not comparableChatGPTn/aChatGPT is a closed, hosted proprietary product with no evidence of a self-hostable core model or server; self-hosting is not a fair axis for this SaaS product category (it does not ship open weights or an installable core).
team-adminManage members, permissions, and data policies for my organization's workspace
weight 2 · not comparableOnly a single explicit doc line notes that a workspace administrator can control access to ChatGPT Work, plugins, browser use, and network access, plus general mentions of GPT builders using 'approved workspace apps' and data controls applying to processed content. There is no evidence of a full admin console covering member invitation/removal, granular role-based permissions, or explicit data retention/training-opt-out policy controls. missing for 10: admin console/member management UI, granular role/permission settings, explicit data-retention and training-opt-out policy controls, independent corroboration of these admin features.
- [claimed-docs] “Your workspace administrator can control access to ChatGPT Work, plugins, browser use, and network access.”
- [claimed-docs] “Allow GPT builders to use approved workspace apps or configure actions that interact with permitted third-party APIs.”
- [claimed-docs] “configure actions that interact with permitted third-party APIs”
- [claimed-docs] “Your ChatGPT data controls apply to content processed through ChatGPT, including screenshots taken by Computer Use.”