Skip to content

Codex wins · 438 (15 drawn)

Agenticness — how well agents can access and operate the productAgenticness

How well agents can access and operate the product

Agent access

  1. ai-native userPoint an agent at llms.txt or agent-oriented docs

    weight 2 · round to Slate

    A dedicated llms.txt file is absent (404 at platform.openai.com/llms.txt), but Codex does publish machine-readable markdown docs (learn.chatgpt.com/docs/codex/cli.md) confirmed reachable by probe, which is an agent-friendly doc format an AI agent could be pointed at. Missing for 10: a standard llms.txt manifest, evidence of agents actually being pointed at these docs, and confirmation across all doc pages (docs/codex.md also 404s).

    • [probe] PROBE llms.txt: HTTP 404 at https://platform.openai.com/llms.txt
    • [probe] PROBE docs-md: HTTP 404 at https://platform.openai.com/docs/codex.md
    • [probe] official CLI documented at https://learn.chatgpt.com/docs/codex/cli
    Slatefullprobed8/10

    A probe confirms Slate's docs site serves a valid llms.txt at the root with links to actual docs pages, directly satisfying the ability to point an agent at agent-oriented docs. Missing for 10: no independent/community confirmation that agents successfully consume this llms.txt in practice, and no broader agent-oriented doc format beyond the single file.

    • [probe] PROBE llms.txt: HTTP 200 at https://docs.randomlabs.ai/llms.txt # Slate ## Docs - [Introduction](https://docs.randomlabs.ai/en/getting-sta…
  2. ai-native userRun the product headlessly / in CI for automation

    weight 2 · round to Codex

    Codex CLI explicitly documents non-interactive execution via `codex exec` for use in repeatable workflows, scripts, and CI/CD pipelines (codex-docs-19, codex-docs-32), and permissions/sandbox controls can be configured for unattended runs (codex-docs-17, codex-docs-39). Missing for 10: no independent case study or CI provider (e.g. GitHub Actions) integration example, and no explicit exit-code/output-format spec for CI parsing.

    • [claimed-docs] Run a non-interactive command in a repeatable workflow.
    • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.
    • [claimed-docs] Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.
    • [claimed-docs] Set the boundaries for each run — /permissions: Choose when Codex can edit files or run commands without asking, and inspect the active sand…
    Slatenone0/10

    The docs describe Slate as an interactive terminal agent (onboarding, slash commands, hotkeys, subagent cards) with a permission-bypass flag (--dangerously-skip-permissions/--yolo), but there is no mention of a headless mode, non-interactive CLI flags, exit-code/scriptable output, or CI integration examples. Missing for 10: documented headless/non-interactive invocation, CI pipeline examples, scriptable output format, and any evidence of automation use outside the interactive TUI.

    • [claimed-docs] We support `--dangerously-skip-permissions` (alias: `--yolo`) to bypass permission prompts.
    • [claimed-docs] npm i -g @randomlabs/slate
    • [claimed-docs] Use `/sessions` to switch between existing sessions
  3. ai-native userPlug MCP servers into this product so it can use their tools

    weight 3 · round to Codex

    Codex CLI explicitly supports adding local/remote MCP servers via `codex mcp add`, inspecting available tools before use, and viewing active servers via `/mcp`; this configuration is shared across ChatGPT desktop app, CLI, and IDE extension. Docs also describe using MCP to connect to third-party tools like browsers or Figma. Missing for 10: independent hands-on verification of MCP tool usage in a real session beyond first-party docs.

    • [claimed-docs] Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.
    • [claimed-docs] The ChatGPT desktop app, Codex CLI, and IDE extension share this configuration. Once you configure your MCP servers, you can switch among th…
    • [claimed-docs] Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…
    • [claimed-docs] Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…
    • [claimed-docs] codex mcp add <server-name> --env VAR1=VALUE1 --env VAR2=VALUE2 -- <stdio server-command>
    • [claimed-docs] In the `codex` TUI, use `/mcp` to see your active MCP servers.
    Slatenone0/10

    No evidence in the pack mentions MCP servers or integrating external tool providers into Slate; the docs cover skills, permissions, orchestration, and CLI usage but never MCP support. Missing for 10: any documentation or claim about connecting/plugging in MCP servers, configuring MCP tool sources, or using MCP-provided tools.

    • ai-native userUse an official CLI

      weight 2 · round to Codex
      Codexfullprobed9/10

      Codex ships an official, well-documented CLI (npm install -g @openai/codex) with rich agentic capabilities: local repo editing, exec/non-interactive scripting, MCP support, subagents, image input, sandbox/permissions control, cloud task delegation, and shell completions — all first-party documented and confirmed via GitHub repo and docs. Missing for 10: independent hands-on benchmarking specifically of CLI workflows (community evidence focuses mostly on model quality/UX rather than CLI mechanics) and some Linux-specific gaps noted by users.

      • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
      • [github] npm install -g @openai/codex
      • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.
      • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.
      • [claimed-docs] Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…
      • [claimed-docs] Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…
      • [claimed-docs] Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.
      • [claimed-docs] Install the Codex CLI with the standalone installer for macOS and Linux.
      • [probe] official CLI documented at https://learn.chatgpt.com/docs/codex/cli
      Slatefullprobed8/10

      Slate is delivered as an official CLI (npm-installed, terminal-based) with rich first-party docs covering install, sessions, hotkeys, shell execution, and configuration — squarely matching the 'official CLI' story for an AI-native user. Missing for 10: independent/hands-on corroboration of the CLI experience itself (community evidence found only relates to unrelated porting-quality claims, not CLI usage).

      • [claimed-docs] npm i -g @randomlabs/slate
      • [claimed-docs] Use `/sessions` to switch between existing sessions
      • [claimed-docs] Press Tab to queue the current message so it runs after the current turn finishes.
      • [claimed-docs] Execute shell commands directly with `!`
      • [claimed-docs] Ctrl+X then N New session
      • [probe] official CLI documented at https://docs.randomlabs.ai/en/getting-started/quickstart
    • ai-native userDrive the product through a documented public API

      weight 3 · round to Codex

      Codex documents multiple programmatic entry points — an MCP server interface for JSON-RPC control (though explicitly marked deprecated/experimental in favor of an undocumented 'app server'), a non-interactive `codex exec` mode for scripts/CI, and 'API key' usage — but these come with real caveats: API-key use 'requires additional setup', the flagship gpt-5.3-codex model was reportedly not yet available via API, and the primary MCP server route is deprecated rather than a stable first-class API. missing for 10: a single stable, non-deprecated documented public API surface, confirmation that the current model is API-accessible, and independent corroboration that third parties successfully drive Codex via this API.

      • [github] You can also use Codex with an API key, but this requires additional setup.
      • [github] Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…
      • [claimed-docs] codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.
      • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.
      • [community] gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'
      Slatenone0/10

      Slate is documented as a CLI/terminal agent with configuration, skills, and hotkeys, but there is no evidence of a documented public API for programmatic/agentic access—the OpenAPI probe returned 404 across all candidate paths and no SDK or REST/API docs are mentioned anywhere in the pack.

      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
      • [probe] PROBE llms.txt: HTTP 200 at https://docs.randomlabs.ai/llms.txt # Slate ## Docs - [Introduction](https://docs.randomlabs.ai/en/getting-sta…
    • ai-native userIssue scoped/least-privilege API credentials for an agent

      weight 2 · round to Codex

      OpenAI's platform docs describe RBAC and project/org-scoped API keys/custom roles, and Codex can authenticate via an API key (codex-gh-4), so scoped credentials are technically available to a Codex-using account. However, none of the evidence ties this RBAC/API-key scoping specifically to configuring or restricting a Codex agent's own permissions — missing for 10: Codex-specific docs on issuing least-privilege keys for agent sessions, guidance on scoping credentials per-task/per-repo, and independent confirmation that this RBAC applies to Codex's own execution rather than just general API access.

      • [claimed-docs] Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…
      • [github] You can also use Codex with an API key, but this requires additional setup.
      • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
      Slatenone0/10

      Slate is a coding-agent CLI; its evidence only covers permission settings (allow/ask/deny) for tool actions, not issuance of scoped/least-privilege API credentials or tokens for agents. No mention of credential/token scoping, API key generation, or IAM-style access control.

      • [claimed-docs] Each permission key maps to an action ("allow", "ask", or "deny"), or a pattern object for fine-grained control.
      • [claimed-docs] We support `--dangerously-skip-permissions` (alias: `--yolo`) to bypass permission prompts.
    • ai-native userBuild against official SDKs

      weight 2 · round to Codex

      Codex is a coding agent, but the evidence shows a genuine SDK-adjacent surface: the underlying OpenAI Responses API has an official OpenAPI spec and multi-language code samples (Python, TypeScript, Go, Ruby, Java, HTTP, CLI), and Codex integrates via CLI/MCP for programmatic extension. However, there is no evidence of an official Codex-specific SDK (as opposed to the general OpenAI API SDK), and API access for the Codex model itself is explicitly noted as not yet available. missing for 10: a dedicated Codex SDK/library distinct from the general OpenAI Responses API, confirmation that Codex agent capabilities (not just chat completions) are exposed via SDK, independent developer corroboration of building against these SDKs.

      • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
      • [github] A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.
      • [community] gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'
      • [claimed-docs] Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.
      Slatenone0/10

      The evidence pack covers Slate's CLI, skills, configuration, and orchestration features but contains no mention of an official SDK (Python/TypeScript/etc.) for building applications on top of Slate, and the OpenAPI probe returned 404s across all candidate paths. Missing for 10: any documented SDK package, API reference, or programmatic interface for building against Slate.

      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
      • [claimed-docs] npm i -g @randomlabs/slate

    Agentic features

    1. ai-native userGet AI-generated insights and suggestions from my data inside the product

      weight 2 · round to Codex

      Codex generates AI-driven insights and suggestions specifically about code: it produces prioritized review findings, diffs, and summaries during automated reviews and delegated tasks (codex-docs-5, codex-docs-10, codex-docs-41, codex-docs-45), and can delegate to subagents for deeper investigation (codex-docs-35). However, this is scoped to code/repository data rather than general business or product data insights. Missing for 10: evidence of insight generation over non-code data sources, dashboards, or analytics-style summaries beyond code review findings.

      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
      • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
      • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
      • [claimed-docs] Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…

      Slate is a coding-agent CLI whose evidence shows it can analyze a codebase and produce suggestions (e.g., generating an ARCH.md with improvement ideas), which maps loosely to 'AI-generated insights from data' but only in the narrow sense of source code, not general data analysis. Community evidence also raises skepticism about the real quality of generated output (e.g., criticism of a ported-code example as low quality/unverified). Missing for 10: evidence of insights/suggestions over non-code datasets, dashboards or analytics-style outputs, and independent validation of suggestion quality.

      • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
      • [community] Blog post claimed porting a library with one sentence, but critic noted it was JS->TS (trivial rename) not Python->TS, excluded tests/exampl…
      • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…
    2. ai-native userSet up automations that run autonomously in the background

      weight 2 · round to Codex

      Codex cloud supports delegating longer tasks that run in isolated cloud environments in parallel, triggered from GitHub, GitLab, Linear, or Slack, and returning results (diff/PR) when ready — a clear background-automation workflow, and the CLI also supports non-interactive/repeatable workflows for scripted automation. Missing for 10: no documentation of scheduled/cron-style recurring triggers, and no independent/hands-on confirmation that long unattended background runs work reliably (community commentary focuses on interactive model quality/UX rather than background automation specifically).

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.

      Slate supports background subagents, parallel task orchestration, and built-in workflows like goal/deep-research that run while the user keeps interacting, which shows some autonomous background execution. However, this is task-level parallelism within an active session, not scheduled or trigger-based automations that run independently of user presence. missing for 10: evidence of scheduled/cron-like automations, persistent background jobs surviving session end, or trigger-based (event-driven) autonomous runs without an active user session.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.
    3. ai-native userDelegate tasks to a built-in AI assistant inside the product

      weight 3 · round to Codex

      Codex documents explicit task delegation to its built-in agent, both for long-running cloud tasks ('Delegate a longer task and return when it is ready') and for sub-agent delegation within a session ('Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session'), backed by detailed CLI/cloud docs. Missing for 10: independent hands-on verification specifically of the subagent delegation flow (community evidence discusses general agent quality/UX but not this feature directly).

      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.
      • [claimed-docs] Split up a larger investigation — subagents: Ask Codex to delegate focused work to specialized agents, then bring their findings back into t…
      • [claimed-docs] Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …
      • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.

      Slate is a CLI-based AI assistant where users delegate whole tasks (e.g., 'review architecture and write ARCH.md') and it spins up parallel subagents, orchestration programs like goal/deep-research, and long multi-hour sessions per first-party docs. Community evidence (comm-1/2/3) raises skepticism about output quality/novelty but does not contradict the core delegation mechanism itself. Missing for 10: independent hands-on validation that delegated multi-agent tasks reliably complete as advertised.

      • [claimed-docs] Parallelize working and orchestration of many tasks at once.
      • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.
      • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…
    4. ai-native userOperate the product with natural-language commands

      weight 2 · round to Codex

      Codex CLI, IDE extension, cloud, and web surfaces are all operated by natural-language prompts/chats — e.g. starting tasks from prompts, resuming chats, delegating subagents, pasting images into the composer, and non-interactive `codex exec` for scripted natural-language instructions — all documented as the primary interaction mode across surfaces. Community threads corroborate heavy real-world use of this conversational/agentic workflow, even amid quality complaints about model performance. missing for 10: independent benchmarking specifically of natural-language command comprehension/robustness (community evidence is about overall agent quality/speed, not NL parsing specifically).

      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
      • [claimed-docs] Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.
      • [claimed-docs] `codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
      • [claimed-docs] Bring visual context into the prompt — codex --image: Pass an error screenshot, architecture diagram, or design reference with the first pro…
      • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
      • [community] Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…

      Docs show Slate is driven primarily via natural-language prompts (e.g. the quickstart example 'Please review the architecture of my entire codebase...') alongside slash-commands, shell escapes, and file references, indicating natural-language is the core interaction mode for an AI-native agent CLI. Missing for 10: independent/hands-on confirmation that complex natural-language commands are reliably parsed and executed as intended (community evidence only discusses code-porting quality, not NL command usage itself).

      • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
      • [claimed-docs] Execute shell commands directly with `!`
      • [claimed-docs] Use `@filename` references
      • [claimed-docs] Use `/sessions` to switch between existing sessions
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.

    Api quality

    1. ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

      weight 2 · round to Codex

      OpenAI publishes a machine-readable OpenAPI 3.1 spec for its REST API (codex-gh-9) and Codex can be used via that API (codex-gh-4), but the evidence never confirms this spec explicitly covers or is dedicated to Codex-specific endpoints, nor is there a direct 'download spec' link tied to Codex docs. missing for 10: a Codex-specific OpenAPI/spec file, explicit download instructions, or confirmation the general OpenAI OpenAPI spec includes Codex CLI/agent endpoints.

      • [github] A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.
      • [github] You can also use Codex with an API key, but this requires additional setup.
      • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
      Slatenone0/10

      Slate's docs site was directly probed for an OpenAPI/swagger spec at standard locations and all returned 404, and no documentation anywhere mentions a machine-readable API spec for AI-native consumption.

      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
    2. ai-native userRely on versioned APIs with a documented deprecation policy

      weight 2 · round to Codex

      There is a documented OpenAPI 3.1 spec and API reference (codex-gh-9, codex-docs-29) and one concrete example of a deprecation notice (codex mcp-server deprecated in favor of the Codex app server, codex-docs-23), showing some practice of versioning and deprecation. However, there is no comprehensive, documented deprecation policy (timelines, notice periods, version numbering scheme) covering the Codex/OpenAI API generally. Missing for 10: an explicit deprecation policy document, API version numbering scheme, and independent corroboration of adherence to it.

      • [github] A machine-readable description of the OpenAI REST API, authored in OpenAPI 3.1.
      • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
      • [claimed-docs] codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …
      Slatenone0/10

      Slate is a CLI coding agent product; no evidence of any versioned public API, API reference, or deprecation policy documentation exists—openapi probes returned 404 and no docs mention API versioning or deprecation. Absence of evidence for this applicable axis (a product could plausibly document API stability) yields 'none'.

      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…

    Automation depth — how much of the product can run unattendedAutomation depth

    How much of the product can run unattended

    1. ai-native userPerform bulk operations across many items at once

      weight 2 · round drawn

      Codex supports running multiple cloud tasks in parallel across repos (codex-docs-1, codex-docs-3, codex-docs-6) and delegating focused work to specialized sub-agents within a session (codex-docs-13), which gives some bulk/parallel automation capability. However, there's no explicit evidence of a bulk operation primitive (e.g., batch-apply an action across many files/items/tickets in one command) — the parallelism described is task-level (multiple independent runs) rather than a documented 'operate over N items at once' feature. Missing for 10: explicit bulk/batch API or CLI verb for acting across many items in one invocation, and independent confirmation of large-scale parallel throughput in practice.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.

      Docs describe running many agents in parallel and orchestrating multiple tasks simultaneously (grid of subagent cards, spinning up more agents to parallelize work), which supports bulk-style operations across many items. However, there's no independent corroboration of this working at scale and no detail on failure handling, limits, or item-level bulk operations (e.g., bulk file edits/refactors) beyond task orchestration. Missing for 10: independent/hands-on verification of large-scale parallel task execution, concrete examples of bulk item processing (files/records), and failure/error handling details at scale.

      • [claimed-docs] Parallelize working and orchestration of many tasks at once.
      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.
    2. ai-native userDefine rules that trigger actions automatically on events

      weight 3 · round drawn
      Codexnone0/10

      Codex supports triggering tasks from external events (GitHub/GitLab/Linear/Slack) and running non-interactive workflows, but there is no evidence of a user-defined rules engine that lets users specify arbitrary trigger conditions and automated actions (e.g., 'on X event, do Y') — this is closer to integration hooks than a rules/automation framework. missing for 10: evidence of a rules/trigger definition interface, conditional logic configuration, or event-to-action mapping system that users can author themselves.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.
      Slatenone0/10

      Slate's docs describe agent rules for permissions/behavior ordering (docs-15, docs-20) and orchestration of parallel agents (docs-12, docs-13), but there is no evidence of user-defined rules that trigger actions automatically on external events (e.g., file changes, webhooks, schedule, git events). This is a plausible axis for a coding agent (many support hooks/triggers), so absence of evidence yields none rather than na.

      • [claimed-docs] Slate by default respects agent rules in the following order
      • [claimed-docs] Each permission key maps to an action ("allow", "ask", or "deny"), or a pattern object for fine-grained control.
    3. ai-native userSchedule recurring jobs or workflows

      weight 2 · round drawn
      Codexnone0/10

      The evidence shows Codex can run in CI/scripts (codex exec), be triggered from GitHub/GitLab/Slack, and run cloud tasks, but there is no mention of a native recurring/scheduled job or cron-like trigger mechanism within Codex itself. Automation is triggered by external events or manual invocation, not scheduled recurrence.

      • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.
      • [claimed-docs] Run a non-interactive command in a repeatable workflow.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …
      Slatenone0/10

      Slate is a coding-agent CLI with orchestration/parallel-agent features and sessions, but nothing in the evidence describes scheduling recurring jobs or workflows (e.g., cron-like triggers, timed recurring runs). Orchestration docs cover on-demand parallelization, not recurrence.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.
    4. ai-native userVersion, review, and roll back my automations

      weight 1 · round to Codex

      Codex's CLI includes a dedicated review command that inspects diffs/commits without modifying the working tree (codex-docs-10, codex-doces-41/45), and it operates within git repos so changes are inherently versioned and revertible via git; skills/plugins can be packaged as reusable automations (codex-docs-20/42). However, there is no documented mechanism to version, review, or roll back the automations/skills/workflows themselves (e.g., skill version history, rollback of a plugin config, audit trail for automation changes) — only code diffs are reviewed. Missing for 10: explicit versioning of skills/automations, a rollback UI/command for automation configs, and independent evidence of this workflow in practice.

      • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
      • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
      • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
      • [claimed-docs] Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.
      • [claimed-docs] Use skills and plugins: Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without l…
      Slatenone0/10

      Evidence shows session management (/sessions, /workspace) and built-in 'programs' like goal/deep-research, but nothing about versioning automations, reviewing history of changes, or rolling back to prior states of an automation/workflow. Missing for 10: any documentation of version history, diffing, or rollback mechanisms for automations/workflows.

      • [claimed-docs] Use `/sessions` to switch between existing sessions
      • [claimed-docs] Use `/workspace` to open the workspace manager, where you can review and remove workspace directories.
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.

    Autonomy agents — stories about autonomy agents in this arenaAutonomy agents

    Stories about autonomy agents in this arena

    Background execution

    1. ai-native userHave a cloud agent build, test, and demo a feature end-to-end for my review

      weight 2 · round to Codex

      Codex cloud lets users delegate tasks that run in isolated cloud environments, inspect summaries/diffs, request follow-ups, and open pull requests for review, effectively building/testing/demoing changes end-to-end for user review (codex-docs-1,5,6,7,37). Community commentary corroborates real-world agentic task completion, though with performance/reliability caveats. Missing for 10: independent hands-on verification specifically of the cloud (not CLI) workflow's demo/test artifacts, and no explicit mention of a 'demo' step (e.g., live preview) beyond diff/PR review.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Start and review work from the web or Codex CLI.
      • [claimed-docs] Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …
      • [community] Often Claude Code Opus 4.6, on hard enough problems, can do the impression of acting fast without really making progress. Then you spin the …
      • [community] Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…

      Slate's docs claim orchestration of parallel background subagents and being 'one of the few agents capable of performing integration tests manually,' suggesting it could build and test a feature autonomously, but no docs mention a 'demo' output or cloud-hosted execution environment. Community hands-on critique of an actual Slate-produced port directly contradicts the build/test claim: reviewers found the work excluded tests/examples and provided no verifiable repo, undermining confidence that Slate reliably builds+tests end-to-end for review. Missing for 10: evidence of cloud/remote execution infra, an explicit demo-generation feature, and independent confirmation that test suites are actually run and pass.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] Slate is one of the few agents capable of performing integration tests manually.
      • [community] Blog post claimed porting a library with one sentence, but critic noted it was JS->TS (trivial rename) not Python->TS, excluded tests/exampl…
      • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…
    2. developerDelegate longer-running coding tasks to run in the background in an isolated cloud environment

      weight 3 · round to Codex

      OpenAI's docs describe a dedicated Codex cloud mode that runs tasks in isolated cloud environments, in parallel, triggered from web/GitHub/GitLab/Linear/Slack, with configurable repo setup and a workflow to inspect diffs/PRs on completion, plus a CLI command (`codex cloud`) to submit and later pull results locally — squarely matching the story of delegating longer background tasks to an isolated cloud environment. missing for 10: independent or hands-on community corroboration specifically validating the cloud/background execution feature (community evidence in the pack discusses CLI/app UX and model quality, not the cloud delegation flow itself).

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …
      • [github] If you are looking for the cloud-based agent from OpenAI, Codex Web, go to chatgpt.com/codex.
      • [github] If you are looking for the <em>cloud-based agent</em> from OpenAI, <strong>Codex Web</strong>, go to <a href="https://chatgpt.com/codex">cha…
      Slatenone0/10

      Slate's docs describe subagents running 'in the background' locally while you keep chatting and orchestration/parallelization of tasks, but there is no mention of an isolated cloud environment, remote execution sandbox, or delegation to a hosted service — everything described appears to run within the local CLI session. This axis is applicable to coding agent tools generally, but no evidence supports a cloud-isolated background execution capability for Slate.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] Slate works with you across long, multi-hour sessions.
      • [claimed-docs] Parallelize working and orchestration of many tasks at once.
    3. developerConfigure a reproducible cloud environment with the dependencies and setup steps my repository needs

      weight 2 · round to Codex

      Codex Cloud docs state you can configure the dependencies, tools, variables, and setup steps each repository needs for isolated cloud environments, directly matching the story. However, there is no detail on how reproducibility is guaranteed (e.g., container images, caching, version pinning) or independent hands-on confirmation of this setup workflow. Missing for 10: concrete configuration file/schema details, reproducibility guarantees, and independent verification of the setup working as documented.

      • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      Slatenone0/10

      Slate's docs describe a local CLI agent (npm install, terminal sessions, permissions, skills, orchestration) but contain no mention of provisioning or configuring a reproducible cloud environment, dependency setup, or devcontainer-style configuration for a repository. This is a fair capability to ask of an autonomous coding agent, but no evidence shows Slate supports it.

      • [claimed-docs] npm i -g @randomlabs/slate
      • [claimed-docs] Slate by default respects agent rules in the following order
      • [claimed-docs] Each permission key maps to an action ("allow", "ask", or "deny"), or a pattern object for fine-grained control.

    Parallel agents

    1. ai-native userLaunch fleets of autonomous agents that work in parallel on different tasks for hours or days

      weight 2 · round to Slate

      Codex Cloud supports running multiple tasks in parallel in isolated cloud environments, triggered from GitHub/GitLab/Linear/Slack, and delegating longer tasks to return to later, which covers parallel/async agent work. However, there is no explicit evidence of orchestrating large 'fleets' of many simultaneous agents, no stated duration limits confirming multi-day autonomous runs, and community feedback highlights usage-limit throttling that would constrain sustained parallel/long-running fleets. missing for 10: evidence of fleet-scale orchestration (many concurrent agents), confirmed multi-day autonomous run duration, and independent confirmation that parallel tasks aren't throttled by usage limits.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
      • [community] Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…
      • [community] The main issue I have with Codex is that the best model is insanely slow, except at nights and weekends when Silicon Valley goes to bed... I…

      Docs describe genuine parallel agent orchestration (grid of subagent cards, spinning up more agents while others run in background) and explicitly support multi-hour sessions, matching much of the story. However, evidence only confirms 'multi-hour' not multi-day autonomy, and community commentary raises skepticism about output quality/novelty without directly refuting the parallel-agent mechanics themselves. Missing for 10: confirmation of multi-day unattended runs, independent hands-on validation of fleet-scale parallel task completion.

      • [claimed-docs] Slate works with you across long, multi-hour sessions.
      • [claimed-docs] Parallelize working and orchestration of many tasks at once.
      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [community] Commenter compared the approach to 'Ralph as a service' referencing an existing agentic coding technique (ghuntley.com/ralph), suggesting Sl…
    2. developerRun several task attempts in parallel and compare results before choosing one

      weight 1 · round drawn

      Docs confirm Codex cloud can run tasks in parallel in isolated cloud environments without tying up the local machine, and results can be inspected (summary/diff) before choosing to follow up or open a PR — this covers running multiple attempts and reviewing outcomes. However, there's no explicit documentation of a dedicated 'compare multiple attempts side-by-side' UI/workflow, and no independent/community evidence confirming this parallel-comparison workflow works well in practice. missing for 10: explicit side-by-side comparison UI documentation, independent hands-on confirmation of comparing parallel attempts.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Delegate a longer task and return when it is ready.

      Slate's orchestration docs show multiple subagents running in parallel as a grid of cards while the user keeps working, directly supporting parallel task execution (docs-12, docs-13). However, there's no explicit documentation of a compare/diff view or a 'choose winning attempt' workflow for reconciling multiple parallel results into one choice. Missing for 10: explicit comparison/selection UI or workflow for multiple attempts of the same task, and independent/hands-on confirmation of this specific use case.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.

    Scheduled automation

    1. ai-native userSet up always-on agents that run on schedules or triggers to maintain and fix my software autonomously

      weight 2 · round to Codex

      Codex cloud supports starting tasks from external triggers (GitHub/GitLab issues & PRs, Linear issues, Slack messages) and running them in parallel isolated environments, which covers the 'triggers' half of the story, but there's no evidence of a true schedule/cron-based always-on agent that proactively maintains a repo without an external event. Missing for 10: explicit scheduled/cron execution, evidence of continuous unattended monitoring/maintenance loops, and independent confirmation these triggers reliably run autonomous fixes end-to-end.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Run tasks in parallel without tying up your local machine.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      Slatenone0/10

      Slate's docs describe parallel subagent orchestration within a live session (background agents you keep talking to, spin up more agents to parallelize tasks) but there is no evidence of scheduling, event/webhook triggers, or persistent always-on agents that run autonomously outside an active session to maintain/fix software over time.

      • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
      • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
      • [claimed-docs] `goal` and `deep-research` are built-in programs. They are user-visible workflows, not something you need to author before using Slate.

    Code generation — quality of generated code — correctness, style, fit to the codebaseCode generation

    Quality of generated code — correctness, style, fit to the codebase

    Debugging

    1. developerDebug issues and troubleshoot using natural-language queries

      weight 2 · round to Codex

      Codex CLI docs show clear natural-language debugging workflows: exploring unfamiliar code, running local tools, passing error screenshots for context, and running dedicated code review that reports prioritized findings (codex-docs-8, codex-docs-9, codex-docs-10, codex-docs-12). However, community evidence shows mixed real-world reliability on agentic/coding tasks and no independent confirmation specifically validating debugging accuracy. Missing for 10: hands-on validation of debugging/troubleshooting accuracy, and independent case studies showing successful root-cause diagnosis via NL queries.

      • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
      • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
      • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
      • [claimed-docs] Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.
      • [community] Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.
      • [community] Often Claude Code Opus 4.6, on hard enough problems, can do the impression of acting fast without really making progress. Then you spin the …

      Slate's docs show it operates via natural-language prompts, executes shell commands (`!`), references files (`@filename`), and can run integration tests and review codebase architecture in NL form, which implies it could be used for debugging and troubleshooting queries. However there is no explicit example, workflow, or documentation section dedicated to debugging/troubleshooting via natural language, and community evidence is skeptical/unrelated to this specific capability. Missing for 10: explicit debugging-focused examples or docs, independent verification that NL-based debugging works well, dedicated troubleshooting workflow beyond generic agent capabilities.

      • [claimed-docs] Execute shell commands directly with `!`
      • [claimed-docs] Use `@filename` references
      • [claimed-docs] Slate is one of the few agents capable of performing integration tests manually.
      • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.

    Feature implementation

    1. developerTurn a tracked issue into a complete pull request end-to-end

      weight 3 · round to Codex

      Codex explicitly supports starting work from a tracked issue (GitHub, GitLab, Linear) in cloud environments, running the task, inspecting the diff/summary, and opening a pull request when done, covering the full issue-to-PR loop. missing for 10: independent hands-on confirmation of a full issue-to-merged-PR workflow succeeding end-to-end, and detail on how issue context/acceptance criteria are actually parsed.

      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
      • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
      • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
      • [claimed-docs] Delegate a longer task and return when it is ready.
      Slatenone0/10

      The evidence pack covers Slate's session management, orchestration, skills, and permissions but contains no mention of issue-tracker integration (e.g., GitHub issues) or automated pull-request creation/submission. Without evidence of ingesting a tracked issue and producing a PR end-to-end, this capability is unshown.

      • developerDescribe a feature or bug in plain language and have the agent implement or fix it across multiple files

        weight 3 · round to Codex

        Codex CLI and cloud docs describe the core loop of natural-language task description leading to autonomous file inspection, editing, running local tools, and producing a diff/PR (codex-docs-30, codex-docs-9, codex-docs-6, codex-docs-5), and community commentary corroborates it does real multi-file edits ('it edits the files and I can use the proper tooling', 'shockingly good... no worse than average L3-L4 engs') alongside some negative UX complaints that don't dispute the core capability. Missing for 10: independent benchmark/case-study evidence specifically confirming complex multi-file refactors across large codebases, and some community reports of it 'struggling with almost anything' create mild quality tension without rising to a concrete dispute.

        • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
        • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
        • [claimed-docs] Delegate a longer task and return when it is ready.
        • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
        • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
        • [community] Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…
        • [community] Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…
        • [community] Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.

        Docs imply broad multi-file code work (e.g. the quickstart example asks Slate to review an entire codebase and produce ARCH.md, plus orchestration features for parallelizing tasks across files/agents), suggesting Slate can act on plain-language requests across a codebase. However, independent community scrutiny of a specific real-world claim (a 'ported library' from one sentence) found it was actually a trivial JS->TS rename, excluded tests, lacked a verifiable repo, and drew explicit skepticism about the quality/usefulness of the generated code — concretely contradicting the marketed multi-file code-generation capability. Missing for 10: first-party documentation of a genuine multi-file bug-fix/feature-implementation workflow with verifiable before/after results, and independent hands-on confirmation that resolves the community dispute.

        • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
        • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
        • [claimed-docs] Parallelize working and orchestration of many tasks at once.
        • [community] Blog post claimed porting a library with one sentence, but critic noted it was JS->TS (trivial rename) not Python->TS, excluded tests/exampl…
        • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…

      Maintenance automation

      1. developerHave the agent write tests, fix lint errors, resolve merge conflicts, and update dependencies for me

        weight 3 · round to Codex

        Codex CLI/cloud docs describe a general-purpose coding agent that can inspect code, edit files, run local dev tools, automate repeatable work, and review diffs before PRs — capabilities broad enough to plausibly cover writing tests, fixing lint issues, resolving conflicts, and updating dependencies (codex-docs-8, codex-docs-9, codex-docs-30, codex-docs-41). However, none of the docs explicitly name test-writing, lint-fixing, merge-conflict resolution, or dependency updates as supported workflows, and community feedback is mixed on real-world reliability for complex agentic tasks (codex-comm-3, codex-comm-13). missing for 10: explicit documentation/examples of test generation, lint-fix automation, merge-conflict resolution, and dependency-update workflows, plus hands-on confirmation these specific tasks succeed.

        • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
        • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
        • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
        • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
        • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
        • [community] Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…
        • [community] Having used codex a fair bit I find it really struggles with … almost anything. However using the equivalent chat gpt model is fantastic.

        Slate is documented as a general-purpose coding agent with shell execution, file editing, permissioning, and orchestration of multiple sub-agents (random-labs-docs-9, random-labs-docs-13, random-labs-docs-16), which implies it could perform tasks like running tests or lint/dependency commands, but the evidence never explicitly documents test-writing, lint-fixing, merge-conflict resolution, or dependency updates as capabilities. Community commentary raises quality concerns about generated code but doesn't specifically address these tasks. Missing for 10: explicit documentation or examples of writing/fixing tests, resolving lint errors, resolving merge conflicts, and updating dependencies.

        • [claimed-docs] Execute shell commands directly with `!`
        • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
        • [claimed-docs] Slate is one of the few agents capable of performing integration tests manually.
        • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…

      Multimodal generation

      1. ai-native userGenerate a working app from a sketch, image, or PDF design

        weight 2 · round to Codex

        Codex supports passing images (error screenshots, architecture diagrams, design references) into prompts, which is a partial building block for generating apps from a sketch/image, but there's no evidence of dedicated PDF-to-app workflows, multi-page design ingestion, or documented end-to-end 'sketch/image to working app' generation feature. missing for 10: explicit PDF design ingestion, dedicated image/design-to-app pipeline or template, independent hands-on demonstration of generating a full app from a design artifact.

        • [claimed-docs] Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.
        Slatenone0/10

        The evidence describes Slate as a terminal-based CLI agent for coding sessions, orchestration, skills, and permissions, but nothing in the docs or community evidence mentions accepting sketches, images, or PDF designs as input to generate an app. missing for 10: any mention of image/sketch/PDF input, multimodal design-to-code capability, or UI mockup ingestion.

        • [claimed-docs] npm i -g @randomlabs/slate
        • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
        • [claimed-docs] Skills are markdown instruction packages that give the agent domain-specific knowledge and behavior.
        • [claimed-docs] description: "Create distinctive, production-grade frontend interfaces."

      Codebase understanding — how deeply the tool maps your repo — cross-file context, architecture awareness, historyCodebase understanding

      How deeply the tool maps your repo — cross-file context, architecture awareness, history

      Codebase mapping

      1. developerUnderstand how a codebase fits together to find where to start making changes

        weight 3 · round to Codex

        Codex CLI docs explicitly mention exploring unfamiliar code and planning changes within a repository, and it can inspect code, run local dev tools, and review diffs/commits — supporting codebase orientation. However, there's no dedicated codebase-mapping/visualization feature, no evidence of dependency-graph or architecture-summary generation, and community feedback focuses on agentic task execution rather than comprehension aids. Missing for 10: dedicated codebase-map/architecture-overview feature, independent hands-on evidence of effectively onboarding to unfamiliar large codebases, and richer navigation/search tooling beyond terminal chat resume.

        • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
        • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
        • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
        • [claimed-docs] Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.

        The quickstart example explicitly shows Slate producing an ARCH.md architecture review of an entire codebase with improvement suggestions, directly supporting codebase-understanding use cases, and @filename references plus workspace management help navigate a repo. However there's no dedicated codebase-mapping/search feature (e.g., symbol index, dependency graph) documented beyond this one example, and no independent evidence confirming quality of such architecture summaries. missing for 10: dedicated code-navigation/search tooling, independent validation of architecture-summary accuracy, more than a single example of codebase-understanding workflow.

        • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
        • [claimed-docs] Use `@filename` references
        • [claimed-docs] Use `/workspace` to open the workspace manager, where you can review and remove workspace directories.
      2. developerHave the agent map and explain an entire unfamiliar codebase without manually selecting context files

        weight 3 · round to Slate

        Codex CLI docs explicitly state it can be started in a repository 'to explore unfamiliar code, plan a change, edit files, and run your local development tools' (codex-docs-9), implying the agent autonomously navigates the codebase rather than requiring manual file selection, and codex-gh-1 confirms it runs as an autonomous coding agent locally. However, there's no detailed documentation of how it builds a whole-codebase map/summary, no explicit 'explain codebase' feature, and no independent hands-on evidence confirming this works well on large unfamiliar repos. Missing for 10: dedicated codebase-mapping/summarization feature documentation, evidence of handling very large repos, and independent user reports validating this specific capability.

        • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
        • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.

        Docs show Slate's quickstart example explicitly demonstrates asking it to 'review the architecture of my entire codebase' and generate an ARCH.md without manual file selection, and it has orchestration/subagent features for broad exploration. However there's no independent/hands-on verification that this codebase-mapping actually works well on large unfamiliar repos, and community evidence raises quality skepticism about other generated outputs. missing for 10: independent hands-on validation of full-codebase mapping accuracy, evidence of handling very large/unfamiliar codebases without manual curation, detail on how context is auto-selected under the hood.

        • [claimed-docs] Please review the architecture of my entire codebase creating an ARCH.md and then give me ways I can improve it.
        • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
        • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
        • [community] "Why trumpet code that is so ready for the garbage that you wouldn't even bother to publish it" - skepticism about the quality/usefulness of…

      Context management

      1. developerHave the agent build and recall memory automatically across sessions

        weight 2 · round to Codex

        Codex CLI supports `codex resume` to reopen or search past local chats in a repository, giving a limited form of session recall, but this requires manual user action rather than automatic memory building/recall across sessions. Missing for 10: evidence of automatic persistent memory (learned facts, preferences, or context) that Codex builds unprompted and recalls without explicit resume/search commands, and any cross-session synthesis beyond raw chat transcripts.

        • [claimed-docs] Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
        • [claimed-docs] `codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
        Slatenone0/10

        Docs describe session switching (/sessions), long multi-hour session support, and diagnostic context attachment, but there is no evidence of automatic cross-session memory building or recall — sessions appear to be manually selected/switched contexts, not an automatic memory system. Missing for higher verdict: any documentation of persistent memory storage, automatic recall of past codebase context, or memory summarization across sessions.

        • [claimed-docs] Slate works with you across long, multi-hour sessions.
        • [claimed-docs] Use `/sessions` to switch between existing sessions
        • [claimed-docs] Slate automatically attaches relevant diagnostic information (OS, version, session context) to your report.
      2. developerInclude multiple project directories in a single session for broader context

        weight 2 · round to Slate
        Codexnone0/10

        No evidence in the pack describes Codex supporting multiple project directories or repositories being combined in a single session/context; documentation focuses on single-repository sessions, cloud tasks, and per-repository setup steps.

          Docs mention a `/workspace` manager for reviewing and removing 'workspace directories' (plural), implying support for multiple project directories in one session, but there's no detailed documentation on how directories are added or how context is merged across them, and no independent/hands-on confirmation. Missing for 10: explicit instructions/examples for adding multiple directories, and independent verification that broader multi-directory context actually works in practice.

          • [claimed-docs] Use `/workspace` to open the workspace manager, where you can review and remove workspace directories.
        • developerAdd a project instructions file to set coding standards and conventions the agent follows

          weight 3 · round to Slate
          Codexnone0/10

          The evidence pack covers Codex's CLI, cloud, MCP, and review features but contains no mention of a project-level instructions/config file (e.g., AGENTS.md or similar) for setting coding standards or conventions the agent should follow. This is a plausible and common capability for coding agents, but nothing in the pack documents or demonstrates it.

            Docs confirm Slate 'respects agent rules' in a defined precedence order and supports Skills (markdown instruction packages, including Claude Code-compatible `.claude/skills/` paths), which cover project-level conventions/instructions, but there's no explicit example of a single top-level 'instructions file' analogous to AGENTS.md/CLAUDE.md being demonstrated end-to-end. missing for 10: explicit naming/format of the project instructions file, a worked example showing the agent following custom conventions from it, and independent/community confirmation it works as documented.

            • [claimed-docs] Slate by default respects agent rules in the following order
            • [claimed-docs] Skills are markdown instruction packages that give the agent domain-specific knowledge and behavior.
            • [claimed-docs] `.claude/skills/` | Claude Code compatibility

          Issue diagnosis

          1. developerReproduce issues, narrow down root causes, and verify fixes

            weight 3 · round to Codex

            Codex CLI docs explicitly describe exploring unfamiliar code and running local dev tools to investigate issues, passing error screenshots for context, delegating focused investigation to subagents, and running dedicated reviews against uncommitted changes/commits/base branches to verify fixes before committing — covering reproduce, narrow-down, and verify steps. Missing for 10: no explicit 'reproduce a bug' walkthrough or first-hand/independent account of successfully diagnosing and fixing a real bug end-to-end with Codex.

            • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
            • [claimed-docs] Pass an error screenshot, architecture diagram, or design reference with the first prompt, or paste an image into the interactive composer.
            • [claimed-docs] Ask Codex to delegate focused work to specialized agents, then bring their findings back into the main terminal session.
            • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
            • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
            • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
            • [claimed-docs] Run a non-interactive command in a repeatable workflow.

            Slate documents shell execution (`!`), file references, and being 'one of the few agents capable of performing integration tests manually,' which are plausible building blocks for debugging workflows, but there's no explicit documentation of a reproduce→diagnose→verify-fix workflow. Missing for 10: explicit debugging/root-cause-analysis workflow documentation, evidence of test-driven verification loops, and independent hands-on confirmation that Slate helps developers actually reproduce and fix bugs.

            • [claimed-docs] Slate is one of the few agents capable of performing integration tests manually.
            • [claimed-docs] Execute shell commands directly with `!`
            • [claimed-docs] Use `@filename` references
            • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…

          Ecosystem — integrations, plugins, and third-party ecosystem storiesEcosystem

          Integrations, plugins, and third-party ecosystem stories

          Marketplace

          1. developerEquip the agent with custom skills to perform specialized tasks

            weight 1 · round drawn

            Codex CLI docs explicitly describe packaging repeatable instructions as "skills" and adding plugins to connect Codex to team tools/data from the CLI, directly matching the custom-skills story. Missing for 10: independent hands-on validation of skill creation/usage, and deeper documentation on skill authoring format/lifecycle beyond a single mention.

            • [claimed-docs] Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.

            Slate has a documented Skills system: markdown instruction packages that give the agent domain-specific knowledge/behavior, with example skill definitions and compatibility with Claude Code's `.claude/skills/` format, letting developers equip the agent with custom specialized capabilities. Missing for 10: independent/hands-on verification that custom skills work as documented, and more detail on skill authoring/discovery beyond the single example.

            • [claimed-docs] Skills are markdown instruction packages that give the agent domain-specific knowledge and behavior.
            • [claimed-docs] description: "Create distinctive, production-grade frontend interfaces."
            • [claimed-docs] `.claude/skills/` | Claude Code compatibility
          2. engineering-leadIntegrate third-party partner-built agent apps into my workflows

            weight 1 · round to Codex

            Codex documents integration points for third-party ecosystem tools — triggering work from GitHub, GitLab, Linear, and Slack (partner platforms), and connecting to third-party MCP servers, plugins, and skills that give access to tools like Figma or a browser — which supports embedding partner-built capabilities into engineering workflows. However, the evidence is framed around Codex consuming tools/data sources rather than a curated marketplace of partner-built 'agent apps,' and there's no independent case study of a partner agent integration working end-to-end. Missing for 10: evidence of a partner/agent-app marketplace or certified third-party agent integrations, and independent verification of such integrations working in practice.

            • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
            • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
            • [claimed-docs] Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.
            • [claimed-docs] Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.
            • [claimed-docs] Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…
            • [claimed-docs] Use skills and plugins: Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without l…
            • [claimed-docs] Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…
            Slatenone0/10

            Slate is a coding CLI agent focused on subagents, skills, sessions, and model orchestration; there is no evidence of an ecosystem for integrating third-party partner-built agent apps (e.g., a marketplace, app store, or partner integration framework). Skills compatibility with Claude Code is mentioned but that is file-format compatibility, not partner app integration.

            • [claimed-docs] Skills are markdown instruction packages that give the agent domain-specific knowledge and behavior.
            • [claimed-docs] `.claude/skills/` | Claude Code compatibility

          Team knowledge

          1. engineering-leadCreate a shared workspace from my docs and repos as a common source of truth for the team

            weight 1 · round drawn
            Codexnone0/10

            Codex documents repo-level cloud environments, RBAC, and MCP connections to team tools, but no evidence describes a shared 'workspace' feature that unifies docs and repos into a common source of truth for a team; this is a plausible ask for an engineering tool but Codex's evidence only covers per-task cloud environments and repo configuration, not a persistent shared knowledge/workspace layer.

            • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
            • [claimed-docs] Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…
            • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
            Slatenone0/10

            Slate is a CLI coding agent focused on individual sessions, workspaces (local directories), skills, and orchestration of subagents—there is no evidence of a shared team workspace or collaborative source-of-truth feature built from docs and repos. The 'workspace' concept here refers to local directory management (/workspace), not a shared team hub.

            • [claimed-docs] Use `/workspace` to open the workspace manager, where you can review and remove workspace directories.
            • [claimed-docs] Slate by default respects agent rules in the following order
            • [claimed-docs] Skills are markdown instruction packages that give the agent domain-specific knowledge and behavior.

          Tool integration

          1. developerConnect the agent to workflow tools like Jira, Slack, and Google Drive to extend its context

            weight 3 · round to Codex

            Codex explicitly supports starting work from Slack (and GitHub/GitLab/Linear) and lets users add local or remote MCP servers to connect to third-party tools/docs (e.g. Figma, browser), giving a generic mechanism to extend context to workflow tools. However, there is no explicit documentation of native Jira or Google Drive connectors—only Slack is named among the story's specific tools, with Jira/Google Drive requiring the generic (and for one variant, deprecated/experimental) MCP server pathway. Missing for 10: named Jira integration, named Google Drive integration, and confirmation that the current (non-deprecated) MCP mechanism is broadly used for these specific SaaS tools.

            • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
            • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
            • [claimed-docs] Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.
            • [claimed-docs] Package repeatable instructions as skills, then add plugins to connect Codex to your team's tools and data without leaving the CLI.
            • [claimed-docs] Use it to give ChatGPT or Codex access to third-party documentation, or to let it interact with developer tools like your browser or Figma.
            • [claimed-docs] Model Context Protocol (MCP) connects models to tools and context. Use it to give ChatGPT or Codex access to third-party documentation, or t…
            • [claimed-docs] Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…
            • [github] Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…
            • [claimed-docs] codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …
            Slatenone0/10

            No evidence anywhere in the docs pack mentions integrations with Jira, Slack, Google Drive, or any workflow tools/MCP connectors; the docs focus on CLI usage, sessions, skills, and permissions with no mention of external tool connectivity.

            • developerKick off agent tasks directly from GitHub, GitLab, Linear, or Slack

              weight 2 · round to Codex

              First-party docs explicitly state Codex cloud tasks can be started from GitHub pull requests, GitLab merge requests/issues, Linear issues, or Slack channels/threads, matching the story directly. Missing for 10: independent/hands-on verification of these specific integrations working in practice (community evidence covers CLI/app UX but not the GitHub/GitLab/Linear/Slack kickoff flows specifically).

              • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
              • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
              • [claimed-docs] Delegate a longer task and return when it is ready.
              Slatenone0/10

              No evidence of any integration with GitHub, GitLab, Linear, or Slack for triggering agent tasks; Slate's documentation covers CLI usage, sessions, skills, and configuration but nothing about ecosystem/platform triggers.

              Ide terminal integration — meeting you in the IDE and terminal — extensions, inline flows, contextIde terminal integration

              Meeting you in the IDE and terminal — extensions, inline flows, context

              Cross device continuity

              1. developerStart a task on one device and continue it later from another device or browser

                weight 2 · round to Codex

                Codex supports starting tasks in the cloud from web/GitHub/GitLab/Linear/Slack, working in parallel cloud environments, and later resuming or continuing work from the CLI via 'codex cloud' (browse active/completed chats, submit/apply results) or 'codex resume' to reopen local chats, plus a shared MCP config across ChatGPT desktop, CLI, and IDE extension enabling cross-client continuity. missing for 10: independent hands-on confirmation of seamless state sync across devices/browsers, and no explicit mention of resuming a cloud-started task from a different physical device's browser session.

                • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
                • [claimed-docs] Delegate a longer task and return when it is ready.
                • [claimed-docs] Start and review work from the web or Codex CLI.
                • [claimed-docs] Move work to Codex cloud — codex cloud: Browse active and completed chats, submit work to a configured environment, and apply the result to …
                • [claimed-docs] `codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
                • [claimed-docs] The ChatGPT desktop app, Codex CLI, and IDE extension share this configuration. Once you configure your MCP servers, you can switch among th…
                Slatenone0/10

                Docs show session management within Slate (e.g. `/sessions` to switch sessions, `Ctrl+X N` for new session) but only describe local session switching, not any cloud sync or cross-device/browser continuation mechanism. Slate appears to be a terminal-only CLI tool with no mention of a browser interface or account-based sync for resuming tasks elsewhere.

              Ide integration

              1. developerChat with the coding assistant directly inside my IDE for contextual help

                weight 3 · round to Codex

                Codex explicitly offers an IDE extension for VS Code, Cursor, and Windsurf, plus a CLI usable within the terminal in your repo, both providing contextual chat/help with the codebase (edit files, run commands, review diffs). Community evidence confirms real-world usage of Codex CLI/app for editing and testing files in context, though some note UX friction compared to competitors. Missing for 10: deeper first-party documentation/screenshots of the IDE extension's chat UI specifically, and stronger independent hands-on corroboration of in-IDE chat quality.

                • [github] If you want Codex in your code editor (VS Code, Cursor, Windsurf), install in your IDE.
                • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
                • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
                • [community] Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…

                Slate is documented as a terminal-based coding agent with session management, `@filename` references, shell execution, and workspace context — providing contextual chat help that developers can run alongside their editor in a terminal. However, there is no evidence of a native IDE extension/panel (e.g., VS Code/JetBrains plugin) that embeds Slate directly inside the IDE UI itself. missing for 10: dedicated IDE extension/panel integration, evidence of in-editor chat UI beyond terminal, independent corroboration of IDE workflow usage.

              Session management

              1. developerReview diffs visually and run multiple sessions side by side in a desktop app

                weight 2 · round to Codex

                Codex ships a desktop app ("codex app"/Codex App page) and documents parallel task execution plus diff/summary inspection before merging, suggesting the underlying pieces exist, but the evidence never shows the desktop app UI actually presenting a visual diff viewer or multiple sessions arranged side by side. Community notes even flag basic desktop-app reliability issues (stuck on 'Loading projects...', Mac-only availability). Missing for 10: concrete documentation/screenshots of the desktop app's diff viewer, explicit multi-session/side-by-side UI description, and independent confirmation it works smoothly.

                • [github] If you want the desktop app experience, run <code>codex app</code> or visit the Codex App page.
                • [github] If you want the desktop app experience, run <code>codex app</code> or visit <a href="https://chatgpt.com/codex?app-landing-page=true">the Co…
                • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
                • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                • [claimed-docs] Run tasks in parallel without tying up your local machine.
                • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
                • [community] Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…
                • [community] Mac only. Again. Apple is great but this is OpenAI devs showing their disconnect from the mainstream.
                Slatenone0/10

                Slate is documented as a terminal/CLI tool (npm install, terminal-background onboarding, hotkeys, `/sessions` switching, subagent grid) with no mention of a desktop GUI or visual diff review; session switching is terminal-based, not side-by-side desktop windows. Missing for 10: any evidence of a desktop application, a visual diff viewer, or GUI-based side-by-side session comparison.

                • [claimed-docs] npm i -g @randomlabs/slate
                • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                • [claimed-docs] Use `/sessions` to switch between existing sessions
                • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
                • [claimed-docs] Ctrl+X then N New session
              2. engineering-leadManage multiple agent-driven coding sessions from one unified workspace

                weight 2 · round to Slate

                Codex documents cloud parallel task execution across multiple repos/environments (codex-docs-1,3,6), a web/CLI dashboard to browse active and completed chats and apply results locally (codex-docs-15), and resuming/searching across sessions (codex-docs-11,24), which together support managing multiple concurrent agent sessions from a unified interface. However, evidence is vendor-documentation only with no independent hands-on confirmation of a true 'unified workspace' UX for an engineering-lead managing many sessions simultaneously, and some community comments note UX rough edges (codex-comm-9,18). Missing for 10: independent/hands-on verification of multi-session management at scale, and clearer detail on cross-session visibility/coordination for a lead overseeing a team's agents.

                • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                • [claimed-docs] Run tasks in parallel without tying up your local machine.
                • [claimed-docs] Delegate a longer task and return when it is ready.
                • [claimed-docs] Browse active and completed chats, submit work to a configured environment, and apply the result to your local repository from the terminal.
                • [claimed-docs] Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
                • [claimed-docs] `codex resume`: Reopen a recent chat from the current repository, or search across local chats when you need to return to older work.
                • [community] I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …

                Docs describe first-class multi-session/multi-agent workspace features: `/sessions` to switch sessions, `/workspace` manager, new-session hotkey, and orchestration showing a grid of inline subagent cards while continuing to chat, queue tasks, or spin up more parallel agents — directly matching the engineering-lead's need to manage multiple concurrent agent sessions from one place. Missing for 10: independent/hands-on verification of this workspace at scale and any lead-specific team-management features beyond individual session switching.

                • [claimed-docs] Use `/sessions` to switch between existing sessions
                • [claimed-docs] Use `/workspace` to open the workspace manager, where you can review and remove workspace directories.
                • [claimed-docs] Those agents show up as a grid of inline subagent cards, one per agent.
                • [claimed-docs] While one or more agents run in the background, you can keep talking with Slate: plan next steps, queue up additional tasks, or spin up more…
                • [claimed-docs] Ctrl+X then N New session

              Terminal workflow

              1. developerRun a coding agent locally from my terminal

                weight 3 · round to Codex

                Codex CLI is explicitly documented as a coding agent that runs locally in the terminal, with npm/standalone install, working against the local repository, editing files, running commands, and offering interactive TUI plus non-interactive exec mode — well corroborated by first-party docs and GitHub README, with community usage discussion confirming real-world use. Missing for 10: independent hands-on verification specifically of pure local terminal usage (most community commentary discusses model quality/UX rather than the local-run mechanics) and some caveats about performance/limits reported by users.

                • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
                • [github] npm install -g @openai/codex
                • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
                • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
                • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
                • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.
                • [claimed-docs] Install the Codex CLI with the standalone installer for macOS and Linux.
                • [community] Codex is my favorite UX for anything as it edits the files and I can use the proper tooling to adjust and test stuff... However lately the l…
                • [community] Genuinely excited to try this out. I've started using Codex much more heavily in the past two months and honestly, it's been shockingly good…
                Slatefullprobed8/10

                Slate ships as a global npm CLI (`npm i -g @randomlabs/slate`) that runs interactively in the terminal, with documented terminal-native features like hotkeys, shell command execution (`!`), file references (`@filename`), session management (`/sessions`), and configuration via `slate.json` — all consistent with a locally-run terminal coding agent. Missing for 10: independent hands-on confirmation of local terminal usage (community evidence only discusses porting-quality skepticism, not terminal operation itself) and no evidence of offline/non-terminal fallback limitations.

                • [claimed-docs] npm i -g @randomlabs/slate
                • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                • [claimed-docs] Use `/sessions` to switch between existing sessions
                • [claimed-docs] Press Tab to queue the current message so it runs after the current turn finishes.
                • [claimed-docs] Execute shell commands directly with `!`
                • [claimed-docs] Use `@filename` references
                • [claimed-docs] Ctrl+X then N New session
                • [probe] official CLI documented at https://docs.randomlabs.ai/en/getting-started/quickstart
              2. developerRun the agent non-interactively in scripts for workflow automation

                weight 2 · round to Codex

                Docs explicitly describe running 'a non-interactive command in a repeatable workflow' and automating repeatable work without leaving the terminal, plus support for submitting work to configured environments from scripts (codex exec-style usage implied). Missing for 10: independent hands-on confirmation of non-interactive/CI usage and detailed exit-code/output-format documentation for scripting.

                • [claimed-docs] Run a non-interactive command in a repeatable workflow.
                • [claimed-docs] Inspect code, make changes, run commands, and automate repeatable work without leaving your terminal.
                • [claimed-docs] Browse active and completed chats, submit work to a configured environment, and apply the result to your local repository from the terminal.
                • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
                Slatenone0/10

                The evidence shows Slate is a CLI-based interactive agent (npm install, onboarding, in-session commands like /sessions, !, @filename) but nowhere documents a non-interactive/headless mode, flags for scripted execution, or CI/automation usage; --dangerously-skip-permissions bypasses prompts but is not shown as enabling scripted/non-interactive invocation. Missing for 10: documentation of a non-interactive/print/exec mode, exit-code or piping behavior, or any CI/scripting examples.

                • [claimed-docs] npm i -g @randomlabs/slate
                • [claimed-docs] We support `--dangerously-skip-permissions` (alias: `--yolo`) to bypass permission prompts.
                • [probe] official CLI documented at https://docs.randomlabs.ai/en/getting-started/quickstart

              Openness — open source, data portability, and self-hosting storiesOpenness

              Open source, data portability, and self-hosting stories

              1. ai-native userDo everything through the API that I can do in the UI

                weight 2 · round to Codex

                Codex ships rich CLI/UI-only capabilities (cloud tasks, resume/review, skills, plugins, MCP client integration) with no evidence these are exposed via a dedicated Codex API, and the general OpenAI API (RBAC, Responses API) is not shown to cover Codex-specific workflows; community evidence even confirms the latest gpt-5.3-codex model 'isn't available on the API yet,' a documented parity gap. Missing for 10: documented API endpoints for cloud task delegation, chat/session resume, MCP tool orchestration, and confirmation that current models/features are API-accessible at parity with CLI/UI.

                • [github] You can also use Codex with an API key, but this requires additional setup.
                • [community] gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'
                • [claimed-docs] Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…
                • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
                • [claimed-docs] codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …
                Slatenone0/10

                No evidence of any public API for Slate — the openapi.json/swagger.json probes returned 404s and no docs reference programmatic endpoints; Slate is documented purely as a CLI/terminal agent with slash-commands, hotkeys, and config files, not an API-driven product with UI/API parity.

                • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
                • [probe] official CLI documented at https://docs.randomlabs.ai/en/getting-started/quickstart
                • [claimed-docs] npm i -g @randomlabs/slate
              2. ai-native userExport all of my data in open formats and leave

                weight 3 · round drawn
                Codexnone0/10

                No evidence of any data export feature or open-format export mechanism for chat history, project data, or configurations; Codex works with local files/git repos but there's no documented export/portability capability for user data (e.g., conversation history, settings) to leave the platform. Missing for 10: any documented data export tool, open-format export (JSON/Markdown dump), or data portability statement.

                  Slatenone0/10

                  No evidence in the docs or elsewhere describes any data export functionality, open-format export, or data portability mechanism for Slate. Sessions, workspace history, and configurations appear stored locally but no documented export/leave path is mentioned. Missing for 10: any documentation of export commands, data format specifications, or account/data portability guarantees.

                  • ai-native userRead the product's source under an open license

                    weight 2 · round to Codex

                    The Codex CLI source lives in a public GitHub repo (openai/codex) and a community comment implies its openness lets users 'diagnose and file an issue' the way they can't with the closed-source Codex App, suggesting at least the CLI's code is publicly viewable. However, no evidence pack item states an explicit open-source license, and the App/cloud components are explicitly described as closed. missing for 10: explicit license file/name (MIT, Apache, etc.), confirmation the full product (not just CLI) is open, and independent corroboration beyond one forum remark.

                    • [github] Codex CLI is a coding agent from OpenAI that runs locally on your computer.
                    • [community] I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …
                    Slatenone0/10

                    No evidence of an open-source license or public source repository for Slate; the CLI is distributed via npm install with no mention of source availability. missing for 10: open-source license declaration, public source repo link, license file/OSS registry evidence.

                    • ai-native userSelf-host the core product

                      weight 3 · round drawn
                      Codexnone0/10

                      Codex CLI runs locally but requires signing into a ChatGPT account or OpenAI API key, and the core inference/model and cloud environments are OpenAI-hosted only; there is no self-hosted backend option. A commenter explicitly wishes the Codex App were open source, implying it is not, which forecloses self-hosting the core product.

                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] You can also use Codex with an API key, but this requires additional setup.
                      • [community] I wish Codex App was open source. I like it, but there are always a bunch of little paper cuts that, if you were using codex cli, you could …
                      Slatenone0/10

                      No evidence anywhere in the docs of Slate being open-source or offering a self-hosted deployment option; it's installed via npm as a CLI that connects to model subscriptions/credits, implying a hosted/service model rather than self-hostable core infrastructure. Missing for 10: any mention of self-hosting instructions, open-source repo, or on-prem deployment option.

                      • [claimed-docs] npm i -g @randomlabs/slate
                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…

                    Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits

                    Free-tier ceilings, usage caps, and rate limits before you have to pay

                    Authentication

                    1. developerAuthenticate with an API key instead of an account login

                      weight 2 · round to Codex

                      GitHub docs confirm Codex CLI supports API key authentication as an alternative to ChatGPT account login, but note it 'requires additional setup,' and the account-login flow (Sign in with ChatGPT) is the recommended default. Missing for 10: detailed API-key setup documentation, first-party quickstart parity with account login, and independent confirmation that API-key auth is fully feature-equivalent (e.g. codex-comm-5 shows some newer models aren't even available via API yet).

                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] You can also use Codex with an API key, but this requires additional setup.
                      • [community] gpt-5.3-codex isn't available on the API yet — 'We are working to safely enable API access soon.'
                      Slatenone0/10

                      No evidence pack mentions API key authentication as an alternative to account login; onboarding docs only describe choosing a model source (ChatGPT/Codex, SuperGrok, or Slate credits subscription), not API-key auth. No mention of an API key mechanism anywhere, and the openapi probe returned 404s, giving no indication of an API-key based auth path.

                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
                    2. engineering-leadAuthenticate through an enterprise identity or cloud platform for compliance and scalability

                      weight 2 · round to Codex

                      Codex supports signing in with a ChatGPT Business/Enterprise/Edu account (codex-gh-3, codex-gh-7) and OpenAI's platform offers RBAC to scope access at org/project level (codex-docs-28), suggesting enterprise-grade authentication and access control exist. However, there is no explicit documentation of SSO/SAML/OIDC federation with enterprise identity providers (e.g., Okta, Azure AD) specific to Codex, nor details on how ChatGPT Enterprise auth ties into RBAC for Codex usage. Missing for 10: explicit SSO/SAML/OIDC integration docs, enterprise IdP federation details, and independent confirmation of compliance-grade auth flows for Codex specifically.

                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…
                      • [claimed-docs] Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…
                      Slatenone0/10

                      No evidence of SSO/SAML/OIDC enterprise identity integration or cloud-platform authentication for compliance; onboarding only mentions choosing a model source (ChatGPT/Codex, SuperGrok, or Slate credits), not enterprise identity federation.

                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                    3. developerSign in with my existing product subscription plan to use the coding agent

                      weight 2 · round to Codex

                      GitHub docs explicitly recommend signing in with ChatGPT to use Codex under existing Plus, Pro, Business, Edu, or Enterprise subscription plans, with API key as an alternative for those without such plans, directly confirming subscription-based sign-in. missing for 10: independent hands-on confirmation of the sign-in flow itself (evidence focuses on capability descriptions rather than a walkthrough).

                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…
                      • [github] Run codex and select Sign in with ChatGPT. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, …
                      • [github] You can also use Codex with an API key, but this requires additional setup.

                      Docs explicitly state onboarding lets you choose your model source as your existing ChatGPT/Codex subscription or SuperGrok subscription (in addition to Slate credits), directly matching the story of signing in with an existing subscription plan to use the agent. Missing for 10: independent/hands-on confirmation that subscription sign-in actually works end-to-end and any detail on limitations of that mode vs credits.

                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                    4. developerSign in with a personal account to get free-tier access without managing API keys

                      weight 1 · round to Codex

                      Codex CLI explicitly recommends signing in with a ChatGPT account (Plus/Pro/Business/Edu/Enterprise) to use Codex without an API key, with API key usage noted as an alternative requiring additional setup. This directly matches the story of personal-account sign-in without managing API keys, though the exact free-tier scope/limits aren't detailed. Missing for 10: explicit confirmation of a genuinely free tier (vs. paid ChatGPT plans) and independent corroboration of the login flow's simplicity.

                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] You can also use Codex with an API key, but this requires additional setup.
                      • [github] Run `codex` and select **Sign in with ChatGPT**. We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Busi…

                      Docs show onboarding lets users choose a model source including an existing ChatGPT/Codex or SuperGrok subscription instead of managing API keys, implying account-based auth is supported, but there's no explicit mention of a free tier or of signing in with a personal Slate account for free credits without a paid subscription. Missing for 10: explicit free-tier account sign-in flow, confirmation that 'Slate credits' option requires no payment, and any account-based (not subscription-based) login mechanism.

                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…

                    Model choice

                    1. developerLet the tool automatically pick the best model for each task

                      weight 1 · round to Slate
                      Codexnone0/10

                      Evidence shows Codex lets users manually choose the model and reasoning effort ('Stay in control: Choose the model, reasoning effort, permissions...') rather than any automatic best-model-per-task selection; no docs or community evidence describe an automatic model-routing/selection feature tied to cost or task type.

                      • [claimed-docs] Stay in control: Choose the model, reasoning effort, permissions, and commands that fit the task.
                      • [community] The main issue I have with Codex is that the best model is insanely slow, except at nights and weekends when Silicon Valley goes to bed... I…
                      • [community] First thoughts using gpt-5.3-codex-spark in Codex CLI: Blazing fast but it definitely has a small model feel... It has to be prompted to do …

                      Docs explicitly claim Slate 'automatically selects the right model for the job' and also allow developers to set preferred default models per slot via `/models` or `slate.json`, suggesting a hybrid automatic+manual approach relevant to pricing/limits tradeoffs. However, there's no detail on the selection logic, cost-awareness, or independent verification that auto-selection actually optimizes for task/price. Missing for 10: independent hands-on confirmation of auto-selection quality, explanation of selection criteria (cost vs capability), and evidence of pricing-limit awareness in model choice.

                      • [claimed-docs] Slate automatically selects the right model for the job.
                      • [claimed-docs] Set preferred default models for each slot with the `/models` dialog or `slate.json` under `models`.
                    2. developerChoose which underlying AI model powers my session from multiple providers

                      weight 2 · round to Slate
                      Codexnone0/10

                      Docs confirm Codex lets users 'Choose the model, reasoning effort, permissions' (codex-docs-31), but this refers to selecting among OpenAI's own Codex/GPT models, not switching between different AI providers (e.g., Anthropic, Google). No evidence shows Codex supports plugging in or selecting non-OpenAI models/providers within a session.

                      • [claimed-docs] Stay in control: Choose the model, reasoning effort, permissions, and commands that fit the task.
                      • [github] We recommend signing into your ChatGPT account to use Codex as part of your Plus, Pro, Business, Edu, or Enterprise plan.
                      • [github] You can also use Codex with an API key, but this requires additional setup.

                      Docs confirm model source can be chosen at onboarding (ChatGPT/Codex, SuperGrok, or Slate credits) and that default models per 'slot' can be set via `/models` or slate.json, showing multi-provider flexibility. However, this is framed around subscription/credit sources rather than a clear list of many independent model providers, and there's no independent/hands-on verification of switching providers mid-session. missing for 10: independent corroboration of provider switching, a full list of supported model providers, and confirmation this works reliably in practice.

                      • [claimed-docs] Onboarding asks for your terminal background, multiline input preference, and model source: your ChatGPT/Codex subscription, SuperGrok subsc…
                      • [claimed-docs] Set preferred default models for each slot with the `/models` dialog or `slate.json` under `models`.

                    Privacy posture — data-handling and privacy storiesPrivacy posture

                    Data-handling and privacy stories

                    1. ai-native userChoose where my data is stored (region/residency)

                      weight 2 · round drawn
                      Codexnone0/10

                      No evidence in the pack mentions data residency, regional storage options, or geographic controls for where Codex data is stored; the pack covers RBAC, MCP, CLI features, and cloud task execution but nothing about choosing a storage region.

                        Slatenone0/10

                        Slate is a CLI coding agent tool; the evidence pack contains no mention of data residency, region selection, or storage location controls. Missing for 10: any documentation of data residency options, regional storage configuration, or compliance controls.

                        • ai-native userPrevent my data from being used to train AI models

                          weight 3 · round drawn
                          Codexnone0/10

                          The evidence pack contains no mention of data-training opt-out controls, enterprise data usage policies, or privacy settings for excluding user data from model training; it covers CLI features, MCP, RBAC, and community sentiment but nothing about training-data exclusion.

                            Slatenone0/10

                            No evidence in the pack addresses data-training opt-out, privacy policy, or any control over model training use; the documentation covers CLI usage, orchestration, and skills but nothing about data privacy posture. Missing for 10: any privacy policy statement, opt-out settings, or data usage terms regarding AI training.

                            • ai-native userControl data retention and deletion

                              weight 2 · round drawn
                              Codexnone0/10

                              No evidence pack items address data retention controls, deletion policies, or configurable retention windows for Codex; RBAC docs address access control, not retention/deletion. Missing for 10: any documentation of data retention settings, deletion APIs/workflows, or retention policy configuration.

                                Slatenone0/10

                                No evidence pack items mention data retention policies, deletion controls, or privacy settings for user data/sessions; docs cover workspace management and permissions but not data retention/deletion. Missing for 10: any documentation of data retention periods, deletion mechanisms, or export/erase controls.

                                • ai-native userOpt out of telemetry and usage tracking

                                  weight 2 · round drawn
                                  Codexnone0/10

                                  No evidence in the pack mentions telemetry, usage tracking, data collection settings, or an opt-out mechanism for Codex; the docs and community threads cover features like MCP, CLI usage, and performance but never privacy/telemetry controls.

                                    Slatenone0/10

                                    No evidence pack item mentions telemetry, usage tracking, analytics, or an opt-out setting anywhere in Slate's docs or community coverage; the closest item (diagnostic attachment on bug reports) doesn't address general telemetry opt-out. Missing for 10: any mention of telemetry collection, a privacy policy, or a documented opt-out flag/setting.

                                    • [claimed-docs] Slate automatically attaches relevant diagnostic information (OS, version, session context) to your report.

                                  Review safety — keeping generated changes safe — diffs, approvals, guardrailsReview safety

                                  Keeping generated changes safe — diffs, approvals, guardrails

                                  Data governance

                                  1. engineering-leadOpt out of having my code and prompts used for AI model training

                                    weight 1 · round drawn
                                    Codexnone0/10

                                    No evidence in the pack addresses data usage or training opt-out policies for code/prompts; RBAC and MCP docs are unrelated to this axis. Missing for 10: any enterprise data-usage/training opt-out policy documentation, admin controls for opting out, or third-party confirmation of such a policy.

                                      Slatenone0/10

                                      No evidence in the pack addresses data usage, training opt-out, or privacy policy for prompts/code submitted to Slate or its model providers.

                                      Pr review

                                      1. developerHave the agent stage changes, write commit messages, create branches, and open pull requests

                                        weight 3 · round to Codex

                                        Codex docs explicitly describe inspecting diffs and opening a pull request when cloud work is ready (codex-docs-5), and CLI docs note reviewing changes 'before you commit or open a pull request' (codex-docs-45), implying git workflow integration. However, staging changes, writing commit messages, and creating branches are not explicitly documented as first-class agent actions — they are only implied via general local repo access and command execution (codex-docs-9, codex-docs-30, codex-docs-17). Missing for 10: explicit documentation of commit-message generation, branch creation, and staging as named agent capabilities, plus independent hands-on confirmation of full PR workflow automation.

                                        • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
                                        • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                        • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
                                        • [claimed-docs] Start Codex in a repository to explore unfamiliar code, plan a change, edit files, and run your local development tools.
                                        • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
                                        Slatenone0/10

                                        No evidence in the pack mentions git operations like staging, committing, branching, or opening pull requests; documentation covers sessions, orchestration, skills, permissions, and CLI setup but not any git/PR workflow. Absence of evidence for this applicable capability means the verdict is none.

                                        • developerInspect diffs and run checks to catch problems before merging

                                          weight 3 · round to Codex

                                          Codex CLI has a dedicated review command that inspects diffs against uncommitted changes, a commit, or a base branch, reporting prioritized findings without modifying the working tree, plus cloud/web flows to inspect summaries and diffs before opening a PR. Missing for 10: independent/hands-on corroboration of the review command's accuracy and any CI-integrated check-running beyond exec/scripts.

                                          • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                          • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
                                          • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                          • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
                                          • [claimed-docs] Compose with scripts and CI: Use Codex interactively or call codex exec from repeatable workflows and pipelines.

                                          Slate's docs mention it can perform integration tests manually (random-labs-docs-16), implying some check-running capability, but there is no evidence of diff inspection, git diff review, PR-style change summaries, or pre-merge validation workflows. missing for 10: diff/change inspection UI or command, explicit pre-merge check/test running workflow, and any corroborating hands-on evidence of catching problems before merge.

                                          • [claimed-docs] Slate is one of the few agents capable of performing integration tests manually.

                                        Safe execution

                                        1. engineering-leadControl which external tools and integrations the agent is allowed to access

                                          weight 2 · round drawn

                                          Codex documents fine-grained control over external tool access at the session/repo level: engineers can add/remove local or remote MCP servers, inspect available tools before they're used, and set permission boundaries for edits/commands via /permissions (codex-docs-16, codex-docs-38, codex-docs-39, codex-docs-46). This gives an engineer meaningful control over which integrations the agent can reach, and RBAC exists for org/project-level API access (codex-docs-28), but that RBAC is about API/dashboard permissions, not specifically about restricting agent tool/integration access org-wide for a lead managing a team's Codex usage. Missing for 10: evidence of centralized, lead-enforced policy that restricts which MCP servers/tools individual developers can enable (vs. per-session self-configuration), and independent confirmation this control actually prevents unauthorized tool access in practice.

                                          • [claimed-docs] Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.
                                          • [claimed-docs] Connect external tools with MCP — codex mcp: Add local or remote MCP servers, authenticate when needed, and inspect the tools available to t…
                                          • [claimed-docs] Set the boundaries for each run — /permissions: Choose when Codex can edit files or run commands without asking, and inspect the active sand…
                                          • [claimed-docs] In the `codex` TUI, use `/mcp` to see your active MCP servers.
                                          • [claimed-docs] Role-based access control (RBAC) lets you decide who can do what across your organization and projects—both through the API and in the Dashb…

                                          Slate's configuration docs describe a permission system where each permission key maps to allow/ask/deny actions or fine-grained pattern objects, which supports controlling what tools/actions the agent can perform, and a `--yolo` flag exists to bypass these prompts entirely. However there's no explicit documentation of controlling specific external integrations (e.g., MCP servers, API connectors) or org/team-level lockdown for an engineering lead specifically. Missing for 10: explicit external-integration/MCP allowlist docs, engineering-lead/team-level enforcement (vs individual config), and independent verification that permission enforcement can't be trivially bypassed.

                                          • [claimed-docs] Each permission key maps to an action ("allow", "ask", or "deny"), or a pattern object for fine-grained control.
                                          • [claimed-docs] We support `--dangerously-skip-permissions` (alias: `--yolo`) to bypass permission prompts.
                                        2. engineering-leadHave the agent operate inside a sandbox when interacting with code, tools, and network resources

                                          weight 2 · round to Codex

                                          First-party docs explicitly describe sandboxed execution: Codex lets you 'choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots' (codex-docs-17), and cloud tasks run in 'isolated cloud environments' with configurable dependencies/tools (codex-docs-1, codex-docs-4). This directly matches the engineering-lead's need for sandboxed code/tool interaction, though network-resource sandboxing specifics are not spelled out and there's no independent hands-on verification of sandbox robustness (a community comment raises but does not concretely confirm a sandbox-bypass issue). Missing for 10: explicit documentation of network-level sandbox controls, and independent/hands-on confirmation that the sandbox reliably contains tool/network access.

                                          • [claimed-docs] Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.
                                          • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                                          • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
                                          • [community] Do people really want codex to have control over their computer and apps? I'm still paranoid about keeping things securely sandboxed.
                                          Slatenone0/10

                                          The evidence shows a permission system (allow/ask/deny actions) and a --yolo flag to bypass prompts, but there is no mention of sandboxed execution, containerization, or network isolation for the agent's code/tool interactions. missing for 10: any documentation of sandbox/container execution, network isolation controls, or filesystem confinement mechanisms.

                                          • [claimed-docs] Each permission key maps to an action ("allow", "ask", or "deny"), or a pattern object for fine-grained control.
                                          • [claimed-docs] We support `--dangerously-skip-permissions` (alias: `--yolo`) to bypass permission prompts.

                                        Security checks

                                        1. engineering-leadSee license and public-code matching references for AI-suggested code

                                          weight 1 · round drawn
                                          Codexnone0/10

                                          No evidence anywhere in the pack mentions license detection, public-code matching, or provenance references for AI-suggested code; Codex's review features (codex-docs-10, -41, -45) only cover code quality/prioritized findings, not license/public-code attribution.

                                            Slatenone0/10

                                            No evidence anywhere in the pack mentions license compliance checks, public-code/OSS matching, or provenance references for AI-suggested code; the docs cover orchestration, skills, permissions, and CLI usage but nothing about code provenance/license scanning.

                                            • developerGet contextual explanations and automatic fixes for security vulnerabilities

                                              weight 2 · round to Codex

                                              Codex CLI has a dedicated review command that inspects uncommitted changes, commits, or branches and reports 'prioritized findings' (codex-docs-10, codex-docs-41, codex-docs-45), which could surface security issues, and as a general coding agent it can edit files/run commands. However, the review feature explicitly reports findings 'without modifying your working tree,' meaning it does not auto-fix, and no evidence specifically frames this as security-vulnerability detection/explanation with automatic remediation. missing for 10: explicit security-vulnerability scanning/explanation feature, evidence of automatic fix application (vs. just flagging), and any independent confirmation that Codex reliably identifies/fixes security issues.

                                              • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                              • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
                                              • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                              • [claimed-docs] Work against your local repository: Let Codex inspect files, make edits, and run the tools already installed on your machine.
                                              Slatenone0/10

                                              No evidence in the pack mentions security vulnerability detection, explanations, or automatic fixes; documentation covers session management, orchestration, skills, and configuration but nothing about security review or vulnerability remediation. Missing for 10: any mention of vulnerability scanning, security explanations, or auto-fix capability.

                                              Not comparable on these axes

                                              1. ai-native userConnect an agent via an official MCP server

                                                weight 3 · not comparable

                                                Codex explicitly supports running itself as an MCP server (codex mcp-server) so other MCP clients can connect, but OpenAI's own docs mark this interface 'experimental' and now 'deprecated', pointing users to a newer 'Codex app server' as the recommended replacement. This is a genuine server-mode capability (not just Codex-as-MCP-client), but the deprecation and lack of independent hands-on confirmation of the replacement's stability keep it from a full verdict. Missing for 10: independent corroboration that the current 'Codex app server' MCP mode works reliably in production, and clearer first-party documentation of its interface now that the original is deprecated.

                                                • [github] Codex MCP Server Interface [experimental]: a JSON-RPC API that runs over the Model Context Protocol (MCP) transport to control a local Codex…
                                                • [claimed-docs] codex mcp-server is deprecated. Use the Codex app server instead. ... This page documents the deprecated command for existing integrations. …
                                                • [claimed-docs] Add local or remote MCP servers, authenticate when needed, and inspect the tools available to the current session before Codex uses them.
                                                Slaten/a

                                                Slate is itself a coding agent (CLI-based, with sessions, subagents, skills); serving as an MCP server for other agents to connect to is a different product role. No evidence shows Slate exposing an official MCP server endpoint, so the axis does not apply per the agent-role exception.

                                                • ai-native userSubscribe to events via webhooks

                                                  weight 2 · not comparable
                                                  Codexnone0/10

                                                  No evidence in the pack mentions webhooks or event subscription capabilities for Codex; the product exposes MCP servers, CLI, and cloud task integrations but nothing about outbound webhook events for AI-native consumers.

                                                    Slaten/a

                                                    Slate is a CLI-based coding agent, not a service/platform exposing an event system; webhook subscriptions are outside its product category, and no evidence pack item references webhooks or event subscriptions at all.

                                                    • ai-native userExplore an interactive API reference with runnable examples

                                                      weight 2 · not comparable

                                                      The evidence shows OpenAI's general API reference (developers.openai.com) has runnable, per-language code samples with live examples, which an AI-native user could explore. However this is the general OpenAI Responses API reference, not a Codex-specific interactive API reference, and Codex itself is documented as a CLI/agent product rather than an API with its own dedicated reference docs. Missing for 10: a Codex-specific API reference page, evidence of interactivity beyond code-sample selection (e.g., live sandbox execution), and any Codex-specific documentation of this reference.

                                                      • [claimed-docs] Create a model response — request/response reference with runnable code samples selectable per language: HTTP, Python, TypeScript, Go, Ruby,…
                                                      • [github] You can also use Codex with an API key, but this requires additional setup.
                                                      Slaten/a

                                                      Slate is a CLI coding agent, not an API/service product with its own API reference; the probe explicitly found no OpenAPI spec, confirming this axis is a category mismatch rather than a missing feature.

                                                      • [probe] PROBE openapi: all candidate paths 404 (https://docs.randomlabs.ai/openapi.json, https://docs.randomlabs.ai/swagger.json, https://docs.rando…
                                                    • ai-native userTest against a sandbox environment without touching production data

                                                      weight 1 · not comparable

                                                      Codex offers isolated cloud task environments and CLI sandbox controls (writable roots, permission gating) that keep agent actions contained rather than acting directly on a live system, which functions as a sandbox layer for testing changes. However, there's no explicit documentation of test-vs-production data separation, and a community report raises unresolved concerns about the sandbox reading sensitive filesystem data without asking. Missing for 10: explicit production-data isolation guarantees, first-party documentation addressing the raised sandbox-safety concern, and independent verification that isolated environments never touch real prod data.

                                                      • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                                                      • [claimed-docs] Configure the dependencies, tools, variables, and setup steps each repository needs.
                                                      • [claimed-docs] Choose when Codex can edit files or run commands without asking, and inspect the active sandbox and writable roots before you continue.
                                                      • [community] Does that version of Codex still read sensitive data on your file system without even asking? Just curious. [links to github.com/openai/code…
                                                      Slaten/a

                                                      Slate is a coding agent CLI tool, not a service with a sandbox/production data separation model; there's no evidence of a hosted environment with production data that would need a sandbox testing mode. This axis is a category error for this kind of local developer tool.

                                                      • developerReceive inline code completions and next-edit suggestions as I type

                                                        weight 3 · not comparable
                                                        Codexn/a

                                                        Codex is an agentic coding assistant (CLI, cloud tasks, IDE extension) focused on delegated task completion, code review, and terminal-based editing, not on inline autocomplete-style completions or next-edit suggestions as you type. This story targets IDE-style inline autocomplete tooling, a different axis than Codex's agent-driven workflow model.

                                                          Slaten/a

                                                          Slate is a terminal/CLI-based agentic coding assistant that operates via chat sessions, orchestration, and shell commands, not an IDE-integrated editor extension providing inline completions or next-edit suggestions as the user types. This story targets an IDE-autocomplete category error for Slate's product type.

                                                          • developerDebug a live running web application directly from my coding assistant

                                                            weight 1 · not comparable
                                                            Codexn/a

                                                            Codex is a coding agent focused on code generation, editing, review, and CLI/cloud task automation; there is no evidence of any capability to attach to or inspect a live running web application (e.g., browser DevTools integration, runtime debugging, log/network inspection of a live app). Debugging a live running app is a different axis (runtime observability/dev-tools) than code editing and static review, which is what this product's evidence covers.

                                                              Slatenone0/10

                                                              No evidence in the pack mentions debugging live running applications, attaching to running processes, browser/runtime debugging, or any live-app inspection capability; Slate's docs focus on codebase review, shell commands, orchestration, and skills, none of which address live debugging.

                                                              • developerView interactive diffs and share selected code as context from within my JetBrains IDE

                                                                weight 1 · not comparable
                                                                Codexnone0/10

                                                                Evidence shows Codex's IDE extension explicitly targets VS Code, Cursor, and Windsurf (codex-gh-2), with no mention of JetBrains IDEs, interactive diff viewing within an IDE, or a 'share selected code as context' feature. The axis (IDE integration) is clearly applicable to Codex as a coding agent, but JetBrains-specific support and the described interactive-diff/context-sharing workflow are simply absent from the evidence pack.

                                                                • [github] If you want Codex in your code editor (VS Code, Cursor, Windsurf), install in your IDE.
                                                                Slaten/a

                                                                Slate is a terminal/CLI-based coding agent (npm-installed CLI, terminal UI, hotkeys), with no evidence of a JetBrains IDE plugin, interactive diff viewer inside an IDE, or IDE-based context sharing. This story targets IDE-native integration, which is a different product surface than Slate's terminal-first design.

                                                              • developerGet automatic code review with contextual feedback on every pull request

                                                                weight 3 · not comparable

                                                                Codex CLI/cloud ships a dedicated 'review' capability that inspects uncommitted changes, a commit, or a base branch and reports prioritized findings without touching the working tree, and cloud tasks can be kicked off from GitHub PRs and later opened as PRs. However, there is no evidence of an automatic, PR-triggered review bot that comments on every pull request without manual invocation. missing for 10: evidence of automatic triggering on every PR (e.g., GitHub App/webhook auto-review), evidence of inline PR comments, independent confirmation of review quality on real PRs.

                                                                • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                                                • [claimed-docs] Review changes before they ship: Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized f…
                                                                • [claimed-docs] Run a dedicated review against uncommitted changes, a commit, or a base branch. Codex reports prioritized findings without modifying your wo…
                                                                • [claimed-docs] Run tasks in isolated cloud environments, work in parallel, and start work from the web, GitHub, GitLab, Linear, or Slack.
                                                                • [claimed-docs] Start work in Codex cloud from GitHub pull requests, GitLab merge requests and issues, Linear issues, or Slack channels and threads.
                                                                • [claimed-docs] Inspect the summary and diff, request a follow-up, or open a pull request when the result is ready.
                                                                Slaten/a

                                                                Slate is a terminal-based coding agent CLI (session management, orchestration, skills, permissions) with no evidence of PR/VCS integration or automated code review on pull requests. Automatic PR review is a GitHub/CI-integration feature category, not something this agentic CLI tool is positioned to do — no docs mention PR hooks, CI integration, or review workflows tied to pull requests, making this a category mismatch rather than a gap in an applicable feature.