Skip to content

Anthropic Skills vs Skills for Real Engineers

open-source · free-tier

·

open-source · free-tier

Anthropic Skills wins · 1511 (14 drawn)

Agenticness
AGENT-READY18/100BUILT-IN AI26/100Skills for Real Engineers

Agent workflows — stories about agent workflows in this arenaAgent workflows

Stories about agent workflows in this arena

Agent ops

  1. ai-native userMy agent can author and package a new skill end to end by following the project's own spec, template, or meta-skill

    weight 2 · round to Anthropic Skills
    Anthropic Skillsfullclaimed9/10

    Anthropic ships a dedicated 'skill-creator' meta-skill for creating new skills and iteratively improving/benchmarking them, plus an official template SKILL.md and the open agentskills.io specification detailing required frontmatter, description rules, and folder structure (scripts/references/assets) — together these let an agent author and package a skill end-to-end per spec. Missing for 10: independent hands-on account of an agent successfully using skill-creator to author a new skill from scratch (community evidence only discusses skill triggering/invocation issues, not authoring/packaging).

    • [claimed-docs] A skill for creating new skills and iteratively improving them.
    • [claimed-docs] benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy
    • [claimed-docs] Create new skills, modify and improve existing skills, and measure skill performance.
    • [claimed-docs] Replace with description of the skill and when Claude should use it.
    • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
    • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
    • [claimed-docs] name: skill-name description: A description of what this skill does and when to use it.
    • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
    • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
    • [claimed-docs] add a skill, and test it locally using the \`--plugin-dir\` flag
    Skills for Real Engineersnone0/10

    The evidence pack describes numerous existing skills (tdd, grilling, triage, setup-matt-pocock-skills) and packaging/distribution mechanics (npx skills, Claude plugin marketplace), but nothing documents a meta-skill, spec, or template that an agent would follow to author and package a brand-new SKILL.md from scratch within this project's own conventions.

    • ai-native userMy coding agent can install a skill by itself — a non-interactive, promptless install path an agent can run headlessly end to end

      weight 3 · round drawn
      Anthropic Skillspartialclaimed4/10

      Docs show scriptable install commands (`/plugin marketplace add`, `/plugin install ...`) and a Skills API for uploading/managing skills programmatically, which could in principle be run non-interactively by automation. However, none of the evidence explicitly documents a promptless, fully headless end-to-end install flow (e.g., a CLI flag or API call an agent invokes autonomously without any human-run slash command or confirmation step). Missing for 10: explicit non-interactive/headless install documentation, evidence of an agent autonomously choosing and installing a skill without human-issued commands, and confirmation that no interactive confirmation/prompt is required during install.

      • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
      • [claimed-docs] /plugin install quality-review-plugin@my-plugins
      • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
      • [claimed-docs] Upload and manage through the [Skills API]
      • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
      • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
      • [github] /plugin marketplace add anthropics/skills
      Skills for Real Engineerspartialclaimed4/10

      The docs describe simple CLI install commands (`npx skills`, `claude plugins install mattpocock-skills`, `npx skills update`) that could in principle be scripted, but the accompanying setup skill explicitly interviews the user ('Ask you which issue tracker you want to use') and other examples show a human typing a slash command, not a fully headless agent-run flow. missing for 10: explicit non-interactive/CI flag or documented flow for an agent to run install end-to-end without any prompts, and confirmation that the interactive setup step can be skipped or automated.

      • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
      • [claimed-docs] Install the ones you want, then type a slash command.
      • [claimed-docs] Ask you which issue tracker you want to use (GitHub, Linear, or local files)
      • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
      • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`

    Agenticness — how well agents can access and operate the productAgenticness

    How well agents can access and operate the product

    Agent access

    1. ai-native userPoint an agent at llms.txt or agent-oriented docs

      weight 2 · round to Anthropic Skills
      Anthropic Skillsfullprobed8/10

      Anthropic's docs site serves a working llms.txt (HTTP 200) plus .md variants of docs pages (e.g. skills.md), explicitly designed for agent/LLM consumption, and Claude Code skills follow the agentskills.io open spec. missing for 10: no independent/community confirmation that agents were actually pointed at llms.txt and successfully used it end-to-end.

      • [probe] PROBE llms.txt: HTTP 200 at https://code.claude.com/llms.txt # Claude Code Docs > Official documentation for Claude Code, Anthropic's agent…
      • [probe] PROBE docs-md: HTTP 200 at https://code.claude.com/docs/en/skills.md > ## Documentation Index > Fetch the complete documentation index at: h…
      • [claimed-docs] Claude Code skills follow the [Agent Skills](https://agentskills.io) open standard, which works across multiple AI tools.
      Skills for Real Engineerspartialprobed5/10

      The product's entire mechanism is agent-oriented markdown docs (SKILL.md, CONTEXT.md files) explicitly written for agents to read and act on, and it documents raw.githubusercontent URLs an agent could be pointed at directly. However, there is no evidence of a dedicated llms.txt for this product itself — the probe shows only GitHub's own generic llms.txt (unrelated to this project) and a 404 for skills.md, so the specific llms.txt convention is not supported, only the broader 'agent-readable docs' pattern. missing for 10: a product-specific llms.txt file, and confirmation that agents can be pointed at a single canonical docs entry point rather than individual SKILL.md files.

      • [claimed-docs] Test only at pre-agreed seams. Before writing any test, write down the seams under test and confirm them with the user.
      • [claimed-docs] Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**
      • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
      • [github] Here's an example [`CONTEXT.md`]... Which one is easier to read?
      • [probe] PROBE llms.txt: HTTP 200 at https://github.com/llms.txt # GitHub > GitHub is a developer platform for building, shipping, and maintaining s…
      • [probe] PROBE docs-md: HTTP 404 at https://github.com/mattpocock/skills.md
    2. ai-native userRun the product headlessly / in CI for automation

      weight 2 · round to Anthropic Skills
      Anthropic Skillspartialclaimed5/10

      Skills can be invoked programmatically via the Messages API (`container` parameter, Skills API for upload/management), which is inherently headless and scriptable, implying CI/automation use is possible. However, there is no explicit documentation or example of running Skills in a CI pipeline, headless mode, or automated build system — missing for 10: explicit CI/headless workflow examples, CLI flags for non-interactive automation, and independent evidence of production CI usage.

      • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
      • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
      • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
      • [claimed-docs] You can include up to 20 Skills for each request.
      • [claimed-docs] This guide shows you how to use both pre-built and custom Skills with the Claude API.
      Skills for Real Engineersnone0/10

      The evidence describes skill files consumed by interactive coding agents (Claude Code, Cursor, etc.) that rely on human interviewing/grilling and manual slash-command invocation, with no mention of a CLI flag, non-interactive mode, or CI/automation pipeline usage. Nothing in the docs, GitHub, or community evidence discusses running the skills headlessly or in a CI pipeline.

      • ai-native userUse an official CLI

        weight 2 · round to Skills for Real Engineers
        Anthropic Skillspartialprobed5/10

        Skills are used and managed through Claude Code, which is described as a terminal-based agentic CLI tool, via slash commands like `/skill-name`, `/plugin marketplace add`, and `--plugin-dir` flags for local testing; the `skill-creator` skill also supports building/testing skills. However, there is no evidence of a dedicated standalone 'skills' CLI binary or command set (e.g., `skills create`, `skills validate`) separate from Claude Code's general slash-command interface, and API-based skill management (Skills API) is not CLI-based at all. Missing for 10: a purpose-built skills CLI tool, independent hands-on confirmation of CLI-based skill workflows, and CLI support outside the Claude Code product.

        • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
        • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
        • [claimed-docs] add a skill, and test it locally using the \`--plugin-dir\` flag
        • [claimed-docs] A skill for creating new skills and iteratively improving them.
        • [probe] PROBE llms.txt: HTTP 200 at https://code.claude.com/llms.txt # Claude Code Docs > Official documentation for Claude Code, Anthropic's agent…
        Skills for Real Engineersfullclaimed7/10

        The product ships an official CLI (`npx skills`, with `npx skills update` to pull latest changes) and is also installable via the Claude Code plugin CLI route (`claude plugins install mattpocock-skills`), matching the AI-native CLI story. Missing for 10: independent hands-on verification of the CLI's full command set and behavior beyond first-party docs.

        • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
        • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
        • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
        • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
        • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
      • ai-native userBuild against official SDKs

        weight 2 · round to Anthropic Skills
        Anthropic Skillsfullclaimed7/10

        Docs describe a dedicated Skills API and Messages API `container` parameter for programmatically attaching Skills (with `skill_id`, versioning, up to 20 per request), plus upload/management endpoints, which constitutes an official API/SDK surface to build against. Missing for 10: explicit language-specific SDK code samples (Python/TypeScript) and independent hands-on developer reports confirming building production integrations against this API.

        • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
        • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
        • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
        • [claimed-docs] You can include up to 20 Skills for each request.
        • [claimed-docs] Upload and manage through the [Skills API]
        • [claimed-docs] This guide shows you how to use both pre-built and custom Skills with the Claude API.
        Skills for Real Engineersnone0/10

        The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

        Agentic features

        1. ai-native userGet AI-generated insights and suggestions from my data inside the product

          weight 2 · round to Skills for Real Engineers
          Anthropic Skillspartialcommunity4/10

          Skills can package data-handling capabilities (PDF, xlsx, docx extraction/manipulation) and are pitched for 'analyzing data using your organization's specific workflows,' giving Claude a path to generate insights from user data, but this is a general extensibility framework rather than a built-in insights/dashboard feature, and community reports show skills are frequently not invoked or unreliable in practice. missing for 10: a dedicated insights/analytics feature, proactive suggestion UI, and evidence that skills reliably surface unsolicited insights rather than requiring explicit triggering.

          • [claimed-docs] Claude already knows a lot about understanding PDFs, but is limited in its ability to manipulate them directly (e.g. to fill out a form). Th…
          • [claimed-docs] Skills extend Claude’s capabilities by packaging your expertise into composable resources for Claude, transforming general-purpose agents in…
          • [github] whether that's creating documents with your company's brand guidelines, analyzing data using your organization's specific workflows, or auto…
          • [github] whether that's creating documents with your company's brand guidelines, analyzing data using your organization's specific workflows, or auto…
          • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
          • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
          Skills for Real Engineerspartialclaimed6/10

          The skill bundle includes several skills that produce AI-generated insights/suggestions from a user's own project data — a visual refactor-worthiness report (docs-18), bug diagnosis from a repro (docs-19), diff review against standards (docs-14), and issue triage/sorting (docs-30, docs-34) — but these are discrete slash-command skills rather than a unified 'insights' surface, and there's no evidence of a dashboard or proactive analytics view. missing for 10: a consolidated insights UI/report aggregating findings, independent hands-on evidence of the insight-generating skills actually producing useful output, and evidence these insights update automatically rather than being invoked per-skill.

          • [claimed-docs] Find the modules worth refactoring, as a visual report.
          • [claimed-docs] Diagnose a hard bug, starting from a repro that fails.
          • [claimed-docs] Review a diff against your standards and against the spec.
          • [claimed-docs] Sort raw issues into work someone can pick up.
        2. ai-native userSet up automations that run autonomously in the background

          weight 2 · round drawn
          Anthropic Skillsnone0/10

          Skills are packaged instructions/capabilities that Claude loads and uses during a session (invoked automatically or via /skill-name), but the evidence pack contains no mention of scheduling, triggers, or background/autonomous execution outside an active user session. Plugins and marketplaces cover distribution, not autonomous background automation.

          • [claimed-docs] Create a `SKILL.md` file with instructions, and Claude adds it to its toolkit.
          • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
          • [claimed-docs] If Claude thinks the skill is relevant to the current task, it will load the skill by reading its full `SKILL.md` into context.
          Skills for Real Engineersnone0/10

          The skills are built around interactive, human-in-the-loop workflows (grilling/interviewing the user, confirming test seams, triage state machines) rather than unattended background automation; the project's own philosophy explicitly rejects processes that 'take away your control' in favor of user-confirmed steps. No evidence of scheduling, background triggers, or autonomous execution without human interaction is present in the pack.

          • [github] They help you align with the agent before you get started, and think deeply about the change you're making. Use them _every_ time you want t…
          • [claimed-docs] Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**
          • [claimed-docs] Verify the claim. Before any grilling, check that the claim holds up. For a bug, reproduce it from the reporter's steps.
          • [github] Approaches like GSD, BMAD, and Spec-Kit try to help by owning the process. But while doing so, they take away your control
        3. ai-native userDelegate tasks to a built-in AI assistant inside the product

          weight 3 · round to Anthropic Skills
          Anthropic Skillsdisputedcontradicted5/10

          Anthropic's docs claim Claude will autonomously discover and load relevant Skills to perform delegated work ('Claude uses skills when relevant... transforms general-purpose agents into specialized agents'), which matches the story of delegating tasks to a built-in assistant. However, hands-on community reports directly contradict reliable automatic delegation: a Vercel eval found the skill was never invoked in 56% of cases despite being available, and multiple users report invocation is 'hit or miss' even with proper YAML triggers, often requiring explicit manual pointers. Missing for 10: independent benchmarks showing consistent autonomous task delegation, and resolution of the documented invocation-reliability gap.

          • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
          • [claimed-docs] Skills extend Claude’s capabilities by packaging your expertise into composable resources for Claude, transforming general-purpose agents in…
          • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
          • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
          • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
          Skills for Real Engineersnone0/10

          The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

          • ai-native userOperate the product with natural-language commands

            weight 2 · round to Skills for Real Engineers
            Anthropic Skillsdisputedcontradicted5/10

            Anthropic's docs explicitly promise natural-language operation: Claude 'uses skills when relevant' by matching the task to a skill's description, in addition to explicit `/skill-name` invocation (anthropic-skills-docs-2, -3, -25, -34). However, hands-on community reports concretely contradict this: a Vercel eval found skills were never invoked in 56% of cases despite being available, and multiple practitioners report skills are 'hit or miss' or require explicitly telling Claude to use them even when asked (anthropic-skills-comm-7, -8, -9). Missing for 10: reliable first-party benchmark of natural-language trigger accuracy, and resolution of the discovery/triggering inconsistency reported by users.

            • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
            • [claimed-docs] Skills add optional features: a directory for supporting files, frontmatter to [control whether you or Claude invokes them]... and the abili…
            • [claimed-docs] If Claude thinks the skill is relevant to the current task, it will load the skill by reading its full `SKILL.md` into context.
            • [claimed-docs] frontmatter to [control whether you or Claude invokes them](#control-who-invokes-a-skill)
            • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
            • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
            • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
            Skills for Real Engineersfullclaimed8/10

            The skills are designed to be invoked and operated conversationally: users type slash commands (e.g., /grill-with-docs) and the agent then interviews/grills the user in natural language to reach shared understanding before acting (docs-43, gh-6, docs-6/26/33). This natural-language interaction model is central and repeated across multiple skill docs (grilling, triage, TDD flows). Missing for 10: independent hands-on confirmation that the natural-language command flow works smoothly in practice (community evidence only critiques prose quality, not the interaction mechanism), and no demonstration of free-form (non-slash) natural language command parsing beyond the interview pattern.

            • [claimed-docs] Install the ones you want, then type a slash command.
            • [github] getting the agent to ask you detailed questions about what you're building
            • [claimed-docs] Interview the user relentlessly until you reach a shared understanding. Map this as a **design tree**
            • [claimed-docs] Interview the user relentlessly until you reach a shared understanding.
            • [github] They help you align with the agent before you get started, and think deeply about the change you're making. Use them _every_ time you want t…
            • [claimed-docs] Turn an agreed conversation into a written spec.

          Api quality

          1. ai-native userTest against a sandbox environment without touching production data

            weight 1 · round to Anthropic Skills
            Anthropic Skillspartialclaimed4/10

            Skills invoked via the Messages API run inside Anthropic's 'code execution environment' (a sandboxed container), and plugin docs mention testing skills locally with the `--plugin-dir` flag before sharing/distribution, which implies some separation from a live/production setup. However, there is no explicit documentation of a dedicated sandbox/staging environment for testing skills against non-production data, no discussion of data isolation guarantees, and no hands-on validation of this specific safety property. Missing for 10: explicit sandbox/staging environment documentation, data-isolation guarantees, and independent confirmation that local/test skill runs cannot touch production data.

            • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
            • [claimed-docs] add a skill, and test it locally using the \`--plugin-dir\` flag
            • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
            Skills for Real Engineersnone0/10

            The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

            • ai-native userRely on versioned APIs with a documented deprecation policy

              weight 2 · round drawn
              Anthropic Skillsnone0/10

              Evidence covers Skills' structure, plugins, marketplaces, and API usage, but there is no mention of API versioning schemes or a documented deprecation policy for Skills/Claude API. Missing for 10: any documentation of API version numbers, backward-compatibility guarantees, or deprecation timelines/policy.

                Skills for Real Engineersnone0/10

                The evidence describes an update mechanism (`npx skills update`, 'nothing updates behind your back') but there is no documentation of semantic versioning, an API surface, or any deprecation policy for skills as they evolve or are removed. Missing for 10: any explicit version numbering scheme, changelog, or documented deprecation/backwards-compatibility policy for the skill files or plugin.

                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

              Automation depth — how much of the product can run unattendedAutomation depth

              How much of the product can run unattended

              1. ai-native userPerform bulk operations across many items at once

                weight 2 · round to Anthropic Skills
                Anthropic Skillspartialcommunity3/10

                Skills can bundle scripts that operate on multiple files (e.g., the PDF skill 'merges multiple PDFs' and fills forms across documents), suggesting some batch/bulk-processing capability, and skills may include arbitrary scripts/executable code for such tasks. However, there is no explicit documentation or example demonstrating bulk operations across many items (e.g., batch-processing hundreds of records/files) as a core Skills feature, and community feedback focuses on skill-triggering reliability rather than bulk-processing performance. Missing for 10: explicit bulk/batch-processing examples or docs, evidence of scale (many items processed reliably), and independent confirmation that bulk workflows work as intended.

                • [claimed-docs] Claude already knows a lot about understanding PDFs, but is limited in its ability to manipulate them directly (e.g. to fill out a form). Th…
                • [claimed-docs] skills can bundle additional files within the skill directory and reference them by name from `SKILL.md`.
                • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
                • [community] If you can write a bash or python script, or an API or MCP to do what you want, then write it and include it in the skill. Keep top-level co…
                Skills for Real Engineersnone0/10

                The skill set is built around structured, one-at-a-time workflows (grilling, triage state machine, TDD, spec-to-ticket) rather than any documented bulk/batch processing across many items simultaneously; no evidence describes running a skill across multiple issues, files, or specs in one operation. Missing for 10: any documented bulk/batch command or workflow, evidence of parallel/multi-item processing, and independent confirmation of such usage.

                • [claimed-docs] Move issues on the project issue tracker through a small state machine of triage roles.
                • [claimed-docs] Move issues on the project issue tracker through a small state machine of triage roles, categorise, verify, grill if needed, and write agent…
                • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
              2. ai-native userDefine rules that trigger actions automatically on events

                weight 3 · round to Anthropic Skills
                Anthropic Skillsdisputedcontradicted4/10

                Anthropic's docs claim skills auto-trigger ('Claude uses skills when relevant... loads it automatically') based on description matching, which is the closest analog to event-driven rule triggering in this product, but this is relevance-based context loading, not true event/webhook/schedule triggers. Hands-on community reports directly contradict reliability of this claimed automation: Vercel's eval found skills were never invoked in 56% of cases despite being applicable, and multiple users report 'hit or miss' triggering even with proper YAML triggers, sometimes requiring explicit manual invocation. Missing for 10: genuine event-based triggers (webhooks, schedules, file-watchers), reliable automatic invocation without manual prompting, and independent confirmation that trigger accuracy is dependable in production.

                • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
                • [claimed-docs] If Claude thinks the skill is relevant to the current task, it will load the skill by reading its full `SKILL.md` into context.
                • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
                • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
                • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
                Skills for Real Engineersnone0/10

                The product ships a library of manually-invoked skill files triggered by slash commands or explicit user direction ('Install the ones you want, then type a slash command'; 'Use them every time you want to make a change'), not an event-driven rule/automation engine. No evidence describes defining rules that fire automatically on repo events, webhooks, or triggers without user invocation.

                • [claimed-docs] Install the ones you want, then type a slash command.
                • [github] They help you align with the agent before you get started, and think deeply about the change you're making. Use them _every_ time you want t…
                • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
              3. ai-native userVersion, review, and roll back my automations

                weight 1 · round drawn
                Anthropic Skillspartialclaimed4/10

                Docs mention plugin marketplaces provide 'version tracking' and 'versioned releases', and the API lets you specify an optional `version` for skill_id, implying some versioning support. However there is no documented review/approval workflow or explicit rollback mechanism — skills are just files/folders, so any versioning or rollback would rely on external git tooling not described as a first-class feature. missing for 10: explicit rollback command/feature, in-product review or approval workflow for skill changes, changelog/diff tooling, independent confirmation of version tracking in practice.

                • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
                • [claimed-docs] Plugins (self-contained directories with skills, agents, hooks, or a `.claude-plugin/plugin.json` manifest) | `/plugin-name:hello` | Sharing…
                Skills for Real Engineerspartialclaimed4/10

                The product offers two install modes—a locked, read-only bundle that only updates via an explicit `npx skills update` pull, or an editable copy where skills become 'ordinary files you own' in your repo—which gives some control over when changes land and implies normal git-based versioning/rollback, but there is no explicit changelog, diff view, or rollback command for the skills themselves. Missing for 10: dedicated version history or diff/review UI for skill changes, an explicit rollback mechanism, and any evidence of reviewing skill updates before applying them (beyond opting into `update`).

                • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

              Cross agent portability — stories about cross agent portability in this arenaCross agent portability

              Stories about cross agent portability in this arena

              Portability

              1. developerInstall the same collection into multiple different coding agents — Claude Code, Codex, Cursor, and others — with per-harness instructions

                weight 3 · round to Skills for Real Engineers
                Anthropic Skillspartialclaimed4/10

                Docs state Claude Code skills follow the open 'Agent Skills' standard which 'works across multiple AI tools' (agentskills.io spec), implying cross-agent portability, and the spec itself defines a tool-agnostic SKILL.md format. However, there is no concrete evidence of per-harness install instructions or documented support for Codex, Cursor, or other named agents — all install/marketplace instructions (plugin marketplace, /plugin install, container skill_id) are Claude-specific. Missing for 10: explicit Codex/Cursor installation docs, per-harness setup instructions, and independent confirmation that the same skill collection actually runs unmodified in non-Anthropic tools.

                • [claimed-docs] Claude Code skills follow the [Agent Skills](https://agentskills.io) open standard, which works across multiple AI tools.
                • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                Skills for Real Engineerspartialclaimed6/10

                Docs explicitly claim the collection 'works with any agent' — listing Claude Code, Cursor, Codex, Copilot — and show at least one harness-specific install path (Claude Code plugin marketplace vs. generic `npx skills` copy/subscribe modes), supporting cross-agent installability. However, evidence lacks concrete per-harness instructions for Cursor, Codex, or Copilot individually (only Claude Code's plugin route is documented in detail), and no hands-on confirmation that the same skill set actually functions identically across harnesses. Missing for 10: explicit install/config steps for Cursor, Codex, and Copilot, and independent verification that per-harness behavior matches claims.

                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
              2. developerSkills are plain markdown files and folders I can read, copy, and carry to another harness — not a proprietary binary format

                weight 2 · round to Anthropic Skills
                Anthropic Skillsfullclaimed9/10

                Docs confirm skills are just plain folders with a SKILL.md file (YAML frontmatter + markdown) plus optional scripts/references/assets, explicitly following the open Agent Skills standard (agentskills.io) that 'works across multiple AI tools,' and GitHub examples show these as ordinary files/folders anyone can read, copy, or fork. This directly supports the portability claim of non-proprietary, cross-harness markdown format. missing for 10: no independent hands-on report of someone actually carrying a skill folder to a different (non-Anthropic) harness and confirming it works unmodified.

                • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
                • [claimed-docs] Claude Code skills follow the [Agent Skills](https://agentskills.io) open standard, which works across multiple AI tools.
                • [github] Skills are simple to create - just a folder with a `SKILL.md` file containing YAML frontmatter and instructions.
                • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                Skills for Real Engineersfullclaimed8/10

                Docs explicitly state skills are written as ordinary markdown files into the repo that you own and can edit, not a proprietary format, and work across multiple harnesses (Claude Code, Cursor, Codex, Copilot). This directly matches the story's plain-file, portable-across-harness claim. Missing for 10: independent hands-on confirmation that files are literally copy-pasteable markdown (only vendor docs/ADR cited) and no explicit demonstration of moving skills to a different harness in practice.

                • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

              Discovery distribution — stories about discovery distribution in this arenaDiscovery distribution

              Stories about discovery distribution in this arena

              Discovery

              1. developerBrowse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anything

                weight 2 · round drawn
                Anthropic Skillspartialclaimed6/10

                Anthropic provides a discoverable catalog via the official `anthropics/skills` GitHub repo (browsable list of example skills) and a formal 'plugin marketplace' concept described as providing 'centralized discovery, version tracking, automatic updates' that developers can add and install from (`/plugin marketplace add`, `/plugin install`). However, there's no evidence of a dedicated searchable registry UI, ratings/leaderboard, or cross-marketplace search — discovery is limited to browsing a GitHub repo or manually adding marketplace sources one at a time. Missing for 10: a searchable/filterable registry UI, popularity or leaderboard signals, and evidence of a unified cross-marketplace search experience.

                • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                • [github] These example skills are all already available to paid plans in Claude.ai.
                • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
                Skills for Real Engineerspartialclaimed6/10

                The aihero.dev/skills page functions as a lightweight catalog listing all 25 skills with one-line descriptions (docs-11 through docs-23, docs-30/35/37, docs-42/43), and the project is also listed in Claude Code's official plugin marketplace (docs-10), letting a developer browse before installing. However there's no evidence of search, filtering, ratings, or a leaderboard-style comparison across skills/authors. Missing for 10: searchable/filterable registry UI, ratings or usage leaderboard, independent confirmation of the marketplace listing's browsability.

                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                • [claimed-docs] Install the ones you want, then type a slash command.
                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                • [claimed-docs] Find out which skill to use for the situation you are in.
                • [claimed-docs] Turn an agreed conversation into a written spec.
                • [claimed-docs] Sort raw issues into work someone can pick up.

              Distribution

              1. engineering-leadDistribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the project

                weight 2 · round drawn
                Anthropic Skillsfullclaimed8/10

                Anthropic Skills supports distribution via plugin marketplaces (git repos or local paths), shared GitHub repos (e.g., anthropics/skills registered as a marketplace), and files committed directly to a project (.claude/skills/SKILL.md), covering all three distribution channels named in the story. Community feedback confirms real-world use of shared skill repos and marketplace installs, though adoption/triggering reliability is debated. missing for 10: independent case study of an engineering team formally rolling out a standardized skill set org-wide, and more detail on permissions/governance controls for enforcing a 'standard' team-wide set.

                • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
                • [claimed-docs] A file at `.claude/commands/deploy.md` and a skill at `.claude/skills/deploy/SKILL.md` both create `/deploy` and work the same way.
                • [claimed-docs] Plugins (self-contained directories with skills, agents, hooks, or a `.claude-plugin/plugin.json` manifest) | `/plugin-name:hello` | Sharing…
                Skills for Real Engineersfullclaimed8/10

                The product ships via Claude Code's official plugin marketplace (docs-10), as an npx-installed, updatable managed bundle (docs-1, docs-4/31, docs-41), or as editable files committed directly into the project repo (docs-2, docs-24), and is agent-agnostic/MIT-licensed so a lead can standardize it across a whole team's tools (docs-42). This covers all three named distribution paths (marketplace, subscribe/shared-source, and in-repo files). Missing for 10: no explicit 'shared git repo' distribution mode distinct from marketplace/npx, and no independent/community confirmation that team-wide rollout works smoothly in practice beyond first-party docs.

                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT

              Triggering

              1. developerInstalled skills trigger automatically from task context, with descriptions engineered so the agent activates the right skill at the right moment

                weight 3 · round to Anthropic Skills
                Anthropic Skillsdisputedcontradicted5/10

                Anthropic's docs explicitly describe automatic activation via engineered description fields (e.g. docs-17, docs-25, docs-32, docs-44) and even ship a skill-creator tool to optimize descriptions for 'triggering accuracy' (docs-14). However, hands-on community reports directly contradict reliable auto-triggering: Vercel's eval found the skill was never invoked in 56% of cases with no improvement over baseline, and multiple users report skills are 'hit or miss' or 'incredibly hard' to get invoked even with proper YAML descriptions (comm-7, comm-8, comm-9). missing for 10: independent benchmark showing consistent correct auto-activation, and resolution of the documented reliability gap.

                • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                • [claimed-docs] If Claude thinks the skill is relevant to the current task, it will load the skill by reading its full `SKILL.md` into context.
                • [claimed-docs] benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy
                • [claimed-docs] description: Extracts text and tables from PDF files, fills PDF forms, and merges multiple PDFs. Use when working with PDF documents or when…
                • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
                • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
                • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
                Skills for Real Engineersnone0/10

                The evidence describes manual invocation — 'Install the ones you want, then type a slash command' (docs-43) and 'Use them every time you want to make a change' (gh-1/gh-4) — rather than automatic, context-triggered activation via engineered descriptions. There is a 'find out which skill to use' meta-skill (docs-23) but it's itself a skill you must invoke, not evidence of automatic description-based dispatch. No documentation or independent evidence shows the agent autonomously selecting/activating skills from task context.

                • [claimed-docs] Install the ones you want, then type a slash command.
                • [claimed-docs] Find out which skill to use for the situation you are in.
                • [github] They help you align with the agent before you get started, and think deeply about the change you're making. Use them _every_ time you want t…
                • [github] They help you align with the agent before you get started, and think deeply about the change you're making. Use them every time you want to …

              Docs onboarding — stories about docs onboarding in this arenaDocs onboarding

              Stories about docs onboarding in this arena

              Onboarding

              1. developerA quickstart takes me from nothing to a working installed skill in under five minutes

                weight 3 · round drawn
                Anthropic Skillspartialcommunity6/10

                Docs and GitHub show a genuinely lightweight path — a skill is 'just a folder with a SKILL.md file containing YAML frontmatter and instructions' (anthropic-skills-gh-3), installable via a single 'plugin marketplace add' + 'plugin install' command (anthropic-skills-gh-1, anthropic-skills-gh-4), which plausibly fits a five-minute window. However there's no dedicated timed 'quickstart' tutorial artifact, and community feedback shows friction getting Claude to actually invoke/use a newly installed skill reliably (anthropic-skills-comm-8, anthropic-skills-comm-9), meaning 'installed and working' isn't fully guaranteed in five minutes. Missing for 10: an explicit timed quickstart doc/tutorial, and independent hands-on confirmation of sub-5-minute install-to-working success.

                • [github] Skills are simple to create - just a folder with a `SKILL.md` file containing YAML frontmatter and instructions.
                • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
                • [claimed-docs] Create a `SKILL.md` file with instructions, and Claude adds it to its toolkit.
                • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
                • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
                Skills for Real Engineerspartialclaimed6/10

                Docs show a simple two-step install path ("Install the ones you want, then type a slash command" and `claude plugins install mattpocock-skills`), plus npx-based add/update commands, suggesting a fast setup, but there is no explicit quickstart walkthrough or timed benchmark confirming a five-minute install-to-working-skill experience, and no independent hands-on report timing the process. missing for 10: an actual quickstart guide/tutorial with step timings, independent/hands-on confirmation of install speed, and evidence that a first skill run succeeds quickly without extra config.

                • [claimed-docs] Install the ones you want, then type a slash command.
                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
              2. developerEvery skill documents what it does and when it activates, so I can predict my agent's new behavior before it surprises me

                weight 2 · round to Skills for Real Engineers
                Anthropic Skillsdisputedcontradicted5/10

                Anthropic's spec strongly documents that every SKILL.md must include a description of what the skill does and when to use it (frontmatter is designed to be both machine- and human-readable for predicting activation), and Claude Code/Claude API docs describe controllable invocation triggers. However, multiple hands-on community reports (Vercel eval: skill never invoked in 56% of cases despite being documented and available; users reporting 'hit or miss' triggering even with well-written descriptions) concretely contradict the promise that documented descriptions reliably predict when the agent will actually activate a skill. missing for 10: consistent, benchmarked activation reliability matching documented triggers, and independent evidence resolving the invocation unpredictability reported by users.

                • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                • [claimed-docs] name: skill-name description: A description of what this skill does and when to use it.
                • [claimed-docs] description: Extracts text and tables from PDF files, fills PDF forms, and merges multiple PDFs. Use when working with PDF documents or when…
                • [claimed-docs] Skills add optional features: a directory for supporting files, frontmatter to [control whether you or Claude invokes them]... and the abili…
                • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
                • [community] I have an incredibly hard time getting them to use Skills at all, even when asked. I saw someone's analysis finding their agents were more a…
                • [community] Same, I have a bunch of skills defined with proper YAML headers and semantic triggers... it's hit or miss if it picks up on the skill -- usu…
                Skills for Real Engineersfullcommunity7/10

                Each skill ships a SKILL.md/one-line description stating its purpose and trigger (e.g., TDD skill's red-green rules, triage skill's state machine, grilling skill's interview process, and the aihero.dev list of 25 skills each with a one-line 'what it does' summary), and activation is explicit and user-controlled via install + slash command rather than silent background changes ('Nothing updates behind your back', 'Install the ones you want, then type a slash command'). A community thread does critique the prose quality/clarity of some SKILL.md files as having 'little utility,' which tempers confidence but doesn't concretely contradict that each skill documents what/when it activates. missing for 10: independent verification that every one of the 25 skills' docs clearly states activation triggers (not just a sample), and resolution of the community critique about jargon-heavy or low-utility prose in some skill docs.

                • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
                • [claimed-docs] Test only at pre-agreed seams. Before writing any test, write down the seams under test and confirm them with the user.
                • [claimed-docs] Red before green. Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
                • [claimed-docs] Move issues on the project issue tracker through a small state machine of triage roles.
                • [claimed-docs] Move issues on the project issue tracker through a small state machine of triage roles, categorise, verify, grill if needed, and write agent…
                • [claimed-docs] Turn an agreed conversation into a written spec.
                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                • [claimed-docs] Install the ones you want, then type a slash command.
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [community] Critique of Matt Pocock's skill files: he has a hand-wavy approach to engineering in a silo, then uses that output as a foundation for tools…

              Install experience — stories about install experience in this arenaInstall experience

              Stories about install experience in this arena

              Install

              1. developerInstall a skill collection with one documented command — a package-manager one-liner, CLI, or in-agent marketplace command — and it is active in my next session

                weight 3 · round drawn
                Anthropic Skillspartialclaimed6/10

                Anthropic documents a clear CLI install path — `/plugin marketplace add anthropics/skills` followed by `/plugin install document-skills@anthropic-agent-skills` — which registers and installs a skill collection inside Claude Code, and docs imply skills become part of Claude's toolkit thereafter. However, this is a two-step sequence rather than a single one-liner, and there is no independent/hands-on confirmation that the installed skill persists and is reliably active in the very next session (community reports focus on activation/triggering reliability, not install itself). Missing for 10: a true single-command one-liner, and first-party or community confirmation of session-persistence after install.

                • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
                • [github] /plugin marketplace add anthropics/skills
                • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                Skills for Real Engineerspartialclaimed6/10

                First-party docs give a genuine one-liner (`claude plugins install mattpocock-skills`) now that the pack is in Claude Code's official marketplace, plus an `npx skills update` command for the managed-bundle mode, both documented as first-party. However there's no independent/hands-on confirmation that the skill set is actually active in the next session, and the exact single-command install syntax for the other supported agents (Cursor, Codex, Copilot) beyond Claude Code isn't shown — only 'install the ones you want, then type a slash command' is vague. missing for 10: independent hands-on confirmation of post-install activation, explicit one-liner install commands for non-Claude agents.

                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                • [claimed-docs] Install the ones you want, then type a slash command.
              2. developerChoose install scope — project-local files committed with my repo, or user-global across all projects

                weight 2 · round drawn
                Anthropic Skillspartialclaimed4/10

                Docs show skills live in project-level directories like `.claude/skills/deploy/SKILL.md` and can be packaged into plugins shareable 'across projects and teams,' implying some notion of local vs shared scope, but there is no explicit documentation of a user-global (e.g. home-directory) skill install path or a direct project-vs-user scope toggle. missing for 10: explicit documentation of a user-global skill directory/location, explicit contrast between project-committed vs user-global install scope, and confirmation that both scopes are simultaneously supported and selectable by the developer.

                • [claimed-docs] A file at `.claude/commands/deploy.md` and a skill at `.claude/skills/deploy/SKILL.md` both create `/deploy` and work the same way.
                • [claimed-docs] Plugins let you extend Claude Code with custom functionality that can be shared across projects and teams. This guide covers creating your o…
                • [claimed-docs] Plugins let you extend Claude Code with custom functionality that can be shared across projects and teams.
                • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                Skills for Real Engineerspartialclaimed4/10

                The docs describe two install modes—copying editable skill files into a repo (project-local, committed) versus subscribing to a managed, read-only, auto-updating bundle via the Claude Code plugin marketplace—but never explicitly frame the second mode as 'user-global across all projects' vs project-local; scope (per-project vs per-user) is never directly addressed. Missing for 10: explicit documentation of a user-global/home-directory install option and confirmation that the plugin/subscribe mode applies across all projects rather than just being non-editable.

                • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own
                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
              3. developerInstall only the specific skills I want from a collection instead of taking the whole bundle

                weight 2 · round to Anthropic Skills
                Anthropic Skillspartialclaimed6/10

                Docs show marketplace-based installs where a user can add a marketplace and then install just one named plugin (e.g. `/plugin install document-skills@anthropic-agent-skills` vs `example-skills@anthropic-agent-skills`, or `quality-review-plugin@my-plugins`) rather than every plugin in the marketplace, and `disableBundledSkills` can turn off built-in skills en masse. This demonstrates selecting a specific plugin/skill bundle out of a larger marketplace, but there's no documented mechanism to cherry-pick an individual skill from inside a single plugin bundle. Missing for 10: explicit per-skill (sub-bundle) install granularity, and independent confirmation that partial installs work as described.

                • [github] /plugin install document-skills@anthropic-agent-skills /plugin install example-skills@anthropic-agent-skills
                • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] the [`disableBundledSkills`](/docs/en/settings-reference#disablebundledskills) setting, which disables every bundled skill except `/doctor`.
                • [claimed-docs] To turn bundled skills off, use the [`disableBundledSkills`](/docs/en/settings-reference#disablebundledskills) setting, which disables every…
                Skills for Real Engineerspartialclaimed5/10

                The docs explicitly state "Install the ones you want, then type a slash command" (mattpocock-skills-docs-43), indicating selective installation is possible, and the file-copy mode writes only the skills you choose as editable files (mattpocock-skills-docs-24). However, the marketplace/plugin route is described as installing 'the whole set as a managed, read-only bundle' (mattpocock-skills-docs-1, docs-41), so whole-bundle install remains the primary documented path and there's no detailed CLI flag or command example showing per-skill selection. Missing for 10: a concrete CLI command/flag demonstrating selecting individual skills, and independent/hands-on confirmation that partial installs work as described.

                • [claimed-docs] Install the ones you want, then type a slash command.
                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

              Lifecycle

              1. developerList what is installed and remove skills cleanly, without orphaned files or lingering instructions

                weight 1 · round drawn
                Anthropic Skillsnone0/10

                The evidence pack shows how to create, add, invoke, and disable skills (e.g., disableBundledSkills, /plugin install), but there is no documentation of a command or workflow to list all installed skills or to cleanly uninstall/remove a skill and its plugin files, hooks, or references without leftovers.

                • [claimed-docs] the [`disableBundledSkills`](/docs/en/settings-reference#disablebundledskills) setting, which disables every bundled skill except `/doctor`.
                • [claimed-docs] To turn bundled skills off, use the [`disableBundledSkills`](/docs/en/settings-reference#disablebundledskills) setting, which disables every…
                • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                • [claimed-docs] A file at `.claude/commands/deploy.md` and a skill at `.claude/skills/deploy/SKILL.md` both create `/deploy` and work the same way.
                Skills for Real Engineersnone0/10

                Evidence covers installation modes (editable copy vs. read-only subscribed bundle) and updating via `npx skills update`, but there is no mention of any command or mechanism to list installed skills or cleanly uninstall/remove them without leftover files.

                Openness — open source, data portability, and self-hosting storiesOpenness

                Open source, data portability, and self-hosting stories

                1. ai-native userDo everything through the API that I can do in the UI

                  weight 2 · round to Anthropic Skills
                  Anthropic Skillspartialclaimed5/10

                  Docs confirm Skills can be invoked and managed programmatically via the Messages API (container parameter, skill_id/type/version, up to 20 skills per request) and via a dedicated Skills API for upload/management, giving real API-level parity for core skill usage. However, other capabilities visible in the Claude Code/Claude.ai UI — plugin marketplaces, automatic relevance-based skill loading, /plugin and /skill-name invocation, bundled-skill toggling — are documented only as CLI/UI features with no evidence of an equivalent API path. Missing for 10: API equivalents for plugin marketplace distribution, automatic skill discovery/loading parity, and confirmation that all UI-configurable settings (e.g., disableBundledSkills) are reachable via API.

                  • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
                  • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
                  • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
                  • [claimed-docs] You can include up to 20 Skills for each request.
                  • [claimed-docs] Upload and manage through the [Skills API]
                  • [claimed-docs] Availability | Available to all users | Private to your workspace
                  • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                  • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                  Skills for Real Engineersnone0/10

                  The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                  • ai-native userExport all of my data in open formats and leave

                    weight 3 · round to Skills for Real Engineers
                    Anthropic Skillsnone0/10

                    The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                      Skills for Real Engineerspartialclaimed5/10

                      The product ships skills as plain, MIT-licensed markdown files copied directly into the user's repo (mattpocock-skills-docs-2, docs-24, docs-42), meaning the artifacts themselves are already open, human-readable, and fully owned/editable with no proprietary lock-in or vendor updates without consent (docs-4, docs-31). However, there is no explicit 'export my data' feature or documentation addressing exporting configuration state (issue tracker settings, triage labels, ADRs) generated while using the skills, nor any statement about a formal data-portability/leave process. missing for 10: explicit data-export tooling/documentation, evidence about exporting generated artifacts (ADRs, triage state, configs) beyond the skill files themselves, and any independent confirmation of portability.

                      • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                      • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                      • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                      • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                      • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                    • ai-native userRead the product's source under an open license

                      weight 2 · round to Skills for Real Engineers
                      Anthropic Skillsnone0/10

                      While the anthropics/skills GitHub repo makes example skill files (SKILL.md, templates) publicly readable, none of the evidence cites an open-source license for Skills, Claude Code, or the underlying product; core Claude Code/Skills functionality itself is closed, proprietary tooling with no license grant shown.

                        Skills for Real Engineersfullclaimed8/10

                        The product's source lives in a public GitHub repo (github.com/mattpocock/skills) and is explicitly described as MIT-licensed with 25 skills, confirming both open-source hosting and license terms; docs also emphasize files are 'ordinary files you own and can edit,' reinforcing readability/openness of the source. Missing for 10: no explicit LICENSE file citation or independent third-party confirmation of the license text.

                        • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                        • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                        • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own

                      Privacy posture — data-handling and privacy storiesPrivacy posture

                      Data-handling and privacy stories

                      1. ai-native userControl data retention and deletion

                        weight 2 · round drawn
                        Anthropic Skillsnone0/10

                        The evidence pack covers how Skills are created, invoked, packaged, and distributed, but contains no documentation about data retention policies, deletion controls, or privacy settings for skill data/usage. No mention of retention windows, user-initiated deletion, or data handling controls exists in this pack.

                          Skills for Real Engineersnone0/10

                          The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na". (na/none harmonized at arena bring-up — see pipeline/scripts/na-harmonize.ts.)

                          Safety review — stories about safety review in this arenaSafety review

                          Stories about safety review in this arena

                          Review

                          1. engineering-leadReview exactly what instructions and scripts a skill will add — list contents before installing and read every file afterward

                            weight 3 · round drawn
                            Anthropic Skillspartialclaimed6/10

                            Docs confirm skills are transparent, plain-text directories (SKILL.md plus optional scripts/references/assets folders) that a lead can browse in a repo/marketplace and read after installation, and plugins/marketplaces are just git repos or local paths a lead could inspect. However, there is no documented feature for listing or previewing a skill's full file contents specifically before installation (e.g., a dry-run/manifest-diff command), so the 'before install' half of the story is only implicit via manual repo browsing. missing for 10: a dedicated pre-install content-listing/manifest command, any audit/review tooling, and independent confirmation that installs can't run additional undisclosed files.

                            • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
                            • [claimed-docs] skills can bundle additional files within the skill directory and reference them by name from `SKILL.md`.
                            • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                            • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                            • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                            • [claimed-docs] add a skill, and test it locally using the \`--plugin-dir\` flag
                            Skills for Real Engineerspartialclaimed6/10

                            The skills ship as plain, open-source Markdown SKILL.md files (visible directly via raw GitHub links quoted in evidence) and are described as 'ordinary files you own and can edit' with no silent updates, which supports post-install readability and transparency. However, there is no documented pre-install 'list contents' or dry-run command shown in the evidence — inspection relies on browsing the public GitHub repo rather than a built-in review step. Missing for 10: an explicit CLI/list command to preview a skill's files before installing, and independent confirmation that all installed files match what's shown pre-install.

                            • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                            • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                            • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                            • [claimed-docs] Test only at pre-agreed seams. Before writing any test, write down the seams under test and confirm them with the user.
                            • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
                            • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

                          Trust

                          1. engineering-leadThe project documents its security posture — what skills can execute, the trust model for third-party skills, and any telemetry or data collection

                            weight 2 · round to Skills for Real Engineers
                            Anthropic Skillsnone0/10

                            The evidence pack shows extensive functional docs (SKILL.md structure, plugin marketplaces, execution environment, disableBundledSkills toggle) but no dedicated security-posture documentation addressing what skills can execute (sandboxing, permissions), a trust model for vetting third-party/marketplace skills, or telemetry/data-collection disclosures tied to skills usage.

                              Skills for Real Engineerspartialclaimed4/10

                              The docs give some trust-relevant transparency—skills are shipped as plain, ownable files that 'nothing updates behind your back' and can be pulled explicitly via `npx skills update`, plus acceptance into Claude Code's official marketplace as a vetted distribution channel—but there is no explicit security-posture document covering what skills can execute (tool/permission scope), a formal trust model for arbitrary third-party skills, or any statement on telemetry/data collection. missing for 10: explicit execution/permission model for skills, documented telemetry or data-collection policy, formal third-party skill vetting/trust framework beyond marketplace acceptance.

                              • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                              • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                              • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                              • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                              • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

                            Skill authoring — stories about skill authoring in this arenaSkill authoring

                            Stories about skill authoring in this arena

                            Authoring

                            1. developerAuthor a new skill from a documented template — a SKILL.md with name and description frontmatter — without reverse-engineering existing skills

                              weight 3 · round to Anthropic Skills
                              Anthropic Skillsfullcommunity8/10

                              There is a documented open specification (agentskills.io) detailing required frontmatter (name, description, length constraints), file structure (SKILL.md plus optional scripts/references/assets folders), and even a template file (template/SKILL.md) with placeholder text, plus a dedicated skill-creator skill for authoring/improving skills without needing to reverse-engineer existing ones. Community evidence (comm-1) shows some users still had to ask others for example SKILL.md files, suggesting the template isn't universally discovered/used. missing for 10: independent hands-on confirmation that a developer successfully authored a skill purely from the template without consulting other examples, and more visibility/discoverability of the template in mainline docs.

                              • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                              • [claimed-docs] Replace with description of the skill and when Claude should use it.
                              • [claimed-docs] name: skill-name description: A description of what this skill does and when to use it.
                              • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                              • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
                              • [claimed-docs] A skill for creating new skills and iteratively improving them.
                              • [claimed-docs] Create new skills, modify and improve existing skills, and measure skill performance.
                              • [community] Github MCP: 39 tools, 30K tokens - had to disable it. Asked if anyone has a good SKILLS.md file to study.
                              Skills for Real Engineersnone0/10

                              The evidence describes an existing bundle of pre-built skills (tdd, grilling, triage, setup) that you can install, subscribe to, or copy as editable files, but there is no documented template, generator, or guide specifically for authoring a brand-new SKILL.md with frontmatter — the closest thing to a 'template' would be reverse-engineering the shipped example skills, which the story explicitly excludes. Community commentary even critiques the prose quality of existing skill files rather than pointing to any authoring template.

                              • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                              • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                              • [claimed-docs] Scaffold the per-repo configuration that the engineering skills assume: **Issue tracker**... **Triage labels**... **Domain docs**
                              • [community] Critique of Matt Pocock's skill files: he has a hand-wavy approach to engineering in a silo, then uses that output as a foundation for tools…
                            2. developerThe collection ships a meta-skill or tool that guides my agent through writing, improving, and packaging new skills

                              weight 2 · round to Anthropic Skills
                              Anthropic Skillsfullclaimed8/10

                              Anthropic ships a dedicated "skill-creator" meta-skill that explicitly guides creation, iteration, and improvement of skills, including benchmarking performance and optimizing description triggers for accuracy, and there's a template SKILL.md and open spec to follow. This directly matches the story of a meta-skill guiding authoring/improving/packaging skills, backed by first-party GitHub and docs evidence. Missing for 10: independent hands-on validation of the skill-creator workflow itself (community evidence only discusses skills generally, not this meta-skill specifically) and no evidence of a dedicated 'packaging for distribution' step within skill-creator beyond plugin/marketplace mechanisms.

                              • [claimed-docs] A skill for creating new skills and iteratively improving them.
                              • [claimed-docs] benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy
                              • [claimed-docs] Create new skills, modify and improve existing skills, and measure skill performance.
                              • [claimed-docs] Replace with description of the skill and when Claude should use it.
                              • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                              • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                              Skills for Real Engineersnone0/10

                              The evidence describes many individual skills (TDD, triage, grilling, spec-writing, setup-per-repo config) but none of them describe a meta-skill that guides writing, improving, or packaging new skills for the collection itself — 'setup-matt-pocock-skills' only scaffolds per-repo config for existing skills, not skill authoring.

                              Spec

                              1. developerSkills follow the open Agent Skills specification so the same skill folder is valid beyond this one vendor's tooling

                                weight 2 · round to Anthropic Skills
                                Anthropic Skillsfullclaimed8/10

                                Anthropic explicitly states Claude Code skills follow the open Agent Skills standard (agentskills.io) that works across multiple AI tools, and the same SKILL.md folder format (frontmatter + markdown, optional scripts/references/assets dirs) is documented both on the vendor docs and the independent agentskills.io spec site. Missing for 10: no third-party/independent tool (outside Anthropic) is shown actually consuming the same folder, and community commentary questions skill invocation reliability rather than spec portability.

                                • [claimed-docs] Claude Code skills follow the [Agent Skills](https://agentskills.io) open standard, which works across multiple AI tools.
                                • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                                • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                                • [claimed-docs] scripts/ # Optional: executable code references/ # Optional: documentation assets/ # Optional: templates, resources
                                • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                                • [claimed-docs] name: skill-name description: A description of what this skill does and when to use it.
                                Skills for Real Engineerspartialclaimed6/10

                                Evidence shows skill folders are plain, editable files (SKILL.md) that work across multiple agents (Claude Code, Cursor, Codex, Copilot) and were accepted into Claude Code's official plugin marketplace, implying broad cross-tool portability consistent with an open skill format. However, none of the evidence explicitly names or cites conformance to the 'Agent Skills specification' itself, so spec-adherence is inferred rather than documented. Missing for 10: explicit reference to the Agent Skills spec, independent confirmation that the folder validates against that spec outside vendor claims.

                                • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                                • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                                • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…

                              Testing quality — stories about testing quality in this arenaTesting quality

                              Stories about testing quality in this arena

                              Maintenance

                              1. developerThe collection is actively maintained — recent releases, triaged issues, and accepted community contributions

                                weight 2 · round to Skills for Real Engineers
                                Anthropic Skillsnone0/10

                                Evidence pack covers docs, specification, and usage patterns for Skills, but contains no information about release cadence, issue triage, or acceptance of community contributions to the anthropics/skills repository. No changelog, release notes, contributor stats, or issue-response evidence is present.

                                  Skills for Real Engineerspartialcommunity3/10

                                  There's some indirect signal of active development — an ADR documenting a recent architectural decision (shipping as a Claude Code plugin) and acceptance into Claude Code's official marketplace, plus an update mechanism (`npx skills update`) implying ongoing releases — but no direct evidence of a release cadence, an issue tracker for the project itself being triaged, or accepted external community contributions/PRs. Community commentary (comm-1/2/3) shows engagement and critique of content quality but says nothing about maintenance cadence or contribution acceptance. missing for 10: changelog/release history, evidence of external PRs being merged, evidence of issues on the repo itself being triaged.

                                  • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                                  • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own
                                  • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                                  • [community] Matt Pocock is still a nice guy with reasonable opinions and shares a lot with us. I've personally learned from reading his skills.
                                  • [community] Critique of Matt Pocock's skill files: he has a hand-wavy approach to engineering in a silo, then uses that output as a foundation for tools…

                                Testing

                                1. developerThe collection maintains tests or evals for its skills so changes are verified against regressions rather than shipped on vibes

                                  weight 2 · round drawn
                                  Anthropic Skillsnone0/10

                                  The skill-creator skill offers ad-hoc 'benchmark skill performance' tooling for authors (anthropic-skills-docs-14, -30), but there is no evidence of a maintained test suite, CI pipeline, or regression eval framework for the official skills collection itself. The only concrete eval-style evidence (Vercel's finding that skills were never invoked in 56% of cases) is a third-party community critique, not Anthropic's own maintained regression testing.

                                  • [claimed-docs] benchmark skill performance with variance analysis, or optimize a skill's description for better triggering accuracy
                                  • [claimed-docs] Create new skills, modify and improve existing skills, and measure skill performance.
                                  • [community] Vercel found: In 56% of eval cases, the skill was never invoked. The agent had access to the documentation but didn't use it. Adding the ski…
                                  Skills for Real Engineersnone0/10

                                  The evidence describes TDD/testing as a discipline the skills teach users to apply to their own code, but nothing shows the maintainer running automated tests, evals, or CI against the skill prompts themselves to catch regressions. A community critique even suggests the SKILL.md prose lacks rigor and 'should be checked... by another LLM,' implying no such verification pipeline exists.

                                  • [claimed-docs] Red before green. Write the failing test first, then only enough code to pass it. Don't anticipate future tests or add speculative features.
                                  • [claimed-docs] Tests verify behavior through public interfaces, not implementation details.
                                  • [community] Critique of Matt Pocock's skill files: he has a hand-wavy approach to engineering in a silo, then uses that output as a foundation for tools…

                                Versioning updates — stories about versioning updates in this arenaVersioning updates

                                Stories about versioning updates in this arena

                                Pinning

                                1. engineering-leadControl when skill changes reach my team — pinned versions or a lockfile rather than silent behind-the-back updates

                                  weight 1 · round to Skills for Real Engineers
                                  Anthropic Skillspartialclaimed5/10

                                  The Skills API explicitly supports pinning a specific version via the `container` parameter's optional `version` field when invoking a skill, and plugin manifests include a version field with marketplaces offering 'version tracking' (docs-10, docs-23, docs-27, docs-36). However, for Claude Code's locally-installed/bundled skills there is no lockfile or team-wide pinning mechanism described — marketplaces are instead touted for 'automatic updates', which is the opposite of controlled rollout, and no evidence shows a way to pin or freeze skill versions across a team's Claude Code installs. missing for 10: lockfile or pinned-version mechanism for Claude Code skill/plugin installs, evidence that automatic marketplace updates can be disabled or gated per-team, and any hands-on confirmation that version pinning works as documented.

                                  • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
                                  • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
                                  • [claimed-docs] The manifest file at `.claude-plugin/plugin.json` defines your plugin's identity: its name, description, and version.
                                  • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                                  Skills for Real Engineerspartialclaimed7/10

                                  Docs explicitly state skills don't auto-update and changes only land when the user runs `npx skills update`, and offer a 'copy' mode where files become editable local copies you fully own — both give an engineering lead control over rollout timing. However, there is no mention of an actual lockfile, pinned semantic versions, or per-team version pinning mechanism, so the control is manual/all-or-nothing rather than granular version pinning. Missing for 10: explicit lockfile/version-pin mechanism, ability to pin to a specific historical version rather than just delaying `update`, independent confirmation of update-control behavior in practice.

                                  • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                                  • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                                  • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                                  • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                                  • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                                  • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own

                                Updates

                                1. developerThere is a documented update path — marketplace auto-updates or an explicit update command — so I get fixes without reinstalling from scratch

                                  weight 3 · round to Skills for Real Engineers
                                  Anthropic Skillspartialclaimed6/10

                                  Docs explicitly state plugin marketplaces provide 'centralized discovery, version tracking, automatic updates' for plugins that can bundle skills, and marketplace/plugin install commands are documented (e.g. /plugin marketplace add, /plugin install). However, this update path is scoped to the plugin-marketplace distribution mechanism rather than a general update command for individually-created or hand-copied skills (e.g. skills dropped into .claude/skills/), and there's no explicit 'update' CLI verb or independent confirmation of update behavior in practice. missing for 10: an explicit 'update' command example/output, confirmation this works for non-marketplace skills, independent/hands-on evidence of auto-update actually firing.

                                  • [claimed-docs] A **plugin marketplace** is a catalog that lets you distribute plugins to others. Marketplaces provide centralized discovery, version tracki…
                                  • [claimed-docs] Marketplaces provide centralized discovery, version tracking, automatic updates, and support for multiple source types, including git reposi…
                                  • [claimed-docs] /plugin marketplace add ./my-marketplace /plugin install quality-review-plugin@my-plugins
                                  • [github] You can register this repository as a Claude Code Plugin marketplace by running the following command in Claude Code: /plugin marketplace a…
                                  • [claimed-docs] A plugin marketplace is a catalog that lets you distribute plugins to others.
                                  Skills for Real Engineersfullclaimed8/10

                                  Docs explicitly document an update path: `npx skills update` pulls latest changes on demand (nothing auto-updates behind your back) for the copy-into-repo mode, and a separate managed/marketplace mode (Claude Code plugin, `claude plugins install mattpocock-skills`) that updates as the author ships. Both paths are documented first-party. missing for 10: independent/hands-on confirmation that `npx skills update` or marketplace auto-update actually works in practice, and no changelog/version-diff evidence showing successful update history.

                                  • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                                  • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                                  • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                                  • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                                  • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own
                                  • [claimed-docs] installs the whole set as a managed, read-only bundle that updates when I ship, so you subscribe rather than fork
                                2. developerReleases ship with notes or a changelog so I can see what changed in the skills before I take an update

                                  weight 1 · round drawn
                                  Anthropic Skillsnone0/10

                                  Evidence shows plugin marketplaces support 'version tracking' and skills/plugins have version fields in manifests, but there is no evidence of actual release notes or a changelog documenting what changed between skill versions before update.

                                    Skills for Real Engineersnone0/10

                                    Evidence describes an update mechanism (`npx skills update`, subscribing to a read-only bundle) but nothing about release notes, a changelog, or version history that would let a developer see what changed before updating.

                                    Not comparable on these axes

                                    1. ai-native userPlug MCP servers into this product so it can use their tools

                                      weight 3 · not comparable
                                      Anthropic Skillspartialcommunity5/10

                                      Docs show that Claude Code plugins — the packaging mechanism used alongside Skills — can bundle MCP servers together with skills, agents, and hooks (anthropic-skills-docs-6, -7, -21, -35), implying an ai-native user could add an MCP server via the plugin/marketplace system. However, Skills themselves are a separate mechanism (plain SKILL.md instructions), and community commentary explicitly notes skills and MCP are distinct, sometimes competing approaches with skills lacking MCP's tool-calling functionality (anthropic-skills-comm-3, -10). Missing for 10: concrete first-party guide/example of installing an MCP server via a skill or plugin, hands-on confirmation that MCP tools become usable once added this way, and clarity on whether Skills (as opposed to Claude Code plugins broadly) directly expose MCP tool use.

                                      • [claimed-docs] Create custom plugins to extend Claude Code with skills, agents, hooks, and MCP servers.
                                      • [claimed-docs] Plugins (self-contained directories with skills, agents, hooks, or a `.claude-plugin/plugin.json` manifest) | `/plugin-name:hello` | Sharing…
                                      • [claimed-docs] Plugins let you extend Claude Code with custom functionality that can be shared across projects and teams. This guide covers creating your o…
                                      • [claimed-docs] Plugins let you extend Claude Code with custom functionality that can be shared across projects and teams.
                                      • [community] A skill is a markdown & yaml file on your filesystem... but skills don't cover a lot of the functionality that MCP offers.
                                      • [community] If you can write a bash or python script, or an API or MCP to do what you want, then write it and include it in the skill. Keep top-level co…
                                      Skills for Real Engineersn/a

                                      This product is a library of skill/prompt files installed into external coding agents (Claude Code, Cursor, Codex, Copilot); it is not itself an agent or platform that consumes or hosts MCP servers, so plugging MCP servers into it is a category error — that capability belongs to the host agents, not to this skills package.

                                      • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                                      • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                                    2. ai-native userConnect an agent via an official MCP server

                                      weight 3 · not comparable
                                      Anthropic Skillsn/a

                                      Anthropic Skills is a packaging/format for capabilities (SKILL.md files, plugins, marketplaces) consumed by Claude/Claude Code, not a service or agent that itself exposes an official MCP server endpoint for other agents to connect to. Evidence mentions plugins can *bundle* MCP servers as one of several extension types, but this is about extending Claude Code, not Skills serving as an MCP server itself — the axis is a category mismatch for this product type.

                                        Skills for Real Engineersnone0/10

                                        Evidence shows distribution via npm CLI, Claude Code plugin marketplace, and file-copy installation, but nowhere does it mention an MCP server for agents to connect to. Since this is a skills/plugin package (not itself an agent), the axis is a fair question, but there is no evidence of an official MCP server offering.

                                        • [claimed-docs] `mattpocock-skills` was accepted into **Claude Code's official marketplace**... `claude plugins install mattpocock-skills` is now the docume…
                                        • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own
                                        • [claimed-docs] Works with any agent Claude Code · Cursor · Codex · Copilot · 25 skills · MIT
                                        • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`.
                                      • ai-native userDrive the product through a documented public API

                                        weight 3 · not comparable
                                        Anthropic Skillsfullclaimed8/10

                                        Skills can be driven via the documented Messages API `container` parameter with `skill_id`, uploaded/managed through a dedicated Skills API, and invoked with up to 20 skills per request, giving AI-native users a clear programmatic path distinct from the chat UI. missing for 10: independent/hands-on corroboration of the API workflow itself (community evidence only covers Claude Code skill-triggering reliability, not the Messages/Skills API), and no evidence of API rate limits, SDKs, or error handling specifics.

                                        • [claimed-docs] You specify Skills in the `container` parameter with a `skill_id`, `type`, and optional `version`, and they run in the code execution enviro…
                                        • [claimed-docs] Skills are specified using the `container` parameter in the Messages API. You can include up to 20 Skills for each request.
                                        • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
                                        • [claimed-docs] You can include up to 20 Skills for each request.
                                        • [claimed-docs] Upload and manage through the [Skills API]
                                        • [claimed-docs] This guide shows you how to use both pre-built and custom Skills with the Claude API.
                                        Skills for Real Engineersn/a

                                        This product is a static collection of skill/markdown files distributed via CLI installers (npx skills, claude plugins) and consumed inside a host agent's context — it is not a service or platform that exposes its own public API for programmatic control. The 'driven through a documented public API' axis is a category error for a skills-file bundle rather than an applicable-but-unmet capability.

                                        • ai-native userIssue scoped/least-privilege API credentials for an agent

                                          weight 2 · not comparable
                                          Anthropic Skillsn/a

                                          Anthropic Skills is about packaging instructions/resources for Claude to load dynamically, not about credential/permission scoping or API key issuance for agents; no evidence pack content addresses scoped credential issuance.

                                            Skills for Real Engineersn/a

                                            This product is a collection of AI agent 'skills'/prompt workflows for coding tasks, not an identity/access-management or credentialing system; issuing scoped API credentials is entirely outside its category.

                                            • ai-native userSubscribe to events via webhooks

                                              weight 2 · not comparable
                                              Anthropic Skillsn/a

                                              Anthropic Skills is a mechanism for packaging instructions/scripts that agents load into context, not a service with an event system; webhook subscription is a wrong axis for this product type and no evidence suggests otherwise.

                                                Skills for Real Engineersn/a

                                                This product is a skills/prompt library for AI coding agents, not an event-driven platform or service with a webhook subscription mechanism; nothing in the evidence pertains to webhooks or event subscriptions, and the concept doesn't fit this product's category.

                                                • ai-native userExplore an interactive API reference with runnable examples

                                                  weight 2 · not comparable
                                                  Anthropic Skillsn/a

                                                  Anthropic Skills is a mechanism for packaging agent capabilities (SKILL.md files, plugins), not an API-reference product with an interactive documentation explorer; 'runnable examples in an API reference' is a category mismatch for this product type.

                                                    Skills for Real Engineersn/a

                                                    Skills for Real Engineers is a set of agent skill files/prompts, not an API product with an interactive reference — this axis is a category error for this product type.

                                                    • ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)

                                                      weight 2 · not comparable
                                                      Anthropic Skillspartialclaimed3/10

                                                      The evidence shows an open, machine-readable specification document (agentskills.io/specification.md) defining the SKILL.md format and frontmatter fields, plus references to a 'Skills API' for programmatic upload/management — these are the closest analogs to a downloadable machine-readable spec, but neither is an OpenAPI document nor explicitly offered as a downloadable API spec for the Skills/Messages API itself. Missing for 10: an actual OpenAPI/JSON-schema file for the Skills API endpoints, explicit download link/format, and independent confirmation that AI-native tooling consumes it.

                                                      • [claimed-docs] The required `description` field: Must be 1-1024 characters, Should describe both what the skill does and when to use it
                                                      • [claimed-docs] A skill is a directory containing, at minimum, a `SKILL.md` file
                                                      • [claimed-docs] The `SKILL.md` file must contain YAML frontmatter followed by Markdown content.
                                                      • [claimed-docs] Upload and manage through the [Skills API](https://platform.claude.com/docs/en/api/skills/create)
                                                      • [claimed-docs] This guide shows you how to use both pre-built and custom Skills with the Claude API.
                                                      Skills for Real Engineersn/a

                                                      This product is a collection of AI agent 'skill' files/instructions for coding workflows, not an API or service with an interface to document via OpenAPI. There is no evidence of an API surface that would warrant a machine-readable spec, making this axis a category error for this product type.

                                                      • ai-native userSchedule recurring jobs or workflows

                                                        weight 2 · not comparable
                                                        Anthropic Skillsnone0/10

                                                        The evidence pack describes Skills as on-demand or auto-triggered instruction modules invoked by Claude during a task (via /skill-name, semantic triggering, or API container calls), but nothing describes a scheduler, cron-like trigger, or persistent recurring job mechanism. Automation-depth for scheduling is a fair ask for an agentic tool, but no evidence shows Skills or Claude Code support recurring/scheduled execution.

                                                        • [claimed-docs] Claude uses skills when relevant, or you can invoke one directly with `/skill-name`.
                                                        • [claimed-docs] If Claude thinks the skill is relevant to the current task, it will load the skill by reading its full `SKILL.md` into context.
                                                        • [claimed-docs] `/run` and `/verify` work without setup. They infer the launch from your project type (CLI, server, TUI, browser-driven) and from what's in …
                                                        Skills for Real Engineersn/a

                                                        This product is a library of AI agent skills/prompts for coding workflows (TDD, triage, grilling, etc.), not a scheduler or automation platform that runs recurring jobs/workflows on a schedule. No evidence of cron-like scheduling or persistent job orchestration, and this capability is outside the product's category.

                                                        • ai-native userSelf-host the core product

                                                          weight 3 · not comparable
                                                          Anthropic Skillsn/a

                                                          Anthropic Skills is a packaging/format layer (SKILL.md files, plugins, marketplaces) that runs on top of Claude Code or the hosted Claude API — it has no standalone server/model component of its own to self-host. Self-hosting is a category mismatch for a skills/plugin framework built atop a proprietary hosted LLM service, not an applicable axis for this product type.

                                                            Skills for Real Engineersfullclaimed7/10

                                                            The product is an MIT-licensed, MIT-open GitHub repo of skill files that are copied directly into the user's own repo as editable, ordinary files ('you own and can edit... nothing updates behind your back'), which is effectively full self-hosting since there is no server component to host beyond the files themselves. This is corroborated by both the docs and the ADR describing the plugin/fork model. missing for 10: no independent/hands-on confirmation of a full self-hosted install working end-to-end outside vendor docs, and no explicit statement addressing infrastructure/hosting concerns (e.g., private registries, offline use).

                                                            • [claimed-docs] copies editable skill files into your project, so you can hack on them and make them your own
                                                            • [claimed-docs] It writes the skills into your repo as ordinary files you own and can edit. Nothing updates behind your back
                                                            • [claimed-docs] Nothing updates behind your back; pull my latest changes when you want them with `npx skills update`
                                                            • [claimed-docs] subscribe to the set as a read-only, always-current bundle you don't edit, rather than a fork you own
                                                          • ai-native userChoose where my data is stored (region/residency)

                                                            weight 2 · not comparable
                                                            Anthropic Skillsn/a

                                                            Anthropic Skills is a feature for packaging instructions/scripts for Claude agents; data residency/region storage is an infrastructure/compliance concern of the underlying platform (Claude API/Claude.ai), not something Skills as a capability could expose or configure.

                                                              Skills for Real Engineersn/a

                                                              This product is a set of skill files/prompts for coding agents, not a data-hosting or storage service; data residency/region selection is not an applicable axis for this kind of tool.

                                                              • ai-native userPrevent my data from being used to train AI models

                                                                weight 3 · not comparable
                                                                Anthropic Skillsn/a

                                                                Anthropic Skills is a feature/framework for packaging agent capabilities, not a data-privacy or training-opt-out control; the evidence pack contains no data-training-consent settings and this axis is a category error for this product type.

                                                                  Skills for Real Engineersn/a

                                                                  This product is a collection of AI agent 'skills'/prompt files for coding workflows, not a data-processing or model-training service; it has no data-handling relationship with end users' data being used for AI training, so an AI-training opt-out story is a category error for this kind of product.

                                                                  • ai-native userOpt out of telemetry and usage tracking

                                                                    weight 2 · not comparable
                                                                    Anthropic Skillsn/a

                                                                    Anthropic Skills is a feature/packaging format for extending Claude's capabilities, not a telemetry-collecting service with its own privacy/tracking controls to opt out of — this axis belongs to platform-level privacy settings (e.g., Claude.ai/Claude Code), not the Skills feature itself.

                                                                      Skills for Real Engineersnone0/10

                                                                      The evidence pack covers installation modes, skill content, and community commentary but contains no mention of telemetry, analytics, or usage tracking of any kind, let alone an opt-out mechanism. Since this is a tool a buyer could reasonably ask about data collection, absence of any documentation on the topic means the axis applies but is unaddressed.