Rank #3 of 5 in Agent Skills & Extensions
Access
Showcase


Try itExperimental
See what an agent can do with Skills for Real Engineers before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$npx -y skills add mattpocock/skills --skill tdd -a claude-code -y --copy # in a scratch dir, then print installed SKILL.md frontmatterrecorded session — replayed, not liveVerified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agent workflows — stories about agent workflows in this arenaAgent workflowsevidence →
Stories about agent workflows in this arena
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Cross agent portability — stories about cross agent portability in this arenaCross agent portabilityevidence →
Stories about cross agent portability in this arena
Discovery distribution — stories about discovery distribution in this arenaDiscovery distributionevidence →
Stories about discovery distribution in this arena
Docs onboarding — stories about docs onboarding in this arenaDocs onboardingevidence →
Stories about docs onboarding in this arena
Install experience — stories about install experience in this arenaInstall experienceevidence →
Stories about install experience in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Safety review — stories about safety review in this arenaSafety reviewevidence →
Stories about safety review in this arena
Skill authoring — stories about skill authoring in this arenaSkill authoringevidence →
Stories about skill authoring in this arena
Testing quality — stories about testing quality in this arenaTesting qualityevidence →
Stories about testing quality in this arena
Versioning updates — stories about versioning updates in this arenaVersioning updatesevidence →
Stories about versioning updates in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 1 free · 0 paid · 0 enterprise · 23 not stated in evidence
Follow the green: where the map greys out is where Skills for Real Engineers stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agent workflows — stories about agent workflows in this arenaAgent workflows
Stories about agent workflows in this arena
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
n/an/a
Subscribe to events via webhooks
n/an/a
Build against official SDKs
—–
Issue scoped/least-privilege API credentials for an agent
n/an/a
Connect an agent via an official MCP server
—0/10
Download a machine-readable API spec (OpenAPI or equivalent)
n/an/a
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
n/an/a
CLI & headless
Use an official CLI
✓7/10
unlocks → Headless / CI
Run the product headlessly / in CI for automation
—–
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—–
Operate the product with natural-language commands
✓8/10
unlocks → Autonomous automations
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
~6/10
Set up automations that run autonomously in the background
—0/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Cross agent portability — stories about cross agent portability in this arenaCross agent portability
Stories about cross agent portability in this arena
Discovery distribution — stories about discovery distribution in this arenaDiscovery distribution
Stories about discovery distribution in this arena
Browse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anything
~6/10
Distribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the project
✓8/10
Installed skills trigger automatically from task context, with descriptions engineered so the agent activates the right skill at the right moment
—0/10
Docs onboarding — stories about docs onboarding in this arenaDocs onboarding
Stories about docs onboarding in this arena
Install experience — stories about install experience in this arenaInstall experience
Stories about install experience in this arena
Install
List what is installed and remove skills cleanly, without orphaned files or lingering instructions
—–
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Safety review — stories about safety review in this arenaSafety review
Stories about safety review in this arena
Skill authoring — stories about skill authoring in this arenaSkill authoring
Stories about skill authoring in this arena
Testing quality — stories about testing quality in this arenaTesting quality
Stories about testing quality in this arena
Versioning updates — stories about versioning updates in this arenaVersioning updates
Stories about versioning updates in this arena
Sorted by importance (agentic first) (high → low) · 52/52 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | 0/10 | ||
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
There is a documented update path — marketplace auto-updates or an explicit update command — so I get fixes without reinstalling from scratch C Updates | developer | Versioning updates — stories about versioning updates in this arenaVersioning updates | 3 | full | 8/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullfree | 7/10 | Cclaimed | |
A quickstart takes me from nothing to a working installed skill in under five minutes C Onboarding | developer | Docs onboarding — stories about docs onboarding in this arenaDocs onboarding | 3 | partial | 6/10 | Cclaimed | |
Install a skill collection with one documented command — a package-manager one-liner, CLI, or in-agent marketplace command — and it is active in my next session C Install | developer | Install experience — stories about install experience in this arenaInstall experience | 3 | partial | 6/10 | Cclaimed | |
Install the same collection into multiple different coding agents — Claude Code, Codex, Cursor, and others — with per-harness instructions C Portability | developer | Cross agent portability — stories about cross agent portability in this arenaCross agent portability | 3 | partial | 6/10 | Cclaimed | |
Review exactly what instructions and scripts a skill will add — list contents before installing and read every file afterward C Review | engineering-lead | Safety review — stories about safety review in this arenaSafety review | 3 | partial | 6/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 5/10 | Cclaimed | |
My coding agent can install a skill by itself — a non-interactive, promptless install path an agent can run headlessly end to end C Agent ops | ai-native user | Agent workflows — stories about agent workflows in this arenaAgent workflows | 3 | partial | 4/10 | Cclaimed | |
Author a new skill from a documented template — a SKILL.md with name and description frontmatter — without reverse-engineering existing skills C Authoring | developer | Skill authoring — stories about skill authoring in this arenaSkill authoring | 3 | none | 0/10 | ||
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | none | 0/10 | ||
Installed skills trigger automatically from task context, with descriptions engineered so the agent activates the right skill at the right moment C Triggering | developer | Discovery distribution — stories about discovery distribution in this arenaDiscovery distribution | 3 | none | 0/10 | ||
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | n/a | untested | none yet | |
Distribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the project C Distribution | engineering-lead | Discovery distribution — stories about discovery distribution in this arenaDiscovery distribution | 2 | full | 8/10 | Cclaimed | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | full | 8/10 | Cclaimed | |
Skills are plain markdown files and folders I can read, copy, and carry to another harness — not a proprietary binary format C Portability | developer | Cross agent portability — stories about cross agent portability in this arenaCross agent portability | 2 | full | 8/10 | Cclaimed | |
Every skill documents what it does and when it activates, so I can predict my agent's new behavior before it surprises me C Onboarding | developer | Docs onboarding — stories about docs onboarding in this arenaDocs onboarding | 2 | full | 7/10 | Xcommunity | |
Browse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anything C Discovery | developer | Discovery distribution — stories about discovery distribution in this arenaDiscovery distribution | 2 | partial | 6/10 | Cclaimed | |
Skills follow the open Agent Skills specification so the same skill folder is valid beyond this one vendor's tooling C Spec | developer | Skill authoring — stories about skill authoring in this arenaSkill authoring | 2 | partial | 6/10 | Cclaimed | |
Install only the specific skills I want from a collection instead of taking the whole bundle C Install | developer | Install experience — stories about install experience in this arenaInstall experience | 2 | partial | 5/10 | Cclaimed | |
Choose install scope — project-local files committed with my repo, or user-global across all projects C Install | developer | Install experience — stories about install experience in this arenaInstall experience | 2 | partial | 4/10 | Cclaimed | |
The project documents its security posture — what skills can execute, the trust model for third-party skills, and any telemetry or data collection C Trust | engineering-lead | Safety review — stories about safety review in this arenaSafety review | 2 | partial | 4/10 | Cclaimed | |
The collection is actively maintained — recent releases, triaged issues, and accepted community contributions C Maintenance | developer | Testing quality — stories about testing quality in this arenaTesting quality | 2 | partial | 3/10 | Xcommunity | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
The collection maintains tests or evals for its skills so changes are verified against regressions rather than shipped on vibes C Testing | developer | Testing quality — stories about testing quality in this arenaTesting quality | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | n/a | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
My agent can author and package a new skill end to end by following the project's own spec, template, or meta-skill C Agent ops | ai-native user | Agent workflows — stories about agent workflows in this arenaAgent workflows | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | n/a | untested | none yet | |
The collection ships a meta-skill or tool that guides my agent through writing, improving, and packaging new skills C Authoring | developer | Skill authoring — stories about skill authoring in this arenaSkill authoring | 2 | none | untested | none yet | |
Control when skill changes reach my team — pinned versions or a lockfile rather than silent behind-the-back updates C Pinning | engineering-lead | Versioning updates — stories about versioning updates in this arenaVersioning updates | 1 | partial | 7/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 4/10 | Cclaimed | |
List what is installed and remove skills cleanly, without orphaned files or lingering instructions C Lifecycle | developer | Install experience — stories about install experience in this arenaInstall experience | 1 | none | untested | none yet | |
Releases ship with notes or a changelog so I can see what changed in the skills before I take an update C Updates | developer | Versioning updates — stories about versioning updates in this arenaVersioning updates | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 35 stories with headroom
What would move Skills for Real Engineers’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
Evidence shows distribution via npm CLI, Claude Code plugin marketplace, and file-copy installation, but nowhere does it mention an MCP server for agents to connect to.
Skill authoring — stories about skill authoring in this arenaAuthor a new skill from a documented template — a SKILL.md with name and description frontmatter — without reverse-engineering existing skills
nonemoves PA Scoreimpact 30
The evidence describes an existing bundle of pre-built skills (tdd, grilling, triage, setup) that you can install, subscribe to, or copy as editable files, but there is no documented template, generator, or guide specifically for authoring a brand-new SKILL.md with frontmatter — the closest thing to a 'template' would be reverse-engineering the shipped example skills, which the story explicitly excludes.
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events
nonemoves PA Scoreimpact 30
The product ships a library of manually-invoked skill files triggered by slash commands or explicit user direction ('Install the ones you want, then type a slash command'; 'Use them every time you want to make a change'), not an event-driven rule/automation engine.
Discovery distribution — stories about discovery distribution in this arenaInstalled skills trigger automatically from task context, with descriptions engineered so the agent activates the right skill at the right moment
nonemoves PA Scoreimpact 30
The evidence describes manual invocation — 'Install the ones you want, then type a slash command' (docs-43) and 'Use them every time you want to make a change' (gh-1/gh-4) — rather than automatic, context-triggered activation via engineered descriptions.
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background
nonemoves Built-in AIimpact 30
The skills are built around interactive, human-in-the-loop workflows (grilling/interviewing the user, confirming test seams, triage state machines) rather than unattended background automation; the project's own philosophy explicitly rejects processes that 'take away your control' in favor of user-confirmed steps.
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation
nonemoves agent-readyimpact 30
The evidence describes skill files consumed by interactive coding agents (Claude Code, Cursor, etc.) that rely on human interviewing/grilling and manual slash-command invocation, with no mention of a CLI flag, non-interactive mode, or CI/automation pipeline usage.
Agenticness — how well agents can access and operate the productBuild against official SDKs
nonemoves agent-readyimpact 30
The axis applies to this product kind (peer products hold positive or none verdicts on this story), so lack of evidence for an applicable capability is "none", never "na".
Showing the top 8 of 35 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map3 surfaces · 24 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
GitHub README23 stories
- My coding agent can install a skill by itself — a non-interactive, promptless install path an agent can run headlessly end to end
- Point an agent at llms.txt or agent-oriented docs
- Use an official CLI
- Operate the product with natural-language commands
- Version, review, and roll back my automations
- Install the same collection into multiple different coding agents — Claude Code, Codex, Cursor, and others — with per-harness instructions
- Skills are plain markdown files and folders I can read, copy, and carry to another harness — not a proprietary binary format
- Browse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anything
- Distribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the project
- A quickstart takes me from nothing to a working installed skill in under five minutes
- Every skill documents what it does and when it activates, so I can predict my agent's new behavior before it surprises me
- Install a skill collection with one documented command — a package-manager one-liner, CLI, or in-agent marketplace command — and it is active in my next session
- Choose install scope — project-local files committed with my repo, or user-global across all projects
- Install only the specific skills I want from a collection instead of taking the whole bundle
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Self-host the core product
- Review exactly what instructions and scripts a skill will add — list contents before installing and read every file afterward
- The project documents its security posture — what skills can execute, the trust model for third-party skills, and any telemetry or data collection
- Skills follow the open Agent Skills specification so the same skill folder is valid beyond this one vendor's tooling
- The collection is actively maintained — recent releases, triaged issues, and accepted community contributions
- Control when skill changes reach my team — pinned versions or a lockfile rather than silent behind-the-back updates
- There is a documented update path — marketplace auto-updates or an explicit update command — so I get fixes without reinstalling from scratch
Skills docs14 stories
- My coding agent can install a skill by itself — a non-interactive, promptless install path an agent can run headlessly end to end
- Get AI-generated insights and suggestions from my data inside the product
- Operate the product with natural-language commands
- Install the same collection into multiple different coding agents — Claude Code, Codex, Cursor, and others — with per-harness instructions
- Skills are plain markdown files and folders I can read, copy, and carry to another harness — not a proprietary binary format
- Browse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anything
- Distribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the project
- A quickstart takes me from nothing to a working installed skill in under five minutes
- Every skill documents what it does and when it activates, so I can predict my agent's new behavior before it surprises me
- Install a skill collection with one documented command — a package-manager one-liner, CLI, or in-agent marketplace command — and it is active in my next session
- Install only the specific skills I want from a collection instead of taking the whole bundle
- Export all of my data in open formats and leave
- Read the product's source under an open license
- Skills follow the open Agent Skills specification so the same skill folder is valid beyond this one vendor's tooling
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$npx -y skills add mattpocock/skills --skill tdd -a claude-code -y --copy # in a scratch dir, then print installed SKILL.md frontmatterreproduced$ npx -y skills add mattpocock/skills --skill tdd -a claude-code -y --copy # in a scratch dir, then print installed SKILL.md frontmatter \ ███████╗██╗ ██╗██╗██╗ ██╗ ███████╗ ██╔════╝██║ ██╔╝██║██║ ██║ ██╔════╝ ███████╗█████╔╝ ██║██║ ██║ ███████╗ ╚════██║██╔═██╗ ██║██║ ██║ ╚════██║ ███████║██║ ██╗██║███████╗███████╗███████║ ╚══════╝╚═╝ ╚═╝╚═╝╚══════╝╚══════╝╚══════╝ ┌ skills │ ◇ Source: https://github.com/mattpocock/skills.git │ [ 3 5 m ◒ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◐ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◓ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◑ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◒ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◐ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◓ [ 3 9m C l o n i n g r e p o s i t o r y … [ 3 5 m ◑ [ 3 9m C l o n i n g r e p o s i t o r y …◇ Repository cloned │ ◇ Found 37 skills │ ● Selected 1 skill: tdd │ ◇ Installation Summary ─╮ │ │ │ │ │ │ │ │ │ [ │ │ 1 │ │ m │ │ M │ │ a │ │ t │ │ t │ │ p │ │ o │ │ c │ │ o │ │ c │ │ k │ │ │ │ S │ │ k │ │ i │ │ l │ │ l │ │ s │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ │ │ [ │ │ 3 │ │ 6 │ │ m │ │ . │ │ / │ │ . │ │ a │ │ g │ │ e │ │ n │ │ t │ │ s │ │ / │ │ s │ │ k │ │ i │ │ l │ │ l │ │ s │ │ / │ │ t │ │ d │ │ d │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ c │ │ o │ │ p │ │ y │ │ │ │ → │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ C │ │ l │ │ a │ │ u │ │ d │ │ e │ │ │ │ C │ │ o │ │ d │ │ e │ │ │ ├────────────────────────╯ │ ◇ Security Risk Assessments ─╮ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ G │ │ e │ │ n │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ S │ │ o │ │ c │ │ k │ │ e │ │ t │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ S │ │ n │ │ y │ │ k │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ [ │ │ 3 │ │ 6 │ │ m │ │ t │ │ d │ │ d │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ │ │ │ │ [ │ │ 3 │ │ 2 │ │ m │ │ S │ │ a │ │ f │ │ e │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ [ │ │ 3 │ │ 2 │ │ m │ │ 0 │ │ │ │ a │ │ l │ │ e │ │ r │ │ t │ │ s │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ │ [ │ │ 3 │ │ 2 │ │ m │ │ L │ │ o │ │ w │ │ │ │ R │ │ i │ │ s │ │ k │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ D │ │ e │ │ t │ │ a │ │ i │ │ l │ │ s │ │ : │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ h │ │ t │ │ t │ │ p │ │ s │ │ : │ │ / │ │ / │ │ s │ │ k │ │ i │ │ l │ │ l │ │ s │ │ . │ │ s │ │ h │ │ / │ │ m │ │ a │ │ t │ │ t │ │ p │ │ o │ │ c │ │ o │ │ c │ │ k │ │ / │ │ s │ │ k │ │ i │ │ l │ │ l │ │ s │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ ├─────────────────────────────╯ │ ◇ Installation complete │ ◇ Installed 1 skill ─╮ │ │ │ │ │ │ │ │ │ [ │ │ 1 │ │ m │ │ M │ │ a │ │ t │ │ t │ │ p │ │ o │ │ c │ │ o │ │ c │ │ k │ │ │ │ S │ │ k │ │ i │ │ l │ │ l │ │ s │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ [ │ │ 3 │ │ 2 │ │ m │ │ ✓ │ │ │ │ [ │ │ 3 │ │ 9m │ │ │ │ t │ │ d │ │ d │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ ( │ │ c │ │ o │ │ p │ │ i │ │ e │ │ d │ │ ) │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ │ │ │ │ │ │ [ │ │ 2 │ │ m │ │ → │ │ │ │ [ │ │ 2 │ │ 2m │ │ │ │ . │ │ / │ │ . │ │ c │ │ l │ │ a │ │ u │ │ d │ │ e │ │ / │ │ s │ │ k │ │ i │ │ l │ │ l │ │ s │ │ / │ │ t │ │ d │ │ d │ │ │ ├─────────────────────╯ │ └ Done! Review skills before use; they run with full agent permissions. \--- installed SKILL.md frontmatter --- --- name: tdd description: Test-driven development. Use when the user wants to build features or fix bugs test-first, mentions "red-green-refactor", or wants integration tests.
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
1 of 6 testable claims verified · 0 contradicted → integrity 17/100
28 distinct capability claims found in Skills for Real Engineers’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
1
Verified
5
Unverified
0
Contradicted
18
Undersold
Verified (1)
“Provides a CONTEXT.md-style document that helps agents decode project-specific jargon”
Point an agent at llms.txt or agent-oriented docspartialproof ↗
Unverified (6)
“Offers a read-only, auto-updating bundle install so you subscribe to vendor updates instead of forking the files”
There is a documented update path — marketplace auto-updates or an explicit update command — so I get fixes without reinstalling from scratchfullproof ↗
“Can copy editable skill files directly into your project so you can customize them freely”
Skills are plain markdown files and folders I can read, copy, and carry to another harness — not a proprietary binary formatfullproof ↗
“Provides an explicit `npx skills update` command so nothing changes without your action”
There is a documented update path — marketplace auto-updates or an explicit update command — so I get fixes without reinstalling from scratchfullproof ↗
“Provides an explicit `npx skills update` command so nothing changes without your action”
Control when skill changes reach my team — pinned versions or a lockfile rather than silent behind-the-back updatespartialproof ↗
“Was accepted into Claude Code's official marketplace with a documented `claude plugins install` command”
Install a skill collection with one documented command — a package-manager one-liner, CLI, or in-agent marketplace command — and it is active in my next sessionpartialproof ↗
“Was accepted into Claude Code's official marketplace with a documented `claude plugins install` command”
Browse or search a catalog of available skills — a registry, leaderboard, or marketplace listing — before installing anythingpartialproof ↗
Undersold (18)
My coding agent can install a skill by itself — a non-interactive, promptless install path an agent can run headlessly end to endpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Install the same collection into multiple different coding agents — Claude Code, Codex, Cursor, and others — with per-harness instructionspartialproof ↗
Distribute a standard skill set to my whole team — via a marketplace, a shared repo, or files committed to the projectfullproof ↗
A quickstart takes me from nothing to a working installed skill in under five minutespartialproof ↗
Every skill documents what it does and when it activates, so I can predict my agent's new behavior before it surprises mefullproof ↗
Choose install scope — project-local files committed with my repo, or user-global across all projectspartialproof ↗
Install only the specific skills I want from a collection instead of taking the whole bundlepartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Review exactly what instructions and scripts a skill will add — list contents before installing and read every file afterwardpartialproof ↗
The project documents its security posture — what skills can execute, the trust model for third-party skills, and any telemetry or data collectionpartialproof ↗
Skills follow the open Agent Skills specification so the same skill folder is valid beyond this one vendor's toolingpartialproof ↗
The collection is actively maintained — recent releases, triaged issues, and accepted community contributionspartialproof ↗
Claims outside our story set (23)
Real capability claims found in Skills for Real Engineers’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Setup asks which issue tracker to use — GitHub, Linear, or local files”
source ↗“Ships an alignment skill that gets you and the agent to agree on a change before work starts”
source ↗“Includes a 'grilling' session skill that builds shared language with the AI and records decisions as ADRs”
source ↗“Enforces a red-green-refactor loop where the agent writes a failing test first, then fixes it”
source ↗“Requires tests only at pre-agreed seams, confirmed with the user before any test is written”
source ↗“Includes an interview skill that relentlessly questions the user and maps understanding as a design tree”
source ↗“Moves issues through a triage state machine: categorise, verify, grill if needed, and write agent-ready briefs”
source ↗“Includes a verification step that reproduces a bug from the reporter's steps before grilling the claim”
source ↗“Provides a scaffolding skill that sets up per-repo config the other skills assume (issue tracker, triage labels, domain docs)”
source ↗“Includes a skill that turns an agreed conversation into a written spec”
source ↗“Includes a skill that splits a spec into small tickets an agent can build”
source ↗“Includes a skill that builds a finished spec into code, test-first”
source ↗“Includes a skill that reviews a diff against your standards and the spec”
source ↗“Includes a skill that charts a large effort as a map of decisions and settles them”
source ↗“Includes a skill that answers a design question with throwaway code you then delete”
source ↗“Includes a skill that produces a cited answer read from primary sources”
source ↗“Includes a skill that finds modules worth refactoring and outputs a visual report”
source ↗“Includes a skill that diagnoses a hard bug starting from a failing repro”
source ↗“Includes a skill that resolves a merge or rebase conflict hunk by hunk”
source ↗“Includes a skill that generates a script walking a human through setup”
source ↗“Includes a skill that writes up a long session so another agent can pick it up later”
source ↗“Includes a skill to help you figure out which skill applies to your current situation”
source ↗“Encourages working in vertical slices — one test then one implementation, repeated as tracer bullets”
source ↗
Business model
MIT open-source skill collection, free via the Claude Code marketplace or skills.sh; funded by AI Hero's paid courses and newsletter rather than by the skills themselves.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
