Try itExperimental
See what an agent can do with Undermind before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$mcp-probe → https://mcp.undermind.ai/mcplive — run just now from our edge$ press ▶ run to send one JSON-RPC initialize from our edge
Verified integrations
Connections to other tracked products — hover a chip for the verbatim evidence quote behind it.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharingevidence →
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflowevidence →
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Report output — stories about report output in this arenaReport outputevidence →
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depthevidence →
Stories about research depth in this arena
Source quality — stories about source quality in this arenaSource qualityevidence →
Stories about source quality in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 0 paid · 4 enterprise · 12 not stated in evidence
Follow the green: where the map greys out is where Undermind stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
~5/10
unlocks → Webhooks · Official SDKs · Scoped API keys · Machine-readable spec · Versioning policy · Official CLI · Full data export
Subscribe to events via webhooks
—0/10
Build against official SDKs
—0/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
n/an/a
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—0/10
Operate the product with natural-language commands
✓8/10
Plug MCP servers into this product so it can use their tools
—0/10
Get AI-generated insights and suggestions from my data inside the product
✓8/10
Set up automations that run autonomously in the background
~3/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Automation depth
Define rules that trigger actions automatically on events
~3/10
unlocks → Upload my own PDFs or corpus and have the agent research over them
Schedule recurring jobs or workflows
—0/10
Perform bulk operations across many items at once
~6/10
Version, review, and roll back my automations
n/an/a
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
Agent runs
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
Sorted by importance (agentic first) (high → low) · 43/43 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partialenterprise | 5/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none± | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Xcommunity | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 3/10 | Cclaimed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partialenterprise | 2/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none± | 0/10 | ||
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none± | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | n/a | untested | none yet | |
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 3 | full | 8/10 | Xcommunity | |
See citations for every substantive claim so I can verify it against the underlying source C Citations | researcher | Source quality — stories about source quality in this arenaSource quality | 3 | full | 7/10 | Xcommunity | |
Get a structured report with sections, tables, and a summary that I can share with stakeholders C Reports | analyst | Report output — stories about report output in this arenaReport output | 3 | partial | 5/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 3/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | untested | none yet | |
Search scholarly literature and primary sources, not just the open web C Corpus | researcher | Source quality — stories about source quality in this arenaSource quality | 2 | full | 8/10 | Xcommunity | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partialenterprise | 6/10 | Cclaimed | |
Run a systematic screening and extraction workflow across many papers with consistent criteria C Reviews | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | partial | 6/10 | Xcommunity | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partialenterprise | 5/10 | Tprobed | |
Start a long research job that keeps working unattended and notifies me when the result is ready C Agent runs | analyst | Research depth — stories about research depth in this arenaResearch depth | 2 | partial | 5/10 | Xcommunity | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
See where sources agree and disagree instead of a single unqualified answer C Synthesis | analyst | Source quality — stories about source quality in this arenaSource quality | 2 | none | 0/10 | ||
Understand plan pricing and usage limits before committing G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | none | 0/10 | ||
Upload my own PDFs or corpus and have the agent research over them C Corpus | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | n/a | untested | none yet | |
Share a research session or report with collaborators who can view or build on it C Sharing | analyst | Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing | 2 | none | untested | none yet | |
Set up standing searches or alerts that surface new relevant sources as they appear C Alerts | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 1 | partial | 6/10 | Cclaimed | |
Try the product meaningfully on a free tier or trial G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | disputed | 4/10 | Dcontradicted | |
Steer the depth, effort, and scope of a research run before or while it executes C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 1 | none | 0/10 | ||
Export results to common formats, including documents, spreadsheets, and reference-manager files G Reports | researcher | Report output — stories about report output in this arenaReport output | 1 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | n/a | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 32 stories with headroom
What would move Undermind’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
Missing: any evidence of a native, built-in AI assistant/chat agent inside Undermind's own UI that a user can delegate tasks to.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
All MCP-related evidence describes Undermind acting as an MCP *server* that other clients (Cursor, VS Code, Claude, ChatGPT) can plug into to use Undermind's own tools — the reverse of this story, which asks whether a user can plug external MCP servers into Undermind so it can use their tools.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
No evidence of any data export feature or open-format export capability for user data/papers/workspaces; probes for docs/API endpoints also 404.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence in the pack addresses data privacy, opt-out of AI training, or data usage policies for Undermind; nothing here confirms or denies such a control exists.
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs
nonemoves agent-readyimpact 30
Direct probes show no llms.txt, no docs.md, and no openapi spec (all 404), meaning there is no agent-consumable documentation file for a generic AI agent to fetch.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
Evidence shows an MCP server, API access, and web/ChatGPT app integrations, but there is no mention of an official CLI tool for Undermind anywhere in the docs or probes; llms.txt, docs-md, and openapi probes all 404, and no CLI is documented.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence of scoped/least-privilege API credential issuance for agents; the API is only mentioned generically ('Programmatic queries via API') with no docs on credential scoping, permissions, or key management, and OpenAPI probes returned 404s.
Agenticness — how well agents can access and operate the productBuild against official SDKs
nonemoves agent-readyimpact 30
Evidence only mentions a vague 'Programmatic queries via API' for enterprise customers and an MCP server, but no official SDKs (client libraries, language bindings) are documented; probes for OpenAPI specs and docs (llms.txt, mcp.md, openapi.json) all return 404, indicating no public developer SDK resources exist.
Showing the top 8 of 32 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map6 surfaces · 17 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
MCP docs13 stories
- Connect an agent via an official MCP server
- Drive the product through a documented public API
- Get AI-generated insights and suggestions from my data inside the product
- Operate the product with natural-language commands
- Perform bulk operations across many items at once
- Set up standing searches or alerts that surface new relevant sources as they appear
- Run a systematic screening and extraction workflow across many papers with consistent criteria
- Do everything through the API that I can do in the UI
- Get a structured report with sections, tables, and a summary that I can share with stakeholders
- Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- Start a long research job that keeps working unattended and notifies me when the result is ready
- See citations for every substantive claim so I can verify it against the underlying source
- Search scholarly literature and primary sources, not just the open web
undermind.ai9 stories
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Define rules that trigger actions automatically on events
- Set up standing searches or alerts that surface new relevant sources as they appear
- Do everything through the API that I can do in the UI
- Try the product meaningfully on a free tier or trial
- Start a long research job that keeps working unattended and notifies me when the result is ready
- See citations for every substantive claim so I can verify it against the underlying source
- Search scholarly literature and primary sources, not just the open web
Hacker News7 stories
- Get AI-generated insights and suggestions from my data inside the product
- Run a systematic screening and extraction workflow across many papers with consistent criteria
- Try the product meaningfully on a free tier or trial
- Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- Start a long research job that keeps working unattended and notifies me when the result is ready
- See citations for every substantive claim so I can verify it against the underlying source
- Search scholarly literature and primary sources, not just the open web
Enterprise docs4 stories
OpenAPI spec3 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
7 of 10 testable claims verified · 1 contradicted → integrity 50/100
11 distinct capability claims found in Undermind’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
7
Verified
2
Unverified
1
Contradicted
7
Undersold
Verified (7)
“Provides an official MCP server so any MCP-compatible AI client (Cursor, VS Code, Claude, etc.) can connect to Undermind”
“Runs an autonomous deep literature review, planning its own searches, following citations/authors until no new relevant papers found”
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summaryfullproof ↗
“Runs an autonomous deep literature review, planning its own searches, following citations/authors until no new relevant papers found”
Start a long research job that keeps working unattended and notifies me when the result is readypartialproof ↗
“Reads full-text PDFs in parallel and answers specific questions across many papers at once, including from figures, tables, and equations”
Run a systematic screening and extraction workflow across many papers with consistent criteriapartialproof ↗
“Search engine claimed to deliver 10x better literature search results than Google Scholar”
Search scholarly literature and primary sources, not just the open webfullproof ↗
“Every substantive statement can be traced via in-line citations back to the source paper”
See citations for every substantive claim so I can verify it against the underlying sourcefullproof ↗
“Offers programmatic access to research queries via an API”
Drive the product through a documented public APIpartialproof ↗
Unverified (2)
“Creates and edits Markdown notes, syntheses, and reports in a persistent workspace accessible across sessions”
Get a structured report with sections, tables, and a summary that I can share with stakeholderspartialproof ↗
“Notifies users whenever relevant new papers are published, acting as a standing search alert”
Set up standing searches or alerts that surface new relevant sources as they appearpartialproof ↗
Contradicted (1)
“Higher-tier plans offer higher usage limits and unlimited workspaces, files, and paper libraries”
Understand plan pricing and usage limits before committingnoneproof ↗
Undersold (7)
Run the product headlessly / in CI for automationpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandsfullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Claims outside our story set (2)
Real capability claims found in Undermind’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Lets users curate papers into a folder for long-term reference”
source ↗“Lets users star important papers across the workspace”
source ↗
Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP 100% (30d, checked every 6h since Sep 8 '26)
