Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharingevidence →
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflowevidence →
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Report output — stories about report output in this arenaReport outputevidence →
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depthevidence →
Stories about research depth in this arena
Source quality — stories about source quality in this arenaSource qualityevidence →
Stories about source quality in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 1 free · 0 paid · 0 enterprise · 16 not stated in evidence
Follow the green: where the map greys out is where FutureHouse Platform stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓7/10
unlocks → Webhooks · Scoped API keys · MCP server · Machine-readable spec · Versioning policy · Official CLI · Full data export
Subscribe to events via webhooks
—0/10
Build against official SDKs
~6/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
—0/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
n/an/a
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓7/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
~7/10
unlocks → Autonomous automations
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
✓7/10
Set up automations that run autonomously in the background
—0/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
Agent runs
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
Sorted by importance (agentic first) (high → low) · 43/43 stories · click a row’s chevron for the rationale and evidence
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 7/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none± | 0/10 | ||
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a± | untested | none yet | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial± | 7/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full± | 7/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial± | 6/10 | Tprobed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none± | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | n/a | untested | none yet | |
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 3 | full | 8/10 | Cclaimed | |
See citations for every substantive claim so I can verify it against the underlying source C Citations | researcher | Source quality — stories about source quality in this arenaSource quality | 3 | full | 7/10 | Cclaimed | |
Get a structured report with sections, tables, and a summary that I can share with stakeholders C Reports | analyst | Report output — stories about report output in this arenaReport output | 3 | partial | 6/10 | Cclaimed | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | n/a | untested | none yet | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | untested | none yet | |
Search scholarly literature and primary sources, not just the open web C Corpus | researcher | Source quality — stories about source quality in this arenaSource quality | 2 | full | 8/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 5/10 | Tprobed | |
Run a systematic screening and extraction workflow across many papers with consistent criteria C Reviews | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | partial | 5/10 | Tprobed | |
See where sources agree and disagree instead of a single unqualified answer C Synthesis | analyst | Source quality — stories about source quality in this arenaSource quality | 2 | partial | 5/10 | Cclaimed | |
Start a long research job that keeps working unattended and notifies me when the result is ready C Agent runs | analyst | Research depth — stories about research depth in this arenaResearch depth | 2 | partial | 5/10 | Cclaimed | |
Understand plan pricing and usage limits before committing G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | partial | 4/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Upload my own PDFs or corpus and have the agent research over them C Corpus | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Share a research session or report with collaborators who can view or build on it C Sharing | analyst | Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing | 2 | none | untested | none yet | |
Try the product meaningfully on a free tier or trial G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | fullfree | 7/10 | Cclaimed | |
Export results to common formats, including documents, spreadsheets, and reference-manager files G Reports | researcher | Report output — stories about report output in this arenaReport output | 1 | none | 0/10 | ||
Steer the depth, effort, and scope of a research run before or while it executes C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 1 | none | 0/10 | ||
Set up standing searches or alerts that surface new relevant sources as they appear C Alerts | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 1 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | n/a | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 29 stories with headroom
What would move FutureHouse Platform’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server
nonemoves agent-readyimpact 45
The evidence pack only shows a Python client (edison-client) for calling FutureHouse agents via API key, plus probes confirming no OpenAPI/MCP-related endpoints were found; there is no mention of an official MCP server for connecting agents.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
No evidence of any data export capability, open-format export, or account portability/deletion feature; documentation covers agent/task usage and API access but nothing about exporting user data or leaving with it.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
No evidence in the pack addresses data usage for AI training opt-out, data privacy controls, or any training-data policy; the docs focus entirely on product features and API usage.
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background
nonemoves Built-in AIimpact 30
Missing: scheduling/cron mechanism, event-driven triggers, persistent background job management, and any docs describing recurring or unattended automation setup.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
Evidence shows only a Python client library (edison-client, installed via pip, used programmatically with client.run_tasks_until_done) rather than a command-line interface; no CLI tool, command syntax, or terminal usage is documented anywhere in the pack.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
Evidence shows only a single, account-wide API token creation flow with no mention of scopes, permissions, or least-privilege controls for agents; no evidence of scoped or restricted credential issuance.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
No mention of webhooks, event subscriptions, or callback mechanisms anywhere in the docs; the client is a polling/run-tasks style API and OpenAPI probe returned 404s, giving no evidence of webhook support.
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
Evidence shows only a quickstart guide with basic client code snippets, not an interactive API reference with runnable examples; probes for OpenAPI/swagger specs and doc endpoints all returned 404s, indicating no interactive reference exists.
Showing the top 8 of 29 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map3 surfaces · 17 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Futurehouse cookbook docs17 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Run a systematic screening and extraction workflow across many papers with consistent criteria
- Do everything through the API that I can do in the UI
- Try the product meaningfully on a free tier or trial
- Understand plan pricing and usage limits before committing
- Get a structured report with sections, tables, and a summary that I can share with stakeholders
- Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- Start a long research job that keeps working unattended and notifies me when the result is ready
- See citations for every substantive claim so I can verify it against the underlying source
- Search scholarly literature and primary sources, not just the open web
- See where sources agree and disagree instead of a single unqualified answer
OpenAPI spec4 stories
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
2 of 10 testable claims verified · 2 contradicted → integrity 0/100
14 distinct capability claims found in FutureHouse Platform’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
2
Verified
6
Unverified
2
Contradicted
9
Undersold
Verified (4)
“Provides an installable Python client package (edison-client) for programmatic access”
“Authentication via an API key obtained from the user's profile page”
Drive the product through a documented public APIfullproof ↗
“Jobs can be submitted programmatically with a task name and query (e.g. literature search job)”
Drive the product through a documented public APIfullproof ↗
“SDK provides a call to run a task and block until it completes”
Unverified (7)
“Kosmos agent autonomously reads literature, writes/executes analysis code, generates hypotheses, and produces a cited report”
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summaryfullproof ↗
“Every conclusion can be traced back to the specific code or literature passage that produced it”
See citations for every substantive claim so I can verify it against the underlying sourcefullproof ↗
“Can answer complex scientific questions with cited responses or run deep literature reviews synthesizing conflicting evidence across many papers”
See where sources agree and disagree instead of a single unqualified answerpartialproof ↗
“Can answer complex scientific questions with cited responses or run deep literature reviews synthesizing conflicting evidence across many papers”
Search scholarly literature and primary sources, not just the open webfullproof ↗
“Turns raw data into detailed analyses, statistical results, and publication-ready figures”
Get a structured report with sections, tables, and a summary that I can share with stakeholderspartialproof ↗
“Paid subscription plans offer higher rate limits and additional features”
Understand plan pricing and usage limits before committingpartialproof ↗
“Searches over 175M+ papers, clinical trials, and patents with understanding of citation graphs and journal quality”
Search scholarly literature and primary sources, not just the open webfullproof ↗
Contradicted (2)
“Users can create new API tokens from an 'API Tokens' section for programmatic access”
Issue scoped/least-privilege API credentials for an agentnoneproof ↗
“Specializes in processing complex experimental datasets such as flow cytometry and RNA-seq data”
Upload my own PDFs or corpus and have the agent research over themnoneproof ↗
Undersold (9)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandspartialproof ↗
Run a systematic screening and extraction workflow across many papers with consistent criteriapartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Try the product meaningfully on a free tier or trialfullproof ↗
Start a long research job that keeps working unattended and notifies me when the result is readypartialproof ↗
Claims outside our story set (2)
Real capability claims found in FutureHouse Platform’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Precedent tool searches across fields to determine if a research idea has been tried before and identify gaps”
source ↗“Molecules agent specializes in chemistry-focused molecular design and analysis”
source ↗
Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime llms.txt 100% (30d, checked every 6h since Sep 8 '26)
