AI Research Agents — procurement report
ProductArena · rankings as of 2026-09-16 · evidence as of 2026-09-16 · 6 products · 43 judged requirements · 258 judged cells
Methodology: Every product is judged against a shared taxonomy of user stories using cited evidence — hands-on probes > repository code > independent community sources > vendor claims — never opinion. Full writeup: https://ultrametric.ai/productarena/methodology
Leaderboard
| # | Product | PA Score | Coverage score | Applicable cells | Confidence |
|---|---|---|---|---|---|
| 1 | Elicit | 28.4 | 35.1 | 41/43 | C |
| 2 | Consensus | 19.1 | 23.8 | 37/43 | C |
| 3 | FutureHouse Platform | 18.8 | 26.3 | 38/43 | B |
| 4 | Undermind | 16.1 | 21.1 | 39/43 | B |
| 5 | Sakana Marlin | 8.9 | 18.4 | 36/43 | B |
| 6 | Gemini Notebook (NotebookLM) | 7.2 | 15.3 | 35/43 | D |
PA Score = agent-readiness blend (see methodology). Coverage score = weighted share of judged requirements met. Confidence = how much of the score rests on tested vs claimed evidence (A–D).
Uncertainty note
The current #1/#2 gap in this arena is not close enough to qualify for the multi-judge uncertainty pass — no extra caveat applies beyond the per-product confidence grades above.
Buyer checklist (RFP)
The arena's 43 judged user stories as requirements, grouped by theme. Priorities mirror the story weights our scoring uses (3 = must-have, 2 = should-have, 1 = nice-to-have). Interactive version with per-requirement verdicts for the top products: /arena/ai-research-agents/checklist
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
- ai-native userPlug MCP servers into this product so it can use their toolsmust-have
- ai-native userConnect an agent via an official MCP servermust-have
- ai-native userDrive the product through a documented public APImust-have
- ai-native userDelegate tasks to a built-in AI assistant inside the productmust-have
- ai-native userPoint an agent at llms.txt or agent-oriented docsshould-have
- ai-native userRun the product headlessly / in CI for automationshould-have
- ai-native userUse an official CLIshould-have
- ai-native userIssue scoped/least-privilege API credentials for an agentshould-have
- ai-native userBuild against official SDKsshould-have
- ai-native userSubscribe to events via webhooksshould-have
- ai-native userGet AI-generated insights and suggestions from my data inside the productshould-have
- ai-native userSet up automations that run autonomously in the backgroundshould-have
- ai-native userOperate the product with natural-language commandsshould-have
- ai-native userExplore an interactive API reference with runnable examplesshould-have
- ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)should-have
- ai-native userRely on versioned APIs with a documented deprecation policyshould-have
- ai-native userTest against a sandbox environment without touching production datanice-to-have
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
- ai-native userDefine rules that trigger actions automatically on eventsmust-have
- ai-native userPerform bulk operations across many items at onceshould-have
- ai-native userSchedule recurring jobs or workflowsshould-have
- ai-native userVersion, review, and roll back my automationsnice-to-have
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
- analystShare a research session or report with collaborators who can view or build on itshould-have
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
- researcherUpload my own PDFs or corpus and have the agent research over themshould-have
- researcherRun a systematic screening and extraction workflow across many papers with consistent criteriashould-have
- researcherSet up standing searches or alerts that surface new relevant sources as they appearnice-to-have
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
- ai-native userExport all of my data in open formats and leavemust-have
- ai-native userSelf-host the core productmust-have
- ai-native userDo everything through the API that I can do in the UIshould-have
- ai-native userRead the product's source under an open licenseshould-have
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
- researcherUnderstand plan pricing and usage limits before committingshould-have
- researcherTry the product meaningfully on a free tier or trialnice-to-have
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
- ai-native userPrevent my data from being used to train AI modelsmust-have
- ai-native userChoose where my data is stored (region/residency)should-have
- ai-native userControl data retention and deletionshould-have
- ai-native userOpt out of telemetry and usage trackingshould-have
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
- analystGet a structured report with sections, tables, and a summary that I can share with stakeholdersmust-have
- researcherExport results to common formats, including documents, spreadsheets, and reference-manager filesnice-to-have
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
- researcherPose a research question and get an autonomous multi-step investigation, not just a single-pass summarymust-have
- analystStart a long research job that keeps working unattended and notifies me when the result is readyshould-have
- researcherSteer the depth, effort, and scope of a research run before or while it executesnice-to-have
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
- researcherSee citations for every substantive claim so I can verify it against the underlying sourcemust-have
- researcherSearch scholarly literature and primary sources, not just the open webshould-have
- analystSee where sources agree and disagree instead of a single unqualified answershould-have
Appendix: recorded probes
Hands-on probe recordings — transcripts/videos a human can replay, the strongest evidence tier. Watch them at https://ultrametric.ai/productarena/proofs
- Gemini Notebook (NotebookLM)
curl -sL 'https://support.google.com/gemininotebook/answer/16322204?hl=en' | grep -oiE 'public notebooks?' | sort | uniq -cterminal · recorded 2026-09-14 · exit 0 - Gemini Notebook (NotebookLM)
curl -sL 'https://support.google.com/gemininotebook/answer/16215270?hl=en' | grep -oiE 'Deep Research|Fast Research|sources for your notebook' | sort | uniq -cterminal · recorded 2026-09-14 · exit 0
Cite as: ProductArena by Ultrametric Inc, AI Research Agents arena, rankings as of 2026-09-16 — https://ultrametric.ai/productarena/arena/ai-research-agents
License: © 2026 Ultrametric Inc. Brief quotation of individual verdicts, scores, or evidence excerpts is permitted with attribution to "ProductArena by Ultrametric Inc (ultrametric.ai/productarena)", as is use of the data to evaluate, contest, or contribute corrections. Bulk copying, redistribution, or use to build competing datasets requires prior written permission (see DATA-LICENSE in the repository).
No liability: rankings, verdicts, and scores are research outputs derived from the cited evidence at a point in time, provided "as is", without warranties. Ultrametric Inc accepts no responsibility for procurement, purchasing, or other decisions made in reliance on them — verify against the cited evidence before acting (https://ultrametric.ai/productarena/terms).