AI Research Agents arenaAI Research Agents
Autonomous research agents and literature-review products that plan searches, read sources, and synthesize evidence-backed reports, judged on research depth, citation quality, and how much real research work they can take on unattended.
43 user stories · 258 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Consensus ↗ | 37/100 | 52/100 | 0/100 | 22/100 | 33/100 | 8/27 verified · 2 disputed | 40/100 integrity | |||
| 2 | free-tier vs Elicit ↗ | 17/100 | 57/100 | 0/100 | 14/100 | 18/100 | 4/15 verified | 40/100 integrity | |||
| 3 | 31/100 | 52/100 | 0/100 | 9/100 | 0/100 | 6/17 verified | 0/100 integrity | ||||
| 4 | 17/100 | 40/100 | 0/100 | 12/100 | 18/100 | 11/17 verified · 1 disputed | 50/100 integrity | ||||
| 5 | subscription-flat vs Elicit ↗ | 0/100 | 59/100 | 0/100 | 0/100 | 0/100 | 0/13 verified | 0/100 integrity | |||
| 6 | free-tier vs Elicit ↗ | 0/100 | 34/100 | untested | 11/100 | 0/100 | 6/14 verified · 1 disputed | 20/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 8/10 | fullT 8/10 | fullT 7/10 | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | partialC 6/10 | partialC 4/10 | fullC 7/10 | partialT 2/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | none 0/10 | n/a | n/a | none 0/10 | n/a | none 0/10 |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | fullC 8/10 | none 0/10 | none 0/10 | fullT 8/10 | n/a | n/a |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | fullT 8/10 | partialT 5/10 | fullT 7/10 | partialT 5/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a | n/a |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | partialC 6/10 | none 0/10 | partialT 6/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Agentic features
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | disputedD 5/10 | fullC 8/10 | fullT 7/10 | fullX 8/10 | fullC 8/10 | fullX 8/10 |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | partialC 5/10 | none 0/10 | none 0/10 | partialC 3/10 | partialC 6/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | fullX 8/10 | fullT 7/10 | fullC 8/10 | none 0/10 | fullC 8/10 | partialX 4/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | fullX 7/10 | fullC 7/10 | partialC 7/10 | fullT 8/10 | partialC 5/10 | partialX 6/10 |
Api quality
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a | n/a |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | n/a | n/a | n/a | n/a | n/a | n/a |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | fullC 8/10 | partialC 6/10 | none 0/10 | partialC 6/10 | none 0/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | partialC 3/10 | n/a | n/a | partialC 3/10 | n/a | none 0/10 |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | partialC 4/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | none 0/10 | n/a | n/a | n/a | none 0/10 | n/a |
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Sharing
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Collaboration sharing — stories about collaboration sharing in this arenaShare a research session or report with collaborators who can view or build on it | analyst | fullC 8/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 7/10 |
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Alerts
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Literature workflow — stories about literature workflow in this arenaSet up standing searches or alerts that surface new relevant sources as they appear | researcher | fullC 7/10 | none 0/10 | none 0/10 | partialC 6/10 | none 0/10 | none 0/10 |
Corpus
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Literature workflow — stories about literature workflow in this arenaUpload my own PDFs or corpus and have the agent research over them | researcher | fullX 7/10 | partialC 6/10 | none 0/10 | none 0/10 | none 0/10 | fullX 8/10 |
Reviews
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Literature workflow — stories about literature workflow in this arenaRun a systematic screening and extraction workflow across many papers with consistent criteria | researcher | fullX 8/10 | partialC 6/10 | partialT 5/10 | partialX 6/10 | none 0/10 | partialX 3/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialC 4/10 | partialC 6/10 | partialT 5/10 | partialT 5/10 | none 0/10 | none 0/10 |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | partialC 6/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | partialC 3/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | none 0/10 | n/a | none 0/10 | n/a | none 0/10 | n/a |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | n/a | n/a | n/a | n/a | n/a | n/a |
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Pricing
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payTry the product meaningfully on a free tier or trial | researcher | none 0/10 | none 0/10 | fullC 7/10 | disputedD 4/10 | none 0/10 | partialX 4/10 |
| Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payUnderstand plan pricing and usage limits before committing | researcher | partialC 3/10 | none 0/10 | partialC 4/10 | none 0/10 | partialC 4/10 | partialC 3/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Report output — stories about report output in this arenaReport output
Reports
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Report output — stories about report output in this arenaExport results to common formats, including documents, spreadsheets, and reference-manager files | researcher | fullC 9/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | partialC 5/10 |
| Report output — stories about report output in this arenaGet a structured report with sections, tables, and a summary that I can share with stakeholders | analyst | fullC 8/10 | partialC 5/10 | partialC 6/10 | partialC 5/10 | fullC 7/10 | partialC 6/10 |
Research depth — stories about research depth in this arenaResearch depth
Agent runs
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Research depth — stories about research depth in this arenaPose a research question and get an autonomous multi-step investigation, not just a single-pass summary | researcher | partialX 6/10 | partialC 5/10 | fullC 8/10 | fullX 8/10 | fullC 8/10 | none 0/10 |
| Research depth — stories about research depth in this arenaStart a long research job that keeps working unattended and notifies me when the result is ready | analyst | partialC 4/10 | none 0/10 | partialC 5/10 | partialX 5/10 | partialC 6/10 | none 0/10 |
| Research depth — stories about research depth in this arenaSteer the depth, effort, and scope of a research run before or while it executes | researcher | partialC 7/10 | none 0/10 | none 0/10 | none 0/10 | partialC 4/10 | partialC 4/10 |
Source quality — stories about source quality in this arenaSource quality
Citations
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Source quality — stories about source quality in this arenaSee citations for every substantive claim so I can verify it against the underlying source | researcher | disputedD 5/10 | fullC 8/10 | fullC 7/10 | fullX 7/10 | partialC 5/10 | fullC 9/10 |
Corpus
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Source quality — stories about source quality in this arenaSearch scholarly literature and primary sources, not just the open web | researcher | fullC 9/10 | fullT 9/10 | fullC 8/10 | fullX 8/10 | partialC 6/10 | disputedD 3/10 |
Synthesis
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Source quality — stories about source quality in this arenaSee where sources agree and disagree instead of a single unqualified answer | analyst | partialX 5/10 | fullC 8/10 | partialC 5/10 | none 0/10 | partialC 4/10 | none 0/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.
AI Assistants arenaAI Assistants
9 products · leader: ChatGPT
AI Coding Agents arenaAI Coding Agents
13 products · leader: OpenCode
Web Scraping APIs arenaWeb Scraping APIs
8 products · leader: Apify
shares: Pricing limits
Edge & App Platforms arenaEdge & App Platforms
6 products · leader: Render
shares: Pricing limits