AI Research Agents arenaBuyer checklist
Every requirement we judge ai research agents products against, as a ready-to-send RFP checklist — with each item's priority, why it matters, and how the top-ranked products score on it today.
43 requirements · 10 themes · verdicts for 6 products · updated 2026-09-16 · priorities mirror the story weights our scoring uses (methodology)
Show the markdown export
# AI Research Agents — buyer checklist (RFP) Derived from ProductArena's evidence-graded user-story taxonomy for AI Research Agents: 43 judged requirements. Priorities mirror story weights (3 = must-have, 2 = should-have, 1 = nice-to-have). ## Agenticness - [ ] **[must-have]** Plug MCP servers into this product so it can use their tools - [ ] **[must-have]** Connect an agent via an official MCP server - [ ] **[must-have]** Drive the product through a documented public API - [ ] **[must-have]** Delegate tasks to a built-in AI assistant inside the product - [ ] **[should-have]** Point an agent at llms.txt or agent-oriented docs - [ ] **[should-have]** Run the product headlessly / in CI for automation - [ ] **[should-have]** Use an official CLI - [ ] **[should-have]** Issue scoped/least-privilege API credentials for an agent - [ ] **[should-have]** Build against official SDKs - [ ] **[should-have]** Subscribe to events via webhooks - [ ] **[should-have]** Get AI-generated insights and suggestions from my data inside the product - [ ] **[should-have]** Set up automations that run autonomously in the background - [ ] **[should-have]** Operate the product with natural-language commands - [ ] **[should-have]** Explore an interactive API reference with runnable examples - [ ] **[should-have]** Download a machine-readable API spec (OpenAPI or equivalent) - [ ] **[should-have]** Rely on versioned APIs with a documented deprecation policy - [ ] **[nice-to-have]** Test against a sandbox environment without touching production data ## Automation depth - [ ] **[must-have]** Define rules that trigger actions automatically on events - [ ] **[should-have]** Perform bulk operations across many items at once - [ ] **[should-have]** Schedule recurring jobs or workflows - [ ] **[nice-to-have]** Version, review, and roll back my automations ## Collaboration sharing - [ ] **[should-have]** Share a research session or report with collaborators who can view or build on it ## Literature workflow - [ ] **[should-have]** Upload my own PDFs or corpus and have the agent research over them - [ ] **[should-have]** Run a systematic screening and extraction workflow across many papers with consistent criteria - [ ] **[nice-to-have]** Set up standing searches or alerts that surface new relevant sources as they appear ## Openness - [ ] **[must-have]** Export all of my data in open formats and leave - [ ] **[must-have]** Self-host the core product - [ ] **[should-have]** Do everything through the API that I can do in the UI - [ ] **[should-have]** Read the product's source under an open license ## Pricing limits - [ ] **[should-have]** Understand plan pricing and usage limits before committing - [ ] **[nice-to-have]** Try the product meaningfully on a free tier or trial ## Privacy posture - [ ] **[must-have]** Prevent my data from being used to train AI models - [ ] **[should-have]** Choose where my data is stored (region/residency) - [ ] **[should-have]** Control data retention and deletion - [ ] **[should-have]** Opt out of telemetry and usage tracking ## Report output - [ ] **[must-have]** Get a structured report with sections, tables, and a summary that I can share with stakeholders - [ ] **[nice-to-have]** Export results to common formats, including documents, spreadsheets, and reference-manager files ## Research depth - [ ] **[must-have]** Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary - [ ] **[should-have]** Start a long research job that keeps working unattended and notifies me when the result is ready - [ ] **[nice-to-have]** Steer the depth, effort, and scope of a research run before or while it executes ## Source quality - [ ] **[must-have]** See citations for every substantive claim so I can verify it against the underlying source - [ ] **[should-have]** Search scholarly literature and primary sources, not just the open web - [ ] **[should-have]** See where sources agree and disagree instead of a single unqualified answer --- Source: https://ultrametric.ai/productarena/arena/ai-research-agents (evidence-graded verdicts for 6 products) · methodology: https://ultrametric.ai/productarena/methodology
Chips show the top 5 ranked products' current verdict on each requirement — ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness· 17 items
How well agents can access and operate the product
- must-have
ai-native userPlug MCP servers into this product so it can use their tools
Core requirement — weighs 3× in arena scoring · no product fully delivers this yet
- must-have
ai-native userConnect an agent via an official MCP server
Core requirement — weighs 3× in arena scoring · 2 of 4 products fully deliver this today
- must-have
ai-native userDrive the product through a documented public API
Core requirement — weighs 3× in arena scoring · 2 of 6 products fully deliver this today
- must-have
ai-native userDelegate tasks to a built-in AI assistant inside the product
Core requirement — weighs 3× in arena scoring · 4 of 6 products fully deliver this today
- should-have
ai-native userPoint an agent at llms.txt or agent-oriented docs
Important, not disqualifying — weighs 2× in arena scoring · 3 of 5 products fully deliver this today
- should-have
ai-native userRun the product headlessly / in CI for automation
Important, not disqualifying — weighs 2× in arena scoring · 1 of 6 products fully deliver this today
- should-have
ai-native userUse an official CLI
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userIssue scoped/least-privilege API credentials for an agent
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userBuild against official SDKs
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userSubscribe to events via webhooks
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userGet AI-generated insights and suggestions from my data inside the product
Important, not disqualifying — weighs 2× in arena scoring · 5 of 6 products fully deliver this today
- should-have
ai-native userSet up automations that run autonomously in the background
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userOperate the product with natural-language commands
Important, not disqualifying — weighs 2× in arena scoring · 3 of 6 products fully deliver this today
- should-have
ai-native userExplore an interactive API reference with runnable examples
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userRely on versioned APIs with a documented deprecation policy
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- nice-to-have
ai-native userTest against a sandbox environment without touching production data
Differentiator, not a dealbreaker — weighs 1× in arena scoring
Automation depth — how much of the product can run unattendedAutomation depth· 4 items
How much of the product can run unattended
- must-have
ai-native userDefine rules that trigger actions automatically on events
Core requirement — weighs 3× in arena scoring · no product fully delivers this yet
- should-have
ai-native userPerform bulk operations across many items at once
Important, not disqualifying — weighs 2× in arena scoring · 1 of 6 products fully deliver this today
- should-have
ai-native userSchedule recurring jobs or workflows
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- nice-to-have
ai-native userVersion, review, and roll back my automations
Differentiator, not a dealbreaker — weighs 1× in arena scoring · no product fully delivers this yet
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing· 1 item
Stories about collaboration sharing in this arena
- should-have
analystShare a research session or report with collaborators who can view or build on it
Important, not disqualifying — weighs 2× in arena scoring · 1 of 6 products fully deliver this today
Literature workflow — stories about literature workflow in this arenaLiterature workflow· 3 items
Stories about literature workflow in this arena
- should-have
researcherUpload my own PDFs or corpus and have the agent research over them
Important, not disqualifying — weighs 2× in arena scoring · 2 of 6 products fully deliver this today
- should-have
researcherRun a systematic screening and extraction workflow across many papers with consistent criteria
Important, not disqualifying — weighs 2× in arena scoring · 1 of 6 products fully deliver this today
- nice-to-have
researcherSet up standing searches or alerts that surface new relevant sources as they appear
Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 6 products fully deliver this today
Openness — open source, data portability, and self-hosting storiesOpenness· 4 items
Open source, data portability, and self-hosting stories
- must-have
ai-native userExport all of my data in open formats and leave
Core requirement — weighs 3× in arena scoring · no product fully delivers this yet
- must-have
ai-native userSelf-host the core product
Core requirement — weighs 3× in arena scoring
- should-have
ai-native userDo everything through the API that I can do in the UI
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userRead the product's source under an open license
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits· 2 items
Free-tier ceilings, usage caps, and rate limits before you have to pay
- should-have
researcherUnderstand plan pricing and usage limits before committing
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- nice-to-have
researcherTry the product meaningfully on a free tier or trial
Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 6 products fully deliver this today
Privacy posture — data-handling and privacy storiesPrivacy posture· 4 items
Data-handling and privacy stories
- must-have
ai-native userPrevent my data from being used to train AI models
Core requirement — weighs 3× in arena scoring · no product fully delivers this yet
- should-have
ai-native userChoose where my data is stored (region/residency)
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userControl data retention and deletion
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- should-have
ai-native userOpt out of telemetry and usage tracking
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
Report output — stories about report output in this arenaReport output· 2 items
Stories about report output in this arena
- must-have
analystGet a structured report with sections, tables, and a summary that I can share with stakeholders
Core requirement — weighs 3× in arena scoring · 2 of 6 products fully deliver this today
- nice-to-have
researcherExport results to common formats, including documents, spreadsheets, and reference-manager files
Differentiator, not a dealbreaker — weighs 1× in arena scoring · 1 of 6 products fully deliver this today
Research depth — stories about research depth in this arenaResearch depth· 3 items
Stories about research depth in this arena
- must-have
researcherPose a research question and get an autonomous multi-step investigation, not just a single-pass summary
Core requirement — weighs 3× in arena scoring · 3 of 6 products fully deliver this today
- should-have
analystStart a long research job that keeps working unattended and notifies me when the result is ready
Important, not disqualifying — weighs 2× in arena scoring · no product fully delivers this yet
- nice-to-have
researcherSteer the depth, effort, and scope of a research run before or while it executes
Differentiator, not a dealbreaker — weighs 1× in arena scoring · no product fully delivers this yet
Source quality — stories about source quality in this arenaSource quality· 3 items
Stories about source quality in this arena
- must-have
researcherSee citations for every substantive claim so I can verify it against the underlying source
Core requirement — weighs 3× in arena scoring · 4 of 6 products fully deliver this today
- should-have
researcherSearch scholarly literature and primary sources, not just the open web
Important, not disqualifying — weighs 2× in arena scoring · 4 of 6 products fully deliver this today
- should-have
analystSee where sources agree and disagree instead of a single unqualified answer
Important, not disqualifying — weighs 2× in arena scoring · 1 of 6 products fully deliver this today
Full evidence behind every verdict lives on the arena page and each product page — chips above deep-link straight to the judged story.