Skip to content

AI Research Agents Arena

AI Research Agents arenaBuyer checklist

Every requirement we judge ai research agents products against, as a ready-to-send RFP checklist — with each item's priority, why it matters, and how the top-ranked products score on it today.

43 requirements · 10 themes · verdicts for 6 products · updated 2026-09-16 · priorities mirror the story weights our scoring uses (methodology)

Procurement report →
Show the markdown export
# AI Research Agents — buyer checklist (RFP)

Derived from ProductArena's evidence-graded user-story taxonomy for AI Research Agents: 43 judged requirements. Priorities mirror story weights (3 = must-have, 2 = should-have, 1 = nice-to-have).

## Agenticness

- [ ] **[must-have]** Plug MCP servers into this product so it can use their tools
- [ ] **[must-have]** Connect an agent via an official MCP server
- [ ] **[must-have]** Drive the product through a documented public API
- [ ] **[must-have]** Delegate tasks to a built-in AI assistant inside the product
- [ ] **[should-have]** Point an agent at llms.txt or agent-oriented docs
- [ ] **[should-have]** Run the product headlessly / in CI for automation
- [ ] **[should-have]** Use an official CLI
- [ ] **[should-have]** Issue scoped/least-privilege API credentials for an agent
- [ ] **[should-have]** Build against official SDKs
- [ ] **[should-have]** Subscribe to events via webhooks
- [ ] **[should-have]** Get AI-generated insights and suggestions from my data inside the product
- [ ] **[should-have]** Set up automations that run autonomously in the background
- [ ] **[should-have]** Operate the product with natural-language commands
- [ ] **[should-have]** Explore an interactive API reference with runnable examples
- [ ] **[should-have]** Download a machine-readable API spec (OpenAPI or equivalent)
- [ ] **[should-have]** Rely on versioned APIs with a documented deprecation policy
- [ ] **[nice-to-have]** Test against a sandbox environment without touching production data

## Automation depth

- [ ] **[must-have]** Define rules that trigger actions automatically on events
- [ ] **[should-have]** Perform bulk operations across many items at once
- [ ] **[should-have]** Schedule recurring jobs or workflows
- [ ] **[nice-to-have]** Version, review, and roll back my automations

## Collaboration sharing

- [ ] **[should-have]** Share a research session or report with collaborators who can view or build on it

## Literature workflow

- [ ] **[should-have]** Upload my own PDFs or corpus and have the agent research over them
- [ ] **[should-have]** Run a systematic screening and extraction workflow across many papers with consistent criteria
- [ ] **[nice-to-have]** Set up standing searches or alerts that surface new relevant sources as they appear

## Openness

- [ ] **[must-have]** Export all of my data in open formats and leave
- [ ] **[must-have]** Self-host the core product
- [ ] **[should-have]** Do everything through the API that I can do in the UI
- [ ] **[should-have]** Read the product's source under an open license

## Pricing limits

- [ ] **[should-have]** Understand plan pricing and usage limits before committing
- [ ] **[nice-to-have]** Try the product meaningfully on a free tier or trial

## Privacy posture

- [ ] **[must-have]** Prevent my data from being used to train AI models
- [ ] **[should-have]** Choose where my data is stored (region/residency)
- [ ] **[should-have]** Control data retention and deletion
- [ ] **[should-have]** Opt out of telemetry and usage tracking

## Report output

- [ ] **[must-have]** Get a structured report with sections, tables, and a summary that I can share with stakeholders
- [ ] **[nice-to-have]** Export results to common formats, including documents, spreadsheets, and reference-manager files

## Research depth

- [ ] **[must-have]** Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- [ ] **[should-have]** Start a long research job that keeps working unattended and notifies me when the result is ready
- [ ] **[nice-to-have]** Steer the depth, effort, and scope of a research run before or while it executes

## Source quality

- [ ] **[must-have]** See citations for every substantive claim so I can verify it against the underlying source
- [ ] **[should-have]** Search scholarly literature and primary sources, not just the open web
- [ ] **[should-have]** See where sources agree and disagree instead of a single unqualified answer

---

Source: https://ultrametric.ai/productarena/arena/ai-research-agents (evidence-graded verdicts for 6 products) · methodology: https://ultrametric.ai/productarena/methodology

Chips show the top 5 ranked products' current verdict on each requirement — ✓ full · ~ partial · ! disputed · — none · n/a not applicable.

Agenticness — how well agents can access and operate the productAgenticness· 17 items

How well agents can access and operate the product

Automation depth — how much of the product can run unattendedAutomation depth· 4 items

How much of the product can run unattended

Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing· 1 item

Stories about collaboration sharing in this arena

Literature workflow — stories about literature workflow in this arenaLiterature workflow· 3 items

Stories about literature workflow in this arena

Openness — open source, data portability, and self-hosting storiesOpenness· 4 items

Open source, data portability, and self-hosting stories

Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits· 2 items

Free-tier ceilings, usage caps, and rate limits before you have to pay

Privacy posture — data-handling and privacy storiesPrivacy posture· 4 items

Data-handling and privacy stories

Report output — stories about report output in this arenaReport output· 2 items

Stories about report output in this arena

Research depth — stories about research depth in this arenaResearch depth· 3 items

Stories about research depth in this arena

Source quality — stories about source quality in this arenaSource quality· 3 items

Stories about source quality in this arena

Full evidence behind every verdict lives on the arena page and each product page — chips above deep-link straight to the judged story.