Consensus vs Sakana Marlin
Sakana Marlin
Sakana AI
Consensus wins · 11–8 (15 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to ConsensusA live probe confirms Consensus serves an llms.txt file at its root (HTTP 200) with a structured summary of the product, directly enabling agents to be pointed at agent-oriented docs. Missing for 10: broader agent-oriented doc formats (e.g. .md endpoints) return 404, and no independent third-party confirmation of llms.txt usage exists.
Sakana Marlinnone0/10The llms.txt probe returned a 404, and no other evidence shows agent-oriented docs (like an API spec or agent-readable documentation) for Marlin; the evidence pack is entirely marketing copy about the product's research capabilities, not machine-readable docs.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to ConsensusConsensus offers an API for integrating its search into custom workflows and running automated searches (consensus-docs-1, consensus-docs-15), which implies some programmatic/headless usability. However, there is no explicit documentation of CI integration, headless execution modes, CLI tooling, or automation pipeline examples. Missing for 10: CI/CD integration examples, headless mode documentation, CLI or SDK for automation, and independent evidence of running in automated pipelines.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
Sakana Marlinnone0/10Marlin is presented as a research-report generation web product with a UI and pay-per-use credits, with no CLI, API, SDK, or webhook documentation for headless/CI usage; probes for llms.txt and OpenAPI specs both returned 404, indicating no programmatic interface is exposed.
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
- [claimed-docs] “Add your card. Start right away.”
ai-native userUse an official CLI
weight 2 · round drawnConsensusnone0/10The evidence pack documents a REST API and an MCP server (consensus-docs-16) but no official command-line interface is mentioned anywhere in the docs or probes. Missing for 10: any mention of a CLI tool, CLI installation instructions, or CLI command reference.
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
Sakana Marlinnone0/10No evidence of an official CLI; Marlin appears to be a web-based research tool with a pay-per-credit UI, and probes for API/llms.txt endpoints returned 404s, suggesting no developer-facing interface is exposed.
ai-native userDrive the product through a documented public API
weight 3 · round to ConsensusConsensus advertises a documented API for integrating citations and running automated searches into custom workflows, and its site provides an llms.txt for AI-agent discovery, showing basic public-API and agent-friendliness. However, the evidence pack only shows marketing/landing pages, not actual API reference documentation, authentication, endpoints, or example requests/responses, and there's no independent or hands-on corroboration that the API works as described. Missing for 10: full API reference/spec details, code/SDK examples, and independent verification of API usage.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
Sakana Marlinnone0/10No documented public API is evidenced; probes for llms.txt and OpenAPI/swagger endpoints all returned 404s, and all other evidence describes the product's research capabilities, not a programmatic interface.
ai-native userBuild against official SDKs
weight 2 · round drawnConsensusnone0/10Consensus documents an API for integration (consensus-docs-1, consensus-docs-15) but no evidence pack item mentions official SDKs (Python, JS, etc.) or client libraries for AI-native development — only the raw API and llms.txt discovery file are shown.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
Sakana Marlinnone0/10No evidence of any official SDK, API, or developer library for Marlin; probes for llms.txt and OpenAPI specs both returned 404, and all docs describe an end-user research product with no mention of programmatic/SDK access.
ai-native userSubscribe to events via webhooks
weight 2 · round drawnConsensusnone0/10No evidence anywhere in the pack mentions webhooks or event subscriptions; Consensus's API/MCP surface is described only as REST retrieval/synthesis, not event-driven push notifications.
Sakana Marlinnone0/10No evidence of a webhook or event subscription mechanism; probes for API/OpenAPI specs returned 404s and docs focus only on research report generation. Missing for 10: any webhook documentation, event subscription API, or callback mechanism.
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round drawnConsensus generates AI-driven synthesis, summaries, the Consensus Meter, PICO extraction, and literature review synthesis directly from the papers in its corpus/library, with citations tracing insights back to sources. This is core native functionality (not a bolt-on), covering search, synthesis, and structured insight generation. Missing for 10: independent/hands-on third-party verification of insight quality beyond vendor docs.
- [claimed-docs] “It searches over 200 million academic papers and uses language models to help you find, understand, and synthesize the literature faster.”
- [claimed-docs] “Every response includes citations, so you can trace each insight back to the original source.”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “Extracted population, intervention, comparator, and outcome (PICO) where applicable”
- [claimed-docs] “Consensus is an AI-powered research engine built to speed up literature reviews. Search, screen, extract, and synthesize evidence faster—whi…”
- [claimed-docs] “The Consensus Library brings your entire research library into one searchable, AI-powered workspace.”
Marlin autonomously researches user-provided topics, mapping causal relationships, comparing hypotheses, and generating structured strategic insights and reports with citations, going beyond simple summarization. Missing for 10: independent/hands-on third-party validation beyond vendor-curated testimonials, and no visibility into underlying data/insight quality benchmarks.
- [claimed-docs] “Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into structured strategic options.”
- [claimed-docs] “It does more than summarize. Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into st…”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to Sakana MarlinConsensusnone0/10Consensus is a research search/synthesis engine with an API and MCP server for on-demand retrieval, but there is no evidence of scheduled or event-triggered automations that run autonomously in the background without user invocation. Missing for 10: any scheduling/trigger mechanism, background job execution, or autonomous recurring workflow capability.
Marlin's docs clearly describe a single research task running autonomously for up to ~8 hours without further human input once a topic is set (docs-9, docs-10, docs-16), which matches the 'runs in background autonomously' idea. However, this is a one-shot session, not a recurring/scheduled automation you configure and forget — there's no evidence of triggers, schedules, or multi-run automation management typical of 'set up automations.' Missing for 10: scheduled/recurring automation setup, background job management UI, independent hands-on corroboration of unattended runtime.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round to Sakana MarlinConsensus ships a built-in "Research Agent" that chains citation crawling, DOI lookup, author search and similar-paper discovery on top of its search engine, and its core AI assistant performs search, screen, extract, and synthesize workflows with cited answers — this is essentially delegating research tasks to an in-product AI assistant. missing for 10: independent/hands-on validation of the agent's autonomy and reliability, and more detail on the scope/limits of delegable tasks beyond literature discovery.
- [claimed-docs] “Citation crawling, DOI lookup, author search, similar papers, and more - chained together on top of the worlds best academic search engine.”
- [claimed-docs] “Search, screen, extract, and synthesize evidence faster—while keeping full transparency and scholarly rigor.”
- [claimed-docs] “Consensus is an AI-powered research engine built to speed up literature reviews. Search, screen, extract, and synthesize evidence faster—whi…”
- [claimed-docs] “It searches over 200 million academic papers and uses language models to help you find, understand, and synthesize the literature faster.”
- [claimed-docs] “Every response includes citations, so you can trace each insight back to the original source.”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
Marlin is explicitly designed as a built-in AI agent that users delegate entire research/strategy tasks to, working autonomously for hours with minimal human input beyond initial framing (docs-1, docs-9, docs-10, docs-16, docs-21). This directly matches the story of delegating tasks to a built-in assistant within the product. Missing for 10: independent/hands-on verification beyond vendor testimonials, and detail on interactive control while a task is delegated.
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
ai-native userOperate the product with natural-language commands
weight 2 · round to ConsensusConsensus's core interaction model is natural-language research queries (search, synthesize, Consensus Meter for yes/no questions) rather than rigid query syntax, and it exposes this same NL-driven retrieval/synthesis surface via an MCP server and REST API for programmatic/agentic use. Missing for 10: independent hands-on evidence of natural-language command execution quality, and no detailed example transcripts showing complex multi-step NL commands being interpreted.
- [claimed-docs] “It searches over 200 million academic papers and uses language models to help you find, understand, and synthesize the literature faster.”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “Think of Consensus as an AI-native alternative to Google Scholar with a more-refined corpus.”
Marlin is initiated by giving it a research topic and a brief natural-language exchange to set direction (docs-10), suggesting natural-language input drives its operation, but there is no evidence of a broader natural-language command interface (e.g., chat-style control, follow-up instructions, or command syntax) beyond initial topic-setting. Missing for 10: documentation of ongoing NL command/control during execution, examples of varied NL commands, independent/hands-on confirmation of NL interaction quality.
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
Api quality
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnConsensusnone0/10Consensus documents a REST API and MCP server (consensus-docs-15, consensus-docs-16) but no evidence pack item mentions an OpenAPI spec, Swagger file, or any downloadable machine-readable API schema; the llms.txt probe returns a plain-text description, not an API spec. missing for 10: OpenAPI/Swagger file, machine-readable schema download link, independent confirmation of spec availability.
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
Sakana Marlinnone0/10Probes for OpenAPI/swagger spec files and llms.txt all returned 404, and no documentation mentions a machine-readable API spec. Missing for 10: any published OpenAPI/Swagger file, API reference docs, or equivalent machine-readable spec.
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnConsensusnone0/10There's an API and MCP server mentioned, but no evidence of API versioning scheme or a documented deprecation policy anywhere in the pack. missing for 10: versioning scheme documentation, deprecation policy, changelog/migration guides.
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
Sakana Marlinnone0/10No evidence of any public API, versioning scheme, or deprecation policy; probes for OpenAPI/llms.txt endpoints returned 404s, and all docs describe the research product itself, not a developer API. Missing for 10: any API documentation, versioning scheme, deprecation policy.
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to ConsensusDocs show bulk-style capabilities: one-click import of thousands of papers into a library, an API/MCP server for automated bulk searches, and Deep Searches across many studies — supporting bulk operations for an AI-native/automation persona. missing for 10: independent/hands-on verification of bulk API throughput or rate limits, explicit batch-processing endpoints (e.g., bulk extract/export across many items in one call), and any third-party confirmation of scale performance.
- [claimed-docs] “Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
- [claimed-docs] “Turn your library into a research engine. Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
- [claimed-docs] “Deep Searches (more comprehensive Lit Reviews across many studies)”
- [claimed-docs] “Import from Zotero ... or import from BibTex, PDF, or RIS”
ai-native userSchedule recurring jobs or workflows
weight 2 · round drawnConsensusnone0/10No evidence in the pack mentions scheduling, recurring jobs, alerts, or automated re-running of searches/workflows over time; the API and MCP server are described as on-demand retrieval/synthesis interfaces, not schedulable automation. missing for 10: any scheduling/cron feature, recurring alert or saved-search re-run capability, or workflow automation trigger.
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
Sakana Marlinnone0/10Sakana Marlin is a single-run autonomous research/report tool triggered by a user topic; no evidence of scheduling, recurrence, cron-like triggers, or workflow automation for repeated jobs. Probes also show no API/OpenAPI surface that could support scheduled invocation.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
Sharing
analystShare a research session or report with collaborators who can view or build on it
weight 2 · round drawnConsensusnone0/10No evidence pack items mention sharing sessions, reports, collaborators, team accounts, or collaborative viewing/editing features—only individual research, library import, and API/agent capabilities are documented. missing for 10: any mention of sharing links, collaborator invites, team workspaces, or comment/build-on functionality.
Sakana Marlinnone0/10No evidence describes any collaboration or sharing features—no mention of shared workspaces, links, comments, or multi-user access to reports/sessions; the pack only covers autonomous research generation, pricing tiers, and API probes returning 404s.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
Alerts
researcherSet up standing searches or alerts that surface new relevant sources as they appear
weight 1 · round drawnConsensusnone0/10No evidence of standing searches, saved-search alerts, or notification features when new relevant papers appear; the evidence pack covers search, library import, citation graph, API/MCP retrieval, and literature review synthesis but nothing about recurring/alert-based monitoring of new sources.
Corpus
researcherUpload my own PDFs or corpus and have the agent research over them
weight 2 · round to ConsensusConsensus's Library feature explicitly supports importing PDFs, BibTeX, RIS, and Zotero corpora and turns them into a 'searchable, AI-powered workspace' for finding gaps and using the collection, which matches the story's upload+research intent. However, evidence doesn't detail how deeply the AI synthesis/agent features (Meter, PICO extraction, literature review synthesis) operate specifically over a user's uploaded corpus versus the general 200M-paper index. Missing for 10: explicit documentation of agent-style synthesis/Q&A running directly over an uploaded private corpus, and independent/hands-on confirmation of this workflow.
- [claimed-docs] “Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
- [claimed-docs] “Turn your library into a research engine. Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
- [claimed-docs] “Import from Zotero ... or import from BibTex, PDF, or RIS”
- [claimed-docs] “The Consensus Library brings your entire research library into one searchable, AI-powered workspace.”
- [claimed-docs] “Reference managers are great at saving papers — not so great at helping you use them. The Consensus Library brings your entire research libr…”
Sakana Marlinnone0/10Evidence describes Marlin as an autonomous web-research agent that gathers information via web navigation and generates reports, but there is no mention of uploading a user's own PDFs or corpus for the agent to research over. Missing for 10: any document/file upload feature, corpus ingestion, or evidence of researching over user-supplied materials rather than open web sources.
Reviews
researcherRun a systematic screening and extraction workflow across many papers with consistent criteria
weight 2 · round to ConsensusConsensus offers literature-review features (search, screen, extract, synthesize per docs-8/12), library import at scale, PICO extraction, and filters by study type/year/discipline that support systematic screening with consistent criteria. However, there is no evidence of documented inter-rater reliability, exportable screening decision logs, or PRISMA-style workflow tracking that a systematic review would require. missing for 10: evidence of structured screening criteria configuration/audit trail, PRISMA-compliant workflow support, independent validation of extraction consistency across large paper sets.
- [claimed-docs] “Search, screen, extract, and synthesize evidence faster—while keeping full transparency and scholarly rigor.”
- [claimed-docs] “Consensus is an AI-powered research engine built to speed up literature reviews. Search, screen, extract, and synthesize evidence faster—whi…”
- [claimed-docs] “Turn your library into a research engine. Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
- [claimed-docs] “Filters allow narrowing by study type (RCT, meta-analysis, systematic review, observational), publication year, journal, open-access status,…”
- [claimed-docs] “Extracted population, intervention, comparator, and outcome (PICO) where applicable”
- [claimed-docs] “Deep Searches (more comprehensive Lit Reviews across many studies)”
Sakana Marlinnone0/10Marlin is positioned as an autonomous business/market strategy research agent producing single deep-dive reports, not as a tool for systematic multi-paper screening/extraction with consistent criteria (a literature-review workflow). No evidence describes handling many papers, applying consistent inclusion/extraction criteria, or batch processing across a corpus — the described unit of work is one topic producing one report.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
- [claimed-docs] “putting Marlin to work on real tasks such as strategy formulation, market research, risk analysis, and competitive analysis”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to ConsensusThe API/MCP server is documented to expose 'the same retrieval and synthesis surface that powers the web app' (consensus-docs-16), and supports automated search (consensus-docs-15), suggesting broad parity for core search/synthesis. However, UI-specific workflows like Library import/reference management (Zotero/BibTeX/RIS import), Citation Graph, and Consensus Meter visualizations are not explicitly confirmed as API-accessible endpoints. Missing for 10: explicit API documentation confirming library management, citation graph, and meter features are callable via API, plus independent/hands-on verification of claimed parity.
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “The Consensus Library brings your entire research library into one searchable, AI-powered workspace.”
- [claimed-docs] “Import from Zotero ... or import from BibTex, PDF, or RIS”
- [claimed-docs] “The Consensus Citation Graph turns a single seed paper into a complete map of the work that built it, the work it inspired, and the studies …”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
Sakana Marlinnone0/10No evidence of any public API for Sakana Marlin; probes for llms.txt and OpenAPI spec both returned 404, and all documentation describes only UI/credit-based access.
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
- [claimed-docs] “Add your card. Start right away.”
ai-native userExport all of my data in open formats and leave
weight 3 · round drawnConsensusnone0/10Evidence shows only import capabilities (Zotero, BibTeX, PDF, RIS) into the Consensus Library, with no mention of exporting a user's library, annotations, or account data back out in open formats. Data portability/export is a fair axis for a reference-manager-style product, but no evidence supports it.
- [claimed-docs] “Import from Zotero ... or import from BibTex, PDF, or RIS”
- [claimed-docs] “Import from Zotero”
- [claimed-docs] “The Consensus Library brings your entire research library into one searchable, AI-powered workspace.”
- [claimed-docs] “Turn your library into a research engine. Import thousands of papers in one click - then search, find gaps, and put your collection to work.”
Sakana Marlinnone0/10No evidence of data export, open-format download, or account portability features; probes for llms.txt and OpenAPI both returned 404, and docs only describe generated reports/slides, not export of underlying user data.
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Pricing
researcherTry the product meaningfully on a free tier or trial
weight 1 · round drawnConsensusnone0/10The evidence pack references a pricing page (consensus-docs-17) but only quotes a single line about 'Deep Searches' feature tiering; there is no description of a free tier, trial period, usage caps, or sign-up-free access that a researcher could evaluate. No first-party or independent evidence confirms Consensus offers a meaningful free/trial experience.
- [claimed-docs] “Deep Searches (more comprehensive Lit Reviews across many studies)”
Sakana Marlinnone0/10Marlin is explicitly pay-per-use available only to paid Pro/Team/Enterprise plans, requires adding a card to start, and cancelling mid-run still consumes credits — there is no free tier or trial for researchers to test it meaningfully.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add your card. Start right away.”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
researcherUnderstand plan pricing and usage limits before committing
weight 2 · round to Sakana MarlinConsensusnone0/10No evidence pack items mention pricing plans, tiers, free/paid limits, or usage quotas — the pack is entirely about product features (citation graph, library, API capabilities). Absence of any pricing/limits documentation for an applicable axis yields none.
Marlin's docs mention it's offered as a pay-per-use add-on to Pro/Team/Enterprise plans, credits are consumed even if cancelled mid-run, and additional credits cost ¥98 each, giving a researcher some pricing signal. However, there's no concrete breakdown of how many credits a typical run consumes, no explicit usage caps, and no comparison table of plan tiers — missing for 10: full plan pricing table, credit-consumption-per-task estimates, explicit usage limits, independent/hands-on cost verification.
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
- [claimed-docs] “Add your card. Start right away.”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnConsensusnone0/10No evidence in the pack mentions data residency, regional storage options, or any data-location controls for Consensus; all evidence concerns search, citation, and library features. Missing for 10: any mention of region selection, data residency policy, or storage location controls.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnConsensusnone0/10No evidence pack item addresses data-training opt-out, privacy controls, or AI training data policies for Consensus; all citations concern search, library, and API features. missing for 10: any privacy policy statement, opt-out mechanism, or data usage/training disclosure.
ai-native userControl data retention and deletion
weight 2 · round drawnConsensusnone0/10No evidence in the pack addresses data retention policies, deletion controls, or privacy settings for user data/library content; all citations focus on search, citation, library, and API features. Missing for 10: any documentation on data retention windows, user-initiated deletion, export/erasure workflows, or privacy policy specifics.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnConsensusnone0/10No evidence in the pack mentions telemetry, usage tracking, or any opt-out/privacy settings for Consensus; all citations are about search, citation, and library features unrelated to telemetry controls.
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
Reports
researcherExport results to common formats, including documents, spreadsheets, and reference-manager files
weight 1 · round to Sakana MarlinConsensusnone0/10Evidence only documents importing papers into Consensus (from Zotero, BibTeX, PDF, RIS) but contains no mention of exporting results to documents, spreadsheets, or reference-manager formats. Missing for 10: any export-to-Word/PDF, export-to-CSV/spreadsheet, or export-to-Zotero/EndNote/BibTeX functionality.
- [claimed-docs] “Import from Zotero ... or import from BibTex, PDF, or RIS”
- [claimed-docs] “Reference managers are great at saving papers — not so great at helping you use them. The Consensus Library brings your entire research libr…”
- [claimed-docs] “Import from Zotero”
Marlin auto-generates full reports with references and PowerPoint slides (docs-3, docs-9, docs-20), covering the 'documents' part of the story, but there is no evidence of spreadsheet export or reference-manager file formats (e.g., BibTeX/RIS/EndNote) for citations. missing for 10: spreadsheet export, reference-manager file export (BibTeX/RIS/EndNote), independent corroboration of export formats.
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
analystGet a structured report with sections, tables, and a summary that I can share with stakeholders
weight 3 · round to Sakana MarlinConsensus offers 'Literature Review' and 'Deep Search' features that synthesize evidence across papers, extract structured fields like PICO, and provide citations—suggesting output with some structure and sourcing suitable for sharing. However, there's no explicit evidence of a polished 'report' format with distinct sections, tables, and an executive summary designed for stakeholder sharing (e.g., export to PDF/Word, formatted report templates). Missing for 10: explicit documentation of report formatting/export (sections, tables, summary), evidence of stakeholder-sharing features like PDF export or presentation-ready output, and independent confirmation of report quality.
- [claimed-docs] “Search, screen, extract, and synthesize evidence faster—while keeping full transparency and scholarly rigor.”
- [claimed-docs] “Consensus is an AI-powered research engine built to speed up literature reviews. Search, screen, extract, and synthesize evidence faster—whi…”
- [claimed-docs] “Deep Searches (more comprehensive Lit Reviews across many studies)”
- [claimed-docs] “Extracted population, intervention, comparator, and outcome (PICO) where applicable”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
Docs describe autonomous generation of a full structured report with main body, appendices, references, and presentation slides/executive summary — directly matching sections, tables (implied by structured strategic options), and summary needs for stakeholder sharing. Missing for 10: no independent/hands-on verification of table formatting or actual sample report shown, and no evidence beyond vendor marketing copy.
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
- [claimed-docs] “Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into structured strategic options.”
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
Agent runs
researcherPose a research question and get an autonomous multi-step investigation, not just a single-pass summary
weight 3 · round to Sakana MarlinConsensus advertises a 'Research Agent' that chains citation crawling, DOI lookup, author search, and similar-papers search on top of its search engine, plus a literature-review feature that searches, screens, extracts, and synthesizes evidence — both suggesting multi-step, not single-pass, investigation. However, evidence is limited to marketing feature pages with no walkthrough, example transcript, or independent corroboration of true autonomous multi-step reasoning over a posed question. Missing for 10: a documented end-to-end example of the agent autonomously chaining steps for a specific question, independent/hands-on verification, and detail on how far it goes without user intervention.
- [claimed-docs] “Citation crawling, DOI lookup, author search, similar papers, and more - chained together on top of the worlds best academic search engine.”
- [claimed-docs] “Search, screen, extract, and synthesize evidence faster—while keeping full transparency and scholarly rigor.”
- [claimed-docs] “Consensus is an AI-powered research engine built to speed up literature reviews. Search, screen, extract, and synthesize evidence faster—whi…”
Vendor docs describe autonomous multi-step research: forming hypotheses, gathering info, resolving contradictions, running for hours across thousands of cycles without further human input, producing a full structured report - directly matching the story. This is corroborated by beta-tester quotes praising depth beyond chat-based research tools, though all evidence is vendor-published/testimonial rather than independent hands-on verification. missing for 10: independent third-party evaluation or benchmark of the autonomous multi-step process, technical detail on how contradictions/hypotheses are actually verified
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “Many told us that Marlin was more practical at digging deeply into information than the chat-based research tools they had used before”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
analystStart a long research job that keeps working unattended and notifies me when the result is ready
weight 2 · round to Sakana MarlinConsensusnone0/10Evidence shows Deep Searches/Lit Reviews and a research agent chaining searches, but there is no mention of async job submission, background/unattended execution, or notification when a long-running job completes.
Vendor docs clearly describe long unattended autonomous research runs (~8 hours) producing full reports, which supports the core of the story, but there is no mention of a notification mechanism when results are ready and no independent/hands-on corroboration beyond marketing copy. missing for 10: evidence of a completion notification/alert feature, independent verification of unattended runtime and reliability, and API/technical docs confirming job-control (start/monitor/cancel) beyond the marketing page.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
researcherSteer the depth, effort, and scope of a research run before or while it executes
weight 1 · round to Sakana MarlinConsensusnone0/10The evidence pack describes search, citation graph, library, and research-agent features but nowhere mentions controls for adjusting depth, effort, or scope of a research run before or during execution — no parameters, modes, or configuration options are documented.
Marlin only allows a brief initial exchange to set direction before running fully autonomously for up to ~8 hours with no mid-run steering, and there's no documented control over depth/effort/scope parameters (e.g., report length, time budget, source breadth) beyond the initial topic framing; cancellation is possible but not adjustment. missing for 10: mid-execution steering controls, explicit depth/effort/scope parameters or settings, independent evidence of pre-run configurability beyond a 'brief exchange'.
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
Citations
researcherSee citations for every substantive claim so I can verify it against the underlying source
weight 3 · round to ConsensusConsensus documents that every AI-generated response includes citations tracing back to the original source paper, and features like the Consensus Meter classify individual papers (supporting/refuting) with traceable provenance, directly matching the researcher's need to verify claims against sources. Missing for 10: independent/hands-on verification of citation accuracy and completeness beyond vendor docs.
- [claimed-docs] “Every response includes citations, so you can trace each insight back to the original source.”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “It searches over 200 million academic papers and uses language models to help you find, understand, and synthesize the literature faster.”
Vendor docs claim Marlin generates 'fully referenced reports' grounded in primary sources with appendices and references, and a testimonial praises its higher-quality citations to primary vs secondary sources, suggesting citation support exists. However, there is no independent verification, no example of inline citation format, and no detail on how claims map to sources for auditability. Missing for 10: independent hands-on verification of citation accuracy, example output showing citation linking, and confirmation citations are traceable/clickable to primary sources.
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “他の生成AIと比較して、引用される情報の量と質が高く、二次情報ではなく一次情報を参照できている点に優位性を感じました。”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
Corpus
researcherSearch scholarly literature and primary sources, not just the open web
weight 2 · round to ConsensusConsensus is explicitly built as a scholarly-search engine over 200M+ academic papers, including full-text and paywalled content, positioned as an AI-native alternative to Google Scholar, with citation tracing back to original sources. Missing for 10: independent third-party verification of corpus quality/coverage beyond vendor claims.
- [claimed-docs] “Consensus analyzes the full text, including paywalled papers from major publishers, so you can find the most relevant papers.”
- [claimed-docs] “Think of Consensus as an AI-native alternative to Google Scholar with a more-refined corpus.”
- [claimed-docs] “It searches over 200 million academic papers and uses language models to help you find, understand, and synthesize the literature faster.”
- [claimed-docs] “Every response includes citations, so you can trace each insight back to the original source.”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
Vendor docs claim research is 'grounded strictly in primary sources' and a testimonial notes it references primary rather than secondary information compared to other generative AI tools, but there is no evidence of dedicated scholarly database access (e.g., PubMed, arXiv, JSTOR integration) or citation verification—only general web navigation is described. missing for 10: evidence of scholarly/academic database integration, independent verification of primary-source claim, and details on how it distinguishes scholarly vs open-web sources.
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “他の生成AIと比較して、引用される情報の量と質が高く、二次情報ではなく一次情報を参照できている点に優位性を感じました。”
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
Synthesis
analystSee where sources agree and disagree instead of a single unqualified answer
weight 2 · round to ConsensusThe Consensus Meter explicitly classifies each relevant paper as supporting, refuting, or mixed/inconclusive on a given question and displays the distribution, directly surfacing agreement/disagreement across sources rather than a single answer, and every response includes citations back to originals. Missing for 10: independent/hands-on corroboration of the Meter's accuracy and no worked example showing disagreement handling in practice.
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “The Consensus Meter is a visual aggregator that, for a yes/no/possibly question, classifies each relevant paper as supporting, refuting, or …”
- [claimed-docs] “Every response includes citations, so you can trace each insight back to the original source.”
Marketing copy claims Marlin 'resolves contradictions' and 'compares and evaluates multiple hypotheses' rather than merely aggregating, implying some handling of conflicting sources, but there is no evidence of a UI feature or report section that explicitly surfaces where sources agree/disagree to the analyst. Missing for 10: concrete example of a report showing conflicting source viewpoints, screenshot/description of how disagreement is presented, independent corroboration beyond vendor marketing.
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
Not comparable on these axes
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · not comparableConsensusn/aConsensus is a research/search product, not an AI agent; the evidence shows an API for integration but nothing about MCP server plug-in support to consume external tools. This axis (agent-side MCP client capability) is a category error for this type of product.
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableConsensusnone0/10Consensus is not an agent product itself, so the MCP-server axis applies as an ecosystem/API capability, but evidence only shows a REST API and llms.txt file — no mention of an official MCP server for connecting agents. missing for 10: any documented MCP server endpoint, MCP spec compliance, or third-party confirmation of MCP support.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [probe] “PROBE llms.txt: HTTP 200 at https://consensus.app/llms.txt # Consensus > Consensus is an AI-powered scientific search engine that finds, ra…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · not comparableConsensusnone0/10Evidence shows an API and MCP server exist, but there is no mention of scoped or least-privilege API keys, permission scopes, or credential management for agents — just generic API access. missing for 10: scoped/least-privilege credential issuance, API key permission controls, agent-specific auth documentation.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
ai-native userExplore an interactive API reference with runnable examples
weight 2 · not comparableConsensusnone0/10Evidence confirms Consensus offers a REST API and MCP server (consensus-docs-1, consensus-docs-15, consensus-docs-16), but there is no mention of an interactive API reference, sandbox, or runnable code examples anywhere in the pack.
- [claimed-docs] “Connect the Consensus API within your project to seamlessly integrate up-to-date peer-reviewed citations into your own custom workflow.”
- [claimed-docs] “Save your team hours of manual discovery research and run automated searches with our API to quickly and easily find the most relevant and r…”
- [claimed-docs] “a REST API and an MCP server expose the same retrieval and synthesis surface that powers the web app”
ai-native userTest against a sandbox environment without touching production data
weight 1 · not comparableConsensusn/aConsensus is a research/literature-search engine over academic papers, not a data-producing or transactional system where 'sandbox vs production data' is a meaningful distinction; there is no concept of production data being modified. This axis is a category error for this product type.
ai-native userDefine rules that trigger actions automatically on events
weight 3 · not comparableConsensusn/aConsensus is a research/literature search and synthesis engine, not an automation/workflow-rules platform; there is no concept of user-defined trigger-action rules for events in its product category.
ai-native userVersion, review, and roll back my automations
weight 1 · not comparableConsensusn/aConsensus is a research/literature-search engine, not an automation-building platform; there is no concept of 'automations' to version, review, or roll back. This axis is a category error for this product type.
ai-native userRead the product's source under an open license
weight 2 · not comparableConsensusn/aConsensus is a closed, commercial SaaS research search engine; there is no indication its source code is open-licensed or expected to be, making this axis a category error for this product type.
Sakana Marlinnone0/10No evidence of any open-source license or public source code repository for Marlin; it is presented as a paid SaaS research product, and probes for open API/docs artifacts returned 404s.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
ai-native userSelf-host the core product
weight 3 · not comparableConsensusn/aConsensus is a hosted SaaS research engine/API, not open-source software; self-hosting is a wrong-axis question for this type of product and no evidence suggests otherwise.
Sakana Marlinn/aSakana Marlin is a hosted SaaS research product with pay-per-use/credit pricing, not open-source or self-hostable software; self-hosting is a category mismatch for this type of managed AI service.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
- [claimed-docs] “Add your card. Start right away.”