Elicit vs Sakana Marlin
free-tier · subscription-flat · subscription-per-seat · enterprise-custom
·subscription-flat · enterprise-custom
Elicit wins · 18–6 (12 drawn)
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
Agent access
ai-native userPoint an agent at llms.txt or agent-oriented docs
weight 2 · round to ElicitElicit hosts a live llms.txt at support.elicit.com/llms.txt (HTTP 200, confirmed via probe) listing agent-oriented docs, and it also exposes an official MCP server for agent access, showing genuine agent-oriented documentation infrastructure. missing for 10: no independent/community corroboration of an agent actually consuming llms.txt successfully, and no OpenAPI/agent-doc spec beyond the llms.txt itself.
- [probe] “PROBE llms.txt: HTTP 200 at https://support.elicit.com/llms.txt # Elicit Help Center > Help center for Elicit ## Getting Started - [Getti…”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Sakana Marlinnone0/10The llms.txt probe returned a 404, and no other evidence shows agent-oriented docs (like an API spec or agent-readable documentation) for Marlin; the evidence pack is entirely marketing copy about the product's research capabilities, not machine-readable docs.
ai-native userRun the product headlessly / in CI for automation
weight 2 · round to ElicitElicit documents an API (and MCP server) explicitly for running its search/report/systematic-review capabilities 'from your own code, scripts, and workflows,' which supports headless/automated use outside the UI. However, there's no explicit CI/pipeline example, and API access appears gated as a paid plan feature rather than a fully documented automation-first workflow. Missing for 10: explicit CI/automation examples or tutorials, rate-limit/auth details for unattended use, and independent confirmation that the API works reliably in automated pipelines.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [claimed-docs] “API access”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Sakana Marlinnone0/10Marlin is presented as a research-report generation web product with a UI and pay-per-use credits, with no CLI, API, SDK, or webhook documentation for headless/CI usage; probes for llms.txt and OpenAPI specs both returned 404, indicating no programmatic interface is exposed.
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
- [claimed-docs] “Add your card. Start right away.”
ai-native userUse an official CLI
weight 2 · round drawnElicitnone0/10Evidence shows Elicit offers an API and an MCP server for programmatic/agentic access, but there is no mention anywhere of an official command-line interface (CLI) tool. Since API-based products could plausibly ship a CLI, absence of evidence means this axis is unmet rather than inapplicable.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
Sakana Marlinnone0/10No evidence of an official CLI; Marlin appears to be a web-based research tool with a pay-per-credit UI, and probes for API/llms.txt endpoints returned 404s, suggesting no developer-facing interface is exposed.
ai-native userDrive the product through a documented public API
weight 3 · round to ElicitElicit documents a public API with keys/auth management and specific programmatic endpoints (search 138M+ papers, automated report generation, full systematic review workflow control), plus an MCP server exposing the same functionality for agentic clients. Missing for 10: an actual OpenAPI/swagger spec was not found (404s on candidate paths) and no independent/hands-on developer report validates real-world API usage.
- [claimed-docs] “API access”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
- [probe] “PROBE openapi: all candidate paths 404 (https://support.elicit.com/openapi.json, https://support.elicit.com/swagger.json, https://support.el…”
Sakana Marlinnone0/10No documented public API is evidenced; probes for llms.txt and OpenAPI/swagger endpoints all returned 404s, and all other evidence describes the product's research capabilities, not a programmatic interface.
ai-native userBuild against official SDKs
weight 2 · round to ElicitElicit documents an official API (with managed API keys) for programmatic access to search and automated research reports, plus full API functionality exposed via an MCP server for Claude Desktop/Code integration, which supports agentic, code-driven workflows. However, evidence shows only a REST-style API and API-key docs, not a dedicated official SDK/client library in specific languages, nor code samples or independent developer corroboration. Missing for 10: named client SDKs (e.g., Python/JS packages), quickstart code examples, and independent/hands-on developer confirmation of API reliability.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [claimed-docs] “API access”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
Sakana Marlinnone0/10No evidence of any official SDK, API, or developer library for Marlin; probes for llms.txt and OpenAPI specs both returned 404, and all docs describe an end-user research product with no mention of programmatic/SDK access.
ai-native userSubscribe to events via webhooks
weight 2 · round drawnElicitnone0/10Elicit offers email alerts, an API, and an MCP server, but no evidence anywhere in the pack mentions webhook subscriptions or event-driven callbacks for programmatic integration.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Sakana Marlinnone0/10No evidence of a webhook or event subscription mechanism; probes for API/OpenAPI specs returned 404s and docs focus only on research report generation. Missing for 10: any webhook documentation, event subscription API, or callback mechanism.
Agentic features
ai-native userGet AI-generated insights and suggestions from my data inside the product
weight 2 · round to Sakana MarlinElicitdisputedcontradicted5/10Elicit's Research Agent, columns, chat-with-papers, and Systematic Review reports are documented to generate AI insights and suggestions from ingested papers (elicit-docs-11, elicit-docs-40, elicit-docs-41), and one HN user found topic analysis genuinely useful (elicit-comm-1). However, independent hands-on reports concretely contradict reliability: users found mostly incorrect summaries, missed key papers, and fabricated/hallucinated facts even when directly quoting sources (elicit-comm-3, elicit-comm-4, elicit-comm-7), and Elicit's own docs admit ~10% inaccuracy requiring manual verification (elicit-comm-2). missing for 10: independent corroboration that generated insights are consistently accurate rather than frequently hallucinated, and resolution of the documented factual-error reports.
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “Chat enables you to: Compare and contrast papers - Summarize multiple papers along specific dimensions (like their methodologies) - Cluster …”
- [claimed-docs] “the Research Agent can pull from a wide range of sources (e.g. publications, public filings, press releases), produce flexible outputs, and …”
- [community] “I asked about media bias detection and used the topic analysis feature. A minute or so later, I had a list of concepts with citations and li…”
- [community] “Elicit states: 'assume that around 90% of the information you see in Elicit is accurate... it's very important for you to check the work in …”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
- [community] “I tried abstract summarization with a whitepaper and it came up with completely made-up facts, describing an algorithm acronym incorrectly. …”
Marlin autonomously researches user-provided topics, mapping causal relationships, comparing hypotheses, and generating structured strategic insights and reports with citations, going beyond simple summarization. Missing for 10: independent/hands-on third-party validation beyond vendor-curated testimonials, and no visibility into underlying data/insight quality benchmarks.
- [claimed-docs] “Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into structured strategic options.”
- [claimed-docs] “It does more than summarize. Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into st…”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
ai-native userSet up automations that run autonomously in the background
weight 2 · round to Sakana MarlinElicit's Alerts feature lets users set up a background process that automatically monitors for new relevant papers and notifies them (e.g., via instant email alerts), which is a real autonomous background automation, and the API/MCP server also enables scripted automated report generation from external workflows. However, this is narrow (limited to paper-discovery alerts) rather than a general-purpose scheduling/automation system for arbitrary agentic tasks, and there's no evidence of recurring scheduled jobs, triggers, or workflow orchestration beyond alerts. Missing for 10: evidence of a general automation/scheduling engine, ability to chain multi-step autonomous tasks, and independent confirmation that alerts reliably run unattended over time.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Alerts allow you to stay up to date on the latest research about topics that are relevant to you, so you can add them to your Library for fu…”
- [claimed-docs] “Turn on Instant email alerts to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Marlin's docs clearly describe a single research task running autonomously for up to ~8 hours without further human input once a topic is set (docs-9, docs-10, docs-16), which matches the 'runs in background autonomously' idea. However, this is a one-shot session, not a recurring/scheduled automation you configure and forget — there's no evidence of triggers, schedules, or multi-run automation management typical of 'set up automations.' Missing for 10: scheduled/recurring automation setup, background job management UI, independent hands-on corroboration of unattended runtime.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
ai-native userDelegate tasks to a built-in AI assistant inside the product
weight 3 · round drawnElicit ships a built-in Research Agent that users can delegate tasks to directly (search, screening, extraction, column creation, chat with papers, skills), with effort-level control and iterative outputs, all inside the product interface (elicit-docs-2,7,8,11,12,21,40,42). Community evidence corroborates real task delegation working in practice (elicit-comm-1) though also raises accuracy concerns that temper trust in outputs (elicit-comm-2,3,4). Missing for 10: independent quality benchmarking beyond anecdotal HN threads and more recent hands-on validation of the newer effort-level/skills features.
- [claimed-docs] “Every Research Agent session runs at an effort level, which you set from the slider in the prompt box before you send your question.”
- [claimed-docs] “Projects in Elicit allow you to link together multiple sessions for fluid reasoning across different aspects of your work.”
- [claimed-docs] “A skill is a set of instructions you can hand to Elicit's agent so it approaches a task a particular way.”
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “the Research Agent can pull from a wide range of sources (e.g. publications, public filings, press releases), produce flexible outputs, and …”
- [claimed-docs] “you can filter and export your sources, generate figures, and draft slides, all without leaving the conversation”
- [claimed-docs] “Chat enables you to: Compare and contrast papers - Summarize multiple papers along specific dimensions (like their methodologies) - Cluster …”
- [claimed-docs] “A skill is a set of instructions you hand to Elicit's agent so it approaches a task a particular way.”
- [community] “I asked about media bias detection and used the topic analysis feature. A minute or so later, I had a list of concepts with citations and li…”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
Marlin is explicitly designed as a built-in AI agent that users delegate entire research/strategy tasks to, working autonomously for hours with minimal human input beyond initial framing (docs-1, docs-9, docs-10, docs-16, docs-21). This directly matches the story of delegating tasks to a built-in assistant within the product. Missing for 10: independent/hands-on verification beyond vendor testimonials, and detail on interactive control while a task is delegated.
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
ai-native userOperate the product with natural-language commands
weight 2 · round to ElicitElicit's Research Agent and semantic search are explicitly natural-language driven (e.g., asking questions in plain language, adding columns via natural-language commands like 'Add a column for study type'), and skills let users reference natural-language instructions instead of re-typing prompts. However, community reports raise accuracy/hallucination concerns that temper confidence in reliability of NL command execution. Missing for 10: independent hands-on verification of complex multi-step NL command chains, and no evidence of NL support outside the research/agent workflows (e.g., no broader command-line or API NL interface).
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “With Elicit's semantic search engine, you can ask a question in natural language, and Elicit will find relevant papers, without you having t…”
- [claimed-docs] “Elicit's semantic search means you don't have to know all the right keywords to get relevant results.”
- [claimed-docs] “Instead of writing out the same detailed prompt every time you run a market sizing or a landscape review, you reference the skill and the ag…”
- [claimed-docs] “the Research Agent can pull from a wide range of sources (e.g. publications, public filings, press releases), produce flexible outputs, and …”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
Marlin is initiated by giving it a research topic and a brief natural-language exchange to set direction (docs-10), suggesting natural-language input drives its operation, but there is no evidence of a broader natural-language command interface (e.g., chat-style control, follow-up instructions, or command syntax) beyond initial topic-setting. Missing for 10: documentation of ongoing NL command/control during execution, examples of varied NL commands, independent/hands-on confirmation of NL interaction quality.
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
Api quality
ai-native userDownload a machine-readable API spec (OpenAPI or equivalent)
weight 2 · round drawnElicitnone0/10Elicit documents a REST API (elicit-docs-34,36) and an MCP server (elicit-docs-28,35), but no evidence of a downloadable OpenAPI/Swagger spec exists; direct probes for openapi.json/swagger.json all returned 404 (elicit-probe-3). Missing for 10: any published OpenAPI/Swagger file, machine-readable schema, or API reference page listing such a spec.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [probe] “PROBE openapi: all candidate paths 404 (https://support.elicit.com/openapi.json, https://support.elicit.com/swagger.json, https://support.el…”
Sakana Marlinnone0/10Probes for OpenAPI/swagger spec files and llms.txt all returned 404, and no documentation mentions a machine-readable API spec. Missing for 10: any published OpenAPI/Swagger file, API reference docs, or equivalent machine-readable spec.
ai-native userRely on versioned APIs with a documented deprecation policy
weight 2 · round drawnElicitnone0/10Evidence confirms Elicit has an API and MCP server (elicit-docs-34, elicit-docs-35, elicit-docs-36) but nothing documents API versioning or a deprecation policy, and probes for an OpenAPI/spec file returned 404s (elicit-probe-3).
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [probe] “PROBE openapi: all candidate paths 404 (https://support.elicit.com/openapi.json, https://support.elicit.com/swagger.json, https://support.el…”
Sakana Marlinnone0/10No evidence of any public API, versioning scheme, or deprecation policy; probes for OpenAPI/llms.txt endpoints returned 404s, and all docs describe the research product itself, not a developer API. Missing for 10: any API documentation, versioning scheme, deprecation policy.
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
ai-native userPerform bulk operations across many items at once
weight 2 · round to ElicitElicit's core Systematic Reviews and Tables/Columns workflows explicitly apply extraction and screening operations across many papers at once (docs-11, docs-18, docs-22, docs-45), with bulk import (RIS/BIB, Zotero) and bulk export (CSV/Excel/RIS/BIB) of entire tables (docs-25, docs-38, docs-47), plus API/MCP access to run full systematic reviews programmatically at scale (docs-34, docs-36). Missing for 10: independent/hands-on evidence specifically validating bulk-scale accuracy or performance (community evidence addresses general accuracy, not bulk-operation mechanics).
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
- [claimed-docs] “Tables can be exported in CSV and Excel format. Certain tables of sources can be exported as RIS or BIB files.”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [claimed-docs] “You can import RIS and BIB files into Elicit. This makes it much easier to import titles/abstracts from other tools like EndNote, Mendeley, …”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
ai-native userSchedule recurring jobs or workflows
weight 2 · round to ElicitElicit's Alerts feature lets users get recurring updates when new relevant papers matching a saved search appear (via instant email alerts), which is a limited form of a recurring job, but there is no evidence of general scheduling of arbitrary Research Agent workflows, systematic reviews, or API-driven jobs on a recurring cadence. Missing for 10: ability to schedule/repeat full Research Agent or Systematic Review workflows, cron-like or interval-based automation beyond paper alerts, and confirmation this works via API/MCP for programmatic recurring runs.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Alerts allow you to stay up to date on the latest research about topics that are relevant to you, so you can add them to your Library for fu…”
- [claimed-docs] “Turn on Instant email alerts to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Sakana Marlinnone0/10Sakana Marlin is a single-run autonomous research/report tool triggered by a user topic; no evidence of scheduling, recurrence, cron-like triggers, or workflow automation for repeated jobs. Probes also show no API/OpenAPI surface that could support scheduled invocation.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
ai-native userVersion, review, and roll back my automations
weight 1 · round drawnElicitnone0/10Elicit provides skills, columns, and projects for automation, but there is no evidence of versioning, review history, or rollback capability for these automations/skills/workflows anywhere in the evidence pack.
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
Sharing
analystShare a research session or report with collaborators who can view or build on it
weight 2 · round to ElicitElicit has a documented feature to invite team members to collaborate live in Research Agent sessions, with edit-access collaborators able to ask the agent questions, produce new artifacts, and edit others' work, plus reports/tables can be exported as PDF/Word/CSV for sharing. missing for 10: no independent/hands-on corroboration of the live collaboration feature working smoothly, and no detail on view-only/read-access sharing permissions.
- [claimed-docs] “Collaborators with edit access can ask the agent questions in your session, produce new artifacts (documents, tables, etc.), and edit other …”
- [claimed-docs] “Invite your team members to collaborate live with you in Research Agent sessions.”
- [claimed-docs] “Invite your team members to collaborate live with you in Research Agent sessions. Collaborators with edit access can ask the agent questions…”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
Sakana Marlinnone0/10No evidence describes any collaboration or sharing features—no mention of shared workspaces, links, comments, or multi-user access to reports/sessions; the pack only covers autonomous research generation, pricing tiers, and API probes returning 404s.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
Alerts
researcherSet up standing searches or alerts that surface new relevant sources as they appear
weight 1 · round to ElicitElicit's Alerts feature explicitly lets researchers set up standing topic alerts with instant email notifications when new relevant papers are found, adding them to a Library for future use — directly matching the standing-search/alert story. Missing for 10: independent/hands-on verification of alert accuracy or timeliness, and detail on how alert relevance/topics are configured beyond docs claims.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Alerts allow you to stay up to date on the latest research about topics that are relevant to you, so you can add them to your Library for fu…”
- [claimed-docs] “Turn on Instant email alerts to be immediately alerted whenever a new relevant research paper is found.”
Corpus
researcherUpload my own PDFs or corpus and have the agent research over them
weight 2 · round to ElicitElicit supports importing/uploading a user's own corpus (RIS/BIB import, Zotero integration, Collections) and running research operations (chat, columns/data extraction, systematic reviews) over those uploaded papers, with the browser extension auto-fetching full text for extraction. Community feedback raises accuracy/hallucination concerns about summarization quality, which tempers reliability but does not contradict the upload/research capability itself. Missing for 10: independent hands-on verification of accuracy when researching over a user-uploaded corpus, and clearer documentation of raw multi-PDF drag-and-drop upload versus reference-manager import formats.
- [claimed-docs] “Elicit's Zotero integration helps you bring your paper collections into Elicit, where you can extract data and analyze your papers!”
- [claimed-docs] “You can import RIS and BIB files into Elicit. This makes it much easier to import titles/abstracts from other tools like EndNote, Mendeley, …”
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “You can now organize and manage your papers in the Elicit Library using Collections. Collections help you group related research, arrange it…”
- [claimed-docs] “Elicit will automatically fetch papers during the data extraction phase of your Systematic Reviews”
- [claimed-docs] “you can spend less time downloading full-text PDFs from publisher websites to extract in Elicit – we'll get the papers for you”
- [claimed-docs] “Chat enables you to: Compare and contrast papers - Summarize multiple papers along specific dimensions (like their methodologies) - Cluster …”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
Sakana Marlinnone0/10Evidence describes Marlin as an autonomous web-research agent that gathers information via web navigation and generates reports, but there is no mention of uploading a user's own PDFs or corpus for the agent to research over. Missing for 10: any document/file upload feature, corpus ingestion, or evidence of researching over user-supplied materials rather than open web sources.
Reviews
researcherRun a systematic screening and extraction workflow across many papers with consistent criteria
weight 2 · round to ElicitElicit documents a dedicated Systematic Reviews workflow covering setup, search/gather, title-abstract screening, automated full-text screening, and data extraction via consistent custom columns applied across all papers, with export of screening/extraction tables — directly matching the story. Community evidence raises general accuracy/hallucination concerns about Elicit's paper analysis (not specifically the systematic review pipeline), which tempers confidence in perfect consistency at scale. Missing for 10: independent hands-on validation of the systematic review workflow's accuracy/consistency specifically (vs. general chat/summarization complaints), and no third-party benchmarking of screening reliability across large paper sets.
- [claimed-docs] “The Systematic Reviews workflow provides step-by-step guidance through search, screening, and data extraction, culminating in a research rep…”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
- [claimed-docs] “Full-Text Screening is a distinct, automated step in the Systematic Review workflow that catches these mismatches before you commit to data …”
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, including the screening recommendation tables and the data…”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
- [claimed-docs] “Elicit can help you save up to 80% of time usually spent on a systematic review without sacrificing accuracy.”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
Sakana Marlinnone0/10Marlin is positioned as an autonomous business/market strategy research agent producing single deep-dive reports, not as a tool for systematic multi-paper screening/extraction with consistent criteria (a literature-review workflow). No evidence describes handling many papers, applying consistent inclusion/extraction criteria, or batch processing across a corpus — the described unit of work is one topic producing one report.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
- [claimed-docs] “putting Marlin to work on real tasks such as strategy formulation, market research, risk analysis, and competitive analysis”
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
ai-native userDo everything through the API that I can do in the UI
weight 2 · round to ElicitElicit's API and MCP server expose search and end-to-end Systematic Review functionality (search, screening, extraction, reporting), giving programmatic access to core research capabilities, but there's no evidence that UI-only features like the interactive Research Agent chat, Skills, real-time collaboration, columns customization, alerts, or the browser extension are exposed via the API/MCP surface. Missing for 10: explicit documentation of full feature parity, API/MCP access to Research Agent conversational sessions, skills, collaboration, and alerts.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [claimed-docs] “API access”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
Sakana Marlinnone0/10No evidence of any public API for Sakana Marlin; probes for llms.txt and OpenAPI spec both returned 404, and all documentation describes only UI/credit-based access.
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
- [claimed-docs] “Add your card. Start right away.”
ai-native userExport all of my data in open formats and leave
weight 3 · round to ElicitElicit supports exporting Library, tables, and reports in open formats (RIS, CSV, BIB, PDF, DOCX) and offers API/MCP access for programmatic retrieval, which supports data portability. However, export of core artifacts like screening/extraction tables is gated behind Pro/Scale/Enterprise plans, and there's no evidence of full account data export (e.g., all research agent sessions, projects, skills, chat history) in open formats, nor an explicit 'delete account and take everything' workflow. missing for 10: full-account/session export beyond tables and library, confirmation of free-tier export ability, independent verification of export completeness/fidelity.
- [claimed-docs] “Your Elicit Library can be exported as a .ris file, which you can import into Zotero, Mendeley, EndNote, or another reference manager.”
- [claimed-docs] “Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, including the screening recommendation tables and the data…”
- [claimed-docs] “Tables can be exported in CSV and Excel format. Certain tables of sources can be exported as RIS or BIB files.”
- [claimed-docs] “Export to RIS, CSV, BIB, PDF and DOCX”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
Sakana Marlinnone0/10No evidence of data export, open-format download, or account portability features; probes for llms.txt and OpenAPI both returned 404, and docs only describe generated reports/slides, not export of underlying user data.
ai-native userRead the product's source under an open license
weight 2 · round drawnElicitnone0/10Elicit is a closed, proprietary SaaS research tool; no evidence in the pack mentions an open-source repository, source code availability, or an open license for its codebase. This axis applies (a product could plausibly open its source), but no evidence supports it.
Sakana Marlinnone0/10No evidence of any open-source license or public source code repository for Marlin; it is presented as a paid SaaS research product, and probes for open API/docs artifacts returned 404s.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [probe] “PROBE llms.txt: HTTP 404 at https://sakana.ai/llms.txt”
- [probe] “PROBE openapi: all candidate paths 404 (https://sakana.ai/openapi.json, https://sakana.ai/swagger.json, https://sakana.ai/api/openapi.json, …”
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Pricing
researcherTry the product meaningfully on a free tier or trial
weight 1 · round drawnElicitnone0/10The evidence pack includes pricing-page references (e.g., elicit-docs-4, elicit-docs-6, elicit-docs-31) but none describe a free tier's scope, limits, or a trial period — no content confirms what a researcher could do without paying. Axis clearly applies to a SaaS research tool, but no evidence substantiates a meaningful free/trial experience.
Sakana Marlinnone0/10Marlin is explicitly pay-per-use available only to paid Pro/Team/Enterprise plans, requires adding a card to start, and cancelling mid-run still consumes credits — there is no free tier or trial for researchers to test it meaningfully.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add your card. Start right away.”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
researcherUnderstand plan pricing and usage limits before committing
weight 2 · round to Sakana MarlinEvidence confirms a public pricing page exists (elicit.com/pricing) and reveals plan tier names (Pro, Scale, Enterprise) tied to feature gating like table exports, but no evidence pack content shows actual price points, free-tier limits, or usage caps that a researcher would need to compare plans before committing. Missing for 10: actual price figures per tier, usage/query limits, free-plan restrictions, and any independent confirmation of pricing transparency.
- [claimed-docs] “Import from Zotero”
- [claimed-docs] “API access”
- [claimed-docs] “Export to RIS, CSV, BIB, PDF and DOCX”
- [claimed-docs] “Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, including the screening recommendation tables and the data…”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
Marlin's docs mention it's offered as a pay-per-use add-on to Pro/Team/Enterprise plans, credits are consumed even if cancelled mid-run, and additional credits cost ¥98 each, giving a researcher some pricing signal. However, there's no concrete breakdown of how many credits a typical run consumes, no explicit usage caps, and no comparison table of plan tiers — missing for 10: full plan pricing table, credit-consumption-per-task estimates, explicit usage limits, independent/hands-on cost verification.
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
- [claimed-docs] “Add your card. Start right away.”
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
ai-native userChoose where my data is stored (region/residency)
weight 2 · round drawnElicitnone0/10No evidence in the pack mentions data residency, regional storage, or the ability to choose where data is stored; Elicit's docs cover exports, imports, API, and workflows but not data residency options. missing for 10: any mention of region/data-residency controls, enterprise data-locality options, or storage location settings.
ai-native userPrevent my data from being used to train AI models
weight 3 · round drawnElicitnone0/10No evidence pack items address opt-out from AI training, data usage policies for model training, or any privacy/data-control settings related to training data; all evidence covers product features (search, review workflows, exports, API/MCP) with no mention of training-data privacy controls. missing for 10: any documentation of a training opt-out setting, data usage/privacy policy statement, or enterprise data handling terms addressing model training.
ai-native userControl data retention and deletion
weight 2 · round drawnElicitnone0/10No evidence pack items address data retention policies, deletion controls, or account/data deletion mechanisms; docs cover export/import formats but not retention or deletion of user data.
ai-native userOpt out of telemetry and usage tracking
weight 2 · round drawnElicitnone0/10No evidence in the pack mentions telemetry, usage tracking, or an opt-out mechanism; only unrelated docs about search, exports, and API/MCP features appear. This is a reasonable privacy-posture question for a SaaS AI product, but nothing in the evidence pack supports Elicit offering telemetry opt-out.
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
Reports
researcherExport results to common formats, including documents, spreadsheets, and reference-manager files
weight 1 · round to ElicitElicit documents export of tables/reports to CSV, Excel, PDF, DOCX, RIS, and BIB, covering documents, spreadsheets, and reference-manager formats, and also supports Zotero/EndNote/Mendeley import/export via RIS. Missing for 10: independent hands-on verification of export fidelity beyond vendor docs.
- [claimed-docs] “Tables can be exported in CSV and Excel format. Certain tables of sources can be exported as RIS or BIB files.”
- [claimed-docs] “Export to RIS, CSV, BIB, PDF and DOCX”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
- [claimed-docs] “Your Elicit Library can be exported as a .ris file, which you can import into Zotero, Mendeley, EndNote, or another reference manager.”
- [claimed-docs] “Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, including the screening recommendation tables and the data…”
- [claimed-docs] “You can import RIS and BIB files into Elicit. This makes it much easier to import titles/abstracts from other tools like EndNote, Mendeley, …”
Marlin auto-generates full reports with references and PowerPoint slides (docs-3, docs-9, docs-20), covering the 'documents' part of the story, but there is no evidence of spreadsheet export or reference-manager file formats (e.g., BibTeX/RIS/EndNote) for citations. missing for 10: spreadsheet export, reference-manager file export (BibTeX/RIS/EndNote), independent corroboration of export formats.
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
analystGet a structured report with sections, tables, and a summary that I can share with stakeholders
weight 3 · round to ElicitElicit's Systematic Reviews workflow produces a research report summarizing papers, includes data extraction and screening tables, and both reports and tables can be exported as PDF/Word/CSV/Excel for sharing with stakeholders (elicit-docs-1, elicit-docs-22, elicit-docs-47, elicit-docs-19, elicit-docs-25). The Research Agent can also produce documents, tables, and figures within a session (elicit-docs-9, elicit-docs-21, elicit-docs-33). Missing for 10: independent/hands-on corroboration that the exported report format is polished enough for external stakeholder sharing, and no evidence of customizable report sectioning beyond the standard systematic-review structure.
- [claimed-docs] “The Systematic Reviews workflow provides step-by-step guidance through search, screening, and data extraction, culminating in a research rep…”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
- [claimed-docs] “Research Reports can be exported as a PDF or Word file... Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, …”
- [claimed-docs] “Pro, Scale, and Enterprise subscribers can export tables from Systematic Reviews, including the screening recommendation tables and the data…”
- [claimed-docs] “Tables can be exported in CSV and Excel format. Certain tables of sources can be exported as RIS or BIB files.”
- [claimed-docs] “Collaborators with edit access can ask the agent questions in your session, produce new artifacts (documents, tables, etc.), and edit other …”
- [claimed-docs] “you can filter and export your sources, generate figures, and draft slides, all without leaving the conversation”
Docs describe autonomous generation of a full structured report with main body, appendices, references, and presentation slides/executive summary — directly matching sections, tables (implied by structured strategic options), and summary needs for stakeholder sharing. Missing for 10: no independent/hands-on verification of table formatting or actual sample report shown, and no evidence beyond vendor marketing copy.
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
- [claimed-docs] “Sakana Marlin maps the causal relationships at work in complex business environments and organizes them into structured strategic options.”
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
Agent runs
researcherPose a research question and get an autonomous multi-step investigation, not just a single-pass summary
weight 3 · round to Sakana MarlinElicit's Research Agent and Systematic Review workflows are explicitly documented as multi-step (effort-level slider from Fastest to Smartest, pulling from multiple source types, iterating until output is complete, and separate search/screen/extract/report phases) rather than single-pass answers (elicit-docs-2, elicit-docs-12/13, elicit-docs-21/33, elicit-docs-22/45). However, hands-on community reports describe results arriving quickly and resembling a single pass over topic clusters rather than deep autonomous investigation, and multiple independent accounts report missed papers, hallucinated conclusions, and shallow reasoning that undercut confidence in true multi-step depth (elicit-comm-1, elicit-comm-3, elicit-comm-4, elicit-comm-5). Missing for 10: independent verification that the agent performs genuinely autonomous multi-step reasoning (not just sequential fixed workflow steps) and evidence rebutting the accuracy/depth complaints.
- [claimed-docs] “Every Research Agent session runs at an effort level, which you set from the slider in the prompt box before you send your question.”
- [claimed-docs] “the Research Agent can pull from a wide range of sources (e.g. publications, public filings, press releases), produce flexible outputs, and …”
- [claimed-docs] “click it to open the slider and move between Fastest, Fast, Balanced, Smart, and Smartest”
- [claimed-docs] “you can filter and export your sources, generate figures, and draft slides, all without leaving the conversation”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
- [claimed-docs] “Full-Text Screening is a distinct, automated step in the Systematic Review workflow that catches these mismatches before you commit to data …”
- [community] “I asked about media bias detection and used the topic analysis feature. A minute or so later, I had a list of concepts with citations and li…”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
- [community] “Testing Elicit gave me quite a bit worse results than using PaperQA by futurehouse. While paperqa could understand a bit of the nuance of a …”
Vendor docs describe autonomous multi-step research: forming hypotheses, gathering info, resolving contradictions, running for hours across thousands of cycles without further human input, producing a full structured report - directly matching the story. This is corroborated by beta-tester quotes praising depth beyond chat-based research tools, though all evidence is vendor-published/testimonial rather than independent hands-on verification. missing for 10: independent third-party evaluation or benchmark of the autonomous multi-step process, technical detail on how contradictions/hypotheses are actually verified
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “Many told us that Marlin was more practical at digging deeply into information than the chat-based research tools they had used before”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
analystStart a long research job that keeps working unattended and notifies me when the result is ready
weight 2 · round to Sakana MarlinElicit documents an alerts feature that emails you when new relevant papers are found (elicit-docs-3, elicit-docs-15, elicit-docs-26), and long-running workflows like Systematic Reviews and Research Agent effort levels ('Smartest' mode) imply tasks that can take longer to complete (elicit-docs-2, elicit-docs-13, elicit-docs-22). However, alerts are for ongoing topic monitoring, not notification of a specific job's completion, and there's no evidence of a 'start and walk away, get notified when this specific job is done' async job model. Missing for 10: explicit documentation of background/async execution of a research job plus a completion notification (vs. topic-monitoring alerts), and any independent confirmation this works as described.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Alerts allow you to stay up to date on the latest research about topics that are relevant to you, so you can add them to your Library for fu…”
- [claimed-docs] “Turn on Instant email alerts to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Every Research Agent session runs at an effort level, which you set from the slider in the prompt box before you send your question.”
- [claimed-docs] “click it to open the slider and move between Fastest, Fast, Balanced, Smart, and Smartest”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
Vendor docs clearly describe long unattended autonomous research runs (~8 hours) producing full reports, which supports the core of the story, but there is no mention of a notification mechanism when results are ready and no independent/hands-on corroboration beyond marketing copy. missing for 10: evidence of a completion notification/alert feature, independent verification of unattended runtime and reliability, and API/technical docs confirming job-control (start/monitor/cancel) beyond the marketing page.
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
researcherSteer the depth, effort, and scope of a research run before or while it executes
weight 1 · round to ElicitElicit's Research Agent lets users set an effort level (Fastest→Smartest) via a slider before sending a query, and users can shape scope through columns, skills, and iterative follow-up prompts within a session; the API also exposes control over search strategy, screening criteria, and extraction parameters for Systematic Reviews. However, evidence only shows steering before/between turns, not genuine mid-execution adjustment of an in-flight run. Missing for 10: documentation of pausing/adjusting effort or scope while a run is actively executing, and independent verification that scope/effort controls meaningfully change output depth.
- [claimed-docs] “Every Research Agent session runs at an effort level, which you set from the slider in the prompt box before you send your question.”
- [claimed-docs] “click it to open the slider and move between Fastest, Fast, Balanced, Smart, and Smartest”
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “you can filter and export your sources, generate figures, and draft slides, all without leaving the conversation”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
Marlin only allows a brief initial exchange to set direction before running fully autonomously for up to ~8 hours with no mid-run steering, and there's no documented control over depth/effort/scope parameters (e.g., report length, time budget, source breadth) beyond the initial topic framing; cancellation is possible but not adjustment. missing for 10: mid-execution steering controls, explicit depth/effort/scope parameters or settings, independent evidence of pre-run configurability beyond a 'brief exchange'.
- [claimed-docs] “Sakana Marlin sharpens the direction of the investigation through a brief exchange with the user. Once the course is set, it works without f…”
- [claimed-docs] “Give it a research topic, and Marlin works autonomously for up to roughly eight hours, crafting a detailed strategy report up to a hundred p…”
- [claimed-docs] “You can cancel at any time during execution, but please note that credits will still be consumed.”
- [claimed-docs] “It is designed for use cases where you want to delegate everything up to "issue framing (agenda setting)" prior to decision-making to the AI…”
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
Citations
researcherSee citations for every substantive claim so I can verify it against the underlying source
weight 3 · round to Sakana MarlinElicitdisputedcontradicted5/10Elicit's search/columns features are built around pulling data from papers with links back to sources (elicit-docs-18, elicit-docs-20), and a hands-on user confirms getting 'a list of concepts with citations and links to papers' (elicit-comm-1). However, independent hands-on reports directly contradict the claim that citations reliably let you verify claims: users found Elicit 'hallucinates just as much and then gives a quote that directly contradicts its statement' (elicit-comm-4), produced 'mostly incorrect summaries' and missed key papers (elicit-comm-3), and Elicit itself warns only ~90% accuracy with a need to 'check the work in Elicit closely' (elicit-comm-2). Missing for 10: evidence that citation/quote extraction is reliably accurate, independent verification benchmarks, and resolution of the hallucination-despite-quoting complaints.
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “With Elicit's semantic search engine, you can ask a question in natural language, and Elicit will find relevant papers, without you having t…”
- [community] “I asked about media bias detection and used the topic analysis feature. A minute or so later, I had a list of concepts with citations and li…”
- [community] “Elicit states: 'assume that around 90% of the information you see in Elicit is accurate... it's very important for you to check the work in …”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
Vendor docs claim Marlin generates 'fully referenced reports' grounded in primary sources with appendices and references, and a testimonial praises its higher-quality citations to primary vs secondary sources, suggesting citation support exists. However, there is no independent verification, no example of inline citation format, and no detail on how claims map to sources for auditability. Missing for 10: independent hands-on verification of citation accuracy, example output showing citation linking, and confirmation citations are traceable/clickable to primary sources.
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “他の生成AIと比較して、引用される情報の量と質が高く、二次情報ではなく一次情報を参照できている点に優位性を感じました。”
- [claimed-docs] “From a fully referenced report to PowerPoint slides, everything is automatically generated — at a quality that stands alongside professional…”
- [claimed-docs] “Everything from the main body to appendices, references, and presentation slides is automatically generated — at a quality that stands along…”
Corpus
researcherSearch scholarly literature and primary sources, not just the open web
weight 2 · round to ElicitElicit's docs clearly show it searches scholarly literature (138M+ academic papers via semantic and keyword search), clinical trials, and journal-restricted queries, plus API/MCP access to the same corpus and systematic-review workflows built around paper screening/extraction rather than general web search. This directly matches the story of searching scholarly/primary sources rather than the open web. Missing for 10: independent corroboration specifically about breadth/quality of the scholarly corpus (community evidence addresses answer accuracy/hallucination, not source scope, so it doesn't contradict this particular axis).
- [claimed-docs] “With Elicit's semantic search engine, you can ask a question in natural language, and Elicit will find relevant papers, without you having t…”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “Elicit allows you to search both research papers and clinical trials.”
- [claimed-docs] “Elicit offers both semantic search and keyword search to find the most relevant papers for your systematic review.”
- [claimed-docs] “Add +journal:"INSERT JOURNAL NAME HERE" to your query”
- [claimed-docs] “Advanced search filters are a powerful hidden feature in the Find Papers workflow. You can use them to search within a particular journal, r…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
Vendor docs claim research is 'grounded strictly in primary sources' and a testimonial notes it references primary rather than secondary information compared to other generative AI tools, but there is no evidence of dedicated scholarly database access (e.g., PubMed, arXiv, JSTOR integration) or citation verification—only general web navigation is described. missing for 10: evidence of scholarly/academic database integration, independent verification of primary-source claim, and details on how it distinguishes scholarly vs open-web sources.
- [claimed-docs] “The research was of extremely high quality, grounded strictly in primary sources, resulting in a highly convincing and reliable final report…”
- [claimed-docs] “他の生成AIと比較して、引用される情報の量と質が高く、二次情報ではなく一次情報を参照できている点に優位性を感じました。”
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
Synthesis
analystSee where sources agree and disagree instead of a single unqualified answer
weight 2 · round to ElicitElicit's Chat/Compare feature explicitly lets analysts 'compare and contrast papers' and 'summarize multiple papers along specific dimensions,' and its column/table extraction lets you see each paper's data side-by-side, which supports spotting agreement/disagreement across sources. However, there is no dedicated feature that explicitly flags or highlights when sources conflict versus concur (no consensus/disagreement indicator), and community reports note the model can hallucinate quotes that contradict its own summaries, undermining confidence in cross-source synthesis. Missing for 10: an explicit contradiction/agreement-detection UI, and independent verification that comparisons are reliably accurate rather than hallucination-prone.
- [claimed-docs] “Chat enables you to: Compare and contrast papers - Summarize multiple papers along specific dimensions (like their methodologies) - Cluster …”
- [claimed-docs] “Create a column for each data point you'd like to extract from your papers. Columns can pull data from the papers' body text or from tables …”
- [claimed-docs] “To add a column, simply tell the research agent what column(s) you'd like to add. For example, you can say: "Add a column for study type."”
- [claimed-docs] “Elicit supports the major steps of a systematic review: 1. Set up your review on the Setup page 2. Gather all papers for the systematic revi…”
- [community] “It seems like it should hallucinate less, as it directly quotes, but nope, it still hallucinates just as much and then gives a quote that di…”
- [community] “I gave it a topic I researched in depth recently. It gave me mostly incorrect summaries (one said hypothesis X is confirmed; nope it hasn't)…”
Marketing copy claims Marlin 'resolves contradictions' and 'compares and evaluates multiple hypotheses' rather than merely aggregating, implying some handling of conflicting sources, but there is no evidence of a UI feature or report section that explicitly surfaces where sources agree/disagree to the analyst. Missing for 10: concrete example of a report showing conflicting source viewpoints, screenshot/description of how disagreement is presented, independent corroboration beyond vendor marketing.
- [claimed-docs] “Once you set a research topic, it works without further human input: repeatedly forming hypotheses, gathering information, navigating the we…”
- [claimed-docs] “From ideation and data gathering to resolving contradictions and structuring the report, the AI autonomously executes the entire research pr…”
- [claimed-docs] “Rather than merely aggregating information, it compares and evaluates multiple hypotheses to provide deep insights.”
Not comparable on these axes
ai-native userPlug MCP servers into this product so it can use their tools
weight 3 · not comparableElicitnone0/10Evidence shows Elicit exposes itself AS an MCP server for other clients (e.g., Claude Desktop) to consume its research tools (elicit-docs-28, elicit-docs-35), not the reverse capability of Elicit acting as an MCP client that plugs in external MCP servers to use their tools. No evidence exists that Elicit can connect to and use third-party MCP servers.
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
ai-native userConnect an agent via an official MCP server
weight 3 · not comparableElicit provides an official documented MCP server endpoint (claude mcp add --transport http elicit https://elicit.com/api/mcp) exposing full API functionality for use from Claude Desktop, Claude Code, and other MCP-compatible clients. missing for 10: no independent/hands-on corroboration of the MCP server working, and no detail on auth/tool-list scope beyond the docs.
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
ai-native userIssue scoped/least-privilege API credentials for an agent
weight 2 · not comparableElicitnone0/10Elicit does offer API keys and an MCP server for programmatic/agent access (elicit-docs-34, elicit-docs-35, elicit-docs-28), so the axis of credential management applies, but there is no evidence of scoped or least-privilege permissions, roles, or restricted-scope API keys — the docs only describe managing API keys generically, not limiting their scope.
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “claude mcp add --transport http elicit https://elicit.com/api/mcp”
ai-native userExplore an interactive API reference with runnable examples
weight 2 · not comparableElicitnone0/10Elicit documents an API and MCP server (elicit-docs-6, elicit-docs-34, elicit-docs-35, elicit-docs-36) but there is no evidence of an interactive API reference/playground with runnable examples; probes for OpenAPI/swagger specs all returned 404 (elicit-probe-3), suggesting no such interactive reference exists.
- [claimed-docs] “API access”
- [claimed-docs] “The Elicit API lets you access Elicit's research capabilities programmatically: search 138 million+ academic papers and generate automated r…”
- [claimed-docs] “All API functionality is also available via MCP (Model Context Protocol) server, enabling use from Claude Desktop, Claude Code, and other MC…”
- [claimed-docs] “Systematic Review: run a full review end to end, with control over search strategy, screening criteria, extraction parameters, and reporting…”
- [probe] “PROBE openapi: all candidate paths 404 (https://support.elicit.com/openapi.json, https://support.elicit.com/swagger.json, https://support.el…”
ai-native userTest against a sandbox environment without touching production data
weight 1 · not comparableElicitn/aElicit is a research/literature review tool, not a system with production data pipelines or deployment environments; the sandbox-vs-production testing story is a category error for this product type.
ai-native userDefine rules that trigger actions automatically on events
weight 3 · not comparableElicit's 'Alerts' feature lets a user set a topic and receive an automatic email notification when a new relevant paper is found, which is a narrow event→action automation, but there is no general rule-builder allowing arbitrary triggers/conditions/actions across the product. Missing for 10: user-defined trigger conditions beyond 'new paper found', support for actions besides email alerts, and any workflow/automation engine tying events to custom actions.
- [claimed-docs] “Turn on **Instant email alerts** to be immediately alerted whenever a new relevant research paper is found.”
- [claimed-docs] “Alerts allow you to stay up to date on the latest research about topics that are relevant to you, so you can add them to your Library for fu…”
- [claimed-docs] “Turn on Instant email alerts to be immediately alerted whenever a new relevant research paper is found.”
ai-native userSelf-host the core product
weight 3 · not comparableElicitn/aElicit is a hosted SaaS research product with no evidence of any self-hostable core offering (only API/MCP access to the hosted service is documented); self-hosting is not a plausible axis for this type of cloud-only product, so this is a category mismatch rather than a missing capability.
Sakana Marlinn/aSakana Marlin is a hosted SaaS research product with pay-per-use/credit pricing, not open-source or self-hostable software; self-hosting is a category mismatch for this type of managed AI service.
- [claimed-docs] “We have made Marlin available as a pay-per-use tier to monthly Pro, Team, and Enterprise-tier plans.”
- [claimed-docs] “Add-on credits available (¥98 / credit)”
- [claimed-docs] “Add your card. Start right away.”