Verified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharingevidence →
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflowevidence →
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Report output — stories about report output in this arenaReport outputevidence →
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depthevidence →
Stories about research depth in this arena
Source quality — stories about source quality in this arenaSource qualityevidence →
Stories about source quality in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 1 paid · 0 enterprise · 12 not stated in evidence
Follow the green: where the map greys out is where Sakana Marlin stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
—0/10
Subscribe to events via webhooks
—0/10
Build against official SDKs
—0/10
Issue scoped/least-privilege API credentials for an agent
n/an/a
Connect an agent via an official MCP server
n/an/a
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
n/an/a
Explore an interactive API reference with runnable examples
n/an/a
Docs for agents
Point an agent at llms.txt or agent-oriented docs
—0/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
✓8/10
Operate the product with natural-language commands
~5/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
✓8/10
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing
Stories about collaboration sharing in this arena
Literature workflow — stories about literature workflow in this arenaLiterature workflow
Stories about literature workflow in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits
Free-tier ceilings, usage caps, and rate limits before you have to pay
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Report output — stories about report output in this arenaReport output
Stories about report output in this arena
Research depth — stories about research depth in this arenaResearch depth
Stories about research depth in this arena
Agent runs
Source quality — stories about source quality in this arenaSource quality
Stories about source quality in this arena
Sorted by importance (agentic first) (high → low) · 43/43 stories · click a row’s chevron for the rationale and evidence
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Cclaimed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | n/a | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | n/a | untested | none yet | |
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 3 | full | 8/10 | Cclaimed | |
Get a structured report with sections, tables, and a summary that I can share with stakeholders C Reports | analyst | Report output — stories about report output in this arenaReport output | 3 | full | 7/10 | Cclaimed | |
See citations for every substantive claim so I can verify it against the underlying source C Citations | researcher | Source quality — stories about source quality in this arenaSource quality | 3 | partial | 5/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | n/a | 0/10 | ||
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | n/a | untested | none yet | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | untested | none yet | |
Search scholarly literature and primary sources, not just the open web C Corpus | researcher | Source quality — stories about source quality in this arenaSource quality | 2 | partial | 6/10 | Cclaimed | |
Start a long research job that keeps working unattended and notifies me when the result is ready C Agent runs | analyst | Research depth — stories about research depth in this arenaResearch depth | 2 | partial | 6/10 | Cclaimed | |
See where sources agree and disagree instead of a single unqualified answer C Synthesis | analyst | Source quality — stories about source quality in this arenaSource quality | 2 | partial | 4/10 | Cclaimed | |
Understand plan pricing and usage limits before committing G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 2 | partialpaid | 4/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Run a systematic screening and extraction workflow across many papers with consistent criteria C Reviews | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Share a research session or report with collaborators who can view or build on it C Sharing | analyst | Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing | 2 | none | 0/10 | ||
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
Upload my own PDFs or corpus and have the agent research over them C Corpus | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 2 | none | untested | none yet | |
Export results to common formats, including documents, spreadsheets, and reference-manager files G Reports | researcher | Report output — stories about report output in this arenaReport output | 1 | partial | 4/10 | Cclaimed | |
Steer the depth, effort, and scope of a research run before or while it executes C Agent runs | researcher | Research depth — stories about research depth in this arenaResearch depth | 1 | partial | 4/10 | Cclaimed | |
Try the product meaningfully on a free tier or trial G Pricing | researcher | Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits | 1 | none | 0/10 | ||
Set up standing searches or alerts that surface new relevant sources as they appear C Alerts | researcher | Literature workflow — stories about literature workflow in this arenaLiterature workflow | 1 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 32 stories with headroom
What would move Sakana Marlin’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDrive the product through a documented public API
nonemoves agent-readyimpact 45
No documented public API is evidenced; probes for llms.txt and OpenAPI/swagger endpoints all returned 404s, and all other evidence describes the product's research capabilities, not a programmatic interface.
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
No evidence of data export, open-format download, or account portability features; probes for llms.txt and OpenAPI both returned 404, and docs only describe generated reports/slides, not export of underlying user data.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
Missing: any mention of training opt-out settings, privacy policy, or data retention controls.
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs
nonemoves agent-readyimpact 30
The llms.txt probe returned a 404, and no other evidence shows agent-oriented docs (like an API spec or agent-readable documentation) for Marlin; the evidence pack is entirely marketing copy about the product's research capabilities, not machine-readable docs.
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation
nonemoves agent-readyimpact 30
Marlin is presented as a research-report generation web product with a UI and pay-per-use credits, with no CLI, API, SDK, or webhook documentation for headless/CI usage; probes for llms.txt and OpenAPI specs both returned 404, indicating no programmatic interface is exposed.
Agenticness — how well agents can access and operate the productUse an official CLI
nonemoves agent-readyimpact 30
No evidence of an official CLI; Marlin appears to be a web-based research tool with a pay-per-credit UI, and probes for API/llms.txt endpoints returned 404s, suggesting no developer-facing interface is exposed.
Agenticness — how well agents can access and operate the productBuild against official SDKs
nonemoves agent-readyimpact 30
No evidence of any official SDK, API, or developer library for Marlin; probes for llms.txt and OpenAPI specs both returned 404, and all docs describe an end-user research product with no mention of programmatic/SDK access.
Agenticness — how well agents can access and operate the productSubscribe to events via webhooks
nonemoves agent-readyimpact 30
Missing: any webhook documentation, event subscription API, or callback mechanism.
Showing the top 8 of 32 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map3 surfaces · 13 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Marlin docs13 stories
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Understand plan pricing and usage limits before committing
- Export results to common formats, including documents, spreadsheets, and reference-manager files
- Get a structured report with sections, tables, and a summary that I can share with stakeholders
- Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- Start a long research job that keeps working unattended and notifies me when the result is ready
- Steer the depth, effort, and scope of a research run before or while it executes
- See citations for every substantive claim so I can verify it against the underlying source
- Search scholarly literature and primary sources, not just the open web
- See where sources agree and disagree instead of a single unqualified answer
Marlin release docs10 stories
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Operate the product with natural-language commands
- Understand plan pricing and usage limits before committing
- Export results to common formats, including documents, spreadsheets, and reference-manager files
- Get a structured report with sections, tables, and a summary that I can share with stakeholders
- Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary
- Start a long research job that keeps working unattended and notifies me when the result is ready
- Steer the depth, effort, and scope of a research run before or while it executes
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
0 of 7 testable claims verified · 0 contradicted → integrity 0/100
10 distinct capability claims found in Sakana Marlin’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
0
Verified
7
Unverified
0
Contradicted
6
Undersold
Unverified (9)
“Runs an end-to-end autonomous research process—forming hypotheses, gathering info, resolving contradictions—without further human input”
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summaryfullproof ↗
“Maps causal relationships in complex business situations into structured strategic options rather than just summarizing”
Get a structured report with sections, tables, and a summary that I can share with stakeholdersfullproof ↗
“Auto-generates a complete deliverable: main report, appendices, references, and presentation slides at professional-researcher quality”
Get a structured report with sections, tables, and a summary that I can share with stakeholdersfullproof ↗
“Works autonomously for up to ~8 hours to produce a detailed strategy report up to 100 pages plus executive summary slides”
Start a long research job that keeps working unattended and notifies me when the result is readypartialproof ↗
“Lets the user sharpen the research direction via a brief exchange before it proceeds fully autonomously”
Steer the depth, effort, and scope of a research run before or while it executespartialproof ↗
“Produces research grounded strictly in primary sources with high-quality, verifiable citations”
Search scholarly literature and primary sources, not just the open webpartialproof ↗
“Produces research grounded strictly in primary sources with high-quality, verifiable citations”
See citations for every substantive claim so I can verify it against the underlying sourcepartialproof ↗
“Available as a pay-per-use tier layered on top of monthly Pro, Team, and Enterprise plans”
Understand plan pricing and usage limits before committingpartialproof ↗
“Applied to real business tasks such as strategy formulation, market research, risk analysis, and competitive analysis”
Pose a research question and get an autonomous multi-step investigation, not just a single-pass summaryfullproof ↗
Undersold (6)
Get AI-generated insights and suggestions from my data inside the productfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Delegate tasks to a built-in AI assistant inside the productfullproof ↗
Operate the product with natural-language commandspartialproof ↗
Export results to common formats, including documents, spreadsheets, and reference-manager filespartialproof ↗
See where sources agree and disagree instead of a single unqualified answerpartialproof ↗
Claims outside our story set (2)
Real capability claims found in Sakana Marlin’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Runs thousands of internal automated cycles to dynamically select key arguments and cut irrelevant content”
source ↗“Users can cancel an in-progress run at any time, though credits already used are still consumed”
source ↗
Business model
Paid plans with a free trial by application; pricing is published on the product page in Japanese and English, with enterprise terms negotiated directly.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
