Skip to content

Rank #5 of 6 in AI Research Agents

Sakana Marlin logo

Sakana AI · commercial

no public signals

Verified integrations

No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

20.5/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

0.0/100

Collaboration sharing — stories about collaboration sharing in this arenaCollaboration sharingevidence →

Stories about collaboration sharing in this arena

0.0/100

Literature workflow — stories about literature workflow in this arenaLiterature workflowevidence →

Stories about literature workflow in this arena

0.0/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

0.0/100

Pricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limitsevidence →

Free-tier ceilings, usage caps, and rate limits before you have to pay

16.0/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

0.0/100

Report output — stories about report output in this arenaReport outputevidence →

Stories about report output in this arena

58.5/100

Research depth — stories about research depth in this arenaResearch depthevidence →

Stories about research depth in this arena

56.0/100

Source quality — stories about source quality in this arenaSource qualityevidence →

Stories about source quality in this arena

30.0/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 0 free · 1 paid · 0 enterprise · 12 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 43/43 stories · click a row’s chevron for the rationale and evidence

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10C

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3none0/10

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10C

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10C

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2n/auntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1n/auntestednone yet

Pose a research question and get an autonomous multi-step investigation, not just a single-pass summary C

Agent runs

researcherResearch depth — stories about research depth in this arenaResearch depth3full8/10C

Get a structured report with sections, tables, and a summary that I can share with stakeholders C

Reports

analystReport output — stories about report output in this arenaReport output3full7/10C

See citations for every substantive claim so I can verify it against the underlying source C

Citations

researcherSource quality — stories about source quality in this arenaSource quality3partial5/10C

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3none0/10

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3n/a0/10

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3n/auntestednone yet

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3noneuntestednone yet

Search scholarly literature and primary sources, not just the open web C

Corpus

researcherSource quality — stories about source quality in this arenaSource quality2partial6/10C

Start a long research job that keeps working unattended and notifies me when the result is ready C

Agent runs

analystResearch depth — stories about research depth in this arenaResearch depth2partial6/10C

See where sources agree and disagree instead of a single unqualified answer C

Synthesis

analystSource quality — stories about source quality in this arenaSource quality2partial4/10C

Understand plan pricing and usage limits before committing G

Pricing

researcherPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits2partialpaid4/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2none0/10

Run a systematic screening and extraction workflow across many papers with consistent criteria C

Reviews

researcherLiterature workflow — stories about literature workflow in this arenaLiterature workflow2none0/10

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2none0/10

Share a research session or report with collaborators who can view or build on it C

Sharing

analystCollaboration sharing — stories about collaboration sharing in this arenaCollaboration sharing2none0/10

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2noneuntestednone yet

Upload my own PDFs or corpus and have the agent research over them C

Corpus

researcherLiterature workflow — stories about literature workflow in this arenaLiterature workflow2noneuntestednone yet

Export results to common formats, including documents, spreadsheets, and reference-manager files G

Reports

researcherReport output — stories about report output in this arenaReport output1partial4/10C

Steer the depth, effort, and scope of a research run before or while it executes C

Agent runs

researcherResearch depth — stories about research depth in this arenaResearch depth1partial4/10C

Try the product meaningfully on a free tier or trial G

Pricing

researcherPricing limits — free-tier ceilings, usage caps, and rate limits before you have to payPricing limits1none0/10

Set up standing searches or alerts that surface new relevant sources as they appear C

Alerts

researcherLiterature workflow — stories about literature workflow in this arenaLiterature workflow1noneuntestednone yet

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1noneuntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 32 stories with headroom

What would move Sakana Marlin’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Agenticness — how well agents can access and operate the productDrive the product through a documented public API

    nonemoves agent-readyimpact 45

    No documented public API is evidenced; probes for llms.txt and OpenAPI/swagger endpoints all returned 404s, and all other evidence describes the product's research capabilities, not a programmatic interface.

  2. Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave

    nonemoves PA Scoreimpact 30

    No evidence of data export, open-format download, or account portability features; probes for llms.txt and OpenAPI both returned 404, and docs only describe generated reports/slides, not export of underlying user data.

  3. Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models

    nonemoves PA Scoreimpact 30

    Missing: any mention of training opt-out settings, privacy policy, or data retention controls.

  4. Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs

    nonemoves agent-readyimpact 30

    The llms.txt probe returned a 404, and no other evidence shows agent-oriented docs (like an API spec or agent-readable documentation) for Marlin; the evidence pack is entirely marketing copy about the product's research capabilities, not machine-readable docs.

  5. Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation

    nonemoves agent-readyimpact 30

    Marlin is presented as a research-report generation web product with a UI and pay-per-use credits, with no CLI, API, SDK, or webhook documentation for headless/CI usage; probes for llms.txt and OpenAPI specs both returned 404, indicating no programmatic interface is exposed.

  6. Agenticness — how well agents can access and operate the productUse an official CLI

    nonemoves agent-readyimpact 30

    No evidence of an official CLI; Marlin appears to be a web-based research tool with a pay-per-credit UI, and probes for API/llms.txt endpoints returned 404s, suggesting no developer-facing interface is exposed.

  7. Agenticness — how well agents can access and operate the productBuild against official SDKs

    nonemoves agent-readyimpact 30

    No evidence of any official SDK, API, or developer library for Marlin; probes for llms.txt and OpenAPI specs both returned 404, and all docs describe an end-user research product with no mention of programmatic/SDK access.

  8. Agenticness — how well agents can access and operate the productSubscribe to events via webhooks

    nonemoves agent-readyimpact 30

    Missing: any webhook documentation, event subscription API, or callback mechanism.

Showing the top 8 of 32 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map3 surfaces · 13 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

0 of 7 testable claims verified · 0 contradictedintegrity 0/100

10 distinct capability claims found in Sakana Marlin’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

0

Verified

7

Unverified

0

Contradicted

6

Undersold

Unverified (9)
Undersold (6)
Claims outside our story set (2)

Real capability claims found in Sakana Marlin’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Runs thousands of internal automated cycles to dynamically select key arguments and cut irrelevant content

    source ↗
  • Users can cancel an in-progress run at any time, though credits already used are still consumed

    source ↗
Suggest a story for these →

Business model

subscription-flatenterprise-custom

Paid plans with a free trial by application; pricing is published on the product page in Japanese and English, with enterprise terms negotiated directly.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score9 (Sep 14 '26)9 (Sep 16 '26)
Agent-ready0 (Sep 14 '26)0 (Sep 16 '26)

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data