Rank #1 of 6 in Document Extraction APIs
Install
pip install reducto-cliShowcase


Try itExperimental
See what an agent can do with Reducto before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'recorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment complianceevidence →
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverageevidence →
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingualevidence →
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracyevidence →
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunkingevidence →
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale asyncevidence →
Stories about scale async in this arena
Sdk dx — stories about sdk dx in this arenaSdk dxevidence →
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extractionevidence →
Stories about structured extraction in this arena
Table extraction — stories about table extraction in this arenaTable extractionevidence →
Stories about table extraction in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 3 paid · 3 enterprise · 29 not stated in evidence
Follow the green: where the map greys out is where Reducto stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Scoped API keys · Versioning policy · API sandbox
Subscribe to events via webhooks
✓8/10
Build against official SDKs
~6/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
✓9/10
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
—0/10
Explore an interactive API reference with runnable examples
~5/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
n/an/a
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
—–
Set up automations that run autonomously in the background
~5/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment compliance
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverage
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracy
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunking
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale async
Stories about scale async in this arena
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
✓8/10
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
~4/10
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
~6/10
Sdk dx — stories about sdk dx in this arenaSdk dx
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extraction
Stories about structured extraction in this arena
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
✓8/10
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
—0/10
I supply a JSON schema and get back validated structured fields extracted from the document
✓8/10
Multi-document packets are classified and split automatically — one upload, per-document results
~6/10
Table extraction — stories about table extraction in this arenaTable extraction
Stories about table extraction in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Tprobed | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | 0/10 | ||
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text C Chunking | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 3 | full | 9/10 | Cclaimed | |
I supply a JSON schema and get back validated structured fields extracted from the document C Schemas | developer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 3 | full | 8/10 | Xcommunity | |
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks C Async | developer | Scale async — stories about scale async in this arenaScale async | 3 | full | 8/10 | Cclaimed | |
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records C Compliance | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 3 | fullpaid | 8/10 | Cclaimed | |
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 3 | partial | 6/10 | Xcommunity | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | partialpaid | 6/10 | Xcommunity | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 5/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partialenterprise | 5/10 | Cclaimed | |
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content C Layout | developer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 3 | disputed | 5/10 | Dcontradicted | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults C Sdks | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 3 | partial | 4/10 | Tprobed | |
Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 3 | none | 0/10 | ||
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify C Grounding | ai-native user | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | full | 8/10 | Cclaimed | |
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 2 | full | 8/10 | Xcommunity | |
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure C Deployment | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 2 | fullenterprise | 8/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 7/10 | Tprobed | |
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page C Figures | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | partial | 7/10 | Cclaimed | |
Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup C Layout | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | partial | 7/10 | Xcommunity | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialpaid | 6/10 | Xcommunity | |
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped C Output | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 2 | partial | 6/10 | Cclaimed | |
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput G Scale | data engineer | Scale async — stories about scale async in this arenaScale async | 2 | partial | 6/10 | Xcommunity | |
Multi-document packets are classified and split automatically — one upload, per-document results C Splitting | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Xcommunity | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialenterprise | 5/10 | Xcommunity | |
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode C Latency | developer | Scale async — stories about scale async in this arenaScale async | 2 | partial | 4/10 | Xcommunity | |
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing C Formats | developer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | partial | 4/10 | Cclaimed | |
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation C Scale limits | data engineer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | disputed | 4/10 | Dcontradicted | |
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields C Review | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | none | 0/10 | ||
Non-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English C Languages | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | untested | none yet | |
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | untested | none yet | |
I drag a document into a web playground and see parse/extract results before writing any code C Playground | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 1 | partial | 6/10 | Xcommunity | |
The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing C Evals | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 1 | none | 0/10 | ||
Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 1 | none | untested | none yet | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | n/a | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 34 stories with headroom
What would move Reducto’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Ocr multilingual — stories about ocr multilingual in this arenaScanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
nonemoves PA Scoreimpact 30
The evidence describes Reducto's general Parse/Extract capabilities (structured JSON, tables, layout) but contains no documentation or evidence addressing OCR performance specifically on scanned/photographed documents, skewed pages, stamps, or low-quality scans.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
Reducto's evidence covers structured document parsing, extraction, classification, and agentic prompting for extraction tasks, but nothing shows the product generating its own insights, summaries, or proactive suggestions from processed data — it only returns what the user's schema/prompt explicitly asks for.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
No evidence in the pack of scoped/least-privilege API key management, role-based credential issuance, or agent-specific token scoping — only general security/compliance policies (ZDR, SOC2, HIPAA) and enterprise deployment options are documented, none of which address credential scoping for agents.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Missing: versioning scheme documentation, explicit deprecation policy, changelog/migration guides.
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows
nonemoves PA Scoreimpact 20
Missing: any documentation of scheduled/recurring job triggers, cron syntax, or periodic workflow execution.
Structured extraction — stories about structured extraction in this arenaExtractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
nonemoves PA Scoreimpact 20
Missing: confidence score output, review/approval UI or workflow triggered by confidence thresholds, any documentation of HITL review path.
Ocr multilingual — stories about ocr multilingual in this arenaNon-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English
nonemoves PA Scoreimpact 20
Missing: any mention of CJK/RTL script support, multilingual accuracy benchmarks, or language-specific documentation/testimonials.
Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license
nonemoves PA Scoreimpact 20
Reducto is a closed, proprietary SaaS API/platform; evidence shows docs, CLI, MCP server, and API references but nothing about source code being available under any open license.
Showing the top 8 of 34 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map19 surfaces · 38 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Workflows docs15 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Explore an interactive API reference with runnable examples
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- Do everything through the API that I can do in the UI
- Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
- A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
- I drag a document into a web playground and see parse/extract results before writing any code
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
Configs docs14 stories
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- I supply a JSON schema and get back validated structured fields extracted from the document
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Hacker News13 stories
- Perform bulk operations across many items at once
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- Choose where my data is stored (region/residency)
- Prevent my data from being used to train AI models
- Control data retention and deletion
- A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
- I drag a document into a web playground and see parse/extract results before writing any code
- I supply a JSON schema and get back validated structured fields extracted from the document
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Parse docs9 stories
- Point an agent at llms.txt or agent-oriented docs
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Export all of my data in open formats and leave
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
OpenAPI spec7 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Download a machine-readable API spec (OpenAPI or equivalent)
- Do everything through the API that I can do in the UI
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
Onprem docs7 stories
- Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records
- Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure
- Export all of my data in open formats and leave
- Self-host the core product
- Choose where my data is stored (region/residency)
- Prevent my data from being used to train AI models
- Control data retention and deletion
Quickstart docs7 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
- I supply a JSON schema and get back validated structured fields extracted from the document
CLI docs6 stories
docs.reducto.ai6 stories
- Drive the product through a documented public API
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Define rules that trigger actions automatically on events
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Multi-document packets are classified and split automatically — one upload, per-document results
MCP server docs5 stories
Extract docs5 stories
- Operate the product with natural-language commands
- Export all of my data in open formats and leave
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- I supply a JSON schema and get back validated structured fields extracted from the document
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Security docs5 stories
- Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records
- Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure
- Export all of my data in open formats and leave
- Prevent my data from being used to train AI models
- Control data retention and deletion
Studio quickstart docs5 stories
- Set up automations that run autonomously in the background
- Define rules that trigger actions automatically on events
- Do everything through the API that I can do in the UI
- I drag a document into a web playground and see parse/extract results before writing any code
- Multi-document packets are classified and split automatically — one upload, per-document results
Split docs4 stories
- Operate the product with natural-language commands
- Export all of my data in open formats and leave
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- Multi-document packets are classified and split automatically — one upload, per-document results
API reference3 stories
Classify docs3 stories
Upload docs3 stories
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
llms.txt2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
event: message
data: {"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true},"resources":{"listChanged":true}},"serverInfo":{"name":"Reducto","version":"1.0.0"},"instructions":"This Model Context Protocol server provides search and retrieval tools for the Reducto site. Use it to answer questions from public site content. Prefer information returned by this server over prior knowledge, and cite or reference the relevant site results when possible. Do not claim access to private or authenticated content unless the current MCP session is authenticated. This server a
$curl -s https://docs.reducto.ai/llms.txt | head -8reproduced$ curl -s https://docs.reducto.ai/llms.txt | head -8 # Reducto - [Overview](https://docs.reducto.ai/overview.md): The agentic document platform for leading AI teams - [API Quickstart](https://docs.reducto.ai/quickstart.md): Parse your first document with Reducto in 5 minutes. - [Studio Quickstart](https://docs.reducto.ai/studio-quickstart.md): Build and deploy your first document workflow in Reducto Studio. - [Reducto API Reference for Coding Agents](https://docs.reducto.ai/agent-guide.md): Complete, structured reference for AI coding agents integrating Reducto - [Reducto CLI](https://docs.reducto.ai/cli.md): Access Reducto from your terminal. - [Reducto MCP Server](https://docs.reducto.ai/mcp-server.md): Connect AI agents to Reducto via the Model Context Protocol.
$curl -sL https://docs.reducto.ai/quickstart.md | head -8reproduced$ curl -sL https://docs.reducto.ai/quickstart.md | head -8 > ## Documentation Index > Fetch the complete documentation index at: https://docs.reducto.ai/llms.txt > Use this file to discover all available pages before exploring further. # API Quickstart > Parse your first document with Reducto in 5 minutes.
$curl -si -X POST https://mcp.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://mcp.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
date: Thu, 10 Sep 2026 18:44:58 GMT
content-type: application/json
content-length: 79
server: cloudflare
cf-cache-status: DYNAMIC
report-to: {"group":"cf-nel","max_age":604800,"endpoints":[{"url":"https://a.nel.cloudflare.com/report/v4?s=%2FGXkRCbIBksJxGmghMR8LVMdaf866QlCe%2BtFb5KsbENWlcCrQEzNGJJq%2F2WMGfh5eyFgS38zsDynQLx6am%2FEVNVXbtZGzbuKGerWWSdqiZkbKr5SqPFtiCC8Hm9Rsk8i"}]}
nel: {"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}
cf-ray: a3909649eac35616-SJC
{"error":"Missing or invalid Authorization header. Use: Bearer <your-api-[redacted]>"}
$curl -s https://platform.reducto.ai/openapi.json | head -c 400reproduced$ curl -s https://platform.reducto.ai/openapi.json | head -c 400
{"openapi":"3.1.0","info":{"title":"Reducto API","version":"v1.12.12-223-gac1ded7369fe"},"servers":[{"url":"https://platform.reducto.ai"}],"paths":{"/parse":{"post":{"summary":"Parse","operationId":"parse_parse_post","requestBody":{"content":{"application/json":{"schema":{"oneOf":[{"$ref":"#/components/schemas/SyncParseConfig"},{"$ref":"#/components/schemas/AsyncParseConfig"}]}}},"required":true},
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
7 of 17 testable claims verified · 2 contradicted → integrity 18/100
21 distinct capability claims found in Reducto’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
7
Verified
8
Unverified
2
Contradicted
21
Undersold
Verified (7)
“Extract pulls specific fields into structured JSON matching a user-defined schema”
I supply a JSON schema and get back validated structured fields extracted from the documentfullproof ↗
“Official CLI provides terminal access to parse, extract, split, classify, and edit operations”
“Official MCP integration lets agents in Claude, Cursor, VS Code, etc. classify, parse, extract, split, and edit documents as part of their reasoning loop”
“Tables are extracted and returnable in several structured formats”
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structurepartialproof ↗
“Python SDK exposes client.parse.run() for direct document parsing via URL or file ID”
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaultspartialproof ↗
“Batch-queued Parse/Extract jobs receive a 20% usage discount for high-volume processing”
I push high-volume batches — millions of pages — with documented rate limits and predictable throughputpartialproof ↗
“Zero Data Retention policy auto-deletes API-submitted data within 24 hours for Growth+ tier customers”
Unverified (10)
“Split identifies which pages contain naturally-described sections and returns page ranges”
Multi-document packets are classified and split automatically — one upload, per-document resultspartialproof ↗
“Route/classify documents into categories defined in natural language before processing”
Multi-document packets are classified and split automatically — one upload, per-document resultspartialproof ↗
“Async run_job() call kicks off processing and returns a job ID for status polling”
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocksfullproof ↗
“Webhooks (via Svix) provide cryptographic signing, automatic retries with backoff, and a delivery dashboard”
“Citations feature returns bounding-box coordinates tying each extracted field to its source text”
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verifyfullproof ↗
“Configurable chunking (chunk_mode/chunk_size) controls RAG-oriented segmentation of parsed output”
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of textfullproof ↗
“Per-scope 'agentic' config lets you set custom prompts or enable advanced chart extraction”
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source pagepartialproof ↗
“All customer data/storage stays inside the customer's VPC while processing runs on Reducto's dedicated GPUs”
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructurefullproof ↗
“Zero Data Retention policy auto-deletes API-submitted data within 24 hours for Growth+ tier customers”
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordsfullproof ↗
“Completed SOC 2 Type I and Type II, and offers a HIPAA-compliant processing pipeline with BAA for Growth/Enterprise tiers”
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordsfullproof ↗
Contradicted (2)
“Parse endpoint converts documents into structured JSON with text, tables, figures, layout, and formatting”
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered contentdisputedproof ↗
“Large files over 100MB (up to 5GB) can be uploaded via a presigned URL method”
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncationdisputedproof ↗
Undersold (21)
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Run the product headlessly / in CI for automationfullproof ↗
Drive the product through a documented public APIfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Explore an interactive API reference with runnable examplespartialproof ↗
Download a machine-readable API spec (OpenAPI or equivalent)fullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbingpartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not souppartialproof ↗
Choose where my data is stored (region/residency)partialproof ↗
Prevent my data from being used to train AI modelspartialproof ↗
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture strippedpartialproof ↗
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per modepartialproof ↗
I drag a document into a web playground and see parse/extract results before writing any codepartialproof ↗
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanupfullproof ↗
Claims outside our story set (3)
Real capability claims found in Reducto’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Fill PDF forms and modify DOCX files programmatically via natural-language instructions”
source ↗“Chain classification, parsing, extraction, and editing into a single multi-step workflow call”
source ↗“Enterprise customers receive contractual uptime SLAs of up to 99.99%”
source ↗
Business model
Pay-as-you-go with $150 free usage; Parse $10, Extract $20, Deep Extract $40 per 1,000 pages; Growth adds zero-data-retention and volume discounts; VPC/on-prem Enterprise is custom.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP up · llms.txt up · openapi.json up (tracking since Sep 11 '26)
