Skip to content

Rank #1 of 6 in Document Extraction APIs

Reducto logo

Reducto

YC W24

Reducto, Inc. · commercial

npm 118.8k/wkpypi 49.9k/wk

Showcase

Reducto homepage screenshot
homepage · captured Sep 2026 · view live ↗
Reducto docs screenshot
docs · captured Sep 2026 · view live ↗

Try itExperimental

See what an agent can do with Reducto before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).

$curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'recorded session — replayed, not live
recorded 2026-09-10 · exit 0 · captured verbatim by our probe harness, secrets redacted

Verified integrations

No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.

By theme — the product's score on each story themeBy theme

Agenticness — how well agents can access and operate the productAgenticnessevidence →

How well agents can access and operate the product

52.1/100

Automation depth — how much of the product can run unattendedAutomation depthevidence →

How much of the product can run unattended

20.6/100

Deployment compliance — stories about deployment compliance in this arenaDeployment complianceevidence →

Stories about deployment compliance in this arena

80.0/100

Format coverage — stories about format coverage in this arenaFormat coverageevidence →

Stories about format coverage in this arena

18.0/100

Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingualevidence →

Stories about ocr multilingual in this arena

0.0/100

Openness — open source, data portability, and self-hosting storiesOpennessevidence →

Open source, data portability, and self-hosting stories

26.4/100

Parse accuracy — stories about parse accuracy in this arenaParse accuracyevidence →

Stories about parse accuracy in this arena

26.6/100

Privacy posture — data-handling and privacy storiesPrivacy postureevidence →

Data-handling and privacy stories

26.7/100

Rag chunking — stories about rag chunking in this arenaRag chunkingevidence →

Stories about rag chunking in this arena

68.4/100

Scale async — stories about scale async in this arenaScale asyncevidence →

Stories about scale async in this arena

51.4/100

Sdk dx — stories about sdk dx in this arenaSdk dxevidence →

Stories about sdk dx in this arena

27.0/100

Structured extraction — stories about structured extraction in this arenaStructured extractionevidence →

Stories about structured extraction in this arena

52.4/100

Table extraction — stories about table extraction in this arenaTable extractionevidence →

Stories about table extraction in this arena

53.6/100

Story verdicts — every judged story with its evidenceStory verdicts

What’s free: 0 free · 3 paid · 3 enterprise · 29 not stated in evidence

?

Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence

Drive the product through a documented public API G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full9/10T

Connect an agent via an official MCP server G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3full8/10T

Delegate tasks to a built-in AI assistant inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Plug MCP servers into this product so it can use their tools G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness3n/auntestednone yet

Download a machine-readable API spec (OpenAPI or equivalent) G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Point an agent at llms.txt or agent-oriented docs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full9/10T

Run the product headlessly / in CI for automation G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Subscribe to events via webhooks G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10C

Use an official CLI G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2full8/10T

Build against official SDKs G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10T

Operate the product with natural-language commands G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial6/10T

Explore an interactive API reference with runnable examples G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10T

Set up automations that run autonomously in the background G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2partial5/10C

Rely on versioned APIs with a documented deprecation policy G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness2none0/10

Get AI-generated insights and suggestions from my data inside the product G

Agentic features

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Issue scoped/least-privilege API credentials for an agent G

Agent access

ai-native userAgenticness — how well agents can access and operate the productAgenticness2noneuntestednone yet

Test against a sandbox environment without touching production data G

Api quality

ai-native userAgenticness — how well agents can access and operate the productAgenticness1none0/10

Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text C

Chunking

ai-native userRag chunking — stories about rag chunking in this arenaRag chunking3full9/10C

I supply a JSON schema and get back validated structured fields extracted from the document C

Schemas

developerStructured extraction — stories about structured extraction in this arenaStructured extraction3full8/10X

Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks C

Async

developerScale async — stories about scale async in this arenaScale async3full8/10C

Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records C

Compliance

data engineerDeployment compliance — stories about deployment compliance in this arenaDeployment compliance3fullpaid8/10C

Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure C

Tables

data engineerTable extraction — stories about table extraction in this arenaTable extraction3partial6/10X

Prevent my data from being used to train AI models G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture3partialpaid6/10X

Export all of my data in open formats and leave G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partial5/10C

Self-host the core product G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness3partialenterprise5/10C

The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content C

Layout

developerParse accuracy — stories about parse accuracy in this arenaParse accuracy3disputed5/10D

Define rules that trigger actions automatically on events G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth3partial4/10C

Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults C

Sdks

developerSdk dx — stories about sdk dx in this arenaSdk dx3partial4/10T

Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included C

Ocr

developerOcr multilingual — stories about ocr multilingual in this arenaOcr multilingual3none0/10

Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify C

Grounding

ai-native userStructured extraction — stories about structured extraction in this arenaStructured extraction2full8/10C

I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup C

Tables

data engineerTable extraction — stories about table extraction in this arenaTable extraction2full8/10X

Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure C

Deployment

data engineerDeployment compliance — stories about deployment compliance in this arenaDeployment compliance2fullenterprise8/10C

Do everything through the API that I can do in the UI G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2partial7/10T

Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page C

Figures

ml engineerParse accuracy — stories about parse accuracy in this arenaParse accuracy2partial7/10C

Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup C

Layout

ml engineerParse accuracy — stories about parse accuracy in this arenaParse accuracy2partial7/10X

Control data retention and deletion G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partialpaid6/10X

I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped C

Output

ai-native userRag chunking — stories about rag chunking in this arenaRag chunking2partial6/10C

I push high-volume batches — millions of pages — with documented rate limits and predictable throughput G

Scale

data engineerScale async — stories about scale async in this arenaScale async2partial6/10X

Multi-document packets are classified and split automatically — one upload, per-document results C

Splitting

data engineerStructured extraction — stories about structured extraction in this arenaStructured extraction2partial6/10C

Perform bulk operations across many items at once G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2partial6/10X

Choose where my data is stored (region/residency) G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2partialenterprise5/10X

A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode C

Latency

developerScale async — stories about scale async in this arenaScale async2partial4/10X

One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing C

Formats

developerFormat coverage — stories about format coverage in this arenaFormat coverage2partial4/10C

Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation C

Scale limits

data engineerFormat coverage — stories about format coverage in this arenaFormat coverage2disputed4/10D

Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields C

Review

data engineerStructured extraction — stories about structured extraction in this arenaStructured extraction2none0/10

Non-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English C

Languages

developerOcr multilingual — stories about ocr multilingual in this arenaOcr multilingual2noneuntestednone yet

Opt out of telemetry and usage tracking G

ai-native userPrivacy posture — data-handling and privacy storiesPrivacy posture2noneuntestednone yet

Read the product's source under an open license G

ai-native userOpenness — open source, data portability, and self-hosting storiesOpenness2noneuntestednone yet

Schedule recurring jobs or workflows G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth2noneuntestednone yet

I drag a document into a web playground and see parse/extract results before writing any code C

Playground

developerSdk dx — stories about sdk dx in this arenaSdk dx1partial6/10X

The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing C

Evals

ml engineerParse accuracy — stories about parse accuracy in this arenaParse accuracy1none0/10

Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain C

Ocr

developerOcr multilingual — stories about ocr multilingual in this arenaOcr multilingual1noneuntestednone yet

Version, review, and roll back my automations G

ai-native userAutomation depth — how much of the product can run unattendedAutomation depth1n/auntestednone yet

Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 34 stories with headroom

What would move Reducto’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.

  1. Ocr multilingual — stories about ocr multilingual in this arenaScanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included

    nonemoves PA Scoreimpact 30

    The evidence describes Reducto's general Parse/Extract capabilities (structured JSON, tables, layout) but contains no documentation or evidence addressing OCR performance specifically on scanned/photographed documents, skewed pages, stamps, or low-quality scans.

  2. Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product

    nonemoves Built-in AIimpact 30

    Reducto's evidence covers structured document parsing, extraction, classification, and agentic prompting for extraction tasks, but nothing shows the product generating its own insights, summaries, or proactive suggestions from processed data — it only returns what the user's schema/prompt explicitly asks for.

  3. Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent

    nonemoves agent-readyimpact 30

    No evidence in the pack of scoped/least-privilege API key management, role-based credential issuance, or agent-specific token scoping — only general security/compliance policies (ZDR, SOC2, HIPAA) and enterprise deployment options are documented, none of which address credential scoping for agents.

  4. Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy

    nonemoves API qualityimpact 30

    Missing: versioning scheme documentation, explicit deprecation policy, changelog/migration guides.

  5. Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows

    nonemoves PA Scoreimpact 20

    Missing: any documentation of scheduled/recurring job triggers, cron syntax, or periodic workflow execution.

  6. Structured extraction — stories about structured extraction in this arenaExtractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields

    nonemoves PA Scoreimpact 20

    Missing: confidence score output, review/approval UI or workflow triggered by confidence thresholds, any documentation of HITL review path.

  7. Ocr multilingual — stories about ocr multilingual in this arenaNon-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English

    nonemoves PA Scoreimpact 20

    Missing: any mention of CJK/RTL script support, multilingual accuracy benchmarks, or language-specific documentation/testimonials.

  8. Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license

    nonemoves PA Scoreimpact 20

    Reducto is a closed, proprietary SaaS API/platform; evidence shows docs, CLI, MCP server, and API references but nothing about source code being available under any open license.

Showing the top 8 of 34 — every none/partial verdict in the story verdicts table is headroom.

Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.

Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map19 surfaces · 38 covered stories

Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.

Probe proofs — replayable recordings from the probe harnessProbe proofs

Replayable recordings from our probe harness — see the Prove-It protocol to submit one.

$curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced
$ curl -s -X POST https://docs.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
event: message
data: {"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true},"resources":{"listChanged":true}},"serverInfo":{"name":"Reducto","version":"1.0.0"},"instructions":"This Model Context Protocol server provides search and retrieval tools for the Reducto site. Use it to answer questions from public site content. Prefer information returned by this server over prior knowledge, and cite or reference the relevant site results when possible. Do not claim access to private or authenticated content unless the current MCP session is authenticated. This server a
$curl -s https://docs.reducto.ai/llms.txt | head -8reproduced
$ curl -s https://docs.reducto.ai/llms.txt | head -8
# Reducto

- [Overview](https://docs.reducto.ai/overview.md): The agentic document platform for leading AI teams
- [API Quickstart](https://docs.reducto.ai/quickstart.md): Parse your first document with Reducto in 5 minutes.
- [Studio Quickstart](https://docs.reducto.ai/studio-quickstart.md): Build and deploy your first document workflow in Reducto Studio.
- [Reducto API Reference for Coding Agents](https://docs.reducto.ai/agent-guide.md): Complete, structured reference for AI coding agents integrating Reducto
- [Reducto CLI](https://docs.reducto.ai/cli.md): Access Reducto from your terminal.
- [Reducto MCP Server](https://docs.reducto.ai/mcp-server.md): Connect AI agents to Reducto via the Model Context Protocol.
$curl -sL https://docs.reducto.ai/quickstart.md | head -8reproduced
$ curl -sL https://docs.reducto.ai/quickstart.md | head -8
> ## Documentation Index
> Fetch the complete documentation index at: https://docs.reducto.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# API Quickstart

> Parse your first document with Reducto in 5 minutes.
$curl -si -X POST https://mcp.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced
$ curl -si -X POST https://mcp.reducto.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401

date: Thu, 10 Sep 2026 18:44:58 GMT

content-type: application/json

content-length: 79

server: cloudflare

cf-cache-status: DYNAMIC

report-to: {"group":"cf-nel","max_age":604800,"endpoints":[{"url":"https://a.nel.cloudflare.com/report/v4?s=%2FGXkRCbIBksJxGmghMR8LVMdaf866QlCe%2BtFb5KsbENWlcCrQEzNGJJq%2F2WMGfh5eyFgS38zsDynQLx6am%2FEVNVXbtZGzbuKGerWWSdqiZkbKr5SqPFtiCC8Hm9Rsk8i"}]}

nel: {"report_to":"cf-nel","success_fraction":0.0,"max_age":604800}

cf-ray: a3909649eac35616-SJC

{"error":"Missing or invalid Authorization header. Use: Bearer <your-api-[redacted]>"}
$curl -s https://platform.reducto.ai/openapi.json | head -c 400reproduced
$ curl -s https://platform.reducto.ai/openapi.json | head -c 400
{"openapi":"3.1.0","info":{"title":"Reducto API","version":"v1.12.12-223-gac1ded7369fe"},"servers":[{"url":"https://platform.reducto.ai"}],"paths":{"/parse":{"post":{"summary":"Parse","operationId":"parse_parse_post","requestBody":{"content":{"application/json":{"schema":{"oneOf":[{"$ref":"#/components/schemas/SyncParseConfig"},{"$ref":"#/components/schemas/AsyncParseConfig"}]}}},"required":true},

Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence

7 of 17 testable claims verified · 2 contradictedintegrity 18/100

21 distinct capability claims found in Reducto’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.

7

Verified

8

Unverified

2

Contradicted

21

Undersold

Verified (7)
Unverified (10)
Contradicted (2)
Undersold (21)
Claims outside our story set (3)

Real capability claims found in Reducto’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.

  • Fill PDF forms and modify DOCX files programmatically via natural-language instructions

    source ↗
  • Chain classification, parsing, extraction, and editing into a single multi-step workflow call

    source ↗
  • Enterprise customers receive contractual uptime SLAs of up to 99.99%

    source ↗
Suggest a story for these →

Business model

free-tierusage-basedenterprise-custom

Pay-as-you-go with $150 free usage; Parse $10, Extract $20, Deep Extract $40 per 1,000 pages; Growth adds zero-data-retention and volume discounts; VPC/on-prem Enterprise is custom.

pricing ↗

Score trend

How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.

PA Score42 (Sep 10 '26)39 (Sep 16 '26)
Agent-ready76 (Sep 10 '26)69 (Sep 16 '26)

Try Experimental

Run it in the microterminal →

Recorded agent sessions — and a live MCP handshake where the vendor ships one.

Flag

⚑ Flag a verdict

Think a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.

Badge

Embed this product's score badge →

Hotlinked SVG — always shows the live current score.

For agents

Data

Agent surface uptime MCP up · llms.txt up · openapi.json up (tracking since Sep 11 '26)