Document Extraction APIs arenaDocument Extraction APIs
Document AI and extraction APIs — the parse/OCR/extract layer that turns messy PDFs, scans, spreadsheets, and multi-document packets into LLM-ready markdown and schema-validated JSON — judged on complex-layout and table fidelity, OCR including handwriting and non-English scripts, schema-driven extraction with citations and confidence scores, classification and splitting, format breadth, RAG-ready chunking, async jobs and webhooks at batch scale, SDK quality, and VPC/self-host and zero-retention compliance for documents that cannot leak. The 2026 field verified by live crawl: Reducto ($108M Series B, "the agentic document platform") and Extend (CrowdView, Inc., YC W23) lead the agent-era startups; LlamaCloud was renamed LlamaParse (Feb 2026) with a new llama-cloud SDK; Datalab (Endless Labs) is the Marker/Surya/Chandra company with a commercial API over its open-source models; Unstructured now actively steers agents AWAY from its once-ubiquitous open-source library (its agent-guide tells LLMs not to recommend it) toward the hosted platform; and Mistral OCR grew into Document AI with OCR 4.1 and schema annotations. Tensorlake pivoted to agent sandboxes and was excluded; Chunkr survives but is too small for the founding six; hyperscaler OCR (Textract, Azure Document Intelligence) is out of scope in favor of what agent builders actually integrate.
53 user stories · 318 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16
Leaderboard — every product ranked by evidenceLeaderboard
| 1 | free-tier vs Extend ↗ | 69/100 | 22/100 | 34/100 | 26/100 | 21/100 | npm 118.8k/wkpypi 49.9k/wk | 22/38 verified · 2 disputed | 18/100 integrity | ||
| 2 | free-tier vs Reducto ↗ | 58/100 | 34/100 | 18/100 | 16/100 | 25/100 | npm 44k/wkpypi 48.8k/wk | 12/42 verified | 30/100 integrity | ||
| 3 | free-tier vs Reducto ↗ | 63/100 | 13/100 | 5/100 | 39/100 | 18/100 | ★ 4.3k▲ 1.6k/yrnpm 33.2k/wkpypi 637.6k/wk | 13/37 verified · 3 disputed | 12/100 integrity | ||
| 4 | free-tier vs Reducto ↗ | 50/100 | 12/100 | 12/100 | 42/100 | 19/100 | ★ 39.6k▲ 13.8k/yrpypi 30k/wk | 10/36 verified | 18/100 integrity | ||
| 5 | free-tier vs Reducto ↗ | 46/100 | 17/100 | 9/100 | 25/100 | 20/100 | ★ 15.4k▲ 3.9k/yrpypi 719k/wk | 5/32 verified | 23/100 integrity | ||
| 6 | usage-based vs Reducto ↗ | 30/100 | 29/100 | 0/100 | 18/100 | 24/100 | 14/26 verified · 4 disputed | 8/100 integrity |
Best by user type — persona-weighted winnersBest by user type
Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.
Best for ml-engineer
Mistral Document AI
46/100
Runner-up:
Reducto (34/100)
3 ml-engineer stories scored
Story matrix — every product × every judged storyStory matrix
Agenticness — how well agents can access and operate the productAgenticness
Agent access
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docs | ai-native | fullT 9/10 | fullT 8/10 | fullT 9/10 | fullT 8/10 | fullT 8/10 | partialT 6/10 |
| Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automation | ai-native | fullT 8/10 | fullC 7/10 | partialT 6/10 | fullT 8/10 | fullC 8/10 | partialX 6/10 |
| Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools | ai-native | n/a | n/a | none 0/10 | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productConnect an agent via an official MCP server | ai-native | fullT 8/10 | fullT 8/10 | fullT 8/10 | none 0/10 | fullT 8/10 | n/a |
| Agenticness — how well agents can access and operate the productUse an official CLI | ai-native | fullT 8/10 | partialC 5/10 | fullT 8/10 | fullT 7/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDrive the product through a documented public API | ai-native | fullT 9/10 | fullT 8/10 | fullX 9/10 | fullT 9/10 | fullT 8/10 | fullT 8/10 |
| Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a |
| Agenticness — how well agents can access and operate the productBuild against official SDKs | ai-native | partialT 6/10 | fullC 8/10 | fullC 8/10 | fullT 8/10 | fullC 8/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productSubscribe to events via webhooks | ai-native | fullC 8/10 | fullC 7/10 | fullC 7/10 | fullC 8/10 | none 0/10 | none 0/10 |
Agentic features
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product | ai-native | none 0/10 | none 0/10 | partialC 6/10 | partialC 4/10 | partialC 2/10 | fullX 7/10 |
| Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the background | ai-native | partialC 5/10 | partialT 4/10 | partialC 6/10 | partialC 5/10 | partialC 5/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product | ai-native | n/a | none 0/10 | partialC 5/10 | none 0/10 | none 0/10 | partialC 4/10 |
| Agenticness — how well agents can access and operate the productOperate the product with natural-language commands | ai-native | partialT 6/10 | partialC 6/10 | partialT 6/10 | none 0/10 | partialC 6/10 | partialC 4/10 |
Api quality
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples | ai-native | partialT 5/10 | none 0/10 | partialC 4/10 | partialT 5/10 | partialT 5/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent) | ai-native | fullT 9/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production data | ai-native | none 0/10 | none 0/10 | fullC 8/10 | partialC 4/10 | none 0/10 | none 0/10 |
| Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy | ai-native | none 0/10 | partialC 3/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Automation depth — how much of the product can run unattendedAutomation depth
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Automation depth — how much of the product can run unattendedPerform bulk operations across many items at once | ai-native | partialX 6/10 | partialC 6/10 | partialC 6/10 | partialC 5/10 | fullC 8/10 | partialX 4/10 |
| Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on events | ai-native | partialC 4/10 | partialC 4/10 | partialC 5/10 | partialC 4/10 | none 0/10 | n/a |
| Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | n/a |
| Automation depth — how much of the product can run unattendedVersion, review, and roll back my automations | ai-native | n/a | none 0/10 | partialC 6/10 | partialC 3/10 | none 0/10 | n/a |
Deployment compliance — stories about deployment compliance in this arenaDeployment compliance
Compliance
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Deployment compliance — stories about deployment compliance in this arenaUploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records | data-engineer | fullC 8/10 | partialC 7/10 | fullC 8/10 | partialC 5/10 | partialC 5/10 | none 0/10 |
Deployment
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Deployment compliance — stories about deployment compliance in this arenaRun the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure | data-engineer | fullC 8/10 | fullC 8/10 | partialC 5/10 | fullC 7/10 | partialC 6/10 | partialC 3/10 |
Format coverage — stories about format coverage in this arenaFormat coverage
Formats
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Format coverage — stories about format coverage in this arenaOne API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing | developer | partialC 4/10 | fullX 8/10 | partialC 6/10 | partialX 6/10 | fullC 9/10 | partialX 7/10 |
Scale limits
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Format coverage — stories about format coverage in this arenaThousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation | data-engineer | disputedD 4/10 | partialC 3/10 | partialC 6/10 | partialC 5/10 | partialC 3/10 | disputedD 4/10 |
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual
Languages
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Ocr multilingual — stories about ocr multilingual in this arenaNon-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English | developer | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | disputedD 4/10 |
Ocr
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Ocr multilingual — stories about ocr multilingual in this arenaHandwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain | developer | none 0/10 | partialC 4/10 | partialC 6/10 | none 0/10 | none 0/10 | partialX 6/10 |
| Ocr multilingual — stories about ocr multilingual in this arenaScanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included | developer | none 0/10 | partialX 5/10 | partialX 6/10 | none 0/10 | partialC 5/10 | disputedD 6/10 |
Openness — open source, data portability, and self-hosting storiesOpenness
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UI | ai-native | partialT 7/10 | partialT 6/10 | partialC 6/10 | fullT 8/10 | partialT 6/10 | n/a |
| Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave | ai-native | partialC 5/10 | partialC 6/10 | none 0/10 | partialC 6/10 | partialC 5/10 | partialC 3/10 |
| Openness — open source, data portability, and self-hosting storiesRead the product's source under an open license | ai-native | none 0/10 | none 0/10 | n/a | partialC 5/10 | none 0/10 | n/a |
| Openness — open source, data portability, and self-hosting storiesSelf-host the core product | ai-native | partialC 5/10 | fullC 7/10 | partialC 3/10 | partialC 5/10 | partialC 5/10 | partialC 3/10 |
Parse accuracy — stories about parse accuracy in this arenaParse accuracy
Evals
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Parse accuracy — stories about parse accuracy in this arenaThe vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing | ml-engineer | none 0/10 | none 0/10 | partialX 5/10 | none 0/10 | none 0/10 | none 0/10 |
Figures
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Parse accuracy — stories about parse accuracy in this arenaFigures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page | ml-engineer | partialC 7/10 | partialX 5/10 | partialC 5/10 | none 0/10 | partialC 5/10 | partialX 6/10 |
Layout
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Parse accuracy — stories about parse accuracy in this arenaThe API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content | developer | disputedD 5/10 | disputedD 5/10 | partialX 5/10 | partialC 3/10 | partialC 5/10 | disputedD 6/10 |
| Parse accuracy — stories about parse accuracy in this arenaParsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup | ml-engineer | partialX 7/10 | partialX 6/10 | none 0/10 | partialX 5/10 | partialC 6/10 | fullX 8/10 |
Privacy posture — data-handling and privacy storiesPrivacy posture
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency) | ai-native | partialX 5/10 | partialC 5/10 | none 0/10 | partialC 5/10 | partialC 4/10 | partialC 3/10 |
| Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models | ai-native | partialX 6/10 | none 0/10 | partialC 6/10 | partialC 4/10 | none 0/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesControl data retention and deletion | ai-native | partialX 6/10 | partialC 3/10 | fullC 8/10 | partialC 4/10 | partialC 4/10 | none 0/10 |
| Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage tracking | ai-native | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 | none 0/10 |
Rag chunking — stories about rag chunking in this arenaRag chunking
Chunking
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Rag chunking — stories about rag chunking in this arenaOutput comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text | ai-native | fullC 9/10 | partialX 5/10 | partialC 6/10 | fullC 7/10 | fullC 9/10 | partialC 5/10 |
Output
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Rag chunking — stories about rag chunking in this arenaI get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped | ai-native | partialC 6/10 | partialX 6/10 | partialC 3/10 | partialX 6/10 | partialC 5/10 | fullX 8/10 |
Scale async — stories about scale async in this arenaScale async
Async
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Scale async — stories about scale async in this arenaLong parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks | developer | fullC 8/10 | fullC 8/10 | fullC 7/10 | partialC 6/10 | partialC 6/10 | none 0/10 |
Latency
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Scale async — stories about scale async in this arenaA fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode | developer | partialX 4/10 | none 0/10 | partialC 4/10 | none 0/10 | none 0/10 | none 0/10 |
Scale
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Scale async — stories about scale async in this arenaI push high-volume batches — millions of pages — with documented rate limits and predictable throughput | data-engineer | partialX 6/10 | partialC 4/10 | partialC 4/10 | partialC 6/10 | partialC 4/10 | none 0/10 |
Sdk dx — stories about sdk dx in this arenaSdk dx
Playground
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Sdk dx — stories about sdk dx in this arenaI drag a document into a web playground and see parse/extract results before writing any code | developer | partialX 6/10 | partialX 6/10 | partialC 4/10 | fullC 8/10 | none 0/10 | none 0/10 |
Sdks
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Sdk dx — stories about sdk dx in this arenaOfficial typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults | developer | partialT 4/10 | partialC 6/10 | partialC 6/10 | partialC 4/10 | partialC 5/10 | none 0/10 |
Structured extraction — stories about structured extraction in this arenaStructured extraction
Grounding
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaEvery extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify | ai-native | fullC 8/10 | partialX 6/10 | fullC 8/10 | partialC 6/10 | partialC 3/10 | partialX 7/10 |
Review
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaExtractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields | data-engineer | none 0/10 | none 0/10 | fullC 8/10 | none 0/10 | none 0/10 | partialX 4/10 |
Schemas
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaI supply a JSON schema and get back validated structured fields extracted from the document | developer | fullX 8/10 | fullC 7/10 | fullX 9/10 | fullC 8/10 | fullC 8/10 | partialC 6/10 |
Splitting
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Structured extraction — stories about structured extraction in this arenaMulti-document packets are classified and split automatically — one upload, per-document results | data-engineer | partialC 6/10 | fullC 8/10 | fullX 8/10 | fullC 7/10 | none 0/10 | none 0/10 |
Table extraction — stories about table extraction in this arenaTable extraction
Tables
| Story | Persona | ||||||
|---|---|---|---|---|---|---|---|
| Table extraction — stories about table extraction in this arenaComplex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure | data-engineer | partialX 6/10 | disputedD 5/10 | partialC 3/10 | none 0/10 | partialC 4/10 | partialX 5/10 |
| Table extraction — stories about table extraction in this arenaI turn extracted tables into typed rows/JSON I can load into a database without manual cleanup | data-engineer | fullX 8/10 | disputedD 5/10 | partialX 6/10 | fullC 8/10 | partialC 6/10 | partialX 5/10 |
Adjacent arenas — categories often shopped togetherAdjacent arenas
Shopping this category often means shopping these too.
Web Scraping APIs arenaWeb Scraping APIs
8 products · leader: Apify
AI Search APIs arenaAI Search APIs
5 products · leader: Tavily
Vector Databases & Memory Stores arenaVector Databases & Memory Stores
7 products · leader: Chroma
Browser Automation for Agents arenaBrowser Automation for Agents
7 products · leader: Steel
shares: Structured extraction