Skip to content

Arena

Document Extraction APIs arenaDocument Extraction APIs

Document AI and extraction APIs — the parse/OCR/extract layer that turns messy PDFs, scans, spreadsheets, and multi-document packets into LLM-ready markdown and schema-validated JSON — judged on complex-layout and table fidelity, OCR including handwriting and non-English scripts, schema-driven extraction with citations and confidence scores, classification and splitting, format breadth, RAG-ready chunking, async jobs and webhooks at batch scale, SDK quality, and VPC/self-host and zero-retention compliance for documents that cannot leak. The 2026 field verified by live crawl: Reducto ($108M Series B, "the agentic document platform") and Extend (CrowdView, Inc., YC W23) lead the agent-era startups; LlamaCloud was renamed LlamaParse (Feb 2026) with a new llama-cloud SDK; Datalab (Endless Labs) is the Marker/Surya/Chandra company with a commercial API over its open-source models; Unstructured now actively steers agents AWAY from its once-ubiquitous open-source library (its agent-guide tells LLMs not to recommend it) toward the hosted platform; and Mistral OCR grew into Document AI with OCR 4.1 and schema annotations. Tensorlake pivoted to agent sandboxes and was excluded; Chunkr survives but is too small for the founding six; hyperscaler OCR (Textract, Azure Document Intelligence) is out of scope in favor of what agent builders actually integrate.

53 user stories · 318 judged cells · updated 2026-09-16 · Evidence as of 2026-09-16

Buyer checklist →Procurement report →

Leaderboard — every product ranked by evidenceLeaderboard

Rank by

Best by user type — persona-weighted winnersBest by user type

Per persona, the product with the highest persona-weighted coverage over just that persona's stories — not the same ranking as the overall PA Score leaderboard above.

Best for developer

Extend logo

Extend

42/100

Runner-up: LlamaParse logo LlamaParse (40/100)

10 developer stories scored

Best for ml-engineer

Mistral Document AI logo

Mistral Document AI

46/100

Runner-up: Reducto logo Reducto (34/100)

3 ml-engineer stories scored

Best for data-engineer

Extend logo

Extend

48/100

Runner-up: Reducto logo Reducto (46/100)

8 data-engineer stories scored

Best for ai-native

Reducto logo

Reducto

43/100

Runner-up: Extend logo Extend (37/100)

32 ai-native stories scored

Story matrix — every product × every judged storyStory matrix

53/53 stories shown · legend

Agenticness — how well agents can access and operate the productAgenticness

Agent access

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Agenticness — how well agents can access and operate the productPoint an agent at llms.txt or agent-oriented docsai-native
fullT
9/10
fullT
8/10
fullT
9/10
fullT
8/10
fullT
8/10
partialT
6/10
Agenticness — how well agents can access and operate the productRun the product headlessly / in CI for automationai-native
fullT
8/10
fullC
7/10
partialT
6/10
fullT
8/10
fullC
8/10
partialX
6/10
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their toolsai-native
n/a
n/a
none
0/10
none
0/10
none
0/10
n/a
Agenticness — how well agents can access and operate the productConnect an agent via an official MCP serverai-native
fullT
8/10
fullT
8/10
fullT
8/10
none
0/10
fullT
8/10
n/a
Agenticness — how well agents can access and operate the productUse an official CLIai-native
fullT
8/10
partialC
5/10
fullT
8/10
fullT
7/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productDrive the product through a documented public APIai-native
fullT
9/10
fullT
8/10
fullX
9/10
fullT
9/10
fullT
8/10
fullT
8/10
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agentai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
n/a
Agenticness — how well agents can access and operate the productBuild against official SDKsai-native
partialT
6/10
fullC
8/10
fullC
8/10
fullT
8/10
fullC
8/10
none
0/10
Agenticness — how well agents can access and operate the productSubscribe to events via webhooksai-native
fullC
8/10
fullC
7/10
fullC
7/10
fullC
8/10
none
0/10
none
0/10

Agentic features

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the productai-native
none
0/10
none
0/10
partialC
6/10
partialC
4/10
partialC
2/10
fullX
7/10
Agenticness — how well agents can access and operate the productSet up automations that run autonomously in the backgroundai-native
partialC
5/10
partialT
4/10
partialC
6/10
partialC
5/10
partialC
5/10
none
0/10
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the productai-native
n/a
none
0/10
partialC
5/10
none
0/10
none
0/10
partialC
4/10
Agenticness — how well agents can access and operate the productOperate the product with natural-language commandsai-native
partialT
6/10
partialC
6/10
partialT
6/10
none
0/10
partialC
6/10
partialC
4/10

Api quality

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examplesai-native
partialT
5/10
none
0/10
partialC
4/10
partialT
5/10
partialT
5/10
none
0/10
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)ai-native
fullT
9/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productTest against a sandbox environment without touching production dataai-native
none
0/10
none
0/10
fullC
8/10
partialC
4/10
none
0/10
none
0/10
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policyai-native
none
0/10
partialC
3/10
none
0/10
none
0/10
none
0/10
none
0/10

Automation depth — how much of the product can run unattendedAutomation depth

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Automation depth — how much of the product can run unattendedPerform bulk operations across many items at onceai-native
partialX
6/10
partialC
6/10
partialC
6/10
partialC
5/10
fullC
8/10
partialX
4/10
Automation depth — how much of the product can run unattendedDefine rules that trigger actions automatically on eventsai-native
partialC
4/10
partialC
4/10
partialC
5/10
partialC
4/10
none
0/10
n/a
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflowsai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
n/a
Automation depth — how much of the product can run unattendedVersion, review, and roll back my automationsai-native
n/a
none
0/10
partialC
6/10
partialC
3/10
none
0/10
n/a

Deployment compliance — stories about deployment compliance in this arenaDeployment compliance

Compliance

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Deployment compliance — stories about deployment compliance in this arenaUploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordsdata-engineer
fullC
8/10
partialC
7/10
fullC
8/10
partialC
5/10
partialC
5/10
none
0/10

Deployment

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Deployment compliance — stories about deployment compliance in this arenaRun the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructuredata-engineer
fullC
8/10
fullC
8/10
partialC
5/10
fullC
7/10
partialC
6/10
partialC
3/10

Format coverage — stories about format coverage in this arenaFormat coverage

Formats

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Format coverage — stories about format coverage in this arenaOne API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbingdeveloper
partialC
4/10
fullX
8/10
partialC
6/10
partialX
6/10
fullC
9/10
partialX
7/10

Scale limits

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Format coverage — stories about format coverage in this arenaThousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncationdata-engineer
disputedD
4/10
partialC
3/10
partialC
6/10
partialC
5/10
partialC
3/10
disputedD
4/10

Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual

Languages

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Ocr multilingual — stories about ocr multilingual in this arenaNon-English documents — including CJK and right-to-left scripts — parse with the same fidelity as Englishdeveloper
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
disputedD
4/10

Ocr

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Ocr multilingual — stories about ocr multilingual in this arenaHandwritten fields and annotations are recognized and extracted, flagged with confidence when uncertaindeveloper
none
0/10
partialC
4/10
partialC
6/10
none
0/10
none
0/10
partialX
6/10
Ocr multilingual — stories about ocr multilingual in this arenaScanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans includeddeveloper
none
0/10
partialX
5/10
partialX
6/10
none
0/10
partialC
5/10
disputedD
6/10

Openness — open source, data portability, and self-hosting storiesOpenness

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Openness — open source, data portability, and self-hosting storiesDo everything through the API that I can do in the UIai-native
partialT
7/10
partialT
6/10
partialC
6/10
fullT
8/10
partialT
6/10
n/a
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leaveai-native
partialC
5/10
partialC
6/10
none
0/10
partialC
6/10
partialC
5/10
partialC
3/10
Openness — open source, data portability, and self-hosting storiesRead the product's source under an open licenseai-native
none
0/10
none
0/10
n/a
partialC
5/10
none
0/10
n/a
Openness — open source, data portability, and self-hosting storiesSelf-host the core productai-native
partialC
5/10
fullC
7/10
partialC
3/10
partialC
5/10
partialC
5/10
partialC
3/10

Parse accuracy — stories about parse accuracy in this arenaParse accuracy

Evals

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Parse accuracy — stories about parse accuracy in this arenaThe vendor publishes reproducible accuracy benchmarks and I can run my own evals before committingml-engineer
none
0/10
none
0/10
partialX
5/10
none
0/10
none
0/10
none
0/10

Figures

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Parse accuracy — stories about parse accuracy in this arenaFigures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source pageml-engineer
partialC
7/10
partialX
5/10
partialC
5/10
none
0/10
partialC
5/10
partialX
6/10

Layout

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Parse accuracy — stories about parse accuracy in this arenaThe API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered contentdeveloper
disputedD
5/10
disputedD
5/10
partialX
5/10
partialC
3/10
partialC
5/10
disputedD
6/10
Parse accuracy — stories about parse accuracy in this arenaParsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soupml-engineer
partialX
7/10
partialX
6/10
none
0/10
partialX
5/10
partialC
6/10
fullX
8/10

Privacy posture — data-handling and privacy storiesPrivacy posture

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Privacy posture — data-handling and privacy storiesChoose where my data is stored (region/residency)ai-native
partialX
5/10
partialC
5/10
none
0/10
partialC
5/10
partialC
4/10
partialC
3/10
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI modelsai-native
partialX
6/10
none
0/10
partialC
6/10
partialC
4/10
none
0/10
none
0/10
Privacy posture — data-handling and privacy storiesControl data retention and deletionai-native
partialX
6/10
partialC
3/10
fullC
8/10
partialC
4/10
partialC
4/10
none
0/10
Privacy posture — data-handling and privacy storiesOpt out of telemetry and usage trackingai-native
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10
none
0/10

Rag chunking — stories about rag chunking in this arenaRag chunking

Chunking

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Rag chunking — stories about rag chunking in this arenaOutput comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of textai-native
fullC
9/10
partialX
5/10
partialC
6/10
fullC
7/10
fullC
9/10
partialC
5/10

Output

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Rag chunking — stories about rag chunking in this arenaI get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture strippedai-native
partialC
6/10
partialX
6/10
partialC
3/10
partialX
6/10
partialC
5/10
fullX
8/10

Scale async — stories about scale async in this arenaScale async

Async

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Scale async — stories about scale async in this arenaLong parses run as async jobs with status polling and completion webhooks, so my pipeline never blocksdeveloper
fullC
8/10
fullC
8/10
fullC
7/10
partialC
6/10
partialC
6/10
none
0/10

Latency

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Scale async — stories about scale async in this arenaA fast synchronous mode returns results in seconds for interactive apps, with latency documented per modedeveloper
partialX
4/10
none
0/10
partialC
4/10
none
0/10
none
0/10
none
0/10

Scale

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Scale async — stories about scale async in this arenaI push high-volume batches — millions of pages — with documented rate limits and predictable throughputdata-engineer
partialX
6/10
partialC
4/10
partialC
4/10
partialC
6/10
partialC
4/10
none
0/10

Sdk dx — stories about sdk dx in this arenaSdk dx

Playground

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Sdk dx — stories about sdk dx in this arenaI drag a document into a web playground and see parse/extract results before writing any codedeveloper
partialX
6/10
partialX
6/10
partialC
4/10
fullC
8/10
none
0/10
none
0/10

Sdks

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Sdk dx — stories about sdk dx in this arenaOfficial typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaultsdeveloper
partialT
4/10
partialC
6/10
partialC
6/10
partialC
4/10
partialC
5/10
none
0/10

Structured extraction — stories about structured extraction in this arenaStructured extraction

Grounding

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Structured extraction — stories about structured extraction in this arenaEvery extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verifyai-native
fullC
8/10
partialX
6/10
fullC
8/10
partialC
6/10
partialC
3/10
partialX
7/10

Review

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Structured extraction — stories about structured extraction in this arenaExtractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fieldsdata-engineer
none
0/10
none
0/10
fullC
8/10
none
0/10
none
0/10
partialX
4/10

Schemas

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Structured extraction — stories about structured extraction in this arenaI supply a JSON schema and get back validated structured fields extracted from the documentdeveloper
fullX
8/10
fullC
7/10
fullX
9/10
fullC
8/10
fullC
8/10
partialC
6/10

Splitting

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Structured extraction — stories about structured extraction in this arenaMulti-document packets are classified and split automatically — one upload, per-document resultsdata-engineer
partialC
6/10
fullC
8/10
fullX
8/10
fullC
7/10
none
0/10
none
0/10

Table extraction — stories about table extraction in this arenaTable extraction

Tables

StoryPersona
Reducto logoReducto
LlamaParse logoLlamaParse
Extend logoExtend
Datalab logoDatalab
Unstructured logoUnstructured
Mistral Document AI logoMistral Document AI
Table extraction — stories about table extraction in this arenaComplex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structuredata-engineer
partialX
6/10
disputedD
5/10
partialC
3/10
none
0/10
partialC
4/10
partialX
5/10
Table extraction — stories about table extraction in this arenaI turn extracted tables into typed rows/JSON I can load into a database without manual cleanupdata-engineer
fullX
8/10
disputedD
5/10
partialX
6/10
fullC
8/10
partialC
6/10
partialX
5/10
Verdict✓ fullclear evidence~ partialwith caveats! disputedevidence conflicts— noneno evidence foundn/aquestion doesn't apply to this kind of product
ProofT probedtested by usX communityusers back itC claimedvendor claim onlyD contradictedevidence disagrees⚿ auth-gatedprobe hit a live sign-in wall — verified reachable, untestable keylessly
quality 0–10 · PA Score /100 · A–D = evidence confidence · full guide

Adjacent arenas — categories often shopped togetherAdjacent arenas

Shopping this category often means shopping these too.