Rank #2 of 6 in Document Extraction APIs
Install
Showcase


Try itExperimental
See what an agent can do with Extend before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -s https://api.extend.ai/extractorsrecorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment complianceevidence →
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverageevidence →
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingualevidence →
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracyevidence →
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunkingevidence →
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale asyncevidence →
Stories about scale async in this arena
Sdk dx — stories about sdk dx in this arenaSdk dxevidence →
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extractionevidence →
Stories about structured extraction in this arena
Table extraction — stories about table extraction in this arenaTable extractionevidence →
Stories about table extraction in this arena
Story verdicts — every judged story with its evidenceStory verdicts
Follow the green: where the map greys out is where Extend stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓9/10
unlocks → Scoped API keys · Machine-readable spec · Versioning policy · Full data export
Subscribe to events via webhooks
✓7/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—–
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
—–
Rely on versioned APIs with a documented deprecation policy
—0/10
Test against a sandbox environment without touching production data
✓8/10
Explore an interactive API reference with runnable examples
~4/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓9/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
~5/10
unlocks → MCP client
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
—0/10
Get AI-generated insights and suggestions from my data inside the product
~6/10
Set up automations that run autonomously in the background
~6/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment compliance
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverage
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracy
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunking
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale async
Stories about scale async in this arena
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
✓7/10
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
~4/10
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
~4/10
Sdk dx — stories about sdk dx in this arenaSdk dx
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extraction
Stories about structured extraction in this arena
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
✓8/10
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
✓8/10
I supply a JSON schema and get back validated structured fields extracted from the document
✓9/10
Multi-document packets are classified and split automatically — one upload, per-document results
✓8/10
Table extraction — stories about table extraction in this arenaTable extraction
Stories about table extraction in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 9/10 | Xcommunity | |
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | partial | 5/10 | Cclaimed | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | 0/10 | ||
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 9/10 | Tprobed | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Tprobed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Cclaimed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | untested | none yet | |
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | full | 8/10 | Cclaimed | |
I supply a JSON schema and get back validated structured fields extracted from the document C Schemas | developer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 3 | full | 9/10 | Xcommunity | |
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records C Compliance | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 3 | full | 8/10 | Cclaimed | |
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks C Async | developer | Scale async — stories about scale async in this arenaScale async | 3 | full | 7/10 | Cclaimed | |
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults C Sdks | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 3 | partial | 6/10 | Cclaimed | |
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text C Chunking | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 3 | partial | 6/10 | Cclaimed | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | partial | 6/10 | Cclaimed | |
Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 3 | partial | 6/10 | Xcommunity | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 5/10 | Cclaimed | |
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content C Layout | developer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 3 | partial | 5/10 | Xcommunity | |
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 3 | partial | 3/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 3/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | none | 0/10 | ||
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | full | 8/10 | Cclaimed | |
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify C Grounding | ai-native user | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | full | 8/10 | Cclaimed | |
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields C Review | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | full | 8/10 | Cclaimed | |
Multi-document packets are classified and split automatically — one upload, per-document results C Splitting | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | full | 8/10 | Xcommunity | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Cclaimed | |
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 2 | partial | 6/10 | Xcommunity | |
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing C Formats | developer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | partial | 6/10 | Cclaimed | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation C Scale limits | data engineer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | partial | 6/10 | Cclaimed | |
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page C Figures | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | partial | 5/10 | Cclaimed | |
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure C Deployment | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 2 | partial | 5/10 | Cclaimed | |
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode C Latency | developer | Scale async — stories about scale async in this arenaScale async | 2 | partial | 4/10 | Cclaimed | |
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput G Scale | data engineer | Scale async — stories about scale async in this arenaScale async | 2 | partial | 4/10 | Cclaimed | |
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped C Output | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 2 | partial | 3/10 | Cclaimed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | 0/10 | ||
Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup C Layout | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
Non-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English C Languages | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | n/a | untested | none yet | |
Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 1 | partial | 6/10 | Cclaimed | |
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | partial | 6/10 | Cclaimed | |
The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing C Evals | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 1 | partial | 5/10 | Xcommunity | |
I drag a document into a web playground and see parse/extract results before writing any code C Playground | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 1 | partial | 4/10 | Cclaimed |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 38 stories with headroom
What would move Extend’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productPlug MCP servers into this product so it can use their tools
nonemoves agent-readyimpact 45
Extend documents 'Using Extend via MCP' (docs.extend.ai/mcp), which describes exposing Extend's own tools via MCP to other agents — this is the opposite direction of the story (product consuming external MCP servers as a client).
Openness — open source, data portability, and self-hosting storiesExport all of my data in open formats and leave
nonemoves PA Scoreimpact 30
Extend is a SaaS document-processing platform holding workflows, processors, evaluation sets and extracted data, so data-portability/export is a fair question, but the evidence pack contains no mention of a bulk data-export feature, open-format export of processed data/configs, or account-closure data dump — only retention/ZDR policies which describe deletion, not export.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
Evidence only shows basic API-key authentication (extend-docs-4) and a separate test-vs-production API key for sandboxing (extend-docs-38), but nothing about issuing scoped, role/permission-limited, or least-privilege credentials for individual agents (e.g., granular scopes, RBAC, per-agent key restrictions).
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
Missing: openAPI/Swagger spec file, documented spec download link, machine-readable API schema.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
nonemoves API qualityimpact 30
Evidence shows versioning for internal artifacts (processors, workflows, base models) but no documentation of API endpoint versioning (e.g., v1/v2 paths) or any deprecation/sunset policy for the REST API/SDK itself.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
partialq5/10moves Built-in AIimpact 22.5
Missing: a documented chat/assistant interface, examples of delegating broader tasks beyond document processing, and independent hands-on confirmation of these agent features in use.
Openness — open source, data portability, and self-hosting storiesSelf-host the core product
partialq3/10moves PA Scoreimpact 21
Missing: open-source or downloadable core product, self-hosting setup docs, infrastructure requirements, and independent confirmation that customers can run it fully outside Extend's cloud.
Table extraction — stories about table extraction in this arenaComplex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
partialq3/10moves PA Scoreimpact 21
Missing: explicit handling of merged cells, nested headers, multi-page table continuity, and documented HTML/markdown fidelity examples.
Showing the top 8 of 38 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map19 surfaces · 42 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Pricing docs22 stories
- Get AI-generated insights and suggestions from my data inside the product
- Set up automations that run autonomously in the background
- Delegate tasks to a built-in AI assistant inside the product
- Define rules that trigger actions automatically on events
- Version, review, and roll back my automations
- Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records
- Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- Do everything through the API that I can do in the UI
- Self-host the core product
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Prevent my data from being used to train AI models
- Control data retention and deletion
- A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
- I drag a document into a web playground and see parse/extract results before writing any code
- Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
General docs16 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Test against a sandbox environment without touching production data
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
- A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
- I supply a JSON schema and get back validated structured fields extracted from the document
- Multi-document packets are classified and split automatically — one upload, per-document results
Extraction docs11 stories
- Get AI-generated insights and suggestions from my data inside the product
- Perform bulk operations across many items at once
- Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
- I supply a JSON schema and get back validated structured fields extracted from the document
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Hacker News7 stories
- Drive the product through a documented public API
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- I supply a JSON schema and get back validated structured fields extracted from the document
- Multi-document packets are classified and split automatically — one upload, per-document results
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
extend.ai7 stories
- Run the product headlessly / in CI for automation
- Build against official SDKs
- Get AI-generated insights and suggestions from my data inside the product
- Delegate tasks to a built-in AI assistant inside the product
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
CLI docs6 stories
API quickstart docs6 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Do everything through the API that I can do in the UI
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
API reference6 stories
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- Multi-document packets are classified and split automatically — one upload, per-document results
Sdks docs6 stories
- Run the product headlessly / in CI for automation
- Drive the product through a documented public API
- Build against official SDKs
- Explore an interactive API reference with runnable examples
- Do everything through the API that I can do in the UI
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
Evaluation docs5 stories
- Perform bulk operations across many items at once
- Version, review, and roll back my automations
- Do everything through the API that I can do in the UI
- The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing
- I drag a document into a web playground and see parse/extract results before writing any code
MCP docs4 stories
Agent quickstart docs4 stories
Webhooks docs4 stories
Workflows docs4 stories
Security docs3 stories
Splitting docs3 stories
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Multi-document packets are classified and split automatically — one upload, per-document results
Changelog docs2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -s https://api.extend.ai/extractorsreproduced$ curl -s https://api.extend.ai/extractors
{"code":"UNAUTHORIZED","message":"Authorization header is missing or malformed.","requestId":"req_9K64nYhzoBrFg1OHBjPaj","retryable":false}
$npx -y @extend-ai/cli --versionreproduced$ npx -y @extend-ai/cli --version extend version v0.8.0 (361e468)
$curl -s -X POST https://docs.extend.ai/_mcp/server -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -s -X POST https://docs.extend.ai/_mcp/server -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
event: message
data: {"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"fern-docs-mcp-server","version":"1.0.0"}},"jsonrpc":"2.0","id":1}
$curl -s https://docs.extend.ai/llms.txt | head -8reproduced$ curl -s https://docs.extend.ai/llms.txt | head -8 # Extend > Extend is a platform for building, iterating on, evaluating, and deploying AI-powered document processing infrastructure. We provide the tooling that you need like parsing, extraction, classification, splitting, and editing so that you can focus on building the features that matter for your business. Extend is a platform for building, iterating on, evaluating, and deploying AI-powered document processing infrastructure. We provide the tooling that you need like parsing, extraction, classification, splitting, and editing so that you can focus on building the features that matter for your business. Extend uses a mix of trained-in-house and frontier models to ensure the ideal mix of accuracy, latency, and cost for your use case. Works across 35+ file types, including PDFs, images, spreadsheets, presentations, and emails. ## Instructions for AI agents
$curl -si -X POST https://mcp.extend.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -si -X POST https://mcp.extend.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
HTTP/2 401
date: Thu, 10 Sep 2026 18:45:00 GMT
content-type: application/json; charset=utf-8
content-length: 193
www-authenticate: Bearer error="invalid_[redacted]", error_description="Authorization required", resource_metadata="https://mcp.extend.ai/.well-known/oauth-protected-resource/mcp"
etag: W/"c1-3fjLP0louiKHlRjsgxa7Uf9CD8Y"
{"error":"unauthorized","error_description":"Authorization required. Sign in via the advertised authorization server (the resource_metadata in WWW-Authenticate) to obtain an MCP access [redacted]."}
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
8 of 27 testable claims verified · 0 contradicted → integrity 30/100
32 distinct capability claims found in Extend’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
8
Verified
19
Unverified
0
Contradicted
15
Undersold
Verified (10)
“API supports request authentication for programmatic access”
Drive the product through a documented public APIfullproof ↗
“Provides an official command-line interface”
“Can be used via an MCP server so agents can invoke it as a tool”
“Automatically classifies document types”
Multi-document packets are classified and split automatically — one upload, per-document resultsfullproof ↗
“Automatically splits multi-document packets into individual documents”
Multi-document packets are classified and split automatically — one upload, per-document resultsfullproof ↗
“Supports schema-driven extraction where a JSON schema defines the fields to extract”
I supply a JSON schema and get back validated structured fields extracted from the documentfullproof ↗
“Includes an Evals feature for testing extraction accuracy”
The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committingpartialproof ↗
“Uses agentic OCR to read scanned/photographed documents”
Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans includedpartialproof ↗
“Detects individual document instances within a combined upload”
Multi-document packets are classified and split automatically — one upload, per-document resultsfullproof ↗
“Provides an agent quickstart guide for building document-processing agents”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
Unverified (23)
“Official Python package installable via pip for building against the API”
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaultspartialproof ↗
“Official Python package installable via pip for building against the API”
“Parse output includes a metadata object alongside pre-split chunks”
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of textpartialproof ↗
“Offers official SDKs for building integrations”
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaultspartialproof ↗
“Offers official SDKs for building integrations”
“Extractions include confidence scores for fields”
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fieldsfullproof ↗
“Supports configurable webhooks to notify on job/processing events”
“Supports configurable webhooks to notify on job/processing events”
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocksfullproof ↗
“Supports asynchronous processing of document jobs”
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocksfullproof ↗
“Supports a broad range of file types for document processing”
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbingpartialproof ↗
“Maintains SOC 2 Type II aligned controls plus HIPAA and GDPR documentation”
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordsfullproof ↗
“Offers configurable automatic data-retention policies, including zero data retention (ZDR) at the workspace level, even with AI subprocessors”
“Offers configurable automatic data-retention policies, including zero data retention (ZDR) at the workspace level, even with AI subprocessors”
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordsfullproof ↗
“Includes a Studio environment to try parsing/extraction before writing code”
I drag a document into a web playground and see parse/extract results before writing any codepartialproof ↗
“Provides Workflows for orchestrating document processing steps”
Define rules that trigger actions automatically on eventspartialproof ↗
“Provides a Composer and Review Agent for building and reviewing document workflows with AI assistance”
Delegate tasks to a built-in AI assistant inside the productpartialproof ↗
“Detects tables within documents for structured extraction”
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structurepartialproof ↗
“Detects images/figures within documents”
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source pagepartialproof ↗
“Detects handwriting within documents”
Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertainpartialproof ↗
“Supports processing of documents with 2,000+ pages”
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncationpartialproof ↗
“Offers a fast synchronous processing mode for quicker results”
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per modepartialproof ↗
“Provides citations linking extracted fields back to source locations”
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verifyfullproof ↗
“Provides an API quickstart guide for getting started with the public API”
Explore an interactive API reference with runnable examplespartialproof ↗
Undersold (15)
Run the product headlessly / in CI for automationpartialproof ↗
Get AI-generated insights and suggestions from my data inside the productpartialproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Test against a sandbox environment without touching production datafullproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructurepartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered contentpartialproof ↗
Prevent my data from being used to train AI modelspartialproof ↗
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture strippedpartialproof ↗
I push high-volume batches — millions of pages — with documented rate limits and predictable throughputpartialproof ↗
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanuppartialproof ↗
Claims outside our story set (3)
Real capability claims found in Extend’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“Parses, extracts, and splits complex documents with high accuracy, enabling document agents to be built quickly”
source ↗“Detects checkboxes as structured form elements”
source ↗“Supports agent-driven automatic form filling”
source ↗
Business model
Pay-as-you-go with 10,000 free credits then $0.0125/credit; Scale is $500/mo with 50,000 credits ($0.01/credit overage); Enterprise is custom.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP up · llms.txt up (tracking since Sep 11 '26)
