Rank #3 of 6 in Document Extraction APIs
Showcase


Try itExperimental
See what an agent can do with LlamaParse before you ever sign up. Pick a story: recorded sessions replay real probe-harness transcripts; commands tagged live-capable can re-run against the real endpoint from our edge, right now (▶ run live — the exact same request, live and recorded lines always labeled); the live MCP handshake runs real requests from our edge, right now — including, where the server allows it, one real read-only tool call (bring your own key for auth-gated servers); sandboxed self-drive sessions are designed and gated (docs/TRY-IT.md).
$curl -s -X POST https://api.cloud.llamaindex.ai/api/v1/parsing/uploadrecorded session — replayed, not liveVerified integrations
No integration evidence found in our corpus for this product yet — that means none was found, never that it doesn’t integrate.
By theme — the product's score on each story themeBy theme
Agenticness — how well agents can access and operate the productAgenticnessevidence →
How well agents can access and operate the product
Automation depth — how much of the product can run unattendedAutomation depthevidence →
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment complianceevidence →
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverageevidence →
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingualevidence →
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpennessevidence →
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracyevidence →
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy postureevidence →
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunkingevidence →
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale asyncevidence →
Stories about scale async in this arena
Sdk dx — stories about sdk dx in this arenaSdk dxevidence →
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extractionevidence →
Stories about structured extraction in this arena
Table extraction — stories about table extraction in this arenaTable extractionevidence →
Stories about table extraction in this arena
Story verdicts — every judged story with its evidenceStory verdicts
What’s free: 0 free · 0 paid · 4 enterprise · 31 not stated in evidence
Follow the green: where the map greys out is where LlamaParse stops today. ✓ full · ~ partial · ! disputed · — none · n/a not applicable.
Agenticness — how well agents can access and operate the productAgenticness
How well agents can access and operate the product
API surface
Drive the product through a documented public API
✓8/10
unlocks → Scoped API keys · Machine-readable spec · API sandbox
Subscribe to events via webhooks
✓7/10
Build against official SDKs
✓8/10
Issue scoped/least-privilege API credentials for an agent
—0/10
Connect an agent via an official MCP server
✓8/10
Download a machine-readable API spec (OpenAPI or equivalent)
—0/10
Rely on versioned APIs with a documented deprecation policy
~3/10
Test against a sandbox environment without touching production data
—–
Explore an interactive API reference with runnable examples
—0/10
Docs for agents
Point an agent at llms.txt or agent-oriented docs
✓8/10
Agentic features
Delegate tasks to a built-in AI assistant inside the product
—–
Operate the product with natural-language commands
~6/10
Plug MCP servers into this product so it can use their tools
n/an/a
Get AI-generated insights and suggestions from my data inside the product
—0/10
Set up automations that run autonomously in the background
~4/10
Automation depth — how much of the product can run unattendedAutomation depth
How much of the product can run unattended
Deployment compliance — stories about deployment compliance in this arenaDeployment compliance
Stories about deployment compliance in this arena
Format coverage — stories about format coverage in this arenaFormat coverage
Stories about format coverage in this arena
Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual
Stories about ocr multilingual in this arena
Openness — open source, data portability, and self-hosting storiesOpenness
Open source, data portability, and self-hosting stories
Parse accuracy — stories about parse accuracy in this arenaParse accuracy
Stories about parse accuracy in this arena
Privacy posture — data-handling and privacy storiesPrivacy posture
Data-handling and privacy stories
Rag chunking — stories about rag chunking in this arenaRag chunking
Stories about rag chunking in this arena
Scale async — stories about scale async in this arenaScale async
Stories about scale async in this arena
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
✓8/10
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode
—–
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
~4/10
Sdk dx — stories about sdk dx in this arenaSdk dx
Stories about sdk dx in this arena
Structured extraction — stories about structured extraction in this arenaStructured extraction
Stories about structured extraction in this arena
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
~6/10
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields
—0/10
I supply a JSON schema and get back validated structured fields extracted from the document
✓7/10
Multi-document packets are classified and split automatically — one upload, per-document results
✓8/10
Table extraction — stories about table extraction in this arenaTable extraction
Stories about table extraction in this arena
Sorted by importance (agentic first) (high → low) · 53/53 stories · click a row’s chevron for the rationale and evidence
Connect an agent via an official MCP server G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Drive the product through a documented public API G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | full | 8/10 | Tprobed | |
Delegate tasks to a built-in AI assistant inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | none | untested | none yet | |
Plug MCP servers into this product so it can use their tools G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 3 | n/a | untested | none yet | |
Build against official SDKs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Cclaimed | |
Point an agent at llms.txt or agent-oriented docs G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 8/10 | Tprobed | |
Run the product headlessly / in CI for automation G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Subscribe to events via webhooks G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | full | 7/10 | Cclaimed | |
Operate the product with natural-language commands G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 6/10 | Cclaimed | |
Use an official CLI G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 5/10 | Cclaimed | |
Set up automations that run autonomously in the background G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 4/10 | Tprobed | |
Rely on versioned APIs with a documented deprecation policy G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | partial | 3/10 | Cclaimed | |
Download a machine-readable API spec (OpenAPI or equivalent) G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Explore an interactive API reference with runnable examples G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Get AI-generated insights and suggestions from my data inside the product G Agentic features | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Issue scoped/least-privilege API credentials for an agent G Agent access | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 2 | none | 0/10 | ||
Test against a sandbox environment without touching production data G Api quality | ai-native user | Agenticness — how well agents can access and operate the productAgenticness | 1 | none | untested | none yet | |
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks C Async | developer | Scale async — stories about scale async in this arenaScale async | 3 | full | 8/10 | Cclaimed | |
I supply a JSON schema and get back validated structured fields extracted from the document C Schemas | developer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 3 | full | 7/10 | Cclaimed | |
Self-host the core product G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | fullenterprise | 7/10 | Cclaimed | |
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records C Compliance | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 3 | partialenterprise | 7/10 | Cclaimed | |
Export all of my data in open formats and leave G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 3 | partial | 6/10 | Cclaimed | |
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults C Sdks | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 3 | partial | 6/10 | Cclaimed | |
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 3 | disputed | 5/10 | Dcontradicted | |
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text C Chunking | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 3 | partial | 5/10 | Xcommunity | |
Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 3 | partial | 5/10 | Xcommunity | |
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content C Layout | developer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 3 | disputed | 5/10 | Dcontradicted | |
Define rules that trigger actions automatically on events G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 3 | partial | 4/10 | Cclaimed | |
Prevent my data from being used to train AI models G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 3 | none | 0/10 | ||
Multi-document packets are classified and split automatically — one upload, per-document results C Splitting | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | full | 8/10 | Cclaimed | |
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing C Formats | developer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | full | 8/10 | Xcommunity | |
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure C Deployment | data engineer | Deployment compliance — stories about deployment compliance in this arenaDeployment compliance | 2 | fullenterprise | 8/10 | Cclaimed | |
Do everything through the API that I can do in the UI G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | partial | 6/10 | Tprobed | |
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify C Grounding | ai-native user | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | partial | 6/10 | Xcommunity | |
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped C Output | ai-native user | Rag chunking — stories about rag chunking in this arenaRag chunking | 2 | partial | 6/10 | Xcommunity | |
Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup C Layout | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | partial | 6/10 | Xcommunity | |
Perform bulk operations across many items at once G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | partial | 6/10 | Cclaimed | |
Choose where my data is stored (region/residency) G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partialenterprise | 5/10 | Cclaimed | |
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page C Figures | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 2 | partial | 5/10 | Xcommunity | |
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup C Tables | data engineer | Table extraction — stories about table extraction in this arenaTable extraction | 2 | disputed | 5/10 | Dcontradicted | |
I push high-volume batches — millions of pages — with documented rate limits and predictable throughput G Scale | data engineer | Scale async — stories about scale async in this arenaScale async | 2 | partial | 4/10 | Cclaimed | |
Control data retention and deletion G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | partial | 3/10 | Cclaimed | |
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation C Scale limits | data engineer | Format coverage — stories about format coverage in this arenaFormat coverage | 2 | partial | 3/10 | Cclaimed | |
Extractions carry calibrated confidence scores with a human-in-the-loop review path for low-confidence fields C Review | data engineer | Structured extraction — stories about structured extraction in this arenaStructured extraction | 2 | none | 0/10 | ||
Read the product's source under an open license G | ai-native user | Openness — open source, data portability, and self-hosting storiesOpenness | 2 | none | 0/10 | ||
Schedule recurring jobs or workflows G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 2 | none | 0/10 | ||
A fast synchronous mode returns results in seconds for interactive apps, with latency documented per mode C Latency | developer | Scale async — stories about scale async in this arenaScale async | 2 | none | untested | none yet | |
Non-English documents — including CJK and right-to-left scripts — parse with the same fidelity as English C Languages | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 2 | none | untested | none yet | |
Opt out of telemetry and usage tracking G | ai-native user | Privacy posture — data-handling and privacy storiesPrivacy posture | 2 | none | untested | none yet | |
I drag a document into a web playground and see parse/extract results before writing any code C Playground | developer | Sdk dx — stories about sdk dx in this arenaSdk dx | 1 | partial | 6/10 | Xcommunity | |
Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain C Ocr | developer | Ocr multilingual — stories about ocr multilingual in this arenaOcr multilingual | 1 | partial | 4/10 | Cclaimed | |
The vendor publishes reproducible accuracy benchmarks and I can run my own evals before committing C Evals | ml engineer | Parse accuracy — stories about parse accuracy in this arenaParse accuracy | 1 | none | 0/10 | ||
Version, review, and roll back my automations G | ai-native user | Automation depth — how much of the product can run unattendedAutomation depth | 1 | none | untested | none yet |
Opportunities — the stories that would move this product's scores, from its own judged verdictsOpportunitiestop 8 of 37 stories with headroom
What would move LlamaParse’s scores — derived from its own judged verdicts, biggest headroom first. Each line quotes what the judge found missing; shipping it (or evidencing it publicly) is the fix.
Agenticness — how well agents can access and operate the productDelegate tasks to a built-in AI assistant inside the product
nonemoves Built-in AIimpact 45
LlamaParse's evidence describes it as a document parsing/extraction API (Parse, Extract, Classify, Split, Index) callable via SDKs, CLI, REST, or exposed to external agents via an MCP server — but there is no mention of a built-in AI assistant inside the product itself that a user could converse with or delegate tasks to.
Privacy posture — data-handling and privacy storiesPrevent my data from being used to train AI models
nonemoves PA Scoreimpact 30
The evidence covers enterprise features like SOC2/HIPAA compliance, SSO/RBAC, and self-hosting/BYOC options, but nowhere states an explicit policy or toggle for preventing customer data from being used to train AI models.
Agenticness — how well agents can access and operate the productGet AI-generated insights and suggestions from my data inside the product
nonemoves Built-in AIimpact 30
Missing: any documented insights/suggestions UI or feature, evidence of autonomous analysis surfaced to users, independent confirmation of such a capability.
Agenticness — how well agents can access and operate the productIssue scoped/least-privilege API credentials for an agent
nonemoves agent-readyimpact 30
Docs mention SSO and role-based access controls for managing org/project access (llamaparse-docs-9, llamaparse-docs-18), but there is no evidence of scoped or least-privilege API key/credential issuance specifically for agents (e.g., per-key permission scopes, agent-specific tokens).
Agenticness — how well agents can access and operate the productExplore an interactive API reference with runnable examples
nonemoves API qualityimpact 30
Docs show many static code snippets/examples (Python calls, curl-like usage) but there is no evidence of an interactive, runnable API reference (e.g., Swagger/OpenAPI explorer or live code sandbox); explicit probes for OpenAPI/Swagger endpoints returned 404s, indicating no such interactive reference exists.
Agenticness — how well agents can access and operate the productDownload a machine-readable API spec (OpenAPI or equivalent)
nonemoves API qualityimpact 30
LlamaParse exposes a REST API, but there is no evidence of a downloadable OpenAPI/Swagger spec; explicit probes for common OpenAPI endpoints (openapi.json, swagger.json, etc.) all returned 404.
Agenticness — how well agents can access and operate the productRely on versioned APIs with a documented deprecation policy
partialq3/10moves API qualityimpact 21
Missing: a published API versioning scheme, a formal deprecation policy/timeline, changelog or release notes, and independent confirmation of stability guarantees.
Automation depth — how much of the product can run unattendedSchedule recurring jobs or workflows
nonemoves PA Scoreimpact 20
LlamaParse's evidence covers parsing, extraction, classification, splitting, webhooks for job status, self-hosting, and MCP tool exposure, but nothing describes native scheduling of recurring jobs or workflows (e.g., cron-like triggers or recurring pipeline runs).
Showing the top 8 of 37 — every none/partial verdict in the story verdicts table is headroom.
Think a verdict is wrong? Every verdicts-table row has a Flag link — see the methodology.
Coverage map — which docs area, API section, or community source covers which judged storiesCoverage map7 surfaces · 37 covered stories
Where the cited evidence behind each covered verdict came from — the same citations the verdicts table shows, no extra judging.
Llamaparse docs36 stories
- Point an agent at llms.txt or agent-oriented docs
- Run the product headlessly / in CI for automation
- Use an official CLI
- Drive the product through a documented public API
- Build against official SDKs
- Subscribe to events via webhooks
- Set up automations that run autonomously in the background
- Operate the product with natural-language commands
- Rely on versioned APIs with a documented deprecation policy
- Perform bulk operations across many items at once
- Define rules that trigger actions automatically on events
- Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical records
- Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructure
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncation
- Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- Do everything through the API that I can do in the UI
- Export all of my data in open formats and leave
- Self-host the core product
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- Choose where my data is stored (region/residency)
- Control data retention and deletion
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocks
- I push high-volume batches — millions of pages — with documented rate limits and predictable throughput
- I drag a document into a web playground and see parse/extract results before writing any code
- Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaults
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- I supply a JSON schema and get back validated structured fields extracted from the document
- Multi-document packets are classified and split automatically — one upload, per-document results
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Hacker News11 stories
- One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbing
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of text
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- I drag a document into a web playground and see parse/extract results before writing any code
- Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verify
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
- I turn extracted tables into typed rows/JSON I can load into a database without manual cleanup
Llamaparse docs8 stories
- Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertain
- Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans included
- Export all of my data in open formats and leave
- Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source page
- The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered content
- Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not soup
- I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture stripped
- Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structure
For agents docs4 stories
GitHub README3 stories
OpenAPI spec2 stories
Probe proofs — replayable recordings from the probe harnessProbe proofs
Replayable recordings from our probe harness — see the Prove-It protocol to submit one.
$curl -s -X POST https://api.cloud.llamaindex.ai/api/v1/parsing/uploadreproduced$ curl -s -X POST https://api.cloud.llamaindex.ai/api/v1/parsing/upload
{"detail":"Not authenticated"}
$curl -s -X POST https://developers.llamaindex.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'reproduced$ curl -s -X POST https://developers.llamaindex.ai/mcp -H 'Content-Type: application/json' -d '<jsonrpc initialize>'
event: message
data: {"result":{"protocolVersion":"2025-06-18","capabilities":{"tools":{"listChanged":true}},"serverInfo":{"name":"mcp-typescript server on vercel","version":"0.1.0"},"instructions":"LlamaIndex documentation server. The documentation site is hosted at https://developers.llamaindex.ai. All page URLs returned by these tools are relative to this root. For example, a page at /llamaparse/parse/getting_started can be viewed at https://developers.llamaindex.ai/llamaparse/parse/getting_started."},"jsonrpc":"2.0","id":1}
$curl -s https://developers.llamaindex.ai/llms.txt | head -8reproduced$ curl -s https://developers.llamaindex.ai/llms.txt | head -8 # LlamaIndex Documentation > LlamaIndex is a framework for building LLM-powered applications over your data. It supports Python and TypeScript, with integrations for LlamaCloud managed services. ## Accessing Documentation Programmatically All documentation pages are available as raw Markdown by appending `index.md` to the page URL. For example, the page at `https://developers.llamaindex.ai/llamaparse/parse/getting_started/` has its Markdown source at `https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md`.
$curl -sL https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md | head -8reproduced$ curl -sL https://developers.llamaindex.ai/llamaparse/parse/getting_started/index.md | head -8 --- title: Getting Started | Developer Documentation description: Quick start guide for Parse, covering API [redacted] generation and document parsing using Python, TypeScript, Go, Java, the CLI, the REST API, or the Web UI. --- Using a coding agent? Give your AI agent access to these docs: `claude mcp add llama-index-docs --transport http https://developers.llamaindex.ai/mcp` — or supercharge your agent with LlamaParse [MCP tools and Skills](/for-agents/index.md).
$curl -s https://api.cloud.llamaindex.ai/api/openapi.json | head -c 300reproduced$ curl -s https://api.cloud.llamaindex.ai/api/openapi.json | head -c 300
{"openapi":"3.1.0","info":{"title":"Llama Platform","version":"0.1.0"},"paths":{"/api/v1/data-sinks":{"get":{"tags":["Data Sinks"],"summary":"List Data Sinks","description":"List data sinks for a given project.","operationId":"list_data_sinks_api_v1_data_sinks_get","security":[{"HTTPBearer":[]}],"pa
Claims vs evidence — vendor claims reconciled against independent verdictsClaims vs evidence
9 of 25 testable claims verified · 3 contradicted → integrity 12/100
20 distinct capability claims found in LlamaParse’s own claimed-docs/GitHub materials, reconciled against our judge’s independent verdicts.
9
Verified
13
Unverified
3
Contradicted
12
Undersold
Verified (11)
“Quick start across Python, TypeScript, Go, Java, CLI, REST API, or Web UI”
Drive the product through a documented public APIfullproof ↗
“Agentic, layout-aware OCR turns PDFs, scans, tables and charts into clean markdown, text or JSON”
I get clean markdown/JSON designed for LLM consumption, with noise like repeated headers and page furniture strippedpartialproof ↗
“Official MCP endpoint exposes Parse, Classify, Extract, Split and Index as callable tools for any MCP client”
“Supports over 130 file formats across four categories”
One API handles my whole document mix — PDF, DOCX, PPTX, XLSX, HTML, images, email — without per-format plumbingfullproof ↗
“Extract tables directly into pandas DataFrames with source-page provenance for each value”
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verifypartialproof ↗
“Drag-and-drop a document in the web UI, define a schema, and extract data with no code”
I drag a document into a web playground and see parse/extract results before writing any codepartialproof ↗
“Build a hosted vector search/index pipeline for retrieval-augmented generation”
Output comes pre-chunked for RAG — semantic boundaries, metadata, embedding-ready segments — not a wall of textpartialproof ↗
“Single API key and SDK give access to all five composable products (Parse, Extract, Classify, Split, Index)”
Drive the product through a documented public APIfullproof ↗
“Agent-oriented documentation tools: lexical search, regex grep, and full-page read for AI agents”
Point an agent at llms.txt or agent-oriented docsfullproof ↗
“Enriched forms pass returns each form page as structured JSON with field values, checkbox states and bounding boxes”
Every extracted field carries provenance — page number, bounding box, source snippet — so agents can cite and humans can verifypartialproof ↗
“Parses complex layouts, tables, charts, handwriting, checkboxes and images into clean markdown”
Figures and charts are extracted or described (VLM summaries, image crops) with positions traceable back to the source pagepartialproof ↗
Unverified (16)
“Quick start across Python, TypeScript, Go, Java, CLI, REST API, or Web UI”
“Quick start across Python, TypeScript, Go, Java, CLI, REST API, or Web UI”
“LlamaExtract API pulls structured data out of unstructured PDFs, text files and images”
I supply a JSON schema and get back validated structured fields extracted from the documentfullproof ↗
“Classify automatically categorizes documents into user-defined types using natural-language rules”
Multi-document packets are classified and split automatically — one upload, per-document resultsfullproof ↗
“Split API automatically segments concatenated PDFs into logical document sections”
Multi-document packets are classified and split automatically — one upload, per-document resultsfullproof ↗
“Configurable webhooks notify you when jobs complete, fail, or change state instead of polling”
“Configurable webhooks notify you when jobs complete, fail, or change state instead of polling”
Long parses run as async jobs with status polling and completion webhooks, so my pipeline never blocksfullproof ↗
“Entire platform can be self-hosted (BYOC) within your own infrastructure, keeping data and models under your control”
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructurefullproof ↗
“Entire platform can be self-hosted (BYOC) within your own infrastructure, keeping data and models under your control”
“Drag-and-drop a document in the web UI, define a schema, and extract data with no code”
I supply a JSON schema and get back validated structured fields extracted from the documentfullproof ↗
“Single API key and SDK give access to all five composable products (Parse, Extract, Classify, Split, Index)”
Official typed SDKs for Python and TypeScript cover the full API — parse, extract, jobs — with sensible defaultspartialproof ↗
“Cost Optimizer setting reduces processing cost on long, mixed-complexity documents”
Thousand-page documents and multi-gigabyte files process reliably without timeouts or silent truncationpartialproof ↗
“Parses complex layouts, tables, charts, handwriting, checkboxes and images into clean markdown”
Handwritten fields and annotations are recognized and extracted, flagged with confidence when uncertainpartialproof ↗
“Platform has completed SOC 2 Type II audit and offers a HIPAA BAA for enterprise customers”
Uploaded documents get zero-retention handling with SOC 2 and HIPAA options, so I can process contracts and medical recordspartialproof ↗
“Multiple deployment/residency options — managed SaaS, single-tenant, BYOC, self-hosting, and regional endpoints”
Choose where my data is stored (region/residency)partialproof ↗
“Multiple deployment/residency options — managed SaaS, single-tenant, BYOC, self-hosting, and regional endpoints”
Run the extraction stack in my own VPC or fully self-hosted when documents can't leave my infrastructurefullproof ↗
Contradicted (3)
“Extract tables directly into pandas DataFrames with source-page provenance for each value”
I turn extracted tables into typed rows/JSON I can load into a database without manual cleanupdisputedproof ↗
“Parses complex layouts, tables, charts, handwriting, checkboxes and images into clean markdown”
The API parses complex real-world PDFs — multi-column layouts, headers, footers, footnotes — into clean, correctly ordered contentdisputedproof ↗
“Parses complex layouts, tables, charts, handwriting, checkboxes and images into clean markdown”
Complex tables — merged cells, nested headers, multi-page spans — come out as faithful HTML/markdown structuredisputedproof ↗
Undersold (12)
Run the product headlessly / in CI for automationfullproof ↗
Set up automations that run autonomously in the backgroundpartialproof ↗
Operate the product with natural-language commandspartialproof ↗
Rely on versioned APIs with a documented deprecation policypartialproof ↗
Perform bulk operations across many items at oncepartialproof ↗
Define rules that trigger actions automatically on eventspartialproof ↗
Scanned and photographed documents OCR accurately — skewed pages, stamps, low quality scans includedpartialproof ↗
Do everything through the API that I can do in the UIpartialproof ↗
Export all of my data in open formats and leavepartialproof ↗
Parsed output preserves document hierarchy — headings, sections, reading order — so downstream LLMs see structure, not souppartialproof ↗
I push high-volume batches — millions of pages — with documented rate limits and predictable throughputpartialproof ↗
Claims outside our story set (1)
Real capability claims found in LlamaParse’s own materials, but no story in this arena’s taxonomy covers them yet — that’s feedback on the taxonomy, not a mark against the product.
“SSO and role-based access controls manage access to organizations and projects”
source ↗
Business model
Free tier with 10K credits/month; Starter $50/mo (40K credits), Pro $500/mo (400K); 1,000 credits = $1.25 and basic parsing from 1 credit/page; Enterprise is custom.
pricing ↗Score trend
How this product’s scores have moved as evidence and verdicts are re-derived — a point per change, not per day.
Try Experimental
Run it in the microterminal →Recorded agent sessions — and a live MCP handshake where the vendor ships one.
Flag
⚑ Flag a verdictThink a verdict is wrong? Opens a prefilled GitHub issue — or use the ⚑ next to any verdict above.
For agents
Agent surface uptime MCP up · llms.txt up (tracking since Sep 11 '26)
