Explore Intelligent Document Processing Tools

{{locationDetails}}

{{locationDetails}}

Back to filters

Browse sub-categories

Top 22 Intelligent Document Processing Tools

Facts and pricing verified .

Twenty-two tools carry the same three letters here, and two completely different products are hiding underneath them. One is a sales-led enterprise suite. It sells you an outcome: 99.5% accuracy, straight-through processing, a human review queue, a posting into your enterprise resource planning (ERP) system.

The other is a self-serve document application programming interface (API) that sells you a primitive at a per-page rate and leaves the workflow to your engineers.

The ranking below interleaves them on merit. So what is your first decision? Not which tool. It is which kind of product you're buying.

The further up-market you go, the less you can check before you sign. Eight of these 22 publish no price at all, and eight give you nothing in the review column below.

And the headline accuracy numbers, 99.5%, 99.7%, 99% at the field level, 0.1% error, are almost all vendor-stated. Only two vendors name a benchmark: LandingAI cites DocVQA and Extend cites RealDoc-Bench.

Keep that in mind when a slide promises a number to two decimal places. The suites are not worse products. They just make you test what a slide asserts.

Here is the short version.

  • Microsoft Azure AI Document Intelligence, the strongest all-round start: 500 free pages a month, prebuilt invoice, tax, mortgage and ID models, and deployment to cloud, edge or on-premises.
  • Amazon Textract, the cheapest published rate in the field, from $0.0006 a page at volume.
  • Rossum, from $18,000 a year, purpose-built for high-volume accounts payable (AP).

Got engineers and want to start today?

  • Unstructured, 10,000 free pages to start, no card.
  • Extend, 10,000 free credits a month, every month.

The two tracks, and which one you are actually shopping for

How much you can verify before you sign depends almost entirely on which track you are standing in, so name your track first.

Track one is the enterprise intelligent document processing (IDP) suite. You buy an outcome. In return you get classification, extraction, validation, business rules, a human-in-the-loop (HITL) review interface and connectors into your system of record.

You also get a sales cycle, a custom quote and an implementation project on your calendar. Rossum, Hyperscience, UiPath, Nanonets, Docsumo, Base64.ai, ABBYY Vantage, Ocrolus, Klippa DocHorizon, Instabase, Infrrd and Indico Data live here.

Track two is the document API built for engineering teams. You buy parse, extract, classify and split as callable endpoints, priced per page or per credit, and you build the workflow yourself.

Azure AI Document Intelligence, Amazon Textract, Unstructured, LlamaIndex, Google Document AI, Reducto, LandingAI, Sensible, Mindee and Extend sit on this side.

The trade is simple. Track one costs more and asks less of your engineers. Track two costs less and asks more.

What you're really deciding is where the integration work lives, because somebody is going to write the validation rules and the exception handling, and the only question is whether that somebody sits in your team or in a statement of work.

Rule of thumb: if you can't name the engineer who will own the pipeline in six months, you're shopping in track one, whatever your budget says.

That is no knock on your engineering team. It's a statement about who maintains a schema two quarters after the person who wrote it has moved on to something else.

The evidence ledger

This is the table to screen with, and the one to keep open during your first three calls. Not the feature grid, the evidence grid. It asks every vendor one question. What can you confirm before a contract exists?

A dash means that column has nothing for you to work with, which is itself worth a question on the call.

Read down the pricing and free-path columns first. Those two decide which vendors you can evaluate on your own schedule and which ones need a calendar invite before you learn anything at all.

ToolTrackPricing publishedFree pathThird-party reviewsAccuracy claimHeadline compliance
1. Microsoft Azure AI Document IntelligenceAPIFree tier plus usage ratesPermanent, 500 pages/mo19-GDPR, EU Data Boundary, on-prem containers
2. Amazon TextractAPIFull rate card3-month trial27-HIPAA eligible, PCI, ISO, SOC
3. RossumSuiteEntry tier only14-day trial140Vendor-statedISO 27001, ISO/IEC 42001, SOC 2 Type II, TX-RAMP
4. HyperscienceSuiteQuote only-5499.5%, vendor-statedFedRAMP High, SOC 2 Type II, Cyber Essentials Plus
5. UiPath (Intelligent Xtraction & Processing / Document Understanding)SuiteEntry tier only-16-SOC 2 Type 2, ISO 27001, HITRUST, HIPAA, C5
6. NanonetsSuiteEntry tier onlyOne-time credits101Up to 95% STP, vendor-statedHIPAA, SOC 2 on Enterprise
7. DocsumoSuiteEntry tier only14-day trial999% field-level, vendor-statedSOC 2, GDPR, HIPAA
8. UnstructuredAPIFull rate cardOne-time, 10,000 pages--FedRAMP High, HIPAA, GDPR, SOC 2 Type II
9. LlamaIndex (LlamaCloud / LlamaParse)APIFull rate cardPermanent, 10K credits-Vendor-statedSOC 2 Type II, HIPAA, GDPR, private VPC
10. Google Document AIAPIUsage-based$300 cloud trial credit37-Regional processing: U.S., EU, Asia-Pacific
11. Base64.aiSuiteQuote only-199.7%, vendor-statedISO 27001, SOC 2 Type II, HIPAA, GDPR, ISO 20243
12. ReductoAPIFull rate cardOne-time credits-Vendor-statedZero data retention, BAA, EU/AU residency, air-gapped
13. LandingAI (Agentic Document Extraction)APIPublished through TeamOne-time credits-DocVQA 99.16%SOC 2 Type II, GDPR, HIPAA with BAA, zero retention
14. ABBYY VantageSuiteQuote only-3490% at the start, vendor-statedSOC 2 cloud instances, on-prem or private cloud
15. OcrolusSuiteQuote only--99+%, vendor-statedSOC 2, SOC 3, ISO/IEC 27001, PCI DSS, CSA STAR Level 1
16. Klippa DocHorizonSuiteQuote only-31, vendor-citedUp to 99%, vendor-statedISO 27001, ISAE 3000 Type I, GDPR, EU-hosted
17. SensibleAPIFull rate card14-day trialVendor-cited-SOC 2 Type II, HIPAA, GDPR
18. InstabaseSuiteQuote only-2-SSO, IP whitelisting, 99.5% uptime SLA
19. MindeeAPIFull rate card14-day trial30+, vendor-cited-Data processing localization on Pro
20. InfrrdSuiteQuote only--0.1% error, vendor-statedSOC 2 Type II, GDPR with standard contractual clauses
21. ExtendAPIFull rate cardPermanent, 10,000 credits/mo-RealDoc-Bench 95.7%Zero data retention standard, HIPAA/BAA add-on
22. Indico DataSuiteQuote only---SOC 2, encryption, role-based access, audit logs

What is actually published, in dollars

Read that pricing column again. The published rates sit almost entirely on the API side, and there's a reason for that. The products that sell without a call publish a price. The suites do not publish rates at all.

For instance, here is every dollar figure the 22 vendors put in public, in the form you can copy straight into a budget deck.

ToolPublished rates
Amazon TextractDetect Document Text $0.0015/page for the first 1M pages/month, then $0.0006/page. Analyze Document (Forms + Tables + Queries) $0.070/page, then $0.055. Analyze Expense $0.01/page, then $0.008. Analyze ID $0.025/page for the first 100K, then $0.01. Analyze Lending $0.07/page, then $0.055.
Unstructured10,000 free pages to start, no card required, then $0.015 per page. Business tier custom.
ReductoClassify $7.50, Parse $10, Extract $20, Split $20, Deep Extract $40, Deep Split $40, Edit $60 or $15 pre-filled, all per 1,000 pages. $150 in free credits to start.
LlamaIndexFree $0/month with 10K credits. Starter $50/month with 40K credits. Pro $500/month with 400k credits. 1,000 credits = $1.25.
MindeeStarter $44/month billed annually at $529/year. Pro $116/month billed annually at $1,393/year. Both 6,000 to 300,000 credits/month.
LandingAIExplore pay-as-you-go at $1 per 100 credits, with 1,000 free starter credits. Team $250/month for 25,000 credits/month, unlimited seats.
Nanonets$50 in credits to start, no card required, then $100/month for 100 credits, up to 3 users. Per-run blocks $0.02 to $0.30.
SensibleGrowth $449/month billed annually or $499/month monthly, 750 documents/month, $0.57 per document over. Scale $1,349/month annually or $1,499/month monthly, 3,200 documents/month, $0.53 over.
ExtendPay As You Go free with 10,000 credits/month, then $0.0125 per credit. Scale $500/month with 50,000 credits, then $0.01 per credit.
UiPathBasic $25/month. Document processing is not fully included at this tier.
RossumStarter starting at $18,000 per year, unlimited seats, 12-month archive. One-year minimum contract.

Azure sits just outside that table: its free F0 tier of 0 to 500 pages a month is published, while the paid per-page rates render by region and currency instead of as a fixed list.

Look at what the table does to the field. The highest hard enterprise floor anyone publishes is Rossum's $18,000 a year. Everything above an entry tier at Hyperscience, ABBYY Vantage, Instabase, Indico, Ocrolus, Infrrd, Klippa and Base64.ai is a conversation, not a number.

Eight questions to answer before you book a demo

Answer these on paper first, before you let a vendor frame the conversation for you. Each one cuts your list, and each one maps to something the vendors actually differ on.

  1. What is your monthly page or document volume band? At 500 pages a month Azure is free. At 3,200 documents a month Sensible is $1,349 on the annual plan. Past the first million pages a month Textract's text detection drops from $0.0015 to $0.0006 a page. The band changes the winner.
  2. What is in the mix, and how much of it is handwritten? Textract handles handwriting in English only. Azure's Read model covers handwritten text in 12 languages. Google Document AI covers 200+ languages including handwriting.
  3. How much of it is not in English? Rossum advertises 276-language extraction, though reviewers report inconsistent accuracy on non-English or unusual formats. Mindee states support for any language or alphabet. Infrrd cites 22+ languages.
  4. Where must the data live, and can it leave your network? Air-gapped, on-premises and virtual private cloud (VPC) options exist at Reducto, ABBYY Vantage, Nanonets, LandingAI and Azure containers. If your answer is "it cannot leave", that single line removes most of this list.
  5. What is the downstream system of record? Rossum ships connectors for SAP, Coupa, Oracle NetSuite, Workday, Microsoft Dynamics, Xero and QuickBooks. Indico integrates with Guidewire. Docsumo posts to NetSuite, Encompass and Epic. Match the connector to the system you already run.
  6. Who owns the pipeline? Be honest here. A document API is cheap right up until nobody is left to maintain the schema.
  7. What is your compliance floor? FedRAMP High for federal work narrows the field to Hyperscience and Unstructured. HIPAA for health data opens up Textract, Nanonets, LlamaIndex, Docsumo, Sensible, LandingAI, Base64.ai, Reducto, Extend and UiPath.
  8. What evaluation can you run without signing anything? Three options give you a standing free plan. Several hand you pages or credits once. The rest give you a calendar invite.

Status and watch items

These are procurement facts, not features, and no comparison column holds them for you.

  • Rossum was acquired by Coupa in 2026, positioned as a move toward autonomous spend management. If you're buying Rossum standalone, ask what the roadmap looks like inside Coupa.
  • Klippa DocHorizon is mid-rebrand to Doxis AI.dp, and both names appear across its own current marketing pages.
  • Indico Data's homepage carries a visible notice pointing to a dedicated page about a data security incident. That belongs in your security review, alongside its stated SOC 2, encryption, role-based access and audit-log controls.
  • Instabase Community sign-ups are reportedly discontinued, which narrows the self-serve on-ramp to a demo or a trial request.
  • Azure has published end-of-support dates for its older interfaces: REST API v2.1 on 15 September 2027 and v3.0 on 30 March 2029. Current is Document Intelligence 2024-11-30 v4.0. If you inherit an old integration, you have a migration on the calendar.

The ranking

1. Microsoft Azure AI Document Intelligence

Best all-round start. The F0 tier gives you 0 to 500 pages a month, permanently, across all supported document types. That's enough to run a real proof of concept on your own documents. Only two other tools here give you a standing monthly free plan at all.

Underneath sits a deep prebuilt catalog:

  • Financial and legal: invoice, receipt, bank statement, check, contract, pay stub, credit card.
  • U.S. tax: W-2, 1098, 1099 and 1040.
  • U.S. mortgage: forms 1003, 1004, 1005 and 1008.
  • Personal identity: health insurance card, identity, marriage certificate.
  • Custom neural, template and composed models, trained from as few as five labeled samples.

Deployment is the other reason it leads. Cloud, edge and on-premises Docker containers, which matters if your documents aren't allowed to leave the building. Optical character recognition (OCR) covers handwritten text in 12 languages, including English, Chinese Simplified, Japanese, Korean, Arabic and Russian.

Reviewers on G2 rate it 4.4 out of 5 across 19 reviews and single out extraction accuracy.

They also flag a difficult learning curve around OCR and custom-model setup, plus costs that escalate at high volumes, and you cannot price those in advance, because the pay-as-you-go rates render by region and currency rather than as a fixed published list.

Microsoft states the wider Azure AI services are used by more than 80,000 enterprises and digital natives, including 80% of Fortune 500 companies. If you are already on Azure, prove the whole thing on the free tier before you write a business case.

2. Amazon Textract

Every rate is on the page, which is rarer here than it should be. Two of them tell you most of what you need. Detect Document Text runs $0.0015 per page for the first million pages a month, and $0.0006 beyond that. Analyze Document, for forms, tables and queries, runs $0.070.

Same page, same file. What you ask of it is what moves your bill.

Custom Adapters let you fine-tune the models on your own documents. Asynchronous operations take PDFs and TIFFs up to 500 MB and 3,000 pages, against a one-page limit for synchronous calls.

The limits are worth naming. Handwriting recognition is English only, and reviewers report lower accuracy on varied handwriting. Text detection covers English, French, German, Italian, Portuguese and Spanish, with no vertical-text support. The free tier expires after three months.

What the rate card can't give you is your true cost on a complex multi-page PDF mix, which reviewers specifically say climbs faster than expected.

The compliance story is solid. HIPAA eligible, PCI, ISO and SOC compliant, with VPC PrivateLink and content encrypted at rest in region. You can opt out of machine-learning training use, and deletion requests are honored.

G2 gives it 4.3 out of 5 across 27 reviews with subscores of 9.3 for data extraction and 8.9 for OCR. Clean, printed, English documents at volume? The per-page economics are hard to argue with.

3. Rossum

Best for high-volume accounts payable. Rossum does something few vendors at this end of the market do. It publishes an annual floor: Starter from $18,000 per year, unlimited seats, ingestion by email, API or upload, and a 12-month archive. Business, Enterprise and Ultimate are all custom-quoted, and the minimum term is a year.

What you buy is the whole transactional-document lifecycle, not an extraction endpoint. Multi-channel intake, classification and splitting, 276-language extraction, master-data validation, duplicate detection, approval workflows, e-invoicing and audit dashboards. Ultimate adds multi-document transaction support and an embeddable interface.

The customer evidence is unusually concrete. Wolt processes 100K invoices a year with 44% fewer errors. Port of Rotterdam Authority reached 90% extraction accuracy after only 10 documents and reports saving 810 AP days a year.

Both are customer stories on Rossum's own site rather than a benchmark, so treat them as direction, not proof.

Connectors cover SAP, Coupa, Oracle NetSuite, Workday, Microsoft Dynamics, Xero and QuickBooks.

Certifications run to ISO 27001:2022, SOC 2 Type II and TX-RAMP Level 1. In fact Rossum also holds ISO/IEC 42001:2023 for AI governance, which is still rare enough to be a genuine differentiator in a security review.

Capterra reviewers score ease of use 4.6 out of 5 across 140 reviews and praise support. Three complaints keep coming back: high cost for smaller businesses, inconsistent accuracy on non-English or unusual formats, and tedious initial model training.

Is $18,000 a year a lot? For a small finance team, yes. For an operation running six figures of invoices a year, it is at least a number you can put in a budget before the first call, and that floor is what your budget conversation will turn on.

4. Hyperscience

Best for enterprise and best integrations. Hyperscience states that customers regularly achieve 99.5% accuracy and 98% automation on its Hypercell platform, built on a vision-language model framework it calls ORCA. The market position around that claim is easier to check. Six analyst firms name it a Leader: Gartner, Forrester, IDC, GigaOm, ISG and Everest Group.

The integration surface is why it takes the integrations pick:

  • AWS: S3, Lambda, SageMaker, Redshift.
  • Google Cloud: BigQuery and Vertex AI.
  • Azure: Blob Storage and Synapse.
  • Business systems: SAP, Salesforce, Microsoft 365, IBM FileNet, UiPath.
  • Model providers: Gemini, Claude through Bedrock, OpenAI, NVIDIA Nemotron.

Compliance clears SOC 2 Type II, Cyber Essentials Plus, GDPR and CCPA. It also holds FedRAMP High through a Palantir partnership, and that's the credential that opens federal work.

The customer list runs to Amex, Schwab, MetLife, the VA and the SSA. Research from IDC cited by the vendor puts client returns at 615% over three years with payback in roughly seven months.

G2 rates it 4.6 out of 5 across 54 reviews with 97% at four or five stars, praising extraction on handwritten and messy documents. The recurring complaint is deployment complexity that demands real in-house IT capability, and cost that reviewers call prohibitive for smaller organizations.

This one is for large enterprises and government agencies with the IT bench to stand a complex deployment up and keep it running.

What you can check: FedRAMP High authorization, the named integration blocks, the analyst placements, and a 54-review base.

What you can't: test the headline numbers. The 99.5% accuracy and 98% automation figures are Hyperscience's own, and the 615% return reaches you through the vendor too.

5. UiPath (Intelligent Xtraction & Processing / Document Understanding)

Already running UiPath robots? Then this is the shortest path you have, and the entry below is really about how short.

What you're looking at is not a document tool exactly. It's a document capability inside a robotic process automation (RPA) and agentic platform, and that's the whole basis on which to judge it. Intelligent Xtraction & Processing combines Document Understanding for structured and semi-structured extraction, Generative Extraction using large language models for messier documents, and Communications Mining.

Human-in-the-loop validation and active learning train the models without ML expertise on your side, on Automation Cloud or self-hosted Automation Suite.

The enterprise results are specific, and all three reach you through UiPath's own case studies. Canon USA processes 4,500 invoices a month with 90% straight-through processing (STP), saving 6,000 hours a year. Dexcom reports 200,000 hours saved and an 80% cycle-time reduction. Hiscox reached a 28% automation rate in claims triage with a 300% lead-time reduction.

Compliance is broad enough for most floors: SOC 2 Type 2 and ISO/IEC 27001 company-wide, with ISO 27017, 27018 and 9001, HITRUST, SOC 1, HIPAA and C5 on Automation Cloud Dedicated.

Reviewers score Document Understanding 9.6 for multi-format support and 9.2 for OCR, and note time-consuming setup on complex document types plus weaker accuracy on small training sets.

Read that sentiment lightly: Document Understanding carries 16 reviews of its own, against a platform reputation built on thousands. And if you are not already on the platform, understand that you're buying a platform to get a feature.

6. Nanonets

Nanonets sells on its proprietary OCR-3 model and on straight-through processing of up to 95%. Pricing is where you should slow down. Starter gives you $50 in credits with no card, then converts to $100 a month for 100 credits, capped at 3 users. Individual blocks run $0.02 to $0.30.

Growth is volume-priced with up to a 40% discount, and Enterprise is fully custom. But the credit-and-block structure is the part users find hardest to forecast at scale, however much they like the product itself. Model your credits before you commit.

Coverage is broad. Structured, semi-structured and unstructured documents, meaning invoices, receipts, bank statements, legal agreements and IDs. Handwriting, signature and checkbox detection come with it, across PDF, PNG, GIF, JPEG, TIFF and BMP.

Classification AI, generative AI blocks and custom Python blocks handle the logic. Integrations named on the site include SAP, Oracle, Salesforce and QuickBooks, plus email and cloud storage.

Enterprise adds HIPAA and SOC 2 compliance, role-based access control, private cloud or on-premises deployment, data residency and audit trails.

G2 puts it at 4.7 out of 5 across 101 reviews, 86% of them five-star, with 9.5 for support and 9.6 for data extraction. The vendor states 35% of Fortune 500 companies use it and cites an AP example where invoice processing fell from 6 minutes to 16 seconds.

What you can check: the entry economics, the per-run block range, the Enterprise compliance list, and one of the largest review bases in this field.

What you can't: stand the capability claim up outside the vendor's own materials. The OCR-3 leaderboard result is self-reported.

7. Docsumo

Docsumo is built as a chain, not a single extractor. An Email agent, a Document Collection agent, Classification across 250+ document types, Extraction, Verification, and Case Management that pushes into workflow.

The vendor claims 99% field-level accuracy with confidence scoring and 95%+ straight-through processing. Two customer cases sit behind that. Arbor reports 99% accurate ACORD-form capture across 75k+ insurance claims a year. Cassena Care processes 130k+ Medicaid applications twice as fast at 99.81% accuracy.

The integration list matters more than the claims if you already run a system of record. Yours may be on it:

  • Business systems: Salesforce, QuickBooks, Xero, Yardi.
  • Payments and support: Stripe, Chargebee, Zendesk.
  • Glue: Zapier with its 500+ apps, Google Sheets, OneDrive, Jotforms, Appsheets.
  • Direct: webhooks and API access.

One caution on the pricing page. The tier labeled Free Plan is a 14-day trial: 1,000 pages, 10 user licenses. It doesn't stand.

Compliance covers SOC 2, GDPR and HIPAA, which is the floor for the lending, insurance and healthcare work this is aimed at.

Capterra scores ease of use 4.8 out of 5, though from a base of 9 reviews. Those reviewers raise diverse invoice formats needing manual correction and thin API documentation, with at least one report of missed scheduled calls.

If your workflow starts with an inbox full of ACORD forms and ends in a case file, the agent chain maps to your process more closely than a raw extraction endpoint would.

What you can check: the trial limits, the integration list, the compliance badges, and the named customer results.

What you can't: lean on the market's verdict. Nine reviews is a thin base, and the 99% and 99.81% figures come from the vendor's own case studies.

8. Unstructured

Ten thousand free pages to start, no card required, then $0.015 a page. That grant is one-time rather than monthly, but for an engineering team that needs to prove a pipeline works before requesting budget, it is still a lot of runway before anyone signs anything.

Unstructured positions itself as the extract-transform-load layer between messy documents and whatever consumes them. You get layout-aware partitioning, table structure preservation and OCR across 65+ file types, returned as structured JSON. Drive it code-first through the Python API or through no-code pipelines. Official documentation lists 35+ connectors and 28 named destinations.

No independent review record stands behind this product on G2 or Capterra, so weigh what follows as capability, not experience.

Plus the destinations are the standout:

  • Vector databases: Pinecone, Weaviate, Milvus, Qdrant.
  • Search and documents: Elasticsearch, MongoDB, Azure AI Search.
  • Warehouses and lakehouses: Snowflake, Databricks.

That is most of what a retrieval-augmented generation (RAG) stack needs. Compliance is unusually strong for a company at this price point: FedRAMP High plus HIPAA, GDPR and SOC 2 Type II. Named customers and testimonials include IBM, Bank of America, JPMorgan Chase, Google, Amazon, SAP, Teradata and Intel.

Keep in mind what it isn't. There's no proprietary extraction model being sold here, and the Business tier with dedicated instance, VPC and full data isolation is custom-quoted.

9. LlamaIndex (LlamaCloud / LlamaParse)

LlamaIndex publishes all four tiers, which is worth something in this field. Free is $0 a month with 10K credits. Starter is $50 a month with 40K credits and pay-as-you-go up to 400K. Pro is $500 a month with 400k credits and pay-as-you-go up to $5,000 a month.

Enterprise is custom, with volume discounts and 5x higher rate limits. Credits convert at 1,000 for $1.25, and overages take more forecasting than a flat plan.

Five composable products sit behind one API key: Parse, Extract, Classify, Split and Index, covering 130+ file formats. You get layout-aware parsing, bounding boxes, word-level grounding and multimodal handling for charts, tables and scans.

Software development kits (SDKs) cover Python, TypeScript, Go and Java, alongside a command-line interface, MCP servers for agents, an n8n node and a REST API.

Trust credentials include SOC 2 Type II, HIPAA and GDPR, with encryption in transit and at rest. Cache retention runs 48 hours, and you can opt out. Private VPC deployment runs on AWS and Azure, with marketplace listings on both.

The vendor reports 99.9% uptime, 1B+ documents parsed across 300k+ users, and 25M+ monthly open-source package downloads, and it pitches itself as 5x more accurate than other APIs and 4x cheaper than frontier labs, though those two comparisons are the vendor's own and no third-party review base sits behind them.

Named customers include Jeppesen, a Boeing company, along with KPMG, PepsiCo, Carlyle, Lazard, Guidehouse and Cemex.

10. Google Document AI

If your data already lives in Google Cloud, this is the default and a strong one. The processor library is deep. Enterprise OCR spans 200+ languages including handwriting. A Layout Parser handles PDF, HTML, DOCX and PPTX. A generative-AI Custom Extractor plus Gemini-powered Custom Classifier and Splitter work zero-shot.

Seven pretrained specialized parsers cover invoices, expenses, bank statements, W2s, U.S. driver licenses, pay slips and identity proofing. Processors deploy across eight official regions in the U.S., Europe and Asia-Pacific, which covers most data-residency requirements.

There is no permanent free tier. Your on-ramp is the general $300 Cloud trial credit, and reviewers are blunt about what comes after. Published quota tiers are paid capacity: 120 pages per minute provisioned, 30 pages per minute for Pro versions.

G2 rates it 4.2 out of 5 across 37 reviews and praises OCR power. What none of those reviewers can settle for your security team is the certification question: the dedicated data-governance page for the processors returned a 404 during the research for this piece, so SOC 2, ISO and HIPAA status went unverified here, and regional processing is what the documentation that did load gives you.

The complaints run to costs that hinder adoption for smaller businesses, extraction inconsistencies on complex PDFs and tabular data, and a product that needs real API expertise despite user-friendly marketing. Documentation gets called outdated or unclear.

Documents spanning a dozen languages will probably decide this one for you on coverage alone. Model the per-page cost before you scale.

11. Base64.ai

The model library is the pitch. The features page documents 2,800+ pretrained models for industry-specific documents. Around them sit automatic document-type classification, unlimited custom models, semantic AI with question-answering, PII (personally identifiable information) redaction, signature and facial verification, and RAG-based AI agents.

Base64.ai puts OCR coverage at 165+ languages with handwriting support across 93 file formats. Integration runs through a REST API and hundreds of prebuilt no-code connectors to scanners and RPA platforms. Named ties include AWS, Microsoft Azure, Google Cloud, Snowflake and Databricks.

Investors named on the site include Sequoia Capital, Long Journey Ventures, Gaingels and Zero Prime.

Security is where it scores highest. The platform states ISO 27001, SOC 2 Type II, HIPAA, GDPR and ISO 20243 certification, with cloud or on-premises deployment.

The pricing page names four tiers, 1 cent OCR, Document AI, Advanced Add-ons and Enterprise, and states that annual plans start at 1,000 pages a month.

Getting an outside read on performance is the hard part. The 99.7% accuracy and 5-second processing figures are self-reported, with a single public review behind them.

12. Reducto

Best value. Reducto publishes an exact rate for every operation it sells. Classify is $7.50 and Parse is $10, per 1,000 pages. Extract and Split are $20. Deep Extract and Deep Split are $40. Edit is $60, or $15 pre-filled.

You start with $150 in free credits, or 15,000 credits through Studio.

That lets you price a workload to the dollar before you talk to anyone, then buy only the operations you need instead of a bundle.

Six composable functions cover Classify, Parse, Extract, Split, Edit and Pipelines. You reach them through a no-code Studio, a command-line interface, a REST API, Python and Node SDKs, an MCP server and an OpenAPI 3.1 spec.

Deployment climbs a ladder, and where you stop on it is usually your security team's call:

  • SaaS.
  • Hybrid VPC.
  • Full VPC.
  • On-premises, air-gapped.

Growth and Enterprise add a zero data retention agreement and a business associate agreement (BAA) for HIPAA work, with EU and Australia data residency and single sign-on.

Standard guarantees 200 concurrent pages and up to 5 Studio seats, Growth 350 with unlimited seats, Enterprise 500+. Encryption is AES-256 at rest and TLS 1.2+ in transit, with 24-hour data deletion on Standard.

Customers include Vanta, Toast, JLL, Scale AI and Harvey, and the platform reports 3 billion+ pages processed. Anterior reports 99%+ accuracy on prior authorization and Elysian reviews claims 16x faster. Pedestal AI moved from 70% to 95-96% on handwritten purchase orders, scanned faxes and 300+ page files.

Corroborating those outcomes is not something you can do from outside: they are customer testimonials on Reducto's own site, and no third-party review base sits behind them.

What you can do is work out what a hundred thousand pages will cost on the back of an envelope, before you speak to a salesperson. In this field that is unusual.

13. LandingAI (Agentic Document Extraction)

LandingAI is one of only two vendors here that names a benchmark rather than a number: 99.16% accuracy on DocVQA. Treat that as a claim you can at least locate, which is more than most of this field offers.

The product is vision-first, built for documents that break ordinary parsers: complex tables, dense multi-page layouts. It returns confidence scores and coordinate-level grounding for every extracted field. Six modular APIs cover Parse, Extract, Split, Classify, Section and Build Schema.

No third-party review data surfaced for this product during the research for this piece, and the DocVQA result is LandingAI's own.

Explore is pay-as-you-go at $1 per 100 credits with 1,000 free starter credits and a single seat. The free Explore grant is one-time, not monthly. Team is $250 a month for 25,000 credits with unlimited seats. SDKs are Python and TypeScript only, with no named no-code connectors, so plan for engineering time.

Compliance covers SOC 2 Type II, GDPR and HIPAA with a BAA available, plus a zero-data-retention option. That is why regulated buyers show up on the customer list. Barclays, Morgan Stanley, Deloitte, Intel, Thermo Fisher, Foxconn and Qualcomm sit among the 50+ enterprise customers the platform names, and it reports 1B+ images and documents processed.

For dense tables and long scanned reports, the grounding and confidence scores give your reviewers something concrete to check an extraction against.

14. ABBYY Vantage

Vantage is the low-code veteran play. It ships 150+ pre-trained extraction skills out of the box, plus a no-code skill designer for your own. Classification, splitting and human-in-the-loop review with continuous learning come standard.

ABBYY puts identity-document processing at 10,000+ document types across 190 countries, which is the number to remember if your workflow involves passports and national IDs from everywhere.

Integration leans toward the automation stack you already own. Prebuilt connectors cover Microsoft Power Automate and the major RPA platforms: Blue Prism, UiPath, Automation Anywhere. A REST API handles business process management, ERP and enterprise content management systems.

It runs in SOC2-certified ABBYY cloud instances in Europe, the USA and Australia, or on-premises and private cloud on Microsoft Azure through Docker and Kubernetes.

Carlsberg reports 92% touchless order processing and 140+ hours saved a month in ABBYY's own case study. ABBYY claims 90% accuracy "at the start", improving through continuous learning. That 90% is ABBYY's own figure, and continuous learning is doing the rest of the work in the sentence.

Reviewers across 34 reviews praise structure-preserving table extraction on contracts and invoices. They complain about complex setup for custom workflows, slow feature rollout, and missing documentation on combining scripted rules. If you already run Power Automate or one of the big RPA platforms, Vantage slots straight into that stack.

15. Ocrolus

Are you a lender? Then this specialist is worth your call, and if you're not, it's the wrong shape of product for your documents however good the portal looks.

Ocrolus isn't a general-purpose IDP platform and doesn't pretend to be. It's document AI for lenders, and the modules are named for lending jobs:

  • Detect, for fraud.
  • Capture, for extraction at a claimed 99+% accuracy.
  • Analyze, for cash flow.
  • Encore, for lead monetization.
  • Inspect, for mortgage condition clearing.

Your use case probably sits in here: mortgage, small-business, auto, consumer, legal, Medicaid, tax or tenant screening. Developer documentation covers 2,100+ supported financial document form types with JSON attribute references, webhook event notifications, and direct embedding into a loan origination system (LOS).

Company standing is the first thing a regulated lender asks about. Founded in 2014, with $110M raised from Oak HC/FT, Fintech Collective, QED and Bullpen. It reports 3,700+ active user accounts and 750k credit applications processed monthly. Customers include PayPal, SoFi, Square, Zillow, Better and Brex.

Security is the second. SOC 2, SOC 3, ISO/IEC 27001, PCI DSS and CSA STAR Level 1, with coverage extending to GDPR, CCPA, GLBA, PIPEDA, VCDPA and New York 23 NYCRR 500. The security portal also shows a BitSight score of 760 and A+ Qualys SSL Labs ratings on both dashboard and API.

What you can check: the funding announcement, the certification list, the security scores, and the documented form-type coverage.

What you can't: see how it lands with the people using it. No independent review record on G2, Capterra or TrustRadius stands against that 99+% claim.

16. Klippa DocHorizon

The name is the first thing to know. Klippa DocHorizon is rebranding to Doxis AI.dp, and both appear across its own current marketing pages, so expect the paperwork to change mid-conversation.

Fraud checks are what set the product apart. EXIF analysis, duplicate detection and pixel-level image analysis sit alongside the usual work: OCR extraction, classification, data anonymization and document verification.

Conversion covers JSON, XML, CSV, XLSX, PDF, UBL and TXT. A drag-and-drop workflow builder and a Prompt Builder let you configure extraction rules, with human review on top.

The platform lists 50+ document types, so check yours against the list: invoices and receipts through ID cards, passports, contracts and medical forms. The vendor claims up to 99% data extraction accuracy, 250+ million documents processed and 1,000+ enterprise clients. Named case-study customers include GLS, Trading 212, SNCF, myWorld and WeClapp.

Compliance covers ISO 27001, ISAE 3000 Type I, GDPR, CCPA and SOC-2, on EU-hosted infrastructure, which is why European buyers put it on a list at all. Access is by API and SDK, integrating with ERP, CRM, accounting software and cloud storage, though without a named connector catalog.

This is the shortlist entry for EU-hosted processing and document fraud checks in one product.

What you can check: the certifications, the EU hosting, the fraud-detection features, and the document-type coverage.

What you can't: verify the rating on the platform it comes from. The 4.8 out of 5 from 31 reviews is Klippa citing Capterra on its own homepage.

17. Sensible

Sensible pairs AI extraction with deterministic rules, so your output is guardrailed instead of probabilistic all the way down. Ever had a model quietly change its mind about a field? Then you know why that matters in production. Your documents move through five stages:

  • Classification.
  • Layout-based and out-of-the-box extraction.
  • Extraction-metrics monitoring.
  • Validation rules.
  • Human review.

There are 150+ prebuilt configurations for common document types, across insurance, healthcare, financial services, real estate and logistics. Custom logic means learning SenseML, the proprietary configuration language. Documented integrations stop at a REST API, per-language SDKs and Zapier.

Pricing is fully public and easy to model. Growth is $449 a month billed annually or $499 monthly for 750 documents, with $0.57 per document over. Scale is $1,349 annually or $1,499 monthly for 3,200 documents, with $0.53 over. Enterprise handles 10,000+ documents a month on a custom basis.

Security covers SOC 2 Type II, HIPAA and GDPR, with NIST-approved encryption in transit and at rest and virtual data segregation across AWS, Microsoft and Google cloud environments. The vendor reports 75 million+ documents processed.

If your problem is production reliability rather than raw model quality, that rules layer is your reason to look here.

What you can check: both published tiers, the exact document limits and overage rates, and the security posture.

What you can't: confirm the score users gave it. The 4.9 out of 5 on G2 is Sensible citing it on its own homepage.

18. Instabase

Instabase AI Hub is built around one unit of work: the packet. Its agents are packet-aware, meaning they validate across documents rather than one at a time, applying multi-step business rules to a bundle. For a mortgage file or an insurance submission, that's the difference between reading pages and understanding a case.

Coverage runs to 55+ file formats, documents of 100+ pages, and handwriting. Around that sit classification, splitting and confidence-scored validation, plus a full human-review interface with SLA and queue management, dev-to-production promotion, accuracy benchmarking and an app marketplace.

Connectivity is where your packet has to arrive from:

  • Storage: Instabase Drive, Google Drive, Google Cloud Storage.
  • Object stores: Amazon S3, Azure Blob Storage.
  • Intake: email inboxes.
  • Direct: API and SDK access.

Enterprise controls include single sign-on, IP whitelisting, multi-region replication, optional dedicated instances and a 99.5% uptime SLA. Those are the controls the capability pages name, and those same pages name no compliance certification beside them.

Traction looks real: a place on the Inc. 5000 Fastest-Growing Private Companies list for 2025, recognition in Gartner's Market Guide for Intelligent Document Processing, and customers including Rocket Mortgage, AXA and NatWest.

Well, the reviewer note is small but pointed, from a base of 2 on G2. Navigation and overall product understanding feel unintuitive next to alternatives, and cost is a challenge. Is your unit of work a page or a bundle? If it's a bundle, put the packet-aware design on your demo agenda.

19. Mindee

Mindee is a developer's tool with a developer's price list. Starter is $44 a month billed annually at $529 a year for one seat. Pro is $116 a month billed annually at $1,393 a year with unlimited seats.

Both cover 6,000 to 300,000 credits a month at roughly $0.044 a credit, and Enterprise starts at 500,000+ credits a year on an annual contract.

No permanent free tier. Fourteen days, then you're paying.

The pipeline chains in a single call: split, classify, crop, extract. You get confidence scores and bounding boxes, plus asynchronous processing with webhooks. Files run to 100MB and 200 pages, in any language or alphabet. A Continuous Learning Model uses retrieval-augmented generation to handle edge cases. Data processing localization appears as a Pro feature.

Integration rests on the raw API with Python and Node.js SDKs, with no named RPA, ERP or automation-platform connectors. Named production customers include Spendesk, Lucca, Indy, Payfit and Circula. Mindee cites 4.8 out of 5 on G2 from 30+ reviews and 4.9 on Capterra from 10+ on its own homepage.

What you can check: both published plans, the credit ranges, the file ceilings, and the named customers.

What you can't: check extraction quality against a number, or pair the Pro-tier data localization with a public security page, because none loaded when the research went looking for one.

20. Infrrd

Infrrd states the largest processing volume in this ranking, 100B+ pages, at a 0.1% error rate across 22+ languages and 1,000+ document types. That error figure is the vendor's own, so weigh it accordingly.

The feature set behind it is genuinely broad, whatever your documents look like. Classification and splitting, extraction and validation, image processing for technical drawings, an agentic capability called Deep Worker, no-touch processing, straight-through processing and human-in-the-loop escalation. It aims at mortgage, insurance, financial services, engineering and logistics.

Analyst standing is solid. Leader placement in the Everest Group PEAK Matrix 2026, plus a highlight in the IDC MarketScape 2024 for AI and large-language-model integration.

The customer roster runs to PwC, Rocket Mortgage, Sedgwick, TCS, Oracle and Johnson & Johnson, and the vendor cites a 70% no-touch rate in a mortgage-origination example.

Security covers SOC 2 Type II attestation with continuous control monitoring and annual audits, plus GDPR privacy-by-design controls with standard contractual clauses for EU data subjects.

The tiers are named but not priced. Intelligent Data Processing. Human-in-the-Loop, with service levels from 15 minutes to 24 hours. No-Touch Processing, with 24/7 support. And AI Agents, built on an agent it calls Ally.

Naming the systems you'd connect it to is the awkward part, because what the official pages offer is API integration in the abstract, and no third-party review base sits behind the vendor's figures either.

Technical drawings sitting alongside mortgage or claims paperwork is an unusual enough combination to justify the call, though expect a procurement cycle rather than a sign-up form.

21. Extend

Ten thousand free credits a month, every month, then $0.0125 per credit. Scale is $500 a month for 50,000 credits with overage at $0.01. Enterprise adds custom per-credit pricing, self-hosted deployment, single sign-on and custom agreements.

That free tier is one of only three standing monthly free plans in this ranking, and it is one of the larger allowances here, though the credits are never converted into a page count you could size against.

Extend is the other vendor here that names a benchmark, claiming 95.7% accuracy on Document Q&A against RealDoc-Bench. No third-party review base stands behind that claim, and no public security page for the HIPAA add-on loaded during the research for this piece either.

Five documented capabilities do the work. Parse to markdown and blocks for RAG. Extract to structured JSON. Then Classify, Split for multi-document PDFs, and Edit for form filling and PDF generation. Workflows chain them, and a Studio lets you iterate on schemas and catch regressions before they reach production.

Confidence scoring runs throughout. Zero data retention is standard on the free and Scale tiers, and a HIPAA business associate agreement is available as an add-on at Scale. The public API version is 2026-02-09.

Customers named include Brex, Flatiron Health, Vendr, CH Robinson, Mercury, Opendoor, Square, Checkr and Ironclad, with the Brex case described as powering document workflows across 30,000 customers. For a real pilot this month, a recurring 10,000 credits is enough to run one on your own documents before you talk to anyone.

22. Indico Data

Carrier or managing general agent drowning in submissions? Start here, then weigh the security notice below alongside the fit.

Indico is one of the narrowest tools in this ranking, and for the right buyer that's the point. It's built for commercial insurance intake: loss runs, ACORD forms and exposure schedules, parsed with field-level validation by three layers of purpose-tuned agents:

  • Ingest: Extraction Agents pull the fields.
  • Enrich: Enrichment Agents add the context.
  • Orchestrate: Orchestration Agents route the decision.

The tooling around them is real, and you don't start from zero. An Agent Gallery holds hundreds of pre-built agents and workflows. An Agent Studio lets you define behavior and validation logic, and an Agentic Workflow Canvas handles drag-and-drop multi-agent decisioning.

Integrations go where carriers live: Guidewire, AWS, Microsoft and Salesforce. The customer list reads like an underwriting conference: Allstate, Aspen, Convex, Everest, Markel, HDI, Aviva and Swiss Re. Recognition comes from Gartner, Quadrant Knowledge Solutions, GigaOm and CB Insights.

No outside account of how the platform performs day to day surfaced during the research for this piece, because every review platform checked was blocked or carried no listing.

So the security review has to carry more weight than usual, and there is something specific in it.

Indico's own homepage carries a visible notice pointing to a dedicated page about a data security incident, sitting next to its stated Trust Center, SOC 2 compliance, encryption, role-based access control and audit logs.

Ask about it directly on the first call, and ask what changed afterwards.

Frequently asked questions

These are the questions buyers send back after they have screened the ledger.

Which intelligent document processing tools have a genuinely free tier rather than a trial?

Three, if you mean an allowance that renews. Microsoft Azure AI Document Intelligence gives you 0 to 500 pages a month on its F0 tier. Extend gives you 10,000 credits a month. LlamaIndex has a Free plan at $0 a month with 10K credits.

Unstructured comes close and is worth counting separately: 10,000 free pages with no card, though that grant is a one-time start rather than a monthly refill.

After those, the clock or the one-time grant applies: three months at Amazon Textract, 14 days at Rossum, Docsumo, Sensible and Mindee, a $300 cloud trial credit at Google Document AI, and starter credits at Nanonets, Reducto and LandingAI. The remaining nine, Hyperscience, UiPath, Base64.ai, ABBYY Vantage, Ocrolus, Klippa DocHorizon, Instabase, Infrrd and Indico Data, publish no free path at all, so your evaluation there starts with a calendar invite.

What does document processing actually cost per page, and which vendors publish rates at all?

Eleven of the 22 publish real figures and eight publish nothing at all, and the numbers themselves sit in the pricing table above, in the form you can copy into a budget deck.

Build your estimate on the operation you'll actually run, because a forms-and-tables call can cost many times what a plain text-detection call costs on the very same page.

Which IDP platforms are HIPAA, SOC 2, or FedRAMP compliant?

Screen by your own gate, not by the longest list.

  • FedRAMP High, the federal gate: Hyperscience, which holds it through a Palantir partnership, and Unstructured.
  • HIPAA: Amazon Textract, Nanonets on Enterprise, LlamaIndex, Docsumo, Sensible, LandingAI, Base64.ai, Reducto on Growth and Enterprise, Extend as a Scale add-on, and UiPath on Automation Cloud Dedicated.
  • SOC 2: the widest credential in this field by a distance. Read it off the compliance column in the ledger rather than off a name-check here.

Ocrolus goes further than most for lending work, adding PCI DSS, GLBA and New York 23 NYCRR 500.

Which tools handle handwritten and non-English documents?

Three answers do most of the discriminating.

  • Google Document AI: 200+ languages, handwriting included.
  • Azure: handwritten text in 12 languages, namely English, Chinese Simplified, French, German, Italian, Thai, Arabic, Japanese, Korean, Portuguese, Spanish and Russian.
  • Amazon Textract: the tight one. Handwriting in English only, text detection in six languages, no vertical-text support.

After that: Base64.ai states OCR in 165+ languages with handwriting, Rossum advertises 276-language extraction, Mindee any language or alphabet, Infrrd 22+ languages, and Hyperscience, ABBYY Vantage, Instabase and Nanonets all handle handwriting.

Can intelligent document processing run on-premises or inside my own VPC?

Yes, at a good number of them, and for a regulated buyer this is the question that ends the shortlist. Ask for the deployment diagram in writing, because your auditors will want it long before your users do. Ordered from the most restrictive down, so you can start at your own floor:

  • Air-gapped on-premises: Reducto.
  • On-premises Docker containers: Azure.
  • On-premises or private cloud on Azure: ABBYY Vantage.
  • Private cloud or on-premises with data residency, on Enterprise: Nanonets.
  • Self-hosted: Extend Enterprise, UiPath Automation Suite.
  • VPC or dedicated instance: LandingAI, Instabase, LlamaIndex on AWS and Azure, Unstructured Business.

Should you buy an IDP platform or build on a document API?

Map yourself onto the tracks first. Short of engineering time and long on process complexity, meaning validation rules, review queues, exception handling and posting into an ERP? Buy the suite. Hyperscience, Rossum, ABBYY Vantage, Ocrolus and Indico Data sell that whole chain, and you pay for it in quote cycles and implementation effort.

Have engineers and a clear schema? The API track costs less and moves faster. Textract, Unstructured, Reducto, Extend and Sensible give you published rates and a same-day start. Which of those two paragraphs describes you is the thing to settle before your first demo, and the section below says how to settle it cheaply.

How to use this list

Cut it in two before you cut it at all. Decide whether you're buying an outcome or a primitive, because that single call removes half the field and most of the confusion.

Then, if the answer is the suite, don't start with the suite. Prototype on a free-tier API first and let your own documents tell you what they do to it.

Walk into the sales cycle carrying a page count and an error rate, and you will negotiate against evidence instead of a requirements list. Your engineers will also know exactly where the hard documents are, which is knowledge no vendor deck hands you.

Screen on the evidence ledger, not the feature list. Who publishes a price? Who gives you a path to test without a contract? Whose accuracy number has a benchmark behind it?

Take the survivors, two or three at most, into demos with the eight questions above already answered. Again, the thing to press on is what you can confirm before a contract exists, so ask each of them the question the deck won't volunteer: what exactly can you show me before I sign?

Two focused hours with the ledger and the checklist will save you a month of demos, and it'll put you in the budget conversation with a defensible number instead of a feeling.