Skip to main content
Use review_decision_managed to select the workflow. False or absent preserves your existing legacy endpoints, booleans, and review behavior. True uses explicit review states, terminal cancellation, and read-only decisions. For managed documents, missing or unknown state stops processing. See the state dispatcher.

What is Documind API?

Documind provides a powerful document extraction API designed for automation developers and integration engineers. Upload documents, define extraction schemas, and receive structured JSON data with confidence scores and automatic review flagging.

Key Capabilities

Multi-Model Extraction

Choose between Basic, VLM-based, or Advanced (multi-model ensemble) extraction modes with different accuracy and speed trade-offs.

Schema Flexibility

Use predefined schemas, generate from samples, or create custom schemas for any document type.

Review Workflow

Automatic flagging of low-confidence fields enables human-in-the-loop validation for critical data.

Credit-Based Usage

Transparent per-page pricing with different credit costs for Basic (2-6), VLM (10), and Advanced (15) extraction.

How It Works

1

Upload Documents

Submit PDF, Word, or image files via the /upload endpoint. Returns document IDs for processing.
2

Define Schema

Specify what data to extract using JSON Schema format. Generate schemas automatically or use predefined templates.
3

Extract Data

Process documents with /extract/{document_id}. Choose extraction mode and review threshold.
4

Handle Reviews

Poll /data/extractions to detect when review_status=approved for documents that needed review. Use corrected data in your automation.

Extraction Modes

Single-model extraction with your choice of:
  • google-gemini-3-flash (6 credits)
  • google-gemini-2.5-flash (4 credits)
  • qwen-3-vl (2 credits)
Best for simple documents where extensive validation isn’t required.
Uses native visual data to process content through multiple Vision-Language Models.Best for low-text, high-visual content like scanned documents, images, or forms where layout is crucial.
Multi-model ensemble extraction utilizing document layout, reading order, and OCR’d text.Includes confidence scores and automatic review flagging. Best for structured documents like invoices, forms, and tables requiring high accuracy. Activated by not setting model or extraction_mode parameters.

Use Cases

Common Use Cases

  • Invoice Processing: Extract line items, totals, vendor details
  • Form Data Entry: Digitize paper forms into structured data
  • Document Classification: Identify document types and route accordingly
  • Compliance Checks: Extract specific fields for validation

Integration Examples

Copy poll_for_review from Polling for Reviews into a local review_polling.py module in your integration. Copy the imports and function definition only; keep the module-level usage demonstration in your caller. This is customer code, not an installed SDK. Synchronous extraction responses do not include technical status; the helper reads stored state.

Authentication

All API requests require authentication using API keys passed in the X-API-Key header:
Never commit API keys to version control. Use environment variables or secure credential storage.

Rate Limits & Credits

  • API Calls: Track usage via /usage/current endpoint
  • Credits: Deducted per page/image processed
  • Daily Refresh: Credits refresh based on your subscription tier
  • Insufficient Credits: Returns 402 Payment Required status
Check your current credits:

Next Steps

Quick Start

Get started with your first extraction in 5 minutes

Authentication

Learn how to create and manage API keys

Extraction Flow

Understand the complete extraction workflow

Review Polling

Implement review polling for automation pipelines