Skip to content

Data OpsLumina AI

Chunkr

Open-source document intelligence API for RAG-ready data.

Categories
Data OpsSearch
Pricing
FREEMIUM
Source
Open core
Hosting
Hybrid
Platforms
WebAPI
Models
Self-contained (on-device)
Verified
Jun 8, 2026

A document parsing and intelligence API that turns complex PDFs, slides, Word docs, and images into clean, LLM/RAG-ready chunks. Chunkr runs layout analysis, OCR, reading-order detection, semantic chunking, and schema-based extraction, emitting HTML, Markdown, or JSON. Self-host the open-source pipeline or call the managed cloud API, which includes a free tier of 200 pages with no card required.

Capabilities 3

What it actually does — grouped by capability family.

  • Document parsing (structured) (primary capability)
  • OCR / scanned-document extraction (secondary capability)
  • Structured extraction (secondary capability)

Pros & cons

  • Self-host or call the managed API
  • Layout analysis + OCR + semantic chunking
  • Outputs HTML, Markdown, or JSON
  • Free cloud tier (200 pages, no card)
  • Accuracy below Reducto on hard layouts
  • Lighter compliance coverage than Unstructured
  • Smaller team / younger product

Tags

Further reading

View all Data Ops
  • View Reducto details
    Data OpsFREEMIUM

    Reducto

    Reducto

    Agentic document parsing and extraction for AI teams, via one API.

    A document-intelligence API that parses, splits, extracts, and edits PDFs, images, spreadsheets, and slides into clean, structured output for RAG and AI pipelines. It blends custom in-house models with frontier ones and bills via usage credits, automatically discounting pages it can parse without the heavier pipeline.

    Strong on complex/nested table layouts
    API-only, no app UI
    • document-parsing
    • ocr
    • extraction
    • rag
  • View Unstructured details
    Data OpsFREEMIUMOpen core

    Unstructured

    Unstructured

    ETL for LLMs — turn PDFs, decks, and emails into clean, structured data.

    Ingests 64+ file types and partitions, chunks, enriches, and embeds them into LLM-ready output, handling OCR, tables, and document hierarchy. An open-source library plus a low-code platform and API; a staple preprocessing layer for production RAG.

    64+ file types ingested
    OSS quality trails hosted partition models
    • document-etl
    • preprocessing
    • rag
    • open-source
  • View Docling details
    Data OpsFREEOSS

    Docling

    Docling Project

    Toolkit that turns documents into AI-ready Markdown and JSON.

    A document-processing toolkit that converts PDF, DOCX, PPTX, XLSX, HTML, images, and audio into clean Markdown or JSON for LLM and RAG pipelines. It does advanced PDF understanding — page layout, reading order, table structure, and OCR for scans — and ships a hybrid chunker plus native LangChain and LlamaIndex integrations. Small enough to run on a laptop via a Python API or CLI; MIT-licensed and community-governed.

    Runs on a laptop via Python API or CLI
    Lower accuracy than top hosted parsers
    • document-parsing
    • rag
    • open-source
    • pdf
    • +1
  • View Nanonets details
    Data OpsFREEMIUM

    Nanonets

    Nanonets

    AI agents for document processing and enterprise data extraction.

    Nanonets automates document-heavy workflows — invoices, orders, contracts, and claims — with AI agents that read, extract, and route structured data across ERPs, email, and approval chains. It runs on its own OCR-3 extraction model and can fold in LLMs for agentic pipelines. Offered as managed cloud with VPC, single-tenant, and on-premises deployment options and regional data residency.

    Handles invoices, orders, contracts, claims
    Leaderboard claims are vendor-reported
    • document-ai
    • idp
    • ocr
    • extraction
    • +1