Quickstart
Install Docpipe and run a first document query.
Getting Started
docpipe is an open-source document pipeline for parsing, structured extraction, vector ingestion, and retrieval-augmented generation. Use it as a Python SDK, CLI, or shared HTTP service; clients provide their own vector-store configuration.
Use the parse, extract, ingest, and RAG pipelines independently or together. Install only the optional integrations your application needs, then configure the SDK or service with your model and storage providers.
- Parse — convert documents to Markdown or text with a selected parser
- Extract — produce structured entities with LangExtract or LangChain
- Ingest — chunk, embed, and store vectors in pgvector, TurboVec, or the optional Qdrant adapter
- RAG — query indexed content with retrieval strategies, filters, and optional reranking
import docpipe
config = docpipe.IngestionConfig(
connection_string="postgresql://user:pass@localhost:5432/mydb",
table_name="docs",
embedding_provider="openai",
embedding_model="text-embedding-3-small",
)
docpipe.ingest("invoice.pdf", config=config)
result = docpipe.query("What is the invoice total?", config=docpipe.RAGConfig(
connection_string=config.connection_string,
table_name=config.table_name,
embedding_provider="openai",
embedding_model="text-embedding-3-small",
llm_provider="openai",
llm_model="gpt-4o",
))
print(result.answer)Install
Install the core SDK, then add optional extras or a curated profile for the integrations your workload uses.
pip install docpipe-sdk # Core only
# Curated profiles (recommended)
pip install "docpipe-sdk[profile-slim]" # MarkItDown, fast chunking
pip install "docpipe-sdk[profile-balanced]" # Docling, semchunk, hybrid RAG (default)
pip install "docpipe-sdk[profile-quality]" # GLM-OCR, BGE rerank
pip install "docpipe-sdk[profile-agents]" # Balanced + AutoGen
pip install "docpipe-sdk[profile-mcp]" # Balanced + hosted MCP server
pip install "docpipe-sdk[profile-eval]" # Balanced + RAGAS evaluation
pip install "docpipe-sdk[profile-gpu]" # Quality + MinerU / PaddleOCR
# À la carte
pip install "docpipe-sdk[docling]" # Docling parser
pip install "docpipe-sdk[markitdown]" # Lightweight Office/PDF → Markdown
pip install "docpipe-sdk[glm-ocr]" # GLM-OCR (scanned docs)
pip install "docpipe-sdk[langextract]" # Google LangExtract
pip install "docpipe-sdk[openai]" # OpenAI embeddings & LLM
pip install "docpipe-sdk[pgvector]" # PostgreSQL vector store
pip install "docpipe-sdk[turbovec]" # On-disk vector indices
pip install "docpipe-sdk[rag]" # Hybrid BM25 + vector
pip install "docpipe-sdk[rag-redis]" # Optional Redis RAG response cache
pip install "docpipe-sdk[rerank]" # FlashRank reranker
pip install "docpipe-sdk[qdrant]" # Experimental Qdrant vector adapter
pip install "docpipe-sdk[s3]" # S3-compatible source adapter
pip install "docpipe-sdk[server]" # FastAPI + /admin + Alembic
pip install "docpipe-sdk[mcp-server]" # Optional Streamable HTTP MCP endpoint
pip install "docpipe-sdk[observability]" # OpenTelemetry + Prometheus
pip install "docpipe-sdk[http]" # Python HTTP client
pip install "docpipe-sdk[all]" # Dev/CI onlyFor the HTTP API, operator panel, and database migrations install `pip install "docpipe-sdk[server]"`; add `observability` for OpenTelemetry and Prometheus support.
# Runtime presets (pass on /ingest, /rag/query, /agents/query)
# fast — low latency, recursive chunks, no rerank
# balanced — general-purpose default (Docling + hybrid RAG)
# quality — OCR parsers, semantic chunks, BGE rerank
# agents — tool-using RAG via /agents/query
# MCP uses the separate opt-in profile-mcp extra and bearer-token configuration.
curl -u admin:pass http://localhost:8000/profiles
curl -u admin:pass http://localhost:8000/plugins
docpipe plugins list
docpipe profiles list
docpipe resolve invoice.pdf --goal ingest