Your documents.
Your database.
Grounded answers.
An open-source document pipeline for parsing, extraction, and grounded retrieval — as a Python SDK, shared HTTP service, or authenticated MCP server.
The pipeline connects your pieces. It doesn’t replace them.
Docpipe coordinates document work while your application keeps control of its storage, providers, and deployment boundaries.
Read the integration guide- Storage
- Vectors stay in the store your application configures.
- Providers
- Choose the parser, embedding, and language-model integrations you install.
- Composition
- Use each pipeline stage on its own or connect them into a workflow.
From a document to a grounded answer.
Index documents into your chosen store, then retrieve relevant context when a question arrives. Select a step to see where it fits.
IndexDocument → Parse → Chunk & embed → Your vector store
AnswerQuestion → Retrieve → Optional reranking → Grounded answer with sources
Choose how you want to connect.
Start with the interface that fits your application and deployment.
Python SDK
Compose parsing, extraction, ingestion, and retrieval directly in Python.
Start with the SDKShared HTTP service
Run Docpipe once; each client supplies its own storage and model settings.
Connect an applicationMCP endpoint
Expose parsing and retrieval tools through authenticated Streamable HTTP.
Configure MCPStart simple. Add retrieval depth when you need it.
Begin with vector similarity, then choose hybrid search, query expansion, context expansion, or automatic strategy selection. Reranking and streaming are optional parts of the RAG workflow.
Operators get focused guides, not a wall of setup notes.
Access, authentication, observability, and deployment each have a dedicated guide.