All documentation
Ingest
Index documents for retrieval and downstream workflows.
Ingest
Chunk documents, create embeddings, and write vectors to the selected backend. pgvector is the default production path; TurboVec is a local on-disk option, and Qdrant is an optional experimental adapter. The service does not centralize client vector data.
Python
import docpipe
config = docpipe.IngestionConfig(
connection_string="postgresql://user:pass@localhost:5432/mydb",
table_name="invoices",
embedding_provider="openai",
embedding_model="text-embedding-3-small",
incremental=True,
)
docpipe.ingest("invoice.pdf", config=config)
# Each application supplies its own vector-store configuration.Streaming ingest (SSE)
curl -u admin:secret -N -X POST http://localhost:8000/ingest/stream \
-H "Content-Type: application/json" \
-d '{
"source": "file:///data/doc.pdf",
"connection_string": "postgresql://user:pass@db:5432/mydb",
"table_name": "docs",
"embedding_provider": "openai",
"embedding_model": "text-embedding-3-small",
"preset": "balanced"
}'
# SSE events report resolve, parse, chunk, and completion progress.Incremental ingestion skips unchanged files when supported by the selected backend. DELETE /ingest removes chunks for a source; use the streaming endpoint when progress events are useful.
- CLI: `docpipe ingest report.pdf --db ... --table docs --incremental`
- API: POST /ingest, POST /ingest/stream, DELETE /ingest, and POST /collection/sources
- Embedding integrations include OpenAI, Google, Ollama, and HuggingFace; install the matching extra.
- Qdrant is installed with `docpipe-sdk[qdrant]`; its configuration and operational limits are documented in the Docpipe repository's Qdrant guide.