◆docpipe
Product
Docs
GitHubPyPI v0.7.0
All documentation

Ingest

Index documents for retrieval and downstream workflows.

Ingest

Chunk documents, create embeddings, and write vectors to the selected backend. pgvector is the default production path; TurboVec is a local on-disk option, and Qdrant is an optional experimental adapter. The service does not centralize client vector data.

Source: Integration and storage boundariesSource: Vector-store support matrix
Python
import docpipe

config = docpipe.IngestionConfig(
    connection_string="postgresql://user:pass@localhost:5432/mydb",
    table_name="invoices",
    embedding_provider="openai",
    embedding_model="text-embedding-3-small",
    incremental=True,
)
docpipe.ingest("invoice.pdf", config=config)
# Each application supplies its own vector-store configuration.
Streaming ingest (SSE)
curl -u admin:secret -N -X POST http://localhost:8000/ingest/stream \
  -H "Content-Type: application/json" \
  -d '{
    "source": "file:///data/doc.pdf",
    "connection_string": "postgresql://user:pass@db:5432/mydb",
    "table_name": "docs",
    "embedding_provider": "openai",
    "embedding_model": "text-embedding-3-small",
    "preset": "balanced"
  }'
# SSE events report resolve, parse, chunk, and completion progress.

Incremental ingestion skips unchanged files when supported by the selected backend. DELETE /ingest removes chunks for a source; use the streaming endpoint when progress events are useful.

  • CLI: `docpipe ingest report.pdf --db ... --table docs --incremental`
  • API: POST /ingest, POST /ingest/stream, DELETE /ingest, and POST /collection/sources
  • Embedding integrations include OpenAI, Google, Ollama, and HuggingFace; install the matching extra.
  • Qdrant is installed with `docpipe-sdk[qdrant]`; its configuration and operational limits are documented in the Docpipe repository's Qdrant guide.