All documentation
RAG
Retrieve relevant context and generate grounded answers.
RAG
Query indexed documents with retrieval strategies, optional reranking, history, metadata filters, structured output, and SSE streaming. Choose storage, embedding, and language-model providers explicitly.
Python query
import docpipe
config = docpipe.RAGConfig(
connection_string="postgresql://user:pass@localhost:5432/mydb",
table_name="invoices",
embedding_provider="openai",
embedding_model="text-embedding-3-small",
llm_provider="openai",
llm_model="gpt-4o",
strategy="hyde",
reranker="flashrank",
)
answer = docpipe.query("What is the total amount on the invoice?", config=config)
print(answer.answer)
print(answer.sources)
print(answer.usage) # Token counts are available when returned by the provider.| Strategy | Description |
|---|---|
| naive | Simple cosine similarity search. Fast and reliable for well-formed queries. |
| hyde | LLM generates a hypothetical answer, embeds it for retrieval. Highest accuracy on complex questions. |
| multi_query | Expands query into N variants, merges and deduplicates results. Best for vague or short queries. |
| parent_document | Retrieves seed chunks then expands context window per source. Best for long documents. |
| hybrid | Combines dense vector search with BM25 keyword matching. Best for exact terms, IDs, and proper nouns. |
| auto | LLM classifies the question and dispatches to the optimal strategy automatically. |
- CLI: `docpipe rag query "..." --strategy hyde --reranker flashrank`
- API: POST /rag/query (JSON) and POST /rag/stream (SSE)
- POST /agents/query provides AutoGen agentic RAG when the agents integration is installed.
- Pass history and metadata filters only in the formats supported by the selected query endpoint.