◆docpipe
Product
Docs
GitHubPyPI v0.7.0
All documentation

RAG

Retrieve relevant context and generate grounded answers.

RAG

Query indexed documents with retrieval strategies, optional reranking, history, metadata filters, structured output, and SSE streaming. Choose storage, embedding, and language-model providers explicitly.

Source: RAG usage overviewSource: RAG pipeline implementation
Python query
import docpipe

config = docpipe.RAGConfig(
    connection_string="postgresql://user:pass@localhost:5432/mydb",
    table_name="invoices",
    embedding_provider="openai",
    embedding_model="text-embedding-3-small",
    llm_provider="openai",
    llm_model="gpt-4o",
    strategy="hyde",
    reranker="flashrank",
)
answer = docpipe.query("What is the total amount on the invoice?", config=config)
print(answer.answer)
print(answer.sources)
print(answer.usage)  # Token counts are available when returned by the provider.
StrategyDescription
naiveSimple cosine similarity search. Fast and reliable for well-formed queries.
hydeLLM generates a hypothetical answer, embeds it for retrieval. Highest accuracy on complex questions.
multi_queryExpands query into N variants, merges and deduplicates results. Best for vague or short queries.
parent_documentRetrieves seed chunks then expands context window per source. Best for long documents.
hybridCombines dense vector search with BM25 keyword matching. Best for exact terms, IDs, and proper nouns.
autoLLM classifies the question and dispatches to the optimal strategy automatically.
  • CLI: `docpipe rag query "..." --strategy hyde --reranker flashrank`
  • API: POST /rag/query (JSON) and POST /rag/stream (SSE)
  • POST /agents/query provides AutoGen agentic RAG when the agents integration is installed.
  • Pass history and metadata filters only in the formats supported by the selected query endpoint.