All documentation
RAG response cache
Configure cache-backed reuse for RAG responses.
RAG response cache
The HTTP RAG response cache is opt-in and exact-question only. Redis can share cache entries across processes; bounded in-memory caching is process-local and suitable only for development or a single worker.
Redis configuration
# HTTP response caching is opt-in; install Redis support only when selected.
pip install 'docpipe-sdk[server,rag-redis]'
DOCPIPE_RAG_CACHE_ENABLED=true
DOCPIPE_RAG_CACHE_BACKEND=redis
DOCPIPE_RAG_CACHE_REDIS_URL=rediss://cache-user:<secret>@redis.example:6379/0
DOCPIPE_RAG_CACHE_TTL_SECONDS=300
DOCPIPE_RAG_CACHE_MAX_PAYLOAD_BYTES=262144
DOCPIPE_RAG_CACHE_SOCKET_TIMEOUT_SECONDS=1.0
# For development or one process only:
# DOCPIPE_RAG_CACHE_BACKEND=memory
# DOCPIPE_RAG_CACHE_MAX_ENTRIES=256- Use TLS, network restrictions, ACLs, and a dedicated Redis database or key prefix. Redis availability, durability, replication, and eviction remain operator responsibilities.
- Cache namespaces include query configuration and verified tenant context; caller-supplied tenant headers are not trusted for cache isolation.
- Cache failures are best-effort: the request continues uncached. Expired, malformed, or oversized entries are treated as misses.
- Cached answers and citation payloads are sensitive data. Restrict Redis access, credentials, and backups accordingly.
- There is no distributed stampede lock, invalidation API, or cross-version cache schema guarantee; use a short TTL during upgrades.