All documentation
Parse
Turn a source document into structured content.
Parse
Convert PDFs, Office files, HTML, and images to Markdown or plain text. Choose an installed parser explicitly or use the automatic router; parser availability depends on installed extras and runtime profile.
Python
import docpipe
# Use parser="auto" to route by document type, or choose a parser explicitly.
document = docpipe.parse("invoice.pdf", parser="markitdown")
print(document.markdown)
# OCR for scanned or image-heavy files (install the glm-ocr extra).
scanned = docpipe.parse("scanned-report.pdf", parser="glm-ocr")- CLI: `docpipe parse invoice.pdf --format markdown`
- API: POST /parse with a configured local path or allowed URL source
- Available integrations include MarkItDown, Docling, GLM-OCR, PyMuPDF, MinerU, PaddleOCR, and Unstructured; each has different optional dependencies.