Docling Serve, published as docling-serve, is the official HTTP server for Docling, IBM’s open-source document parser. It wraps Docling’s conversion pipeline in a FastAPI app: you send a PDF, DOCX, PPTX, HTML or image to /v1/convert/... and get Markdown, HTML, JSON or DocTags back. It ships as a pip package and as CPU and CUDA container images with the default models bundled, offers synchronous and asynchronous endpoints, an optional web UI, and Redis-backed queue engines for scaling out. This guide covers running it yourself.
Everything below is checked against the v1.32.0 source tree, released September 1, 2026. That release bundles docling-slim 2.124.0, docling-jobkit 3.5.0 and docling-parse 7.16.0 (changelog). The stable API lives under /v1; anything you find online posting to /v1alpha is dead.
Two warnings up front, because they cause most of the wasted afternoons:
- The published usage docs still show pre-v1 examples in places (
http_sources,file_sources,-F 'parameters=...'). The v1 migration notes and the source are correct; the examples are not. - Almost every app-level setting is environment-variable only. The CLI exposes just
--artifacts-pathand--enable-ui. Configure through env vars and you never hit the mismatch.
Installing Docling Serve
Pip:
pip install "docling-serve[ui]==1.32.0"
docling-serve run --enable-ui
Container, pinned:
docker run --rm -p 127.0.0.1:5001:5001 \
-e DOCLING_SERVE_ENABLE_UI=1 \
quay.io/docling-project/docling-serve-cpu:v1.32.0
The port publish above is bound to loopback deliberately. Docling Serve listens on 0.0.0.0 inside the container and ships with authentication off and CORS wide open, so -p 5001:5001 on a cloud VM publishes an unauthenticated parser to the internet.
Either way you get three things on port 5001:
http://127.0.0.1:5001for the APIhttp://127.0.0.1:5001/docsfor the API reference — ReDoc, with Swagger UI at/swaggerand Scalar at/scalarhttp://127.0.0.1:5001/uifor a browser playground, when the UI is enabled
Images are published to both quay.io/docling-project/... and ghcr.io/docling-project/.... The image variants, sizing and Compose setup are covered in the companion guide, Docling Docker: CPU/GPU images, Compose and model cache.
Configuration
Every app-level setting is a DOCLING_SERVE_* environment variable, read by a pydantic-settings model; a .env file or a YAML/JSON file named in DOCLING_SERVE_CONFIG_FILE works too, with env vars winning. The ones this guide leans on are MAX_SYNC_WAIT, MAX_DOCUMENT_TIMEOUT, MAX_SOURCES_PER_REQUEST, API_KEY, ENG_KIND and ARTIFACTS_PATH. There are close to 180 in total, and the upstream doc lists about half of them. The Docling Serve configuration reference covers all of them with type, default and effect, grouped by purpose, and includes a Compose file that sets the ones that matter in production.
Your first conversion
Two sync endpoints: POST /v1/convert/source for URLs and inline base64, POST /v1/convert/file for multipart uploads.
curl -X POST http://127.0.0.1:5001/v1/convert/source \
-H "Content-Type: application/json" \
-d '{
"sources": [
{ "kind": "http", "url": "https://arxiv.org/pdf/2501.17887" }
],
"options": { "to_formats": ["md"] }
}'
The response shape, with illustrative values:
{
"document": {
"md_content": "...",
"json_content": {},
"html_content": "",
"text_content": "",
"doctags_content": ""
},
"status": "success",
"processing_time": 12.4,
"timings": {},
"errors": []
}
status is one of success, partial_success, skipped, or failure. Format fields are populated only for what you asked for in options.to_formats, which defaults to Markdown alone. Ask for md and html in one request if you need both; converting twice pays the parse cost twice.
If you are porting from a v1alpha example, the payload shape changed. http_sources and file_sources are gone, replaced by a single sources array where each entry carries a kind of http or file.
Uploading a local file
Multipart upload is where stale examples cost the most time. There is no parameters field holding a JSON blob. Docling Serve flattens the options model into individual form fields, so every option is its own -F, and repeated fields build lists:
curl -X POST http://127.0.0.1:5001/v1/convert/file \
-F "files=@document.pdf" \
-F "to_formats=md" \
-F "table_mode=accurate" \
-F "do_ocr=false"
The flattening happens in FormDepends, which builds one form parameter per model field. Nested objects and dicts are the exception: those are accepted as JSON-encoded strings in a single field. Since v1.31.0 an invalid value returns 422 rather than 500, and the error loc points at the form field name.
Or send the file inline as base64 to /v1/convert/source:
{
"sources": [
{
"kind": "file",
"filename": "document.pdf",
"base64_string": "JVBERi0xLjQK..."
}
],
"options": { "to_formats": ["md"] }
}
Options that matter
Most of Docling’s conversion options are exposed through options, subject to whatever the server allows. The knobs worth knowing, with the defaults the server actually applies:
| Option | Default | Notes |
|---|---|---|
to_formats | ["md"] | Also json, html, text, doctags, chunks, latex, and more |
do_ocr | true | Turn it off for text-native PDFs |
force_ocr | false | Replaces existing text layers with OCR output |
ocr_preset | auto | Replaces the deprecated ocr_engine field |
table_mode | accurate | fast trades table fidelity for throughput |
pdf_backend | threaded_docling_parse | |
image_export_mode | placeholder | placeholder forces include_images off |
images_scale | 2.0 | Server-capped at 2.0 by default |
document_timeout | server maximum | See below — the default is not what you want |
page_range | whole document | 1-indexed |
Three of those defaults deserve more than a table row.
do_ocr defaults to on. It doesn’t re-read text a PDF already embeds, but the stage still selects layout regions and OCRs anything that overlaps an image or shape. On a digital-native corpus, turning it off removes that work. For scans that come back incomplete, see the Docling OCR guide. table_mode defaults to accurate, which is the slower of the two table paths. Those two flags are the cheapest latency wins available and neither changes your deployment.
document_timeout is the trap. If a request omits it, the server substitutes DOCLING_SERVE_MAX_DOCUMENT_TIMEOUT, whose default is 604800 seconds — seven days. Set the server maximum to something sane and one bad scan can no longer occupy a worker for a week.
The full option list is generated in the usage docs; read that table for names and allowed values, and the source for defaults.
Limits that return 422
Three server-side limits are easy to trip without knowing they exist:
DOCLING_SERVE_MAX_SOURCES_PER_REQUESTdefaults to 3. A fourth source, or a fourth file on the multipart endpoint, is rejected. Batch bigger jobs client-side.images_scaleaboveDOCLING_SERVE_MAX_IMAGES_SCALE(default 2.0) is rejected rather than clamped.document_timeoutabove the server maximum is rejected, as is any value at or below zero.
DOCLING_SERVE_MAX_FILE_SIZE and DOCLING_SERVE_MAX_NUM_PAGES are unbounded by default. Set both before you accept untrusted uploads.
Authentication
Off by default. Set one shared secret on the server:
DOCLING_SERVE_API_KEY=changeme docling-serve run
Then send it on every API call, not just the submit:
curl http://127.0.0.1:5001/v1/result/$TASK_ID \
-H "X-Api-Key: changeme"
The dependency is attached to the convert, chunk, status-poll and result routes alike, so a client that authenticates the submission and then polls without the header gets a bare 401 on the poll. The websocket endpoint is the odd one out: it takes the key as a query parameter, ?api_key=SECRET, and the source notes that query-parameter keys tend to land in proxy access logs.
The health surface — /health, /livez, /ready, /readyz, /metrics, /version — is unauthenticated so probes work. /version and the /v1/memory/* endpoints can be closed off with DOCLING_SERVE_SHOW_VERSION_INFO=false and by leaving DOCLING_SERVE_ENABLE_MANAGEMENT_ENDPOINTS off.
That is the whole auth story: one shared key, no per-tenant credentials, no rate limiting, no quota tracking. If you need any of that, it goes in front of Docling Serve, not inside it.
Async jobs
Sync requests have a hard ceiling at DOCLING_SERVE_MAX_SYNC_WAIT, default 120 seconds. Anything that might run longer belongs on the async endpoints.
Same payload, different path. With the key from the previous section and jq to pull out the task id:
DOCLING_URL=http://127.0.0.1:5001
DOCLING_API_KEY=changeme
TASK_ID=$(curl -s -X POST "$DOCLING_URL/v1/convert/source/async" \
-H "X-Api-Key: $DOCLING_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"sources": [
{ "kind": "http", "url": "https://arxiv.org/pdf/2501.17887" }
],
"options": { "to_formats": ["md"] }
}' | jq -r .task_id)
POST /v1/convert/file/async is the multipart equivalent, with the same flat form fields as the sync version.
The full response is a task descriptor, shown here with illustrative values:
{
"task_id": "...",
"task_type": "convert",
"task_status": "pending",
"task_position": 1,
"task_meta": null
}
task_status moves through pending, started, then success or failure. Poll it:
curl "$DOCLING_URL/v1/status/poll/$TASK_ID?wait=30" \
-H "X-Api-Key: $DOCLING_API_KEY"
The wait query parameter is worth knowing — it long-polls for up to that many seconds instead of returning immediately, which cuts request volume substantially compared to a tight loop. Or open a websocket at /v1/status/ws/{task_id} and skip polling entirely.
When status is success, fetch the result:
curl "$DOCLING_URL/v1/result/$TASK_ID" \
-H "X-Api-Key: $DOCLING_API_KEY"
Fetch it once. DOCLING_SERVE_SINGLE_USE_RESULTS defaults to true, which schedules the result for removal as soon as it has been read, with DOCLING_SERVE_RESULT_REMOVAL_DELAY (default 300s) as the grace period. A retry that re-fetches a consumed result after that grace period gets a 404; store the first successful response locally.
Which engine runs your jobs
DOCLING_SERVE_ENG_KIND selects the async engine, and the choice is local, rq, or ray:
| Engine | Backing | Survives an API restart | Use when |
|---|---|---|---|
local (default) | In-process, in-memory | No | Single container, low volume |
rq | Redis + RQ workers | Queued tasks and stored results, as far as Redis keeps them | Multiple containers, real throughput |
ray | Redis + a Ray cluster | Queued tasks and stored results, as far as Redis keeps them | You already operate Ray |
An older version of this post claimed Docling Serve has no queue option. That was wrong, and worth correcting: with local you genuinely do lose pending tasks on restart, because they live in the process that accepted them. With rq you set DOCLING_SERVE_ENG_KIND=rq and DOCLING_SERVE_ENG_RQ_REDIS_URL on both sides and run separate docling-serve rq-worker containers; the API process then does no conversion at all. Durability is then whatever your Redis gives you: run it with persistence enabled, expect results to expire after DOCLING_SERVE_ENG_RQ_RESULTS_TTL (default 4 hours), and do not expect a conversion that was mid-flight on a killed worker to resume — it fails and the client resubmits. That distinction also matters for readiness — /ready gates on model load under local, but under RQ it checks the API queue processor and Redis connection, not whether a conversion worker is ready.
Redis pool sizing, per the configuration reference: the default 50 connections covers 1–4 workers, 100 for 5–10, and 150–200 beyond that.
The ray engine is selectable and validated in settings.py — it requires both DOCLING_SERVE_ENG_RAY_REDIS_URL and DOCLING_SERVE_ENG_RAY_ADDRESS — and carries a large set of DOCLING_SERVE_ENG_RAY_* tuning knobs for autoscaling, tenant fairness and page-slice fan-out. It is not covered in the configuration reference, so expect to read the source. If you do not already run Ray, run RQ.
Fixing timeouts
Timeouts are the single most common Docling Serve complaint, and the mechanics are specific enough to be worth spelling out.
What a sync timeout actually does. The sync handler enqueues the task, then polls its own orchestrator every DOCLING_SERVE_SYNC_POLL_INTERVAL (default 2s) until DOCLING_SERVE_MAX_SYNC_WAIT (default 120s) elapses. It then raises 504 with a detail string naming the setting. The task is not cancelled — the source carries a TODO: abort task! at exactly that point. So a client that hammers the sync endpoint and gives up at 120 seconds keeps stacking live conversions behind the scenes, and the 504 body does not include a task_id you could use to recover the work.
Because of that, raising MAX_SYNC_WAIT is the wrong first move. The fixes, in order:
- Move to
/v1/convert/source/asyncand poll. This is the real fix. It takes the long conversion out of your load balancer’s timeout window entirely. - Set a per-document ceiling.
DOCLING_SERVE_MAX_DOCUMENT_TIMEOUTcaps how long any single document may run, and fills indocument_timeoutwhen the caller omits it. Without it, that ceiling is seven days. - Cut the work.
do_ocr: falseon text-native PDFs,table_mode: "fast", or apage_rangefor documents where you only need the front matter. - Only then raise the wait. If a specific caller genuinely needs synchronous behaviour, raise
MAX_SYNC_WAITand the read timeout on every proxy in the path. A 504 from your ingress at 60 seconds is not the same 504 as Docling’s.
Under RQ there is a second timeout. DOCLING_SERVE_ENG_RQ_JOB_TIMEOUT defaults to 4 hours, and the value is baked into each job at enqueue time. That means it must be set on the API containers, not the workers, and jobs already queued keep whatever value they were enqueued with. Setting it to 0 does not disable it; RQ reads 0 as unset and falls back to its own 180-second default, which will look like conversions being killed for no reason. Use -1 to disable.
Troubleshooting
| Symptom | Cause | Fix |
|---|---|---|
404 on /v1alpha/... | Stale tutorial | Move to /v1 and the sources payload shape |
504 from a sync endpoint | Exceeded MAX_SYNC_WAIT; the task keeps running | Use the async endpoints; see above |
422 naming a form field | Multipart options are flat fields, not a parameters JSON blob | One -F per option |
422 on too many sources | MAX_SOURCES_PER_REQUEST defaults to 3 | Raise it, or batch client-side |
401 on poll or result, submit worked | The API key is required on the status and result routes too | Send X-Api-Key on every call; websockets use ?api_key= |
| Result 404s after the removal delay | SINGLE_USE_RESULTS is on by default | Read once, or set it to false |
| Job runs for hours with no progress | document_timeout defaulted to the 7-day server maximum | Set MAX_DOCUMENT_TIMEOUT |
| Conversions killed at ~3 minutes under RQ | ENG_RQ_JOB_TIMEOUT=0 is read by RQ as unset | Use -1 to disable, or a real value |
ImportError: libGL.so.1 | OpenCV’s native dependency missing on a slim base image | The official images include libglvnd-glx; on a custom image install libgl1 or use opencv-python-headless |
| GPU present, still slow | ONNX Runtime fell back to its CPU provider | Check ort.get_available_providers() inside the container; see below |
| Missing-model error after deploy | Empty bind mount hiding the image’s bundled models | See the Docker guide |
GPU acceleration is uneven
Layout and table models pick up a CUDA device automatically once you run a cu* image and set DOCLING_DEVICE=cuda. OCR is the part that disappoints. RapidOCR supports GPU through both its Torch and ONNX Runtime backends, but the ONNX path silently falls back to CPU unless the CUDA execution provider is actually loadable, which the official GPU guide tells you to assert explicitly:
import onnxruntime as ort
assert "CUDAExecutionProvider" in ort.get_available_providers()
A user in docling#2635 reports slow RapidOCR performance with ONNX GPU and better results with Torch-backed EasyOCR. That report is a troubleshooting lead, not a benchmark of the current release. If you need OCR on a GPU, start from a Torch-backed engine.
Throughput numbers to size against
The Docling maintainers publish measured throughput in the GPU support guide. Their standard-pipeline figures, on a 192-page PDF containing 95 tables and with no OCR:
| Hardware | Standard, no OCR | GraniteDocling VLM |
|---|---|---|
AWS g6e.2xlarge (L40S 48GB) | 3.1 pages/s | 2.4 pages/s |
| RTX 5090 | 7.9 pages/s | 3.8 pages/s |
| RTX 5070 | 4.2 pages/s | 2.0 pages/s |
| CPU only, 16 torch threads | 1.2–1.5 pages/s | — |
These come from one document on one set of machines, and the guide itself notes that GPU optimisation is an active topic. Treat them as an order of magnitude for capacity planning, then measure your own corpus — OCR-heavy scans behave nothing like the digital-native PDF used here.
What about drmingler/docling-api?
drmingler/docling-api is the other Docling HTTP wrapper you find on GitHub: a FastAPI app around Docling with Celery, Redis and Flower, exposing /documents/convert, /conversion-jobs and /batch-conversion-jobs. It is actively maintained and worth knowing about if you already run Celery and want batching handed to you.
Now that docling-serve ships RQ and Ray orchestrators of its own, though, the official server is the better default for new work; pick drmingler when its Celery-based batch API fits infrastructure you already operate.
Self-hosting versus a hosted Docling API
Self-host when you have a long-lived host, control over the runtime, and ideally a GPU. Compliance constraints push you here too: Docling Serve can keep document processing inside your VPC when you use local models and keep remote inference disabled. At high enough sustained volume a GPU instance you keep busy can also beat per-page pricing, but only if you keep it busy — run the numbers on your own volume.
Serverless callers are the awkward case, and it is worth being precise about why. It is not that Docling cannot run on managed platforms — Cloud Run supports GPUs with no drivers to install, and IBM’s serverless fleets with GPUs on Code Engine handle batch corpus conversion well. The problem is that a multi-gigabyte image plus model load makes cold starts expensive on request-scoped platforms, and the fix — minimum instances that keep the model resident — means paying for capacity that sits idle between documents. Multi-gigabyte deploy artefacts are also flatly incompatible with some platforms’ package limits.
So the real question is not “can I self-host” but “do I want to own the model cache, GPU drivers, queue workers, readiness probes and autoscaling.” If the answer is no, Parsebridge runs Docling as a managed API.
It is our own product, so here is the honest scope: Parsebridge is PDF to Markdown through its own API, not a drop-in Docling Serve endpoint. The request shape is ours, not Docling’s:
curl -X POST https://api.parsebridge.com/v1/parse/url \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{"url": "https://example.com/document.pdf"}'
{
"success": true,
"content": "# Document title\n\n...",
"pageCount": 12,
"creditsUsed": 12
}
One call, Markdown back, no PyTorch in your deploy artefact and no /ready probe to wire up. If you already self-host Docling Serve happily, stay there — same parser either way. If you are here because of a timeout you cannot pin down, the Docker guide covers the container-side half of these failures.
The API shapes, defaults, limits and error paths above were read directly from the docling-serve v1.32.0 source rather than reproduced from its docs, which lag the v1 API in places. Throughput and memory figures are attributed to their upstream sources; we did not re-run those measurements.