Docling Serve is the official way to run Docling in Docker. It ships container images on Quay and GHCR, and the smallest working command is one line:
docker run --rm -p 127.0.0.1:5001:5001 \
quay.io/docling-project/docling-serve-cpu:v1.32.0
That gets you a FastAPI server with the full Docling pipeline behind it, on loopback. Everything that matters comes after that line: which image you pick, whether you need a model volume at all, how much memory you give it, and which environment variables you set before real traffic arrives.
This is the container-side checklist. v1.32.0, released September 1, 2026, is the reference throughout, and the details are read from the tagged source tree rather than the docs. For the API shapes, async jobs and timeout behaviour, see the companion guide: Docling Serve: REST API, async jobs and timeout fixes.
Picking a Docling Docker image
Four official images, mirrored on quay.io/docling-project/... and ghcr.io/docling-project/...:
| Image | Torch flavour | Arch | Documented size |
|---|---|---|---|
docling-serve | PyPI torch | amd64, arm64 | 8.7 GB / 4.4 GB |
docling-serve-cpu | CPU-only torch wheels | amd64, arm64 | 4.4 GB |
docling-serve-cu128 | CUDA 12.8 wheels | amd64 | 11.4 GB |
docling-serve-cu130 | CUDA 13.0 wheels | amd64, arm64 | not published |
Sizes are the ones stated in the project README; we have not measured them.
Use docling-serve-cpu if you do not have a GPU. The base docling-serve image is roughly twice the size on amd64 because the PyPI torch build pulls CUDA libraries you will not use — dead weight on a CPU host, and dead weight you pay for on every pull.
For GPU, match the tag to your host driver: CUDA 12.8 wheels need a driver supporting the 12.8 runtime, CUDA 13.0 wheels need 13.0. The documented tagging policy is that CUDA images are published only with explicit version tags and main — no latest — specifically so you cannot drift onto a deprecated CUDA build. Tag availability on the registries has not always matched that policy exactly, which is one more reason to pin: docling-serve-cu128:v1.32.0. Use :main only for testing; it tracks unreleased commits.
An AMD ROCm image is supported but not published because of its size. Build it with make docling-serve-rocm-image.
The model cache: usually nothing to do
This is the section most people get wrong, including an earlier version of this post.
The official images already contain the model weights. The Containerfile runs docling-tools models download at build time for layout, tableformer, picture_classifier, rapidocr and easyocr, into the path it also sets as the default DOCLING_SERVE_ARTIFACTS_PATH:
ENV DOCLING_SERVE_ARTIFACTS_PATH=/opt/app-root/src/.cache/docling/models
ARG MODELS_LIST="layout tableformer picture_classifier rapidocr easyocr"
RUN docling-tools models download -o "${DOCLING_SERVE_ARTIFACTS_PATH}" ${MODELS_LIST}
For the standard pipeline there is no runtime download and no volume to mount. Skip the model-cache machinery entirely.
There is a corollary that surprises people: when DOCLING_SERVE_ARTIFACTS_PATH is set, Docling Serve loads models from that path and raises an error if one is missing rather than fetching it, by design, so the container never calls out to Hugging Face on its own. Which means the failure mode is not a slow first request — it is a hard error.
The two ways to break it
Bind-mounting over the cache. A bind mount of an empty host directory onto /opt/app-root/src/.cache (or the models path under it) hides the bundled weights. The container then looks in an empty directory, finds nothing, and fails. This is a real pattern in the wild — the Compose file in docling#2635 does exactly this while chasing a performance problem. If you bind-mount that path, you own populating it.
Named volumes behave differently: Docker copies the image’s existing content into a fresh, empty named volume on first use, so a new named volume at the models path inherits the bundled weights rather than hiding them. That is a convenience, not a guarantee — a volume populated once by an older image tag keeps that older content forever, which is its own quiet failure.
Mixing up the two variables. DOCLING_SERVE_ARTIFACTS_PATH configures Docling Serve; DOCLING_ARTIFACTS_PATH configures the underlying Docling library. Set the one belonging to the layer you are configuring.
When you do need extra models
Enrichment features are the exception. Picture description, code and formula enrichment, and the VLM pipelines need weights that are not in the standard image, and they will error rather than download. Three supported approaches, from docs/models.md:
Bake them into a derived image — best for production, since the result is still a single immutable artefact:
FROM quay.io/docling-project/docling-serve-cpu:v1.32.0
RUN docling-tools models download smolvlm code_formula
Or mount a directory you populated yourself, pointing the server at it. Note that this replaces the bundled set, so download everything you need, including the defaults:
docling-tools models download --all -o ./models
docker run -p 127.0.0.1:5001:5001 \
-v "$(pwd)/models:/opt/app-root/src/models" \
-e DOCLING_SERVE_ARTIFACTS_PATH=/opt/app-root/src/models \
quay.io/docling-project/docling-serve-cpu:v1.32.0
Or, for local development only, clear the path with -e DOCLING_SERVE_ARTIFACTS_PATH="" to re-enable auto-download at startup. That trades a hard error for a slow, network-dependent boot, which is the wrong trade in production.
On Kubernetes the equivalent is a PVC plus a download Job, with the deployment mounting the populated volume and DOCLING_SERVE_ARTIFACTS_PATH set to the mount path. The project ships manifests for exactly that.
Docling Docker Compose setup
Since the models ship in the image, the Compose file is a single service — no one-shot download container, no volume. What it does spend its lines on is the configuration that decides whether the thing survives contact with traffic: worker count, thread caps, request limits, and a readiness probe with a realistic grace period.
Download compose.yaml — the same file as below, with the reasoning inline as comments.
services:
docling:
image: quay.io/docling-project/docling-serve-cpu:v1.32.0
ports:
- '127.0.0.1:5001:5001'
environment:
UVICORN_WORKERS: '1'
DOCLING_NUM_THREADS: '4'
OMP_NUM_THREADS: '4'
DOCLING_SERVE_ENG_LOC_NUM_WORKERS: '2'
DOCLING_SERVE_ENG_LOC_SHARE_MODELS: 'true'
DOCLING_SERVE_MAX_DOCUMENT_TIMEOUT: '1800'
DOCLING_SERVE_MAX_SYNC_WAIT: '120'
DOCLING_SERVE_MAX_FILE_SIZE: '52428800'
DOCLING_SERVE_MAX_NUM_PAGES: '200'
DOCLING_SERVE_LOG_LEVEL: 'INFO'
DOCLING_SERVE_LOG_FORMAT: 'json'
healthcheck:
test:
- CMD
- python
- -c
- "import urllib.request; urllib.request.urlopen('http://127.0.0.1:5001/ready', timeout=5).read()"
interval: 30s
timeout: 10s
retries: 5
start_period: 180s
stop_grace_period: 2m
restart: unless-stopped
deploy:
resources:
limits:
cpus: '4'
memory: 12G
Four decisions in there are worth defending:
The published port is bound to 127.0.0.1. Docling Serve listens on 0.0.0.0 inside the container, ships with authentication off and CORS set to *. -p 5001:5001 on a cloud VM publishes an unauthenticated document parser to the internet. Put a reverse proxy with TLS in front, and set DOCLING_SERVE_API_KEY before anything else can reach it.
start_period: 180s, not 30s. The probe hits /ready, which returns 503 until the models are loaded, and model load on a cold container is slow. A grace period shorter than your real load time marks the container unhealthy before it ever serves a request. Compose only records that status, dependent services using condition: service_healthy may wait, while an orchestrator that restarts unhealthy containers may interrupt startup. Time your own cold start and set this above the observed worst case.
MAX_DOCUMENT_TIMEOUT is set explicitly. Its default is 604800 seconds — seven days — and the server substitutes that default into any request that omits document_timeout. This single line is the difference between a bad scan failing and a bad scan holding a worker until you notice.
stop_grace_period under the service, not at the top level. Conversions in flight are not resumable. Two minutes of drain is the difference between a clean rollout and a wave of errors handed to callers mid-deploy.
The python -c probe rather than curl is deliberate too: the official image is built on a Python base and the interpreter is guaranteed present, whereas curl is not in the OS package list. Exec form avoids nesting shell quotes inside YAML.
Health and readiness
Docling Serve registers these in app.py, all unauthenticated so probes work without a key:
| Endpoint | Meaning |
|---|---|
/health, /livez | Liveness. /livez also fails if the background queue processor has died |
/ready, /readyz | Readiness. 503 until models are loaded and the orchestrator answers |
/metrics | Prometheus |
/version | Package versions, unless SHOW_VERSION_INFO=false |
/livez and /readyz are the Kubernetes-conventional aliases and are hidden from the OpenAPI schema, but they exist.
Point your load balancer at /ready, not /health. /health returns immediately and tells you nothing about whether the models are loaded; admit traffic on that basis and the first requests after a rollout pile up and time out. One caveat: the model-load gate only applies to the default local engine. Under RQ the models live in the worker containers. The API probe checks its background queue processor and Redis connection; it says nothing about whether a conversion worker is ready. Probe or monitor the workers separately.
GPU passthrough
Three things on the host: an NVIDIA driver matching the CUDA wheels in the image, the nvidia-container-toolkit, and a way to hand the device to the container. The project’s NVIDIA deployment example states driver >= 550.54.14 as its requirement; a newer CUDA tag needs correspondingly newer.
docker run --gpus all -p 127.0.0.1:5001:5001 \
-e DOCLING_DEVICE=cuda \
quay.io/docling-project/docling-serve-cu128:v1.32.0
As a Compose override on the file above:
services:
docling:
image: quay.io/docling-project/docling-serve-cu128:v1.32.0
environment:
DOCLING_DEVICE: 'cuda'
NVIDIA_VISIBLE_DEVICES: 'all'
NVIDIA_DRIVER_CAPABILITIES: 'compute,utility'
runtime: nvidia
DOCLING_DEVICE accepts cpu, cuda, mps, or cuda:N to pick a specific GPU; unset means Docling chooses. The official Compose example uses runtime: nvidia and keeps the deploy.resources.reservations.devices block commented out as the Swarm-compatible alternative — either works, and some orchestrators want both.
If nvidia-smi works on the host and inside the container, the device is wired up. If Docling still runs slowly after that, the device is not the problem — see the OCR backend note in the Serve guide, because RapidOCR on ONNX Runtime falls back to CPU silently.
Sizing
Docling is memory-hungry: the model graph is several gigabytes once loaded, and a single OCR-heavy page can spike well past that. Starting points to load-test from, not measurements:
| Workload | Starting point |
|---|---|
| Smoke test | 2 vCPU, 4–8 GB RAM |
| CPU, low concurrency | 4 vCPU, 8–16 GB RAM |
| CPU, OCR-heavy or large PDFs | 8+ vCPU, 16–32 GB RAM |
| GPU | 4–8 vCPU, 16 GB+ RAM, 8 GB+ VRAM |
| VLM pipelines | 16 GB RAM floor; VRAM set by the model |
That is more than the small examples in the docs imply, and the reason is worth reading rather than taking on faith. In docling-serve#366 — still open — a user reports a 4.4 MB PDF on an AWS g5.xlarge (4 vCPU, 16 GB, A10G) driving host memory to about 12 GB during conversion and holding near 11 GB afterwards, on docling-serve-cu128:v1.5.1. That is one report on one release and not something we have reproduced, but it is a concrete reason not to trust a 4 GB limit that looks fine in dev.
Load-test with the largest, ugliest PDFs you actually expect. Dev fixtures will lie to you.
Workers and concurrency
Stay at one Uvicorn worker until you have evidence you need more. Each worker process loads its own copy of the model graph, so two workers means roughly twice the memory for no extra throughput on a single-document bottleneck. Multi-worker also breaks in-memory async task tracking under the default local engine, because each process keeps its own task store — submit to one worker, poll another, get a 404.
The knobs, with defaults from the configuration reference:
| Variable | Default | Effect |
|---|---|---|
UVICORN_WORKERS | 1 | Server processes. Leave at 1 |
DOCLING_NUM_THREADS | 4 | Torch CPU threads inside conversion |
OMP_NUM_THREADS | 4 | OpenMP thread pool |
DOCLING_SERVE_ENG_LOC_NUM_WORKERS | 2 | Concurrent conversions in the local engine |
DOCLING_SERVE_ENG_LOC_SHARE_MODELS | false | Share one model set across those workers |
Set the thread caps at or below the container’s CPU limit; above it, PyTorch oversubscribes and you spend the difference on context switches. And note SHARE_MODELS defaults to false, so the default configuration allocates a separate model set per local worker — turning it on is the single easiest memory saving available.
Production environment variables
The limits that stop a bad client from holding the server hostage:
DOCLING_SERVE_API_KEY: 'replace-me'
DOCLING_SERVE_MAX_FILE_SIZE: '52428800'
DOCLING_SERVE_MAX_NUM_PAGES: '200'
DOCLING_SERVE_MAX_DOCUMENT_TIMEOUT: '1800'
DOCLING_SERVE_MAX_SYNC_WAIT: '120'
DOCLING_SERVE_MAX_SOURCES_PER_REQUEST: '3'
DOCLING_SERVE_SINGLE_USE_RESULTS: 'true'
DOCLING_SERVE_RESULT_REMOVAL_DELAY: '300'
DOCLING_SERVE_CORS_ORIGINS: '["https://app.example.com"]'
DOCLING_SERVE_LOG_LEVEL: 'INFO'
DOCLING_SERVE_LOG_FORMAT: 'json'
MAX_FILE_SIZE and MAX_NUM_PAGES are unbounded by default, and CORS_ORIGINS defaults to ["*"]. Those three are the ones to change before you accept anything untrusted. MAX_SOURCES_PER_REQUEST already defaults to 3, which is low enough to surprise you in the other direction — a fourth file in one multipart request gets a 422.
Every DOCLING_SERVE_* variable, including the many the upstream doc omits, is listed with type and default in our Docling Serve configuration reference.
Use environment variables rather than CLI flags. The app-level CLI surface is only --artifacts-path and --enable-ui, and the configuration reference warns that CLI values do not survive the subprocess spawn under --reload or multiple workers. (v1.32.0 fixed that for those two specific flags by re-exporting them as env vars, which tells you how the settings layer actually wants to be driven.) For anything complex, DOCLING_SERVE_CONFIG_FILE points at a YAML or JSON file; env vars still win over it.
Logs go to stdout/stderr through Python logging and Uvicorn. INFO is the right default and json is worth setting for anything that aggregates logs. DOCLING_SERVE_LOG_HEADER_PREFIX (default X-Docling-Log-) propagates matching request headers into every log line for that request, which is the cheapest request tracing you will get here.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Startup error about a missing model | A bind mount hid the image’s bundled weights | Remove the mount, or populate it with docling-tools models download |
| Enrichment or VLM request errors on a missing model | Those weights are not in the standard image and are not auto-fetched | Bake them into a derived image, or mount a populated directory |
| Cache worked, then went stale after upgrading the tag | A named volume keeps content from whichever image first populated it | Recreate the volume, or drop it and use the bundled weights |
| Model path ignored | Wrong variable: DOCLING_ARTIFACTS_PATH vs DOCLING_SERVE_ARTIFACTS_PATH | Set the one for the layer you mean |
| OOM kill under load that dev never saw | Memory limit sized from a small fixture | Raise the limit, set ENG_LOC_SHARE_MODELS=true, keep UVICORN_WORKERS=1 |
| Memory stays high after a conversion finishes | Reported retention, see docling-serve#366 | Cap concurrency and recycle the container on an RSS threshold while it is open |
| Container flapping between healthy and unhealthy at boot | start_period shorter than real model load | Raise it above your measured cold start |
| First requests after a rollout time out | Traffic admitted on /health instead of /ready | Probe /ready; under RQ, probe the workers too |
/ready instantly OK but conversions never run | RQ engine with no workers running | Start docling-serve rq-worker containers |
nvidia-smi works, Docling stays on CPU | Not a cu* image, or DOCLING_DEVICE unset | Pull a cu* tag and set DOCLING_DEVICE=cuda |
| GPU busy for layout, idle for OCR | RapidOCR’s ONNX backend fell back to the CPU provider | Assert CUDAExecutionProvider, or use a Torch-backed OCR engine |
docker pull gets an unexpected CUDA build | Relying on a floating tag | Pin :v1.32.0 |
Scaling past one container
A single container with the limits above holds up under controlled traffic. Past that, the next step is a queue. DOCLING_SERVE_ENG_KIND selects the async engine: local (default, in-process), rq (Redis Queue), or ray.
RQ is the one to reach for. Set DOCLING_SERVE_ENG_KIND=rq and DOCLING_SERVE_ENG_RQ_REDIS_URL on both the API containers and the workers, then run the workers as separate services from the same image with command: ["docling-serve", "rq-worker"] in Compose (the Kubernetes manifest below splits that into command and args). The API containers stop converting entirely and become thin dispatchers, which means you can scale the two independently and size their memory very differently. The project ships a Kubernetes manifest for this shape. Redis pool sizing: the default 50 connections covers 1–4 workers, 100 for 5–10, 150–200 beyond.
One RQ-specific trap: DOCLING_SERVE_ENG_RQ_JOB_TIMEOUT is baked into each job when it is enqueued, so it must be set on the API containers rather than the workers, and queued jobs keep the value they were enqueued with. Setting it to 0 does not disable it — RQ reads 0 as unset and applies its own 180-second default.
The ray engine also exists, and requires both a Redis URL and a Ray cluster address. It carries a large surface of DOCLING_SERVE_ENG_RAY_* settings for autoscaling and tenant fairness that the configuration reference does not document, so budget time in the source. If you do not already operate Ray, run RQ.
At this point you are not running Docling in Docker so much as running a small distributed system around Docling. That is fine, and the official engines do most of the work. It is also where most teams decide whether they want to keep operating it.
If you would rather not own a model cache, GPU drivers, queue workers, readiness probes and autoscaling, Parsebridge runs Docling as a hosted API — same parser, no Compose file. It is our own product, and the honest scope is PDF to Markdown through our API rather than a drop-in Docling Serve endpoint; the Serve guide shows the request shape.
Image variants, bundled models, defaults and endpoint behaviour above were read from the docling-serve v1.32.0 source and Containerfile. The Compose file is source-reviewed against that release rather than runtime-tested here; memory figures are attributed to the upstream report they come from and were not reproduced by us.