Docker Model Runner Architecture Breakdown
Docker moved local model serving into the container worldview. Ollama is a separate process you install, start, and connect to over its own port. Docker Model Runner (DMR) takes a different route: models become OCI artifacts managed through docker model commands, and the inference surface is exposed as an OpenAI-compatible endpoint on port 12434. That shift looks like a command-line change, but it redraws the boundary of what "running a model locally" means. This walkthrough is built from Docker's official documentation, including the differences between the three engines and the limitations Docker itself documents.
TL;DR:
-The default engine, llama.cpp, is enough for most people: it reads GGUF, runs on all platforms (macOS/Windows/Linux), and is the only engine that supports CPU-only inference. It works with no extra configuration.
-vLLM only matters if you have an NVIDIA GPU: Linux x86_64 only (Windows needs WSL2 plus Docker Desktop 4.54+), Safetensors format only, and Docker's docs state plainly that it does not support CPU-only inference.
-Diffusers is image generation only: it uses the DDUF format and handles Stable Diffusion-style models. It is not a text-generation engine and also has no CPU inference.
-"Connection refused" from a container is usually not a bug: containers must use http://model-runner.docker.internal, not localhost. Host processes use localhost:12434. These are two different addressing conventions.
-Deployment warning: DMR's API has no authentication by default, and Docker's docs say it ignores the Authorization header. Do not expose port 12434 to an untrusted network.
1. What DMR Actually Changes: From Process to Artifact
The fastest way to understand DMR is to look at what it swaps out.
With Ollama, the model is a standalone service. You run ollama run, it listens on its own port, and models live in its own directory. Your application reaches it over the network — and if that application itself runs in a container, you have to configure container-to-host networking.
DMR works differently. Docker's documentation describes models as packaged as OCI artifacts, pulled and run via docker model, with an OpenAI-compatible interface exposed outward. The practical consequence: teams already standardized on Docker no longer operate a separate service.
Three commands cover the basic workflow:
docker model pull ai/llama3.2:3B-Q4_K_M # pull a model
docker model run ai/smollm2 # run it
docker model status # inspect engine state
docker model status reports each engine individually, so you can see at a glance what is installed:
Docker Model Runner is running
Status: llama.cpp: running
llama.cpp version: 34ce48d
mlx: not installed
sglang: sglang package not installed
vllm: ...
For a non-default backend, Docker provides install-runner:
docker model install-runner --backend vllm --gpu cuda
--backend accepts llama.cpp, vllm, or diffusers; --gpu accepts cuda, rocm, vulkan, or metal depending on platform. There is also a command for pushing models to a registry, which matters if your team uses OCI-based distribution:
docker model package --gguf ./model.gguf --push myorg/mymodel:Q4_K_M
2. The Three Engines, Side by Side
This is Docker's own comparison table. I have translated and annotated it but changed no values:
| Feature | llama.cpp | vLLM | Diffusers |
|---|---|---|---|
| Model formats | GGUF | Safetensors, HuggingFace | DDUF |
| Platforms | All (macOS, Windows, Linux) | Linux x86_64 only | Linux (x86_64, ARM64) |
| GPU support | NVIDIA, AMD, Apple Silicon, Vulkan | NVIDIA CUDA only | NVIDIA CUDA only |
| CPU inference | Yes | No | No |
| Quantization | Built-in (Q4, Q5, Q8, etc.) | Limited | Limited |
| Memory efficiency | High (with quantization) | Moderate | Moderate |
| Throughput | Good | High (with batching) | Good |
| Best for | Local dev, constrained environments | Production, high throughput | Image generation |
Three rows deserve individual attention.The CPU inference row is the dividing line.Docker's exact wording for vLLM is: "vLLM requires an NVIDIA GPU with CUDA support. It does not support CPU-only inference." If your machine has no NVIDIA card, installing vLLM is wasted effort. On Apple Silicon, Docker marks macOS as "Not supported" for vLLM outright.Model format determines which engine you can use.Most quantized models in the community are GGUF, which is llama.cpp territory. Original weights on HuggingFace are usually Safetensors, which is vLLM territory. Docker also notes that Safetensors models typically use more memory than quantized GGUF models.Diffusers sits on a different line entirely.It is not part of text generation — Docker classifies it as image generation for Stable Diffusion workloads.
vLLM's platform limits are strict
Docker's vLLM platform support table is explicit:
| Platform | GPU | Status |
|---|---|---|
| Linux x86_64 | NVIDIA CUDA | Supported |
| Windows (WSL2) | NVIDIA CUDA | Supported (Docker Desktop 4.54+) |
| macOS | — | Not supported |
| Linux ARM64 | — | Not supported |
| AMD GPUs | — | Not supported |
vLLM model tags also tend to carry a -vllm suffix:
docker model run ai/smollm2-vllm
3. Why Containers Cannot Reach It: Two Addressing Conventions
This is the most common stumble, and Docker's documentation dedicates a table to it.
| Runtime | Access from | Base URL |
|---|---|---|
| Docker Desktop | Containers | http://model-runner.docker.internal |
| Docker Desktop | Host processes (TCP) | http://localhost:12434 |
| Docker Engine | Containers | http://172.17.0.1:12434 |
| Docker Engine | Host processes | http://localhost:12434 |Using localhost inside a container is wrong.A container has its own network namespace, so localhost points at the container itself, not the host. This is not a misconfiguration — it falls out of how the network model works.
Docker also provides a fallback. If 172.17.0.1 is unavailable in your Compose project (the docs note this interface may not be available by default to containers within a Compose project), add extra_hosts:
extra_hosts:
- "model-runner.docker.internal:host-gateway"
Base URLs differ per API flavor
The same DMR service needs different base URLs for different clients:
| Client type | Base URL |
|---|---|
| OpenAI SDK / clients | http://localhost:12434/engines/v1 |
| Anthropic SDK / clients | http://localhost:12434 |
| Ollama-compatible clients | http://localhost:12434 |
Docker's OpenAI-compatible endpoint list:
| Endpoint | Method | Description |
|---|---|---|
/engines/v1/models | GET | List models |
/engines/v1/models/{namespace}/{name} | GET | Retrieve model |
/engines/v1/chat/completions | POST | Create chat completion |
/engines/v1/completions | POST | Create completion |
/engines/v1/embeddings | POST | Create embeddings |
Note the /engines/ prefix.An OpenAI-compatible client cannot use http://localhost:12434/v1— dropping /engines returns a 404. This is the most frequent failure when porting an Ollama config to DMR.
You can also include the engine name in the path, for example /engines/llama.cpp/v1/chat/completions, which is useful when several engines run side by side.
The Anthropic-compatible endpoints are:
| Endpoint | Method | Description |
|---|---|---|
/anthropic/v1/messages | POST | Create a message |
/anthropic/v1/messages/count_tokens | POST | Count tokens |
One detail in Docker's example is easy to miss:
curl http://localhost:12434/v1/messages \
-H "Content-Type: application/json" \
-d '{
"model": "ai/smollm2",
"max_tokens": 1024,
"messages": [{"role": "user", "content": "..."}]
}'
The path is /v1/messages, not /anthropic/v1/messages. The example and the endpoint table differ on this point, so treat your client configuration as the source of truth. Model names must use the full identifier including the namespace, such as ai/smollm2.
4. Limitations Docker Documents — The Most Important Section
DMR's API reference has a section titled "Limitations and differences from OpenAI." The original table:
| Feature | DMR behavior |
|---|---|
| API key | Not required. DMR ignores the Authorization header. |
| Function calling | Supported with llama.cpp for compatible models |
| Vision | Supported for multi-modal models (e.g., LLaVA) |
| JSON mode | Supported via response_format: {"type": "json_object"} |
| Logprobs | Supported |
| Token counting | Uses the model's native token encoder,which may differ from OpenAI's |
Two of these directly affect migration work.The API key row is a deployment red line.Docker's documentation is unambiguous: DMR does not require an API key and ignores any Authorization header you send. That means any client able to reach port 12434 can submit inference requests, and may be able to pull, load, or run models. So: do not expose that port on an untrusted LAN, a shared CI network, or a publicly routable developer host.Function calling only works on llama.cpp.If your application depends on function calling, you cannot switch to the vLLM backend. This is a hard constraint that is easy to miss during engine selection.Token counting may differ from OpenAI's, so leave headroom if you budget context or meter by tokens.
Troubleshooting: What Docker Lists as Common Failures
DMR's documentation has a Troubleshooting section whose headings are the symptoms themselves.vLLM won't start.Docker's first suggested check is confirming an NVIDIA GPU is available. Combined with the platform table above, this is usually decisive: AMD GPUs, macOS, and Linux ARM64 are unsupported, and no configuration change will fix that.llama.cpp is slow.Docker groups this with out-of-memory errors, which suggests the two often appear together in practice.
For diagnosis, Docker provides:
docker model logs
Check this before guessing.
5. DMR or Ollama: How to Choose
Cross-checking against multiple independent technical sources, the conclusion is that for single-machine local development the two are close, and the real difference is engineering governance rather than raw performance.
Peak memory use is comparable because both run llama.cpp underneath. The meaningful differences are elsewhere:
-Model distribution: DMR uses OCI artifacts and registries; Ollama uses its own model library and Modelfiles. If you need registry-style versioning and audit control, DMR fits better.
-Idle memory: Ollama keeps models resident by default (OLLAMA_KEEP_ALIVE, 5-minute default), giving faster first response at the cost of idle memory. DMR unloads models when idle.
-Authentication: neither should have its port exposed directly. DMR explicitly ignores the Authorization header; Ollama likewise needs network isolation handled deliberately.
-"OpenAI-compatible" does not mean operationally identical: DMR uses engine-aware paths such as /engines/llama.cpp/v1/..., while Ollama's native chat API is /api/chat. Both fit existing clients, but request fields, streaming chunks, error schemas, and usage data should all be tested before switching providers.
One more point worth keeping in mind: for multi-user or production-grade high-concurrency serving, neither tool should be your answer as-is. Independent comparison write-ups recommend moving to a separately deployed vLLM, SGLang, or a Kubernetes serving stack. A laptop benchmark answers developer-experience questions; it does not establish production throughput, tail latency, or multi-tenant isolation.
FAQ
Q1: Which Docker Desktop version do I need?
Docker's docs specify Docker Desktop 4.54+ for the vLLM-on-WSL2 path. DMR overall evolves with Docker Desktop, and the engines and flags available differ by release, so check the docs matching your installed version.Q2: Can I expose DMR to my team?
Not directly. The docs state it has no built-in authentication and ignores the Authorization header. At minimum, scope it at the network level; put an authenticating reverse proxy in front if shared access is unavoidable.Q3: How do I declare a model dependency in Compose?
DMR's Compose integration lets you declare models alongside a service's other dependencies and have Compose inject the model endpoint and identifier into the application container, rather than hardcoding it. Check the docs for your Compose version for the exact fields and version requirements.Q4: What happens if I use the wrong model format?
The pull or the run will fail. llama.cpp takes GGUF, vLLM takes Safetensors, Diffusers takes DDUF. The typical symptom of a wrong format is that the engine does not recognize the artifact.Q5: Does this work with Docker Engine, not just Desktop?
Yes, but the addressing differs: containers use http://172.17.0.1:12434 and host processes use http://localhost:12434. Docker also warns that this interface may not be available by default inside a Compose project, in which case you need extra_hosts.
All endpoint paths, engine capabilities, and limitations in this article come from Docker's official documentation (as of October 11, 2026):
- Model Runner overview: https://docs.docker.com/ai/model-runner/
- Inference engines and comparison table: https://docs.docker.com/ai/model-runner/inference-engines/
- API reference and limitations: https://docs.docker.com/ai/model-runner/api-reference/
- Compose integration: https://docs.docker.com/compose/bridge/use-model-runner/
A note on scope: this article deliberately omits throughput and tokens-per-second figures. Those appear in third-party comparisons, but their test environments (hardware, model, quantization level) differ too much from any reader's setup and cannot be traced back to a reproducible configuration, so quoting them as conclusions would be misleading.
This article contains no affiliate links. All endpoint paths, engine capabilities, and limitations are taken from Docker official documentation; no unverified performance figures are quoted.
👉 Join Xiaomi MiMo Platform: Leading AI model platform with cost-effective inference
👉 Join Aliyun AI: Top AI products with exclusive coupons for business innovation
📌 This article was AI-assisted generated and human-reviewed | TechPassive — An AI-driven content testing site focused on real tool reviews
🔗 Recommended Tools
These are carefully selected tools. Using our affiliate links supports us to keep producing quality content: