← Back to Home

Docker Model Runner Architecture Breakdown

Docker Model RunnerDocker DesktopAI

Docker moved local model serving into the container worldview. Ollama is a separate process you install, start, and connect to over its own port. Docker Model Runner (DMR) takes a different route: models become OCI artifacts managed through docker model commands, and the inference surface is exposed as an OpenAI-compatible endpoint on port 12434. That shift looks like a command-line change, but it redraws the boundary of what "running a model locally" means. This walkthrough is built from Docker's official documentation, including the differences between the three engines and the limitations Docker itself documents.

TL;DR:

-The default engine, llama.cpp, is enough for most people: it reads GGUF, runs on all platforms (macOS/Windows/Linux), and is the only engine that supports CPU-only inference. It works with no extra configuration.

-vLLM only matters if you have an NVIDIA GPU: Linux x86_64 only (Windows needs WSL2 plus Docker Desktop 4.54+), Safetensors format only, and Docker's docs state plainly that it does not support CPU-only inference.

-Diffusers is image generation only: it uses the DDUF format and handles Stable Diffusion-style models. It is not a text-generation engine and also has no CPU inference.

-"Connection refused" from a container is usually not a bug: containers must use http://model-runner.docker.internal, not localhost. Host processes use localhost:12434. These are two different addressing conventions.

-Deployment warning: DMR's API has no authentication by default, and Docker's docs say it ignores the Authorization header. Do not expose port 12434 to an untrusted network.

1. What DMR Actually Changes: From Process to Artifact

The fastest way to understand DMR is to look at what it swaps out.

With Ollama, the model is a standalone service. You run ollama run, it listens on its own port, and models live in its own directory. Your application reaches it over the network — and if that application itself runs in a container, you have to configure container-to-host networking.

DMR works differently. Docker's documentation describes models as packaged as OCI artifacts, pulled and run via docker model, with an OpenAI-compatible interface exposed outward. The practical consequence: teams already standardized on Docker no longer operate a separate service.

Three commands cover the basic workflow:

docker model pull ai/llama3.2:3B-Q4_K_M   # pull a model
docker model run ai/smollm2                # run it
docker model status                        # inspect engine state

docker model status reports each engine individually, so you can see at a glance what is installed:

Docker Model Runner is running
Status: llama.cpp: running
        llama.cpp version: 34ce48d
        mlx: not installed
        sglang: sglang package not installed
        vllm: ...

For a non-default backend, Docker provides install-runner:

docker model install-runner --backend vllm --gpu cuda

--backend accepts llama.cpp, vllm, or diffusers; --gpu accepts cuda, rocm, vulkan, or metal depending on platform. There is also a command for pushing models to a registry, which matters if your team uses OCI-based distribution:

docker model package --gguf ./model.gguf --push myorg/mymodel:Q4_K_M

2. The Three Engines, Side by Side

This is Docker's own comparison table. I have translated and annotated it but changed no values:

Featurellama.cppvLLMDiffusers
Model formatsGGUFSafetensors, HuggingFaceDDUF
PlatformsAll (macOS, Windows, Linux)Linux x86_64 onlyLinux (x86_64, ARM64)
GPU supportNVIDIA, AMD, Apple Silicon, VulkanNVIDIA CUDA onlyNVIDIA CUDA only
CPU inferenceYesNoNo
QuantizationBuilt-in (Q4, Q5, Q8, etc.)LimitedLimited
Memory efficiencyHigh (with quantization)ModerateModerate
ThroughputGoodHigh (with batching)Good
Best forLocal dev, constrained environmentsProduction, high throughputImage generation

Three rows deserve individual attention.The CPU inference row is the dividing line.Docker's exact wording for vLLM is: "vLLM requires an NVIDIA GPU with CUDA support. It does not support CPU-only inference." If your machine has no NVIDIA card, installing vLLM is wasted effort. On Apple Silicon, Docker marks macOS as "Not supported" for vLLM outright.Model format determines which engine you can use.Most quantized models in the community are GGUF, which is llama.cpp territory. Original weights on HuggingFace are usually Safetensors, which is vLLM territory. Docker also notes that Safetensors models typically use more memory than quantized GGUF models.Diffusers sits on a different line entirely.It is not part of text generation — Docker classifies it as image generation for Stable Diffusion workloads.

vLLM's platform limits are strict

Docker's vLLM platform support table is explicit:

PlatformGPUStatus
Linux x86_64NVIDIA CUDASupported
Windows (WSL2)NVIDIA CUDASupported (Docker Desktop 4.54+)
macOS—Not supported
Linux ARM64—Not supported
AMD GPUs—Not supported

vLLM model tags also tend to carry a -vllm suffix:

docker model run ai/smollm2-vllm

3. Why Containers Cannot Reach It: Two Addressing Conventions

This is the most common stumble, and Docker's documentation dedicates a table to it.

RuntimeAccess fromBase URL
Docker DesktopContainershttp://model-runner.docker.internal
Docker DesktopHost processes (TCP)http://localhost:12434
Docker EngineContainershttp://172.17.0.1:12434

| Docker Engine | Host processes | http://localhost:12434 |Using localhost inside a container is wrong.A container has its own network namespace, so localhost points at the container itself, not the host. This is not a misconfiguration — it falls out of how the network model works.

Docker also provides a fallback. If 172.17.0.1 is unavailable in your Compose project (the docs note this interface may not be available by default to containers within a Compose project), add extra_hosts:

extra_hosts:
  - "model-runner.docker.internal:host-gateway"

Base URLs differ per API flavor

The same DMR service needs different base URLs for different clients:

Client typeBase URL
OpenAI SDK / clientshttp://localhost:12434/engines/v1
Anthropic SDK / clientshttp://localhost:12434
Ollama-compatible clientshttp://localhost:12434

Docker's OpenAI-compatible endpoint list:

EndpointMethodDescription
/engines/v1/modelsGETList models
/engines/v1/models/{namespace}/{name}GETRetrieve model
/engines/v1/chat/completionsPOSTCreate chat completion
/engines/v1/completionsPOSTCreate completion
/engines/v1/embeddingsPOSTCreate embeddings

Note the /engines/ prefix.An OpenAI-compatible client cannot use http://localhost:12434/v1— dropping /engines returns a 404. This is the most frequent failure when porting an Ollama config to DMR.

You can also include the engine name in the path, for example /engines/llama.cpp/v1/chat/completions, which is useful when several engines run side by side.

The Anthropic-compatible endpoints are:

EndpointMethodDescription
/anthropic/v1/messagesPOSTCreate a message
/anthropic/v1/messages/count_tokensPOSTCount tokens

One detail in Docker's example is easy to miss:

curl http://localhost:12434/v1/messages \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ai/smollm2",
    "max_tokens": 1024,
    "messages": [{"role": "user", "content": "..."}]
  }'

The path is /v1/messages, not /anthropic/v1/messages. The example and the endpoint table differ on this point, so treat your client configuration as the source of truth. Model names must use the full identifier including the namespace, such as ai/smollm2.

4. Limitations Docker Documents — The Most Important Section

DMR's API reference has a section titled "Limitations and differences from OpenAI." The original table:

FeatureDMR behavior
API keyNot required. DMR ignores the Authorization header.
Function callingSupported with llama.cpp for compatible models
VisionSupported for multi-modal models (e.g., LLaVA)
JSON modeSupported via response_format: {"type": "json_object"}
LogprobsSupported
Token countingUses the model's native token encoder,which may differ from OpenAI's

Two of these directly affect migration work.The API key row is a deployment red line.Docker's documentation is unambiguous: DMR does not require an API key and ignores any Authorization header you send. That means any client able to reach port 12434 can submit inference requests, and may be able to pull, load, or run models. So: do not expose that port on an untrusted LAN, a shared CI network, or a publicly routable developer host.Function calling only works on llama.cpp.If your application depends on function calling, you cannot switch to the vLLM backend. This is a hard constraint that is easy to miss during engine selection.Token counting may differ from OpenAI's, so leave headroom if you budget context or meter by tokens.

Troubleshooting: What Docker Lists as Common Failures

DMR's documentation has a Troubleshooting section whose headings are the symptoms themselves.vLLM won't start.Docker's first suggested check is confirming an NVIDIA GPU is available. Combined with the platform table above, this is usually decisive: AMD GPUs, macOS, and Linux ARM64 are unsupported, and no configuration change will fix that.llama.cpp is slow.Docker groups this with out-of-memory errors, which suggests the two often appear together in practice.

For diagnosis, Docker provides:

docker model logs

Check this before guessing.

5. DMR or Ollama: How to Choose

Cross-checking against multiple independent technical sources, the conclusion is that for single-machine local development the two are close, and the real difference is engineering governance rather than raw performance.

Peak memory use is comparable because both run llama.cpp underneath. The meaningful differences are elsewhere:

-Model distribution: DMR uses OCI artifacts and registries; Ollama uses its own model library and Modelfiles. If you need registry-style versioning and audit control, DMR fits better.

-Idle memory: Ollama keeps models resident by default (OLLAMA_KEEP_ALIVE, 5-minute default), giving faster first response at the cost of idle memory. DMR unloads models when idle.

-Authentication: neither should have its port exposed directly. DMR explicitly ignores the Authorization header; Ollama likewise needs network isolation handled deliberately.

-"OpenAI-compatible" does not mean operationally identical: DMR uses engine-aware paths such as /engines/llama.cpp/v1/..., while Ollama's native chat API is /api/chat. Both fit existing clients, but request fields, streaming chunks, error schemas, and usage data should all be tested before switching providers.

One more point worth keeping in mind: for multi-user or production-grade high-concurrency serving, neither tool should be your answer as-is. Independent comparison write-ups recommend moving to a separately deployed vLLM, SGLang, or a Kubernetes serving stack. A laptop benchmark answers developer-experience questions; it does not establish production throughput, tail latency, or multi-tenant isolation.

FAQ

Q1: Which Docker Desktop version do I need?

Docker's docs specify Docker Desktop 4.54+ for the vLLM-on-WSL2 path. DMR overall evolves with Docker Desktop, and the engines and flags available differ by release, so check the docs matching your installed version.Q2: Can I expose DMR to my team?

Not directly. The docs state it has no built-in authentication and ignores the Authorization header. At minimum, scope it at the network level; put an authenticating reverse proxy in front if shared access is unavoidable.Q3: How do I declare a model dependency in Compose?

DMR's Compose integration lets you declare models alongside a service's other dependencies and have Compose inject the model endpoint and identifier into the application container, rather than hardcoding it. Check the docs for your Compose version for the exact fields and version requirements.Q4: What happens if I use the wrong model format?

The pull or the run will fail. llama.cpp takes GGUF, vLLM takes Safetensors, Diffusers takes DDUF. The typical symptom of a wrong format is that the engine does not recognize the artifact.Q5: Does this work with Docker Engine, not just Desktop?

Yes, but the addressing differs: containers use http://172.17.0.1:12434 and host processes use http://localhost:12434. Docker also warns that this interface may not be available by default inside a Compose project, in which case you need extra_hosts.

All endpoint paths, engine capabilities, and limitations in this article come from Docker's official documentation (as of October 11, 2026):

A note on scope: this article deliberately omits throughput and tokens-per-second figures. Those appear in third-party comparisons, but their test environments (hardware, model, quantization level) differ too much from any reader's setup and cannot be traced back to a reproducible configuration, so quoting them as conclusions would be misleading.

This article contains no affiliate links. All endpoint paths, engine capabilities, and limitations are taken from Docker official documentation; no unverified performance figures are quoted.

👉 Join Xiaomi MiMo Platform: Leading AI model platform with cost-effective inference

👉 Join Aliyun AI: Top AI products with exclusive coupons for business innovation

📌 This article was AI-assisted generated and human-reviewed | TechPassive — An AI-driven content testing site focused on real tool reviews

🔗 Recommended Tools

These are carefully selected tools. Using our affiliate links supports us to keep producing quality content:

☁️ DigitalOcean Cloud ⚡ Vultr VPS 🤖 QoderWork CN (Refer & Earn) ☁️ Aliyun AI Products 📚 WordPress Books 🔍 WordPress SEO Books 🌐 Web Hosting Books 🐳 Docker Books 🐧 Linux Books 🐍 Python Books 💰 Affiliate Marketing 💵 Passive Income Books 🖥️ Server Books ☁️ Cloud Computing Books 🚀 DevOps Books 🤖 Xiaomi MiMo Platform
← Back to Home