n8n Task Runner Architecture and External Mode
TL;DR — n8n's Code node does not run inside the main process. Since 2.x the official docs split it into a separate Task Runner, scheduled by a component called the task broker. I tested two self-hosted setups and confirmed three things. First, the task runner is the only isolation layer between user-provided code and your n8n database, encryption key, and stored credentials — that is the docs' own wording, not my paraphrase. Second, internal mode is deprecated as of n8n 3.0, and instances still running it print a deprecation warning at startup. Third, in queue mode every worker needs its own runner sidecar; pointing workers at the main instance's broker is the most common way this breaks. Below you get a copy-pasteable compose file, five real error messages taken from official docs and a GitHub issue with their root causes, and the hardware math for self-hosting.
Why the Code node deserves its own process
The Code node lets users write arbitrary JavaScript or Python. If that code ran inside the main process, anyone who can edit a workflow could potentially read your database, your encryption key, your stored credentials, and your environment variables. The docs put this in a warning box at the top of the task runner page:
Without them, or with internal mode, anyone who can edit a workflow could potentially read your database, encryption key, stored credentials, and environment variables.
That is why n8n keeps insisting on task runners in production since 2.x. It is not an optional performance tweak — it is a security boundary.
The three-layer architecture
The official architecture description names three roles, and all three are required:
- Task runner — the process that actually executes code, connected to the broker over a websocket connection. Both JS and Python execution happen here.
- Task broker — the coordinator. The n8n instance itself (main and worker) acts as the broker, listening on port
5679by default. - Task requester — whoever submits the task. In n8n's case, the Code node is the task requester.
The flow: the Code node submits a task request to the broker, the broker broadcasts it, an idle runner picks it up over the websocket, executes it inside its sandbox, and returns the result to the requester. The broker only coordinates — it never runs your code.
Two modes and three defaults you must change
Internal mode (deprecated, do not use)
In internal mode n8n launches the runner as a child process that shares the same uid and gid as the main process. The docs flag two things: it is deprecated from n8n 3.0 and n8n logs a deprecation warning at startup even when N8N_RUNNERS_MODE is unset. The docs also state it is insecure by design — code that escapes the runner sandbox has the same access as n8n itself, including stored credentials. Use it only on isolated instances holding nothing but mock data.
External mode (the only production choice)
In external mode a launcher starts runners on demand and manages their lifecycle. In practice this means adding a sidecar container running the n8nio/runners image next to n8n. Three hard constraints:
1. Versions must match exactly — the n8nio/runners version must equal the n8nio/n8n version.
2. n8n must be ≥ 1.111.0 — external mode is unavailable below that.
3. The broker listens on localhost by default — in Docker Compose you must set N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0, or the runner cannot reach it.
The official compose reference, aligned to those constraints:
services:
n8n:
image: n8nio/n8n:1.111.0
environment:
- N8N_RUNNERS_MODE=external
- N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0
- N8N_RUNNERS_AUTH_TOKEN=your-secret-here
- N8N_NATIVE_PYTHON_RUNNER=true
ports:
- "5678:5678"
task-runners:
image: n8nio/runners:1.111.0
environment:
- N8N_RUNNERS_TASK_BROKER_URI=http://n8n-main:5679
- N8N_RUNNERS_AUTH_TOKEN=your-secret-here
depends_on:
- n8n
⚠️ Two official-docs inconsistencies worth knowing. The sample above still carries N8N_RUNNERS_ENABLED=true, yet the environment variable reference marks that variable deprecated from n8n 2.0 (it is only required on 1.x). And the task runner environment variable page refers to the config file as /etc/n8n-task-runners.json in one place and /etc/task-runners.json in another. Trust whatever is actually inside the image for your version — after upgrading, run docker exec to confirm.
Three defaults worth tuning
| Variable | Default | Meaning |
|---|---|---|
N8N_RUNNERS_MAX_CONCURRENCY | 5 | Tasks a single runner executes at once |
N8N_RUNNERS_TASK_TIMEOUT | 300 | Max seconds per task before the runner stops and restarts it |
N8N_RUNNERS_AUTO_SHUTDOWN_TIMEOUT | 15 | Idle seconds before shutdown; the launcher restarts on new work |
A few security-relevant defaults are also worth memorizing: N8N_RUNNERS_MAX_PAYLOAD defaults to 1 073 741 824 bytes (about 1 GB) as the ceiling for broker-to-runner payloads; N8N_RUNNERS_INSECURE_MODE defaults to false and is discouraged in production; N8N_RUNNERS_HEARTBEAT_INTERVAL defaults to 30 seconds, after which a silent runner is treated as dead and restarted.
On queue mode performance, the docs give a baseline of up to 220 workflow executions per second on a single instance, with further scaling by adding instances. The docs' single-instance benchmark ran on an ECS c5a.large with 4GB RAM; the multi-instance comparison used seven 8GB instances arranged as two webhook processors, four workers, one database, and one main running n8n plus Redis. The docs also state that all execution data between workers and main flows through Redis and Postgres, and workers are stateless, so adding replicas is safe.
Troubleshooting: five real errors and their root causes
Every item below is traceable to official documentation, a GitHub issue, or the n8n community forum. None of them are invented "typical" errors.
Error one: Task request timed out after 60 seconds
This is the most common external-runner error. It means the task waited 60 seconds without an idle runner becoming available. Note that the 60 seconds comes from the default of N8N_RUNNERS_TASK_REQUEST_TIMEOUT, not from N8N_RUNNERS_TASK_TIMEOUT (which defaults to 300).
Two root causes appear in real community cases: the worker has no runner sidecar of its own, and N8N_RUNNERS_AUTH_TOKEN differs between main and runner. The latter shows up in logs as Task runner connection attempt failed with status code 403.
Fix — give every worker its own sidecar; confirm both sides use an identical token (while debugging, switch to a simple value like test123 to rule out special-character parsing issues); confirm the broker listen address is 0.0.0.0. Only raise the timeout when the task is genuinely compute-heavy rather than when the runner never connected.
Error two: getaddrinfo ENOTFOUND while the runner waits for the broker
The signature is the runner container repeatedly printing Waiting for task broker to be ready... and Waiting for launcher's task offer to be accepted..., alongside Task runner failed to start { error: Error: getaddrinfo ENOTFOUND ... }, and never receiving a task.
Cause — the hostname in N8N_RUNNERS_TASK_BROKER_URI is wrong. One documented community case had a compose service named n8n while the URI pointed at http://n8n-main:5679.
Fix — the hostname must match the compose service name exactly. Verify from inside the runner container with getent hosts ; if it resolves to an IP, connectivity is fine.
Error three: JS Code nodes hang while Python nodes work
This one is counterintuitive enough to deserve its own section. In GitHub issue #22798 the main-side log shows Task (xxx) deferred until runner is ready immediately followed by Deregistered runner, repeating every 2 to 3 seconds, while the external launcher-javascript registers successfully but never receives an offer. Under the identical configuration, a Python Code node running return 1 + 1 executed fine.
The issue was eventually closed as closed:cant-reproduce, and the reporter then identified the real cause: their own extended runner image, a custom Dockerfile that added extra JS and Python dependencies. A maintainer had noted earlier that this step should be a non-issue, since the environment variables matched on both sides.
Fix — when JavaScript specifically fails on an external runner, fall back to the stock n8nio/runners image to establish a working baseline, then add custom dependencies back one at a time. The docs give a related constraint: extending the image requires at least n8nio/runners:1.121.0.
Error four: Python imports rejected, or out-of-memory
Two separate problems that often get conflated. The docs are explicit that all Python standard library and third-party imports are disabled by default, and that in external mode the allowlists must be set as env-overrides inside the runner image's /etc/n8n-task-runners.json, not passed as container environment variables:
{
"task-runners": [
{
"runner-type": "python",
"env-overrides": {
"PYTHONPATH": "/opt/runners/task-runner-python",
"N8N_RUNNERS_STDLIB_ALLOW": "json",
"N8N_RUNNERS_EXTERNAL_ALLOW": "numpy,pandas"
}
}
]
}
The official out-of-memory messages come in three shapes: Execution stopped at this node (n8n may have run out of memory while executing it), Allocation failed - JavaScript heap out of memory in server logs, and Problem running workflow, Connection Lost, or 503 Service Temporarily Unavailable when the instance has become unavailable. The docs name the Code node and the older Function node as memory hogs, and also flag manual executions, because n8n copies data for the frontend.
Fix — adjust N8N_RUNNERS_MAX_OLD_SPACE_SIZE (it maps to Node's --max-old-space-size); split large batches as the docs suggest (process 200 rows per execution instead of 10,000); note that N8N_RUNNERS_ID requires n8n 2.35.0 or newer, and that two runners sharing one ID keep evicting each other.
Error five: the worker starts but never does anything
The community consensus here is direct: workers that cannot decrypt credentials are suffering from an N8N_ENCRYPTION_KEY that differs between main and worker. And because manual executions run on main by default, you get the classic symptom of a workflow that works in production but fails when you test it — unless OFFLOAD_MANUAL_EXECUTIONS_TO_WORKERS=true is set. Community reports also flag an environmental factor: a worker configured with 0.5 CPU and 1GB RAM is too small, and while Postgres repeatedly logs Database connection timed out, the grant token exchange cannot complete in time. Adding resources resolved it immediately.
Fix — confirm the key is byte-identical; enable offload if needed; give workers at least 1.0 CPU and 2048MB of memory, and check the database container's resources separately.
Hardware and cost math for self-hosting
For n8n plus Postgres plus Redis, official and community guidance agree closely: self-hosting needs roughly 2 cores, 2GB RAM, and 20GB SSD as a minimum (development and testing only); production wants 4+ cores, 8-16GB of RAM, and PostgreSQL. That last point drives every hardware decision that follows.
Take a Beelink EQ14 N150 mini PC (16GB DDR4, 500GB NVMe, dual 2.5GbE) as the example. At the time of writing its listing specifies an Intel Twin Lake N150 (4 cores, 4 threads, up to 3.6GHz), a single 16GB DDR4 3200MHz module, a 500GB NVMe SSD, and dual 2.5GbE NICs, priced at $199.00 with a 4.4-star rating. The dual 2.5GbE suits splitting main and worker onto separate networks, but note the single memory slot caps out at 16GB and runs single-channel — reaching 32GB means replacing the machine, and since 8-16GB is only a starting point for production, memory gets tight fast once queue mode and workers are added.
The second thing people overlook is power. The docs' memory page lists the ways an instance becomes unavailable, and a workflow that dies mid-run leaves executions stuck in a running state. Community advice is to enable Redis persistence and set N8N_GRACEFUL_SHUTDOWN_TIMEOUT=30. A CyberPower CP1500AVRLCD3 UPS (1500VA/900W, 12 outlets, AVR) is a 1500VA/900W unit listed at $199.95 at the time of writing, with a rated runtime of about 3 minutes at full load and about 12 minutes at half load. That is short, but it is enough for the job that matters: giving Postgres time to finish writing and close cleanly during an outage.
One model-number trap to flag: if you encounter the older CP1500AVRLCD, Amazon's own page states it has been replaced by the CP1500AVRLCD3.
The cost comparison is straightforward. A mini PC at the EQ14 level ($199.00) plus a UPS ($199.95) is a one-time outlay, while n8n Cloud starts around $24 per month and prices by execution. Community template pages put self-hosted queue mode on usage-billed platforms like Railway at roughly $3-5 per month for a typical project. Automations that run hard pay back the hardware within a year; a workflow that fires a few times a week is cheaper and simpler on Cloud. That tradeoff depends on your execution volume, not on which option is technically better.
Hardening checklist: four layers of defense
The official hardening page does not say much, but what it says is specific:
1. Use the distroless image — append -distroless to the Docker tag (the docs use 2.4.6-distroless). It contains only the application and its runtime dependencies, with no package manager or shell.
2. Run as the nobody user — uid and gid 65532.
3. Configure a read-only root filesystem — but mount a minimal emptyDir volume at /tmp, since runners still need temporary space.
4. Apply an AppArmor profile — block reads of sensitive /proc files. With this rule, any attempt to read environ or mounts is denied and logged:
audit deny @{PROC}/[0-9]*/{environ,mounts} rwl,
Environment variable defaults add one more layer: Python code cannot read the runner's environment by default (N8N_BLOCK_RUNNER_ENV_ACCESS defaults to true), and Node built-in module imports are disabled by default (NODE_FUNCTION_ALLOW_BUILTIN is empty). When you need to open something up, add it under env-overrides in the config file rather than as a container environment variable.
How this differs from existing n8n articles here
We have already covered migrating from SQLite to PostgreSQL (5 problems swapping SQLite for Postgres), Redis 502s and timeout debugging (n8n Langfuse 502 and Redis in practice), and early self-hosting Docker deployment pitfalls (5 real n8n self-hosted Docker problems). Those articles answer "how do I get it running at all." This one is only about code execution isolation and horizontal scaling once it runs — they are sequential in a debugging workflow, not overlapping in scope.
When you actually need queue mode
Official and community guidance agree, and it is worth stating plainly to avoid over-engineering:
Signs you need queue mode — executions start queueing and finish minutes after they trigger; one long workflow (a large AI batch, a slow API loop) blocks everything else; webhooks time out under burst traffic; you cannot tolerate the main process being a single point of failure.
Signs you do not — a handful of scheduled workflows per day. The community framing is blunt: adding Redis and Postgres at that point buys operational overhead you will not use. Queue mode turns one process into at least three or four (main, worker, Redis, Postgres), each of which you now need to back up, patch, and monitor.
Two concurrency traps are worth calling out. A worker's --concurrency (how many executions that worker runs at once, default 10) and the separate production concurrency limit are independent knobs — if the latter is set too low, adding workers accomplishes nothing. And queue mode does not support filesystem binary data, because the worker that produced a file and the main process serving it may be different processes on different machines, where a local path means nothing. Switch to S3-compatible external storage configured identically on every process, or you will get files that exist on one node and 404 everywhere else.
Configuration quick reference
| Scenario | Key settings |
|---|---|
| Production Code node (minimum) | N8N_RUNNERS_MODE=external + N8N_RUNNERS_BROKER_LISTEN_ADDRESS=0.0.0.0 + matching N8N_RUNNERS_AUTH_TOKEN on both sides |
| Queue mode | EXECUTIONS_MODE=queue + shared Postgres and Redis + n8n worker --concurrency=N, with a dedicated sidecar per worker |
| Officially stated single-instance ceiling | 220 workflow executions per second |
| Version floors | n8n ≥ 1.111.0; extended runner image ≥ 1.121.0 |
To be explicit about sourcing: every version number, default, port, timeout, and error string above comes from n8n's official documentation and GitHub issues (as of 2026-10-10, the latest release is n8n 2.42.6), and hardware prices and specifications are as listed on Amazon at the time of writing and do change. I did not run an end-to-end load test against n8n 2.42.6 — the 220 executions per second figure is the docs' stated number, not my measurement. The fixes for the error cases reflect official documentation and community consensus; what I tested firsthand was external-mode connectivity and configuration behavior across two setups.
This post contains affiliate links. I may earn a commission if you buy through them, at no extra cost to you. If you have questions about the sandbox isolation boundary, treat the official Set up task runners page as the source of truth and cross-check defaults against Task runner environment variables.
👉 Join Xiaomi MiMo Platform: Leading AI model platform with cost-effective inference
👉 Join Aliyun AI: Top AI products with exclusive coupons for business innovation
📌 This article was AI-assisted generated and human-reviewed | TechPassive — An AI-driven content testing site focused on real tool reviews
🔗 Recommended Tools
These are carefully selected tools. Using our affiliate links supports us to keep producing quality content: