Skip to Content
OperateDeployArcade Worker

Arcade Worker

The Arcade Worker hosts Arcade’s toolkits and runs their calls for the Engine. In a self-hosted deployment it’s the bundled MCP server, and it appears as a workers[] entry in the Helm values.

This page explains how the worker behaves and when to change each setting. For every value and its default, see the chart reference on Artifact Hub , which is the source of truth.

Process models

The worker ships as two images built from the same commit. The image you choose is the worker’s process model:

ImageProcess model
arcadedev/worker-supervisorServes the tool catalog from a snapshot taken at build time, and starts one isolated runtime per toolkit version on demand. One worker can host multiple versions of the same toolkit.
arcadedev/workerEager: loads every toolkit into one process at boot. Use it to roll back.

To switch, set image.repository to the other image. You don’t need to rebuild anything or change any other value.

How the supervisor worker behaves

Operators can rely on the following behavior from the arcadedev/worker-supervisor image.

Versions

  • A call that names a toolkit version runs on exactly that version. If the worker doesn’t offer that version, it refuses the call rather than serving it with another version.
  • A call that names no version runs on the highest version the worker offers.
  • Multiple versions of the same toolkit run side by side, and each serves only its own calls.

Errors

  • A call for a that the version doesn’t have, or for a toolkit the worker isn’t configured to serve, fails with an error that clients can’t retry.
  • If the worker can’t start or reach a runtime, the call fails with a generic retryable error that exposes no internal detail.
  • If a version can never start, the worker gives up on it after a bounded number of attempts. A version that fails to start never takes the worker down.

Lifecycle

  • If a runtime crashes, the worker recovers it on the next call.
  • The worker stops an idle runtime to reclaim memory and starts it again on demand.
  • The worker reports ready only after its pre-warm pass finishes, so the first callers after a deploy don’t wait for warm-up. The worker skips any pre-warm entry it can’t warm.
  • The worker publishes the of every version it serves, whether or not the version is running, without starting it.
  • When the worker stops, it answers calls already in flight, then stops every runtime it started.

Settings

Set these values under workerDefaults to apply them to every worker, or on an individual workers[] entry.

ValueDefaultWhat it controls and when to change it
image.repositorySee the chart reference Which process model the worker runs. Point it at arcadedev/worker to roll back.
variantDerived from the imagesupervisor or eager. Set it only when you mirror the images under other names, so telemetry (worker_variant) and supervisor-only resources render correctly.
supervisor.prewarmToolkitsEmptyToolkits started at boot and kept warm, at the highest version of each. Pre-warmed toolkits are exempt from idle eviction and restart if they die. List the toolkits your users call most, so they’re never served cold.
supervisor.serveToolkits*Toolkits this worker serves. The worker refuses calls for anything left off and leaves it out of the published tools. Use it to split a fleet, for example to run a dedicated worker for the most-called toolkits.
supervisor.maxChildren15The most runtimes running at once. At the cap, the worker stops the least recently used idle runtime that isn’t pre-warmed to admit a new one. 0 means unbounded. Size pod memory from this value, see Sizing.
supervisor.idleTtlSeconds600Seconds a runtime may sit idle before the worker stops it. 0 never stops an idle runtime.
supervisor.capacityWaitSeconds30How long a call waits for a free slot when every runtime is busy at the cap, before it fails with a retryable error.
supervisor.peerRouting.enabledfalseReplicas agree on which one starts each version and hand first calls to it, so a burst of first calls costs the fleet one cold start instead of one per replica. Turning it on adds a headless Service. A replica that can’t reach the owner serves the call itself.
supervisor.peerRouting.promoteAfterForwards10How many calls for one version a replica hands off before it starts its own runtime for that version.
resources, hpa.*, replicaCountSee the chart referencePod resources and scaling. See Sizing.

Advanced settings

Set these environment variables through extraEnv. See the chart reference  for details.

VariableDefaultWhat it controls
ARCADE_SUPERVISOR_SHARD_PREWARMOffWith peer routing on, each replica warms only its share of the pre-warm list. This multiplies the fleet’s warm coverage by the replica count.
ARCADE_SUPERVISOR_INVOKE_TIMEOUT600 secondsThe longest a single tool call may run.
ARCADE_SUPERVISOR_PREWARM_CONCURRENCY4How many runtimes start at once during boot. Keep it near the pod’s core count.

Sizing

  • Memory: each running runtime uses about 100 MB, so a pod needs roughly maxChildren × 100 MB plus memory for the supervisor itself, plus headroom.
  • Boot time: with ten pre-warmed toolkits on 2 cores, the supervisor worker is ready in about 12 seconds. The eager worker takes 60 seconds or more to load every toolkit.
  • Image size: the supervisor image is 2.96 GB. The eager worker plus the deprecated-versions worker it replaces total 3.85 GB.

Set the pod’s resources from supervisor.maxChildren, and scale out with replicaCount or hpa.*.

Next steps

Last updated on