AutoGen's code executor is not a service but a library: five interchangeable back ends that take code blocks written by another agent or by a model, drop each one into a working directory, and run it — in a Docker container, in a containerized Jupyter kernel, in a Jupyter kernel inside the application's own process, in a plain host subprocess, or in a remote Azure sandbox session. Everything untrusted arrives on one path, as code blocks in agent messages, and everything consequential happens on another: the container, the host shell, and the shared working directory that all of the back ends read and write. The two surfaces that decide whether this component is safe are the choice of back end — whether generated code lands inside an isolation boundary at all — and the human confirmation step that is supposed to precede every execution.
Neither holds. There is no place in the executor contract for a confirmation to happen; the only approval gate anywhere is an optional callback in the calling agent, and it is off unless a developer supplies one. The factory that picks a back end quietly hands back the host executor whenever Docker is missing or fails to start, so a deployment that believes it is containerized can be running generated shell commands directly on the machine, with the application's full environment, and with a documented dangerous-command filter that does not exist in the code at all. Two of the five back ends run on the host by design or by fallback, and package setup reaches the host interpreter no matter which back end was chosen. Make an unavailable sandbox stop the run rather than fall through to the host, then give the executor contract a confirmation point of its own instead of relying on whoever calls it.
Three concrete paths deserve an operator's attention. A code block from any participant in the conversation can reach the host shell unscreened, because the sender filter is optional and the confirmation gate defaults to off. Output from executed code flows back into the model's context raw — unlabeled, unbounded, unescaped — so text a program prints can steer the next block that gets generated and run, and the container it runs in has root, the default network, and no egress restriction to carry data out. And the containerized notebook server publishes its code-execution port on every host interface while reporting itself as loopback, protected only by a token the client sends over an unencrypted socket even when it was configured for HTTPS. Nothing anywhere records what was executed, its exit code, or its output, so none of this leaves a trace to investigate.
- Coder / assistant agent a code block arrives in a chat message from a peer whose identity nothing verifies — the sender allowlist is optional and off by default
- CodeExecutorAgent loop [PRAX-2026-08-12-004] the block is extracted and executed with no confirmation, because the only approval hook lives in this caller and defaults to off
- Default executor factory [PRAX-2026-08-12-001] the executor it was handed is the host backend, silently substituted when Docker was unavailable or failed to initialize
- Local command-line executor [PRAX-2026-08-12-002] the block runs as a host subprocess with the application's full environment and no dangerous-command screening, despite the docstring that promises it
- Docker daemon and execution container the running block prints attacker-authorable text — content from a fetched page, a data file or a compromised dependency — into the container's stdout
- Docker command-line executor [PRAX-2026-08-12-012] that output is decoded and returned verbatim, unbounded and unlabeled
- CodeExecutorAgent loop [PRAX-2026-08-12-012] it is appended to the model context as an ordinary message with nothing marking it as data rather than instruction
- Chat completion model client the next completion is generated under the injected instruction and emits a new code block
- CodeExecutorAgent loop [PRAX-2026-08-12-004] the loop extracts and runs that block with no confirmation step
- Docker command-line executor the new block is written into the mounted workspace and executed in the container
- Docker daemon and execution container [PRAX-2026-08-12-006] it runs as root on the default bridge with unrestricted egress, carrying workspace contents to any destination the host can reach
- Containerized kernel gateway [PRAX-2026-08-12-003] the kernel gateway port is published on every host interface with kernel listing enabled, so an unauthorized network caller can reach a stateful arbitrary-code endpoint whose auth token crosses the same network in plaintext
- Host working directory [PRAX-2026-08-12-007] code run through that kernel writes into the world-writable host directory mounted read-write as its working directory, persisting files that later blocks and the host itself read back
- Azure dynamic-sessions pool endpoint the session's file listing is produced by the code that ran there, so the names it returns are attacker-authorable rather than operator-configured
- ACA dynamic-sessions executor [PRAX-2026-08-12-009] membership in that listing is the only gate before the name is joined to the working directory and opened for write
- Host working directory [PRAX-2026-08-12-009] a name carrying traversal segments lands the response body on any existing host path the process can reach — including code a later block will execute
B1 Inbound code blocks from the agent runtime (untrusted-ingress) — 1 threats, 2 remit rules · worst: partial
R-15 verified Code blocks supplied by the authorized coder/orchestrator agent, and files explicitly uploaded by the operator to the working directory / session.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| S | ASI03 | Any participant in the conversation can supply executable blocks: the caller matches senders against an optional source-name allowlist that is off by default and verifies nothing about who actually sent the message, so an impersonating or compromised peer reaches the executor on the same path as the authorized coder agent. | partial python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:681remainder: the allowlist defaults to None and compares runtime-assigned source names, so a peer posting under an allowed name is indistinguishable from the authorized coder agent |
B2 Execution output re-entering the agent's context (untrusted-ingress-2) — 1 threats, 0 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM01 | Whatever an executed block writes to stdout and stderr is decoded and concatenated verbatim into the model context and onto the operator's terminal — no size cap, no label marking it as data rather than instruction, and no control-character neutralization on any backend except the Jupyter traceback path. | confirmed PRAX-2026-08-12-012 |
B3 Decision to run a code block (tool-invocation) — 5 threats, 7 remit rules · worst: confirmed
R-13 partial The local host executor (running code directly on the machine hosting the application) MUST NOT be used unless the operator has explicitly authorized it, and then only in a trusted, non-production context.
R-20 partial Human confirmation by the overseer is required before every code-block execution; there is no risk-threshold carve-out (no category of block is exempt), and a denied confirmation blocks execution.
R-21 gap Installing packages into the host interpreter, or otherwise modifying state outside the working directory and outside the configured sandbox boundary, MUST require human confirmation.
R-22 verified The executor MUST NOT execute a code block that the human overseer has denied.
R-23 gap The executor MUST NOT execute agent- or LLM-generated code outside an isolation boundary (container or remote sandbox), except via the explicitly operator-approved local executor.
R-25 gap If the configured sandbox backend (container or remote session) is unavailable, the executor MUST halt and alert the operator rather than falling back to unsandboxed host execution.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | The executor contract defines no confirmation hook — every backend goes from execute_code_blocks straight to writing and running the block — so the only gate in the picture is the calling agent's opt-in approval callback, which defaults to off and which any other caller of an executor bypasses entirely. | confirmed PRAX-2026-08-12-004 |
| E | ASI05 | The default factory downgrades from the container backend to the host backend whenever Docker is unavailable or its constructor raises, so a caller that asked for isolation silently receives a host executor with only a warning where the remit requires a halt and an operator alert. | confirmed PRAX-2026-08-12-001 |
| E | ASI05 | The local backend's documented dangerous-command sanitizer does not exist anywhere in the tree, so bash, sh and PowerShell blocks reach the host unscreened, with a full copy of the parent process environment and PowerShell invoked with execution policy bypassed. | confirmed PRAX-2026-08-12-002 |
| E | ASI05 | The Jupyter backend runs generated code as notebook cells in a kernel launched in the host process context with no container and no remote session, and unlike the local backend it emits no construction-time warning and requires no operator authorization. | confirmed PRAX-2026-08-12-005 |
| E | LLM03 | Function setup reaches the host whichever backend an operator selected — pip install runs against the application's own interpreter when no virtual-env context is supplied, and string-defined function source is compiled and executed in the host process at construction — with no confirmation on either path. | confirmed PRAX-2026-08-12-010 |
B4 Agent to model provider (model-egress) — 1 threats, 0 remit rules · worst: partial
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Execution output — including anything an executed block read out of the host environment or the working directory — is added to the model context and forwarded to the external model provider on the next turn. | partial python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:477remainder: the only output suppressor exists on the Azure backend, defaults to False, and covers just the result field — stdout and stderr from all five backends are forwarded verbatim |
B5 Host working directory and bind mounts (data-at-rest) — 3 threats, 0 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | ASI05 | The Jupyter server chmods the supplied host bind directory to world-writable and mounts it read-write as the kernel's working directory, so any local user or process can read, modify or replace the code and data the kernel subsequently executes or returns. | confirmed PRAX-2026-08-12-007 |
| T | LLM10 | When a Docker-backend block carries no filename comment the generated name interpolates the block's language string straight into the host path and writes the file before the language is validated, so a language value carrying path separators lands a file outside the working directory (the calling agent's markdown extractor happens to constrain the language; a direct caller of the executor does not). | confirmed PRAX-2026-08-12-008 |
| T | LLM10 | The Azure backend's file download gates only on membership in the remote session's own file listing, then joins that remote-supplied name to the working directory and opens it for write, giving an overwrite primitive on any existing host path the process can reach. | confirmed PRAX-2026-08-12-009 |
B6 Persisted workspace re-entering later execution (stored-state) — 1 threats, 0 remit rules · worst: potential
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | ASI05 | The generated functions module and anything else a block leaves in the working directory persist for the life of the session and are imported or run by later blocks, so a block — or, given the world-writable bind directory, a local user — that rewrites that module changes what every subsequent block executes (checked: _common.py:96-111 — path containment is the only check on workspace writes, and setup never re-verifies the functions module after the first block). | potential |
B7 Backend credentials (secret-material) — 2 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | ASI03 | The Azure backend fetches its access token once and returns the cached value on every later call under an acknowledged expiry TODO, so a long-lived executor holds one credential indefinitely and only an explicit restart clears it. | confirmed PRAX-2026-08-12-015 |
| S | — | An unauthenticated kernel gateway would let any caller start and drive kernels; the server generates a 32-byte random token by default, injects it into the container and requires it on every API and websocket call. | mitigated python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:350 |
B8 Published kernel-gateway execution surface (control-plane-exposure) — 2 threats, 3 remit rules · worst: confirmed
R-24 gap The executor MUST NOT expose the code-execution environment, or the Docker daemon socket it relies on, to untrusted or public networks.
R-05 partial Agent runtime message bus (code blocks in, results out) ... Only from agents within the same authorized AutoGen runtime
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | ASI05 | The bundled image starts the kernel gateway on all interfaces with kernel listing enabled and the container is run with all ports published, so any party that can reach the host — not just the authorized runtime — can enumerate and drive a stateful arbitrary-code kernel, while the client-side connection info still reports loopback. | confirmed PRAX-2026-08-12-003 |
| I | LLM02 | The websocket URL is built as plaintext regardless of the HTTPS setting the REST calls beside it honor, so the gateway auth token and every code block cross the network unencrypted even for a remote gateway an operator configured for TLS. | confirmed PRAX-2026-08-12-013 |
B9 Executed code reaching the network and the host (container-sandbox--external-network) — 1 threats, 2 remit rules · worst: confirmed
R-26 gap If executed code attempts to reach resources outside the sandbox and working directory (host filesystem, unauthorized network destinations, privilege escalation), the executor MUST halt and alert the operator.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | The execution container is created with an image, a name, a bind mount and lifecycle flags only — no network mode, user, capability drop, read-only root or resource ceiling — so executed code runs as root on the default bridge with unrestricted egress to anything the host can reach, and nothing detects or halts an attempt to leave the sandbox. | confirmed PRAX-2026-08-12-006 |
B10 Execution environment images (supply-chain) — 1 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM04 | The environment that runs untrusted code is resolved by mutable tag and pulled on demand with no digest pin or signature check, and the bundled Jupyter image builds from a base with no tag at all, so what actually executes is whatever those tags point at on the day — with no inventory to reconstruct a past run. | confirmed PRAX-2026-08-12-014 |
B11 Execution record (telemetry-egress) — 1 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| R | — | No backend records what it executed — not the block, not the exit code, not the output; the only logging covers container lifecycle, image pulls and cancellation errors, so every other gap here is invisible after the fact and an incident leaves nothing to reconstruct. | confirmed PRAX-2026-08-12-011 |
| Component | Kind | Lane | Description | Evidence |
|---|---|---|---|---|
| Coder / assistant agent | entrypoint | user_inputs | The peer agent in the AutoGen runtime whose chat messages carry the markdown code blocks that get executed. | python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:683; python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:312 |
| Human overseer | entrypoint | user_inputs | The operator the remit expects to confirm each block; reachable only through the caller's optional approval callback. | python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:182; python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:458 |
| CodeExecutorAgent loop | orchestrator | agent_core | Read as context, not scored: the generate/execute/reflect loop that ingests messages, extracts code blocks, drives the executor and feeds results back into the model context. | python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:515; python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:613 |
| Chat completion model client | model | agent_core | The model that generates the code blocks when the caller supplies a model_client, and that judges whether to retry after a failure. | python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:821; python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:590 |
| approval_func gate | control | agent_core | Opt-in confirmation callback in the calling agent — a denial does block execution, but it defaults to None and lives outside the executor contract. | python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:691; python/packages/autogen-agentchat/src/autogen_agentchat/agents/_code_executor_agent.py:454 |
| Default executor factory | orchestrator | agent_core | create_default_code_executor — chooses the Docker backend when the daemon answers and otherwise returns a host executor with only a warning. | python/packages/autogen-ext/src/autogen_ext/code_executors/__init__.py:58; python/packages/autogen-ext/src/autogen_ext/code_executors/__init__.py:25 |
| Docker command-line executor | tool | tools_mcp | Writes each block to the working directory and runs it inside a container via exec_run, returning the decoded output. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:405; python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:340 |
| Local command-line executor | tool | tools_mcp | Runs each block as a host subprocess in the working directory with a full copy of the parent process environment. | python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:341; python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:413 |
| Host Jupyter kernel executor | tool | tools_mcp | Runs blocks as notebook cells in an nbclient kernel started in the host process context — no container, no remote session, no construction-time warning. | python/packages/autogen-ext/src/autogen_ext/code_executors/jupyter/_jupyter_code_executor.py:283; python/packages/autogen-ext/src/autogen_ext/code_executors/jupyter/_jupyter_code_executor.py:139 |
| Docker Jupyter executor | tool | tools_mcp | Sends each block as a cell over a websocket to a containerized kernel gateway and collects the streamed outputs and data items. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_docker_jupyter.py:197; python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_docker_jupyter.py:265 |
| ACA dynamic-sessions executor | tool | tools_mcp | Posts Python blocks to an Azure Container Apps dynamic session over HTTPS and can upload and download session files. | python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:417; python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:354 |
| Function-setup pathway | tool | tools_mcp | Compiles supplied function source in the host process at construction and, on the local backend without a virtual-env context, pip-installs its packages into the application's own interpreter. | python/packages/autogen-core/src/autogen_core/code_executor/_func_with_reqs.py:114; python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:285 |
| Workspace containment check | control | tools_mcp | Resolves a block's '# filename:' comment against the working directory and raises when it escapes; it covers only that branch, not generated file names. | python/packages/autogen-ext/src/autogen_ext/code_executors/_common.py:96; python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:346 |
| Language allowlist | control | tools_mcp | Per-backend SUPPORTED_LANGUAGES lists plus lang_to_cmd, which rejects unsupported languages — enforced before the file write everywhere except the Docker backend, where it runs after. | python/packages/autogen-ext/src/autogen_ext/code_executors/_common.py:155; python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:364 |
| Documented dangerous-command sanitizer | control | tools_mcp | STUB — the local executor's class docstring advertises regex sanitization against a list of dangerous commands; no list, function or call site exists anywhere in the tree. | python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:57; python/packages/autogen-ext/src/autogen_ext/code_executors/local/__init__.py:397 |
| Host working directory | datastore | tools_mcp | The host directory every backend writes code blocks, function modules and downloaded files into, and which the container backends bind-mount read-write as the execution working directory. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:546; python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:325 |
| Docker daemon and execution container | external_service | external_deploy | The host Docker daemon that creates and runs the execution container; the container is created with no network, user, capability, filesystem or resource restriction. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:537; python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:296 |
| Containerized kernel gateway | deploy_surface | external_deploy | DockerJupyterServer builds and runs the bundled kernel-gateway image with publish_all_ports, exposing a stateful arbitrary-code endpoint on every host interface. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:363; python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:278 |
| Azure dynamic-sessions pool endpoint | external_service | external_deploy | The operator-configured pool management endpoint that runs the code, holds the session files and returns their listing. | python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:208; python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:303 |
| Azure token provider | secret_store | external_deploy | The operator-supplied credential object whose dynamicsessions scope token the executor caches in process for its lifetime. | python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:152; python/packages/autogen-ext/src/autogen_ext/code_executors/azure/_azure_container_code_executor.py:279 |
| Kernel-gateway auth token | control | external_deploy | A 32-byte random token generated by default and injected into the container as TOKEN, which the gateway enforces on every API and websocket call. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:350; python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:68 |
| Container registries | external_service | external_deploy | The registries that supply the execution environments — the default python:3-slim tag pulled on demand and the untagged quay.io foundation image the bundled Dockerfile builds from. | python/packages/autogen-ext/src/autogen_ext/code_executors/docker/_docker_code_executor.py:519; python/packages/autogen-ext/src/autogen_ext/code_executors/docker_jupyter/_jupyter_server.py:267 |
omissions: Arbitration divergence: finding PRAX-2026-08-12-013 stores owasp_agentic ASI07 (Insecure Inter-Agent Communication), but the kernel gateway is an execution backend, not a peer agent, so the Agentic KB's inter-agent row does not apply; tagged LLM02 (Sensitive Information Disclosure) per the LLM KB for the auth token crossing in plaintext. Findings 002 and 005 list LLM03 ahead of ASI05 in their stored tag order while 001 lists ASI05 first on the same mechanism; the Agentic KB's live-exec-path row makes ASI05 primary for all three, so they are tagged ASI05 here. Coinages: the boundary container-sandbox--external-network is coined because no archetype covers executed code reaching the network and the host filesystem from inside the sandbox. The host working directory is a shared runtime artifact with no single defining file and no fixed literal path; its id takes the work_dir construct over the Docker backend, though the most consequential evidence for it (the world-writable chmod) lives in the docker_jupyter server file. The absent dangerous-command sanitizer has no defining construct to key on, so its id takes the class scope that documents it. Edgeless: the documented sanitizer node carries no edges because the control does not exist; the telemetry-egress boundary has no crossing edges because nothing is written to any sink. Not modelled: the abstract CodeExecutor contract in autogen-core, whose missing approval hook is dispatch-level plumbing folded into the executor nodes and the tool-invocation threats; and the outbound republication of raw execution output to peer agents in a group chat, which belongs to the calling agent rather than the scored executors. The findings JSON's threat coverage matched the code in front of me on every citation.
generated by: Opus 5 (1M context)