OpenHands ships an always-on control plane for an automated software engineer. A Python API server, driven by a browser client and a REST API, provisions sandboxes, holds the operator's git provider tokens and LLM credentials, composes each task's context from user settings and registered skill sources, and then dispatches the work to a separate agent server that actually runs shell commands, edits files and browses the web. Two trust surfaces carry almost all of the risk: the front door, where every API route takes its authentication from a single helper that contributes nothing at all unless one environment variable is set; and the handoff into the sandbox, which is where the server's credentials cross into a process running model-generated code.
The front door is the thing to fix first. In the shipped default, any caller who can reach the published port is served as the operator, and the reference container publishes the port bound to every interface. From there a stranger can start conversations, provision sandboxes, read and rewrite user settings, and drive the git tools using the operator's tokens. Three consequences follow directly. On a host-direct runtime the agent-server child is handed the whole server environment, so a task an outsider started runs with every provider token, JWT secret and cloud credential the server holds. A settings write can register a skill marketplace that is cloned and loaded into the agent's context at the start of every later conversation, which is persistent poisoning that outlives the session that planted it. And the pull-request tool surface is attached to the application through a mount that escapes the route-level check entirely, so it stays open even in a deployment that has correctly set its session key, letting commits and pull requests land on the operator's repositories under the operator's identity. Setting the session key, configuring explicit cross-origin allowlists, and moving the search-proxy mount under the authenticated router are the concrete fixes.
After that, the secrets and the record. Provider tokens and user secrets are written to disk as unencrypted JSON in a directory the reference deployment bind-mounts from the host, even though the codebase already owns the encryption primitive to protect them; and git clone failures are returned verbatim to API callers, where the clone URL those diagnostics echo carries the provider token in it. The monitoring half is detection with no response: rejected logins and rejected sandbox callbacks are logged with good structured context, but nothing counts repeated failures, nothing alerts an operator, and both JSON formatting and on-disk persistence are off by default, so there is no durable audit record of who did what. Worth protecting from regression is everything downstream of the front door, which is genuinely well built: session keys bound to a running sandbox and to its owner, cross-user ownership checks on callbacks, argument-injection-hardened git invocation, a deliberately narrow crypto algorithm allowlist, and a secret-redaction filter actually attached to the log handlers.
- Unauthenticated network caller [PRAX-2026-08-12-001] Untrusted because no credential is required and no identity is established: the router dependency list is empty unless a session key variable is set, and the shipped image binds every interface.
- V1 API surface (/api/v1) [PRAX-2026-08-12-001] The conversation and sandbox endpoints serve the caller with the same authority the operator would have.
- Conversation orchestrator The caller's task text becomes a conversation start request and the orchestrator provisions a sandbox to run it.
- Host-direct sandbox (ProcessSandboxService) [PRAX-2026-08-12-002] On a host-direct runtime the agent server is spawned with a copy of the app server's entire environment, with no allowlist between them.
- Sandboxed agent server [PRAX-2026-08-12-002] Model-generated shell then runs on the host holding the JWT secret, every git provider token and any cloud credential the server was started with.
- Unauthenticated network caller [PRAX-2026-08-12-001] Untrusted because the settings endpoint accepts an anonymous caller: its router derives protection from the same default-empty dependency list.
- V1 API surface (/api/v1) [PRAX-2026-08-12-001] The settings write is accepted and deep-merged into the stored document with only a name-collision check.
- Persisted user settings (settings.json) A marketplace registration with auto-load set is persisted, so the change survives the request, the session and a restart.
- Conversation orchestrator Every conversation started afterwards composes that marketplace and asks the agent server to clone it and load its skills.
- Sandboxed agent server Attacker-authored skill instructions enter the agent's context on each later task and direct its shell, editor and pull-request tools.
- Unauthenticated network caller [PRAX-2026-08-12-001] Untrusted because the caller proves nothing: conversation creation is reachable without any identity.
- V1 API surface (/api/v1) [PRAX-2026-08-12-001] A conversation is started against a repository the operator has authorised, using the operator's stored provider tokens.
- Conversation orchestrator The orchestrator wires the default MCP server into that conversation and dispatches the task.
- Sandboxed agent server The agent follows the attacker's task and reaches for its pull-request tooling.
- OpenHands MCP server (/mcp) [PRAX-2026-08-12-005] The tool surface sits on a mount outside the router dependency, and resolves the operator's provider token server-side for whatever repository the call names.
- Git provider integrations Commits and pull requests land on the operator's git provider under the operator's identity — an external, durable effect.
- Third-party marketplace repository Untrusted because it is third-party content with unlabelled provenance: the registration carries an optional, unverified ref and no signature or hash, so the source can be rewritten after it was approved.
- Sandboxed agent server Its skill files are cloned at conversation start and their instruction text becomes part of the agent's context alongside the operator's own task.
- OpenHands MCP server (/mcp) The redirected agent calls the pull-request tooling, which resolves the operator's provider token server-side.
- Git provider integrations Attacker-chosen changes are pushed and opened as pull requests on the operator's repositories.
B1 Network caller into the control plane (untrusted-ingress) — 4 threats, 3 remit rules · worst: confirmed
R-09 gap any identity not authenticated through that provider is not a trusted operator
R-27 gap Exposing the Agent Server or any agent-run service to an untrusted network without authentication
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| S | ASI03 | Eleven of the thirteen V1 routers take their protection from a helper that contributes no dependency unless SESSION_API_KEY is set, and the default user auth resolves every request to a null user id, so any caller reaching the port is served as the operator. | confirmed PRAX-2026-08-12-001 |
| S | ASI03 | With no CORS origins configured — the shipped default — the middleware reflects every origin while still allowing credentials, so any web page a logged-in operator visits can drive the control plane on their behalf even in a deployment that has correctly set a session key. | confirmed PRAX-2026-08-12-003 |
| D | LLM06 | The only request ceiling is an in-process limiter keyed on the immediate socket address, so behind the reverse proxy this service is normally deployed under every client collapses into one bucket and each replica enforces its own independent limit. | confirmed PRAX-2026-08-12-007 |
| S | ASI03 | A sandbox whose agent has been hijacked could try to impersonate another sandbox's callbacks; the webhook path resolves the caller strictly from the session API key it presents, scopes every subsequent action to that sandbox's owner, and rejects a conversation whose owner or sandbox does not match. | mitigated openhands/app_server/event_callback/webhook_router.py:259 |
B2 Root-mounted surfaces outside the V1 auth path (control-plane-exposure) — 2 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | The status router is declared with no dependencies argument and included at the application root, so /server_info returns host CPU, memory and runtime inventory to unauthenticated callers even where a session key is configured. | confirmed PRAX-2026-08-12-006 |
| E | LLM03 | The MCP surface — five git pull-request tools plus a proxied search namespace holding the operator's search credential — is attached through the FastAPI constructor rather than the V1 router, so router-level dependencies can never apply and the capability stays reachable in a correctly configured deployment. | confirmed PRAX-2026-08-12-005 |
B3 Orchestrator dispatch into the sandboxed agent (tool-invocation) — 2 threats, 1 remit rules · worst: potential
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM01 | Operator task text, the selected repository, and marketplace-sourced skill files are concatenated into one start request with no provenance labelling separating trusted instruction from third-party content before the agent's tool loop sees it (checked: live_status_app_conversation_service.py:1869-1990 and app_conversation_service_base.py:139-165 — no labelling or sanitisation step on the assembly path). | potential |
| T | ASI04 | The sandbox spec resolves to the agent-server image bundled with the installed SDK, and a mismatched image tag is rewritten to the bundled version with a warning rather than honoured, so the executing runtime cannot be silently swapped through configuration. | mitigated openhands/app_server/sandbox/sandbox_spec_service.py:120 |
B4 Provider tokens and secrets at rest and in the sandbox child (secret-material) — 5 threats, 4 remit rules · worst: confirmed
R-19 verified Raw secret values MUST flow only in the SaaS→sandbox direction and MUST NEVER be returned to the SDK / client channel.
R-20 verified Unmasked secrets MUST NEVER be served without the required authentication
R-21 partial Sensitive information MUST NEVER be exposed in error messages, logs, or agent output.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | ASI03 | The host-direct sandbox copies the app server's entire environment into the agent-server child, so model-generated code on a local runtime inherits every provider token, JWT secret and cloud credential the server holds — while the Docker path in the same codebase forwards only a two-prefix allowlist. | confirmed PRAX-2026-08-12-002 |
| I | LLM02 | The default secrets store serialises git provider tokens and user custom secrets with secrets exposed and writes them as unencrypted JSON into the persistence directory the reference compose file bind-mounts from the host, even though the codebase already owns a JWE wrapper that this path never calls. | confirmed PRAX-2026-08-12-004 |
| I | LLM02 | Marketplace git clone and checkout failures are returned to the API caller verbatim, and the authenticated clone URL those diagnostics can echo embeds the provider token in its userinfo; the redaction filter that covers the log path is not applied to HTTP responses. | confirmed PRAX-2026-08-12-010 |
| I | LLM02 | A leaked sandbox session key could be replayed to read raw secret values; the shared validator rejects any key whose sandbox is not RUNNING and returns 403 when the sandbox belongs to a different user, so a key stops working the moment its sandbox pauses. | mitigated openhands/app_server/sandbox/session_auth.py:76 |
| I | LLM02 | A provider token could be carried over cleartext transport to an http:// git host; building an authenticated URL against such a host raises rather than silently downgrading unless the operator explicitly sets the insecure-access variable. | mitigated openhands/app_server/integrations/provider.py:566 |
B5 Persisted settings that re-enter every later conversation (stored-state) — 2 threats, 2 remit rules · worst: potential
R-03 enp MUST NEVER redefine its own goals, security constraints, or approval requirements on the basis of such untrusted content
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | ASI06 | A settings write can register a marketplace with auto-load set, and every conversation started afterwards asks the agent server to clone that source and load its skill instructions into the agent's context — a change that outlives the session that made it (checked: settings_router.py:224-300 and marketplace_composition.py:206-270 — name-collision validation only, no provenance, pin or content check). | potential |
| E | LLM03 | Custom MCP server entries persisted in the same settings document are merged into every conversation's agent configuration, so a stored entry adds tool surfaces the operator never approved for subsequent runs (checked: live_status_app_conversation_service.py:1511-1545 — a merge with no allowlist or approval step). | potential |
B6 Agent-initiated writes to operator git repositories (state-commit) — 2 threats, 3 remit rules · worst: confirmed
R-24 enp Operations that write to or modify state outside the sandboxed workspace or the authorized repositories.
R-25 enp Destructive version-control operations — force-push, branch or repository deletion, and history rewrites.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | The pull-request tools resolve the operator's provider token server-side and act on whatever repository the call names, and because their mount escapes the router dependency the caller need not be the operator — so commits and pull requests can land on the operator's repositories under the operator's identity. | confirmed PRAX-2026-08-12-005 |
| T | — | The destructive version-control operations the remit gates behind human approval — force-push, branch or repository deletion, history rewrite — have no representation in this tree; the exposed tool set is creation-only. | mitigated openhands/app_server/mcp/mcp_router.py:147 |
B7 Dependencies, image build, and third-party skill sources (supply-chain) — 3 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM04 | About eighteen direct dependencies carry no version constraint in an otherwise strictly pinned manifest, no SBOM or provenance is produced for the primary app image, and Dependabot is the only scanner in the repository, so a newly disclosed vulnerability in a pinned transitive dependency surfaces only if Dependabot happens to raise it. | confirmed PRAX-2026-08-12-008 |
| T | ASI04 | A registered marketplace source is cloned and its skill files become agent instructions with no signature, hash or mandatory commit pin, so a source rewritten after registration silently changes what the agent is told to do on the next run (checked: skill_loader.py:511-560 and settings_models.py:742-750 — ref is optional and never verified). | potential |
| T | LLM04 | A compromised third-party GitHub Action could reach the build; every action outside the first-party namespaces is pinned to a full 40-character commit SHA, with Dependabot covering pip, npm, github-actions and docker across six directory entries. | mitigated .github/workflows/npm-publish-ui.yml:69 |
B8 Traces, analytics, and the application log (telemetry-egress) — 5 threats, 3 remit rules · worst: confirmed
R-30 gap Alert on repeated authentication / authorization failures
R-31 partial Record security-relevant events — authentication events, tool invocations, and outbound posts to external services
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Three call paths prefer the user's email over the internal id as the trace identity sent to a third-party observability service, and the same traces carry repository, branch and commit metadata, so both PII and project identity leave the authorized counterparty set whenever the integration is configured. | confirmed PRAX-2026-08-12-009 |
| R | — | Security-relevant events go to the ordinary application logger with JSON formatting and file persistence both default-off, and there is no audit stream distinct from application logging, so a normal deployment keeps only coloured plaintext on stdout. | confirmed PRAX-2026-08-12-012 |
| R | — | Session-key rejections, sandbox ownership mismatches and webhook rejections are all logged with structured context, but nothing counts repeated failures per source and no alert sink is wired to any rejection path, so the remit's alert-without-halting obligation has a detection half and no response half. | confirmed PRAX-2026-08-12-011 |
| I | LLM02 | Credential values could be written into log records; a redaction filter is attached to the handlers and to the logger itself, scrubbing credential-shaped environment values, twelve named key patterns, and API-key literals. | partial openhands/app_server/utils/logger.py:248remainder: the filter derives its value list from the process environment, so a custom secret that exists only in secrets.json or in a request body is never in that list and is not scrubbed |
| I | LLM02 | Product analytics could carry user identity to a third-party ingest host without permission; every capture call returns immediately when the resolved context is not consented, and person profiles are disabled outside SaaS mode. | mitigated openhands/analytics/analytics_service.py:83 |
| Component | Kind | Lane | Description | Evidence |
|---|---|---|---|---|
| Unauthenticated network caller | entrypoint | user_inputs | Any caller that can reach the published port; in the shipped default it is served with the same authority as the operator. | openhands/app_server/utils/dependencies.py:23; containers/app/Dockerfile:105 |
| Self-hosting operator | entrypoint | user_inputs | The human operator driving the Agent Canvas client and REST API; in OSS mode the deployment models exactly one such local identity. | openhands/app_server/user_auth/default_user_auth.py:24; openhands/app_server/user/user_router.py:21 |
| Third-party marketplace repository | entrypoint | user_inputs | External git repositories registered as skill/plugin marketplaces; their skill files are cloned and their instructions loaded into the agent's context. | openhands/app_server/settings/settings_models.py:742; openhands/app_server/user/skills_router.py:284 |
| V1 API surface (/api/v1) | adapter | client_adapters | The aggregate REST surface: thirteen routers for conversations, sandboxes, secrets, settings, users, skills, git, events, webhooks, web-client and config. | openhands/app_server/v1_router.py:24; openhands/app_server/app.py:71 |
| Status / probe router | adapter | client_adapters | Root-mounted liveness, readiness and /server_info endpoints declared with no dependencies argument of any kind. | openhands/app_server/status/status_router.py:5; openhands/app_server/status/status_router.py:28 |
| FastAPI application root | orchestrator | agent_core | The composition root: constructs the app with the MCP mount, includes the V1 and status routers, and installs the middleware stack. | openhands/app_server/app.py:54; openhands/app_server/app.py:71 |
| Conversation orchestrator | orchestrator | agent_core | Assembles each conversation start request — settings, secrets, skills, marketplaces, LLM and MCP configuration, trace metadata — and dispatches it to the sandboxed agent server. | openhands/app_server/app_conversation/live_status_app_conversation_service.py:382; openhands/app_server/app_conversation/live_status_app_conversation_service.py:1481 |
| Router auth dependency helper | control | agent_core | The single helper every V1 router takes its authentication from; contributes no dependency at all unless SESSION_API_KEY is set, and in SaaS mode only a non-failing header declaration. | openhands/app_server/utils/dependencies.py:9; openhands/app_server/utils/dependencies.py:23 |
| HTTP middleware stack | control | agent_core | Request-path middleware: a CORS origin gate that reflects any origin when no allowlist is configured, a cache-control header setter, and an in-process rate limiter keyed on the immediate socket address. | openhands/app_server/middleware.py:37; openhands/app_server/config.py:103 |
| Sandbox session-key validation | control | agent_core | Validates X-Session-API-Key, rejects keys whose sandbox is not RUNNING, and verifies that the sandbox belongs to the calling user. | openhands/app_server/sandbox/session_auth.py:37; openhands/app_server/sandbox/session_auth.py:103 |
| Log secret-redaction filter | control | agent_core | Filter that scrubs credential-shaped environment values and known key patterns from log records; attached to the handlers rather than merely defined. | openhands/app_server/utils/logger.py:248; openhands/app_server/utils/logger.py:392 |
| JWT / JWE token service | control | agent_core | Signs and verifies the scoped access tokens the sandbox presents for secret retrieval, with a deliberately narrow two-element JWE algorithm allowlist; also exposes symmetric encrypt/decrypt helpers the secrets store never calls. | openhands/app_server/services/jwt_service.py:27; openhands/app_server/services/jwt_service.py:253 |
| Persisted user settings (settings.json) | datastore | agent_core | The default per-user settings document — LLM profiles, disabled skills, registered marketplaces, custom MCP configuration — whose contents are re-read into every later conversation's agent configuration. | openhands/app_server/settings/file_settings_store.py:15; openhands/app_server/server_config/server_config.py:17 |
| OpenHands MCP server (/mcp) | mcp_server | tools_mcp | FastMCP server mounted through the FastAPI constructor: five git pull/merge-request creation tools plus a proxied Tavily search namespace holding the server's search key. | openhands/app_server/mcp/mcp_router.py:43; openhands/app_server/mcp/mcp_router.py:49 |
| Sandbox provisioning (Docker / remote) | tool | tools_mcp | The isolated-sandbox provisioning family: creates the agent-server container, mints its 32-byte session key, sets its webhook callback URL and mounts the configured volumes. | openhands/app_server/sandbox/sandbox_service.py:112; openhands/app_server/sandbox/docker_sandbox_service.py:405 |
| Host-direct sandbox (ProcessSandboxService) | tool | tools_mcp | The host-direct runtime: spawns the agent server as a local subprocess whose environment is a copy of the app server's own. | openhands/app_server/sandbox/process_sandbox_service.py:114; openhands/app_server/config.py:333 |
| Agent-container environment allowlist | control | tools_mcp | Deliberate two-prefix allowlist plus an explicit JSON override map governing which environment variables the Docker agent-server container inherits. | openhands/app_server/sandbox/sandbox_spec_service.py:151; openhands/app_server/sandbox/docker_sandbox_spec_service.py:35 |
| Marketplace skills endpoint | tool | tools_mcp | Server-side git clone of registered marketplace repositories to preview their skill metadata; returns the raw git stderr for failed clones to the caller. | openhands/app_server/user/skills_router.py:32; openhands/app_server/user/skills_router.py:284 |
| Sandboxed agent server | tool | tools_mcp | The dispatched agent process or container that runs the model loop and executes shell, editor and browser actions; the app server starts it, sends it the conversation, and receives its callbacks. | openhands/app_server/sandbox/sandbox_models.py:27; openhands/app_server/sandbox/process_sandbox_service.py:130 |
| Secrets store (secrets.json) | secret_store | external_deploy | The configured default at-rest store: git provider tokens and user custom secrets serialised with secrets exposed and written as plain JSON to the persistence directory. | openhands/app_server/secrets/file_secrets_store.py:32; openhands/app_server/server_config/server_config.py:20 |
| Application logger | log_sink | external_deploy | The single logging surface for security-relevant events; JSON formatting and file persistence are both off by default and there is no audit stream distinct from application logging. | openhands/app_server/utils/logger.py:30; openhands/app_server/utils/logger.py:55 |
| Laminar trace service | external_service | external_deploy | Third-party LLM observability service; the conversation start request carries the user's email as the trace identity along with repository, branch and commit metadata. | openhands/app_server/app_conversation/live_status_app_conversation_service.py:387; openhands/app_server/app_conversation/live_status_app_conversation_service.py:1840 |
| Git provider integrations | external_service | external_deploy | The git provider family (GitHub, GitLab, Bitbucket, Bitbucket DC, Azure DevOps, Forgejo) and the handler that builds authenticated clone URLs with the provider token in the URL userinfo. | openhands/app_server/integrations/service_types.py:16; openhands/app_server/integrations/provider.py:510 |
| Reference container deployment | deploy_surface | external_deploy | The shipped compose deployment and the image it builds: publishes the control plane on port 3000, bind-mounts the host Docker socket and the persistence directory, and runs the server as root on all interfaces. | docker-compose.yml:14; containers/app/Dockerfile:100 |
| Dependency and build posture | deploy_surface | external_deploy | Repo supply-chain posture: a mostly exactly-pinned manifest with about eighteen unconstrained direct dependencies, Dependabot as the only scanner, and SBOM/provenance enabled for the enterprise image only. | pyproject.toml:24; .github/dependabot.yml:1 |
omissions: Arbitration divergence: the findings JSON tags PRAX-2026-08-12-005 with an agentic primary of ASI02 (Tool Misuse); under the current KB arbitration this is grant/gate evidence — a reachable ungated capability with no wrongful use evidenced — so the primary here is LLM03 (Excessive Agency), with ASI02 available only as a co-tag. The KB wins; the graph tags LLM03 at both boundaries where that finding is cited. No other stored tag disagreed with the KB. Scope-driven absences: the agentic core (controller / runtime / llm / mcp event loop), enterprise/, frontend/ and kind/ are outside the scan scope, so no model node and no model-egress boundary are emitted — the app server assembles and forwards LLM credentials and model metadata but makes no model call of its own (checked openhands/app_server/utils/llm.py, where LiteLLM is imported for the model catalogue only). For the same reason the remit's HMAC webhook-signature obligation is not attached to any boundary: the only webhook surface in this tree is the sandbox-to-app-server callback authenticated by session API key, while the git-provider webhook receivers that verify HMAC signatures live under the excluded enterprise/ tree. config.template.toml is named in the scan scope but no module under openhands/ reads it in this snapshot (it configures the legacy V0 core), so it is not modelled as a node. The PostHog analytics egress is covered by a threat on the telemetry-egress boundary but is collapsed into that boundary rather than given its own node, the consent gate being its only remit-relevant behaviour. Two nodes are deliberately edgeless: the reference container deployment and the dependency/build posture node, both repo- and deploy-posture evidence rather than data-flow participants.
generated by: Opus 5 (1M context)