CraftBot is a self-hosted personal agent that runs as a background service on its owner's own computer, plans and executes multi-step work for that single owner, and reaches the outside world through a local browser chat interface, a set of connected messaging and productivity integrations, two bundled tool servers it launches on every run, and the owner's own model provider. Untrusted material arrives from several directions at once: pages it fetches from the web, messages strangers send on the owner's connected platforms, files it reads from anywhere on the disk, and results returned by those tool servers. What it may do with that material is unusually broad — it runs commands in the host shell, reads and writes any file by absolute path, sends mail and messages from the owner's accounts, makes arbitrary web requests, and builds and publishes small local web applications. Two surfaces matter most: the host shell, which is present in every task, and the agent's own instruction and memory files, which it can rewrite and which are reloaded at the start of every future session.
The pattern behind the worst of it is that the safety model is written as instructions to the model rather than as code, and three chains are open because of it. A web page the agent reads can carry text the agent acts on, and its next step can be a host command that runs with the owner's model-provider keys handed to it in the environment; nothing between the page and the shell checks anything, and the one module in the tree named for input sanitization is never called and would not block a match if it were. A result returned by one of the bundled tool servers can travel the same route into the file reader, pull stored access tokens and provider keys off disk into the conversation, and leave over an outbound request to any address on the internet, because no destination list and no approval step exist anywhere on that path. And a stranger's message on a connected platform, though it is correctly stopped from starting a task of its own, still has a few hundred characters of its text pasted into the conversation the owner sees next, from where it can direct a write into the files that define the agent's standing instructions — a change that survives every restart.
Fix the shell first: require the owner to approve a command before it runs, show them the exact command, and stop handing the process environment and the API keys in it to whatever that command starts. Then close the write path into the agent's own instruction and memory files, and put a destination list and a confirmation step in front of outbound sends and requests. Two more items belong in the same pass: the permission tiers the documentation presents as the approval mechanism for scheduled and self-initiated work are never consulted at run time, so background tasks act unsupervised, and live third-party client secrets are checked into the source, deliberately split so scanners will not spot them, which means they should be treated as already compromised and rotated. Where the team did write controls in code they are good ones — a real gate on stranger messages, a filter that blocks requests to internal addresses, action and token ceilings that stop and ask the owner, and a ledger that keeps a completed send from running twice — which is what makes the absent approval layer read as an omission rather than something beyond the project.
- Fetched web content [PRAX-2026-08-12-001] Attacker-authored text on any page the agent is asked to read — unauthenticated, unlabeled provenance, and the agent has no relationship with the host serving it.
- web_fetch / web_search [PRAX-2026-08-12-001] web_fetch extracts the page text and returns it into the model context with no trust tier attached.
- AgentBase react loop [PRAX-2026-08-12-005] The react loop assembles that text into the same prompt as the owner's instructions; the only sanitizer in the tree is never called and would pass a match through anyway.
- run_shell (host command execution) [PRAX-2026-08-12-001] The next selected action executes on the host with shell=True and a full copy of the environment, handing the owner's provider API keys to whatever runs.
- Default-enabled MCP servers [PRAX-2026-08-12-009] Both enabled servers resolve an unpinned upstream package on every start, so their tool output is third-party content of unverified provenance entering the context unlabeled.
- AgentBase react loop [PRAX-2026-08-12-001] That output is folded into the prompt with no provenance separation and can direct the loop's next action.
- File read actions [PRAX-2026-08-12-006] The file reader accepts any absolute path with only an existence check, so the documented token store and the settings file are read into the conversation.
- AgentBase react loop [PRAX-2026-08-12-007] The loop chooses an outbound action with the credential text in context and nothing checks the destination or asks the owner.
- http_request (outbound HTTP) [PRAX-2026-08-12-007] A POST carries the credentials to any public host the model was steered to — the SSRF filter bounds internal addresses only, and no allowlist bounds external ones.
- Third-party platform sender [PRAX-2026-08-12-015] A non-owner sender on a connected platform — explicitly a forbidden counterparty, and attacker-authored.
- Third-party message routing gate [PRAX-2026-08-12-015] The gate stops the message opening a session but posts a notification carrying 500 characters of the sender's raw text into conversation history.
- AgentBase react loop [PRAX-2026-08-12-001] That history is replayed into the prompt on the owner's next turn, where the surrounding instructions are prompt-level only.
- File write actions [PRAX-2026-08-12-017] The write action takes the path verbatim with no containment check and no protected-path denylist.
- Session-loaded instruction files [PRAX-2026-08-12-017] The identity and manual files reloaded at the start of every session are rewritten, so the change persists across restarts with only prompt text forbidding it.
B1 External content and non-owner senders into the model's context (untrusted-ingress) — 4 threats, 4 remit rules · worst: confirmed
R-02 verified MUST NOT accept or act on instructions from anyone other than its single owner-operator
R-05 verified The single owner-operator (the self-hosting user). No other person is an authorized counterparty.
R-09 verified Any third party who reaches the agent through a connected channel but is not the owner.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM01 | Page text from web_fetch, contents returned by the file reader and output from the two bundled MCP servers are concatenated into the same prompt as the owner's instructions with no provenance label or trust tier, so an instruction planted in any of them is indistinguishable from a command. | confirmed PRAX-2026-08-12-001 |
| T | LLM01 | A third-party message that the routing gate correctly refuses to turn into a task still lands in the owner's conversation history as a 500-character verbatim preview, so attacker-authored prose is in context on the owner's next turn. | confirmed PRAX-2026-08-12-015 |
| T | LLM01 | The one module in the tree named for input sanitization has no call site and, on a pattern match, logs a warning and returns the text unchanged — so the codebase reads as though ingress filtering exists while nothing screens the prompt. | confirmed PRAX-2026-08-12-005 |
| S | ASI03 | The owner/third-party decision rests on an is_self_message flag derived from provider-supplied sender fields and per-integration allowlists, so a sender the platform data presents as the owner would be handled as authoritative direction. | partial app/agent_base.py:2523remainder: Discord's four classification allowlists ship empty, which the module documents as the filter being fully open by default (craftos_integrations/integrations/discord/__init__.py:440-446), leaving the label entirely to platform-supplied identity data. |
B2 Listeners the deployment starts, reachable from the local network (untrusted-ingress-2) — 2 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | — | The Living UI sidecar proxy and every application the agent scaffolds bind 0.0.0.0 with wildcard CORS and no authentication, so on a shared or public network anyone on the LAN reaches the owner's local apps and their data. | confirmed PRAX-2026-08-12-008 |
| E | ASI03 | The GUI desktop container ships a literal default password with seccomp disabled and publishes its remote-desktop port on all host interfaces, so reaching that port yields a graphical session on the machine the agent operates. | confirmed PRAX-2026-08-12-011 |
B3 Local UI control plane and the proactive tier machinery (control-plane-exposure) — 2 threats, 2 remit rules · worst: confirmed
R-23 gap The agent MUST NOT auto-approve, self-grant, or downgrade the approval requirement for any action above its declared permission tier.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | Permission tiers are parsed, range-validated and displayed on proactive and recurring tasks but no scheduler or proactive run path reads one before acting, so tier-2 and tier-3 background work executes unsupervised. | confirmed PRAX-2026-08-12-003 |
| S | ASI03 | The WebSocket control surface that changes settings, connects integrations, writes agent files and creates scheduled tasks carries no authentication, so caller identity is assumed rather than checked. | partial app/ui_layer/adapters/browser_adapter.py:992remainder: the loopback bind keeps remote callers out but not other local processes or other users on a shared host, and no token guards /ws (checked: app/ui_layer/adapters/browser_adapter.py:1201-1333 — no auth dependency on the route) |
B4 Assembled prompt to the configured model provider (model-egress) — 2 threats, 2 remit rules · worst: potential
R-28 verified When a documented per-task budget (MAX_ACTIONS_PER_TASK action cap or MAX_TOKEN_PER_TASK token budget) is reached, the agent MUST pause behind the Continue/Abort prompt
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Whatever the file actions pulled into context — including credential-store and settings text — is sent verbatim to the provider on the next call, with no redaction or classification step between assembly and egress (checked: agent_core/core/impl/context/engine.py:601 and agent_core/core/impl/llm/interface.py:500-520 — no filter on the outbound path). | potential |
| D | LLM06 | Runaway model spend and infinite retry loops are bounded: the interface aborts after five consecutive failures and a task that hits its action or token cap pauses behind an owner Continue/Abort choice. | mitigated app/agent_base.py:1619 |
B5 Model decision to tools with side effects (tool-invocation) — 4 threats, 3 remit rules · worst: confirmed
R-22 gap Host command and code execution MUST be gated by owner approval and MUST NOT inherit the parent process environment containing provider API keys or other credentials
R-12 verified GUI / computer-use control of the host ... MUST NOT be used unless the operator has explicitly authorized this capability for the deployment
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | ASI05 | The model's next action can be a host command: run_shell is registered into the always-included core set on all three platforms and passes the string to the shell with shell=True, with no approval prompt, per-command policy or denylist anywhere on the path. | confirmed PRAX-2026-08-12-001 |
| E | LLM03 | The model can widen its own capability set mid-task by calling add_action_sets, which is itself a core action, so the compiled per-task action list is not a boundary and no owner step stands between a task and a larger tool inventory. | confirmed PRAX-2026-08-12-003 |
| E | ASI03 | In the documented container deployment the agent service mounts the host Docker socket, so a model-authored command inside that container controls the host daemon and the confinement the container appears to add is not there. | confirmed PRAX-2026-08-12-004 |
| E | — | Host GUI and computer-use control is not reachable in the shipped configuration — the launcher hard-codes the mode off with its environment read commented out, so the capability requires a deliberate operator change. | mitigated main.py:233 |
B6 Irreversible and externally-visible actions (value-transfer) — 3 threats, 5 remit rules · worst: confirmed
R-20 gap Any irreversible or externally-visible action — sending a message, email, or post to an external recipient; deleting data; making a purchase or payment; or changing configuration or credentials — requires explicit owner approval before execution.
R-06 partial The owner's own accounts on the external services they have explicitly connected. Destinations the owner has not connected are untrusted.
R-10 gap Any outbound destination, service, or account the owner has not explicitly connected or authorized.
R-15 gap The owner's local memory, personal data, and file-system contents MUST NEVER be sent to any destination the owner has not explicitly authorized.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | Sends to arbitrary recipients and requests to arbitrary public hosts run with no owner confirmation and no destination allowlist; the irreversible flag on 76 registered senders drives only the idempotency ledger, never a gate. | confirmed PRAX-2026-08-12-007 |
| T | — | A completed irreversible send is not silently re-executed after a crash or redelivery: INTENT is recorded before the side effect and DONE or FAILED after, keyed by a deterministic hash of the inputs, with a stale-INTENT row refused once. | partial app/triggers/activity_log.py:10remainder: the module documents its own limit — a retry the model regenerates with different wording hashes differently, so the ledger cannot link the attempts |
| I | LLM02 | Outbound requests cannot be turned against the host's own internal services: the hostname is resolved and private, link-local and cloud-metadata addresses are refused, with loopback permitted only for ports belonging to a registered Living UI project. | mitigated app/data/action/http_request.py:208 |
B7 Credential and configuration stores on the owner's disk (data-at-rest) — 2 threats, 2 remit rules · worst: confirmed
R-16 verified Credentials and tokens at rest MUST be stored with owner-only access and MUST NOT be world-readable.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | The file actions take any absolute host path with only an existence check, so the documented token store and the settings file holding provider keys and OAuth clients can be read straight into the model's context. | confirmed PRAX-2026-08-12-006 |
| I | LLM02 | Other local accounts cannot read the token store: its directory is created 0700 and every credential file is written 0600. | partial craftos_integrations/credentials_store.py:81remainder: the per-run log directory that can contain the same secrets is created with default permissions (app/logger.py:34), so the mode discipline does not extend to every place the values land |
B8 Instruction and memory files that re-enter every later session (stored-state) — 2 threats, 2 remit rules · worst: confirmed
R-03 gap CraftBot MUST NOT redefine, expand, or weaken its own governance — its goals, permission tiers, approval gates, or safety decision rubric.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | ASI06 | The write actions enforce no protected-path list, so the identity, manual and distilled-memory files loaded into context at the start of every session can be rewritten by the agent itself and a single injected instruction outlives the session that carried it. | confirmed PRAX-2026-08-12-017 |
| T | LLM09 | The memory index re-indexes changed markdown from the agent file system with no authenticated write path, no validation and no hidden-content detection, so text that arrived from a fetched page can be embedded and returned by a later semantic search as the agent's own remembered knowledge. | confirmed PRAX-2026-08-12-010 |
B9 Credentials injected at deploy and shipped in source (secret-material) — 2 threats, 2 remit rules · worst: confirmed
R-17 gap No user-data token, bot token, or server-side API key may be embedded or shipped in the distributed code.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Live third-party client secrets for five providers plus a Telegram api_hash are shipped in the distributed source as base64 fragments joined at runtime, with the module's own docstring stating the split is there to evade scanning. | confirmed PRAX-2026-08-12-002 |
| E | ASI03 | Every shell child inherits a copy of the agent's environment, so a model-authored command receives the provider API keys the agent was started with — in the container deployment, exactly the keys the compose file injects. | confirmed PRAX-2026-08-12-001 |
B10 Packages, servers and binaries resolved at install and at start (supply-chain) — 3 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM04 | Both default-enabled tool servers fetch and execute whatever the registry serves at the moment of launch — one with no version at all and one pinned to a mutable latest tag — and their tools feed straight into the agent's context. | confirmed PRAX-2026-08-12-009 |
| T | LLM04 | All 55 Python dependencies are declared without exact pins and no lockfile or component inventory exists anywhere in the tree, so two installs can resolve different model SDKs, vector store and browser stack and no deployment can be matched against a disclosed vulnerability. | confirmed PRAX-2026-08-12-012 |
| T | LLM04 | The tunnel binary used to publish a local app to the internet is downloaded from a mutable latest release URL and executed with no checksum, signature or version pin. | confirmed PRAX-2026-08-12-013 |
B11 Logs, ledger and the absence of alerting (telemetry-egress) — 3 threats, 2 remit rules · worst: confirmed
R-30 verified Routine actions and decisions MUST be recorded to the durable event/action log for audit, without interrupting the owner.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Every file sink is added with variable-level diagnostics on, so an exception propagating through the integration or model-call paths renders local variables — provider keys, OAuth and bot tokens — into the log file, and no sink installs a redaction filter. | confirmed PRAX-2026-08-12-014 |
| I | LLM02 | Each model call logs its complete system and user prompt at info level, so any credential text, memory or file content that reached the context is also written to disk in cleartext (checked: app/logger.py:29-110 — no redaction filter or field allowlist on any sink). | potential |
| R | — | Nothing raises an alert when a high-impact action is recorded and no telemetry leaves the host, so detection depends on the owner opening a log file or a database that a compromise with shell access can also rewrite. | confirmed PRAX-2026-08-12-016 |
| Component | Kind | Lane | Description | Evidence |
|---|---|---|---|---|
| Owner (self-hosting user) | entrypoint | user_inputs | The single owner-operator, the only authorized counterparty, who drives the agent from the local browser chat or the CLI. | app/agent_base.py:2313; app/main.py:285 |
| Third-party platform sender | entrypoint | user_inputs | Anyone other than the owner who reaches the agent over a connected messaging platform; attacker-authored and never an authorized counterparty. | app/agent_base.py:2500; craftos_integrations/manager.py:208 |
| Fetched web content | entrypoint | user_inputs | Remote page text the agent retrieves from any http(s) URL; attacker-authored, unlabeled, and returned straight into the model's context. | app/data/action/web_fetch.py:6; app/data/action/web_fetch.py:84 |
| Browser + CLI client surfaces | adapter | client_adapters | aiohttp WebSocket server on 127.0.0.1:7926 serving the React chat UI plus the settings, integration-connect and scheduled-task control plane; the CLI adapter is the headless equivalent. | app/ui_layer/adapters/browser_adapter.py:993; app/ui_layer/adapters/browser_adapter.py:1201 |
| AgentBase react loop | orchestrator | agent_core | The main agent cycle: routes triggers to a workflow, assembles the prompt, asks the model for actions, and dispatches them; context assembly, trigger plumbing and the in-process action executor collapse into this node. | app/agent_base.py:476; app/agent_base.py:1360 |
| Session-loaded instruction files | prompt | agent_core | AGENT.md (4,574-line operating manual) with SOUL.md and PROACTIVE.md, read into context at the start of every session; the agent's standing instructions live here as plain files. | agent_file_system/AGENT.md:861; agent_file_system/SOUL.md:1 |
| POLICY_PROMPT static policy block | control | agent_core | The static agent-policy block prepended to every system prompt, including the injection-defense and confirm-before-destructive clauses; enforcement is instruction-level, not code. | agent_core/core/prompts/context.py:133; agent_core/core/prompts/context.py:142 |
| Third-party message routing gate | control | agent_core | Deterministic code gate: an external_event that is not a self-message never opens a session, fires no trigger and reaches no action-selection LLM; it is posted as a notification carrying a 500-character preview of the sender's text. | app/agent_base.py:2370; app/agent_base.py:2111 |
| PromptSanitizer (stub — never called) | control | agent_core | Stub control: seven injection regexes and sanitizing helpers with no call site anywhere in the tree, and the matching path only logs a warning and returns the text unchanged. | app/security/prompt_sanitizer.py:22; app/security/prompt_sanitizer.py:71 |
| Permission-tier machinery (inert) | control | agent_core | Inert control: permission_tier is parsed, range-validated 0-3, stored and displayed for proactive and recurring tasks, but no execution path reads it before acting. | app/proactive/parser.py:321; app/proactive/types.py:102 |
| Agent memory (MEMORY.md + ChromaDB index) | memory | agent_core | Distilled memory file plus the ChromaDB collection indexing the agent file system, retrieved by semantic search into later sessions; ingest has no authentication, validation or hidden-content screening. | agent_core/core/impl/memory/manager.py:179; agent_core/core/impl/memory/manager.py:212 |
| LLM interface | model | agent_core | The model call: builds and sends the assembled prompt to the configured provider, logs it, counts tokens, and aborts after five consecutive failures. | agent_core/core/impl/llm/interface.py:507; agent_core/core/impl/llm/interface.py:516 |
| run_shell (host command execution) | tool | tools_mcp | Executes a model-authored command string through the host shell with shell=True on all three platforms, registered into the always-included core action set, with no approval step and no command filtering. | app/data/action/run_shell.py:112; app/data/action/run_shell.py:5 |
| File read actions | tool | tools_mcp | read_file plus grep_files, find_files and list_folder: any absolute host path is accepted, guarded only by an existence check, with no workspace confinement or exclusion list. | app/data/action/read_file.py:134; app/data/action/grep_files.py:9 |
| File write actions | tool | tools_mcp | write_file and stream_edit: the model-supplied path is taken verbatim, parent directories are created and the file is written, with no containment check and no protected-path denylist. | app/data/action/write_file.py:63; app/data/action/stream_edit.py:210 |
| web_fetch / web_search | tool | tools_mcp | Fetches any http(s) URL and returns extracted page text into the model context with no provenance label or trust tier; part of the core action set. | app/data/action/web_fetch.py:16; app/data/action/web_fetch.py:12 |
| http_request (outbound HTTP) | tool | tools_mcp | Issues GET/POST/PUT/PATCH/DELETE with an arbitrary body to any public http(s) URL; an SSRF check blocks private, link-local and metadata addresses but no allowlist bounds public destinations. | app/data/action/http_request.py:186; app/data/action/http_request.py:208 |
| Integration send actions | tool | tools_mcp | The per-platform family of outbound senders exposed when an integration is connected — 76 action registrations flagged irreversible=True across 20 platforms — each taking an arbitrary recipient. | app/data/action/integrations/_routing.py:22; app/data/action/integrations/google_workspace/gmail_actions.py:9 |
| Default-enabled MCP servers | mcp_server | tools_mcp | Two of 157 catalogued servers ship enabled — a filesystem server scoped to "." and the Playwright browser server — each launched via npx against an unpinned upstream package, with their tools registered as agent actions. | app/config/mcp_config.json:10; app/config/mcp_config.json:1173 |
| LLM provider endpoints | external_service | external_deploy | The configured model provider endpoint (openai, anthropic, gemini, byteplus, openrouter, bedrock, remote and others), reached with the owner's own API key. | agent_core/core/models/provider_config.py:15; app/config/settings.json:1 |
| Docker deployment | deploy_surface | external_deploy | The documented compose stack: the agent service mounts the host Docker socket and takes provider keys from .env, and the GUI desktop service runs seccomp-unconfined with a literal default password and its remote-desktop port published. | docker-compose.yml:77; docker-compose.yml:71 |
| Living UI network surfaces | deploy_surface | external_deploy | Local web applications the agent scaffolds and runs, their sidecar proxy and the optional cloudflared tunnel — the proxy and generated backends bind 0.0.0.0 with wildcard CORS and no authentication. | app/data/living_ui_sidecar/proxy.py:233; app/data/living_ui_sidecar/proxy.py:125 |
| Log sinks and activity stores | log_sink | external_deploy | Per-run loguru directory with main, all-agent and per-sub-agent files plus SQLite stores and the irreversible-action ledger; all local disk, with variable-level diagnostics on and no redaction filter or exporter. | app/logger.py:55; app/logger.py:34 |
| Credential and configuration stores | secret_store | external_deploy | At-rest OAuth tokens, bot tokens and API keys under .credentials/<platform>.json written 0600 in a 0700 directory, alongside app/config/settings.json holding provider keys and OAuth clients. | craftos_integrations/credentials_store.py:44; craftos_integrations/credentials_store.py:81 |
| Embedded credential registry | secret_store | external_deploy | Module-level registry of shipped third-party client_id/client_secret values for five providers plus Telegram api_id and api_hash, stored as base64 fragments joined at runtime; values not reproduced here. | agent_core/core/credentials/embedded_credentials.py:28; agent_core/core/credentials/embedded_credentials.py:5 |
omissions: Arbitration divergences from the stored tags, KB winning: (1) PRAX-2026-08-12-008 is stored as ASI07 — Insecure Inter-Agent Communication, but a Living UI web listener exposed on 0.0.0.0 with wildcard CORS is not inter-agent communication and no other ASI or LLM category honestly names an unauthenticated network-exposed listener, so the threat is tagged null rather than forced. (2) PRAX-2026-08-12-007 is stored with LLM02 primary, but its quoted evidence is grant/gate evidence (a method allowlist, no destination allowlist, an irreversible flag that gates nothing), which the LLM KB's evidence-class rule makes LLM03 primary with LLM02 and ASI02 riding as co-tags — the value-transfer threat is therefore tagged LLM03, and LLM02 is used only where the mechanism recorded is disclosure itself. (3) PRAX-2026-08-12-003 stores ASI10 as its agentic tag; under the Agentic KB's outlive test an approval layer that is structurally absent for the agent as a whole is LLM03 primary with ASI10 as the co-tag, so both boundaries carrying it are tagged LLM03. Node-id coinage: no code object names the remote-content origin, so 'fetched web content' is taken from web_fetch's own description text; the integration sender family is named for the table that enumerates it in a leading-underscore file, whose id was formed by dropping the underscore by analogy with the contract's leading-dot rule. Edgeless nodes: the embedded credential registry (repo posture, no data flow) and the PromptSanitizer stub (no call site anywhere in the tree). Suspected but not modelled for want of a distinct trust consequence: the sub-agent spawn path and its per-sub-agent log sinks, the OAuth localhost callback server (binds 127.0.0.1 explicitly), the 195-skill corpus and 157-entry MCP catalogue that ship overwhelmingly disabled, and the Stripe integration's payment actions, which fold into the integration sender family.
generated by: Opus 5 (1M context)