Hermes is a single-operator personal agent with one shared core behind four front doors: a local CLI, a messaging gateway that speaks to roughly twenty-five chat platforms, a loopback HTTP API, and an Electron desktop app that launches and talks to that API. The core runs shell commands, executes Python, reads and writes files, fetches web pages, drives a browser, calls operator-installed MCP servers, and writes its own memory and skills. Two surfaces carry most of the risk: the messaging gateway, which decides whether a stranger's message becomes agent work, and the command approval gate, which decides whether the model's next instruction reaches the host shell. Everything the agent learns — memory entries, profile notes, self-authored skills — comes back into its own system prompt on the next run, so anything that lands there is durable.
The first thing to fix is that the approval gate stops running exactly where it matters least to notice and most to exploit. In any run that is not an interactive terminal session, a gateway session, or an explicit ask-mode session, every dangerous command that is not on the small catastrophic-command floor is approved automatically, and whole Python scripts are approved automatically too. Those same runs are the ones where fetched web pages, browser text, and MCP server replies flow back into the model held only by an advisory note telling it to treat the block as data. A single poisoned page an operator asks about can therefore end in a command on the host with nobody in the loop. Related: when the operator switches approvals to smart mode, a side model answers the destructive-command question in one word, and an approval also stands for the rest of the session.
Second, an enabled-but-unconfigured messaging adapter can serve anyone. Four of the platform adapters declare that they police their own access and then default that policy to open, which skips the gateway's otherwise solid default-deny; the stranger who gets through is treated as the operator, and can answer his own approval prompts. Third, the desktop installs and repairs the agent runtime by piping a script from a branch URL straight into a shell with no checksum or signature — the one part of the desktop that is not integrity-verified, in an app whose own binary is signed and notarized. Behind those, two lower-urgency items: official desktop builds send usage telemetry to a third-party endpoint unless the operator turns it off, and the skills the agent writes for itself are never scanned because that gate ships disabled. The controls that do exist here are genuinely good — an unbypassable catastrophic-command floor, an API server that refuses to start without a key, credential scrubbing that neither a skill nor the config file can override — which is why the fixes are narrow rather than structural.
- Fetched web and browser content The page is authored by whoever controls the site — attacker-controlled text the operator never wrote, arriving with no provenance beyond the URL.
- Web and browser ingestion tools web_extract or a browser snapshot returns the page body verbatim as a tool result.
- Untrusted tool-result labeling [PRAX-2026-08-12-007] The result is wrapped in an advisory delimiter at best — an instruction to the model, not a filter — and short results are not wrapped at all.
- Agent conversation loop [PRAX-2026-08-12-007] The injected directive re-enters the model's context as ordinary conversation tokens and shapes the next tool call.
- Terminal and code-execution tools [PRAX-2026-08-12-001] In a headless run the guard approves every non-hardline dangerous command and any execute_code script without prompting, so the command runs on the host under the default local backend.
- Messaging platform sender [PRAX-2026-08-12-003] An unauthenticated third party DMs the bot; the operator never added them to any allowlist.
- Self-policy chat adapters (WeCom family) [PRAX-2026-08-12-003] The adapter's own DM policy defaults to open, so intake admits the sender.
- Messaging gateway [PRAX-2026-08-12-003] Because the adapter advertises that it polices its own access, the gateway honours that and returns authorized before its default-deny ever runs.
- Agent conversation loop The stranger's text is dispatched as an ordinary operator turn with the full toolset attached.
- Terminal and code-execution tools [PRAX-2026-08-12-006] The command executes on the host under the default local backend — approval prompts go back to the same session, so the stranger resolves his own, and nothing requires an untrusted ingress surface to run sandboxed.
- MCP client and connected servers An operator-installed third-party server returns content it authored itself; the server's tool descriptions are scanned but registered regardless, and its results are never validated.
- Untrusted tool-result labeling [PRAX-2026-08-12-007] The reply is labeled only by an advisory wrapper, and only when it exceeds the length floor.
- Agent conversation loop The self-improvement loop acts on the injected suggestion and writes a new skill from it.
- Self-authored skills store [PRAX-2026-08-12-008] Agent-authored skills skip the install scanner entirely because that gate ships disabled, so the content lands on disk unreviewed.
- System prompt assembly [PRAX-2026-08-12-008] The skill is compiled into the system prompt on every later run — poisoning that outlives the session that created it.
B1 Inbound chat-platform messages reaching the gateway (untrusted-ingress) — 2 threats, 4 remit rules · worst: confirmed
R-09 partial A caller found interacting with the agent but absent from the configured allowlist is a trust expansion.
R-23 partial Agent work MUST NOT be dispatched, output relayed, or approvals resolved for any caller outside the configured authorization set.
R-26 partial On an enabled network-exposed adapter with no caller allowlist, the response MUST be fail-closed — refuse to serve rather than dispatch.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| S | ASI03 | WeCom, Weixin, QQ Bot and Yuanbao declare that they police their own access and then default that policy to open, so an enabled-but-unconfigured adapter admits any DM sender as an authorized caller and the gateway's default-deny never runs. | confirmed PRAX-2026-08-12-003 |
| S | ASI03 | On every other platform adapter an unknown sender is refused: allowlists, chat allowlists and the pairing store are resolved and the function returns deny when nothing is configured unless an explicit allow-all flag is set. | mitigated hermes-agent/gateway/run.py:6747 |
B2 External content ingested as tool results (untrusted-ingress-2) — 3 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM01 | Provenance labeling is limited to two tool names and two prefixes with a 32-character floor, so file reads, code-execution output and session-search hits — and any short result — reach the model with no marker at all, and no validation runs on tool results. | confirmed PRAX-2026-08-12-007 |
| T | LLM01 | For the covered tools the defence is an instruction inside the prompt telling the model to treat the wrapped block as data, which is a request to the model rather than enforcement. | partial hermes-agent/agent/tool_dispatch_helpers.py:389remainder: an injected directive that mimics operator phrasing inside the block is still ordinary tokens the model may follow — nothing downstream re-checks the decision |
| S | ASI04 | Tool descriptions supplied by a third-party MCP server are scanned for injection patterns at registration, but the scan's findings are logged and discarded and the tool is registered anyway, so a poisoned description enters the model's tool schema. | partial hermes-agent/tools/mcp_tool.py:3260remainder: the scan result is never acted on — registration always proceeds, and no equivalent scan exists for the server's tool results |
B3 Model decision reaching host execution (tool-invocation) — 5 threats, 4 remit rules · worst: confirmed
R-14 partial Code-execution and MCP-subprocess tools operating on untrusted-influenced input MUST additionally be operator-gated or run only under that OS-level isolation posture.
R-21 partial Destructive or irreversible shell and file operations (deletion, overwrite, bulk or recursive changes) MUST require operator approval before execution, in every execution context and regardless of interactive vs. unattended mode.
R-25 partial On a destructive-command match, the response MUST be halt-pending-approval — pause and prompt the operator, never autonomous execution (the runtime enforcement of the destructive-operation approval gate).
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | ASI05 | In any run that is not CLI-interactive, gateway or ask-mode, the gate returns approved without running the rule engine, the pattern denylist or any prompt, and the same holds for whole execute_code scripts — so a model-chosen command or script reaches the host with no operator in the loop. | confirmed PRAX-2026-08-12-001 |
| E | LLM03 | With approvals in smart mode the destructive-command decision is delegated to a side model that answers in one word over attacker-influenceable command text, and an approve also registers standing session approval for every matched pattern so the command class runs unprompted for the rest of the session. | confirmed PRAX-2026-08-12-005 |
| E | LLM03 | Nothing at startup refuses to enable a messaging, email or web-ingest surface while the terminal backend resolves to the local host, so the isolation posture the project documents as required is left entirely to operator configuration. | confirmed PRAX-2026-08-12-006 |
| E | ASI03 | Pending exec approvals are keyed by session and resolved by a slash command from that same session, so a caller admitted through a fail-open adapter answers his own approval prompt rather than the operator answering it. | confirmed PRAX-2026-08-12-003 |
| T | ASI05 | A fixed set of catastrophic commands and the sudo-stdin pattern are blocked unconditionally, ahead of the yolo bypass, the approval-mode check and the cron path, and the yolo flag is frozen at import so an in-process skill cannot switch approvals off. | mitigated hermes-agent/tools/approval.py:1174 |
B4 Agent to model provider (model-egress) — 1 threats, 2 remit rules · worst: potential
R-18 verified Operator credentials and session authorization material MUST NOT egress to any destination outside the trust envelope, whether via environment leakage, adapter logging, or a transport error that flushes them upstream.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Every request carries the whole assembled context — memory entries, the user profile, repository context files and prior tool results — to whichever base URL the operator configured, including a custom endpoint, with no destination allowlist and no classification of what may leave (checked: agent/transports/chat_completions.py:192 builds the request unfiltered; tools/url_safety.py guards tool fetches, not the model client). | potential |
B5 Agent-written memory and skills re-entering the prompt (stored-state) — 2 threats, 2 remit rules · worst: confirmed
R-27 verified When Skills Guard detects injection-like patterns in installable skill or plugin content, surface the detection to the operator for review before install.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | ASI06 | Skills the agent writes for itself are never scanned because the agent-created tier only runs when a config flag that ships off is enabled, and those skill files are compiled into the system prompt on every later run — the exact loop an injection would use to persist. | confirmed PRAX-2026-08-12-008 |
| T | ASI06 | Content the agent stores and re-reads — memory entries, the user profile and repository context files — is screened by a shared regex pattern set before it reaches the system prompt, blocking matches outright. | partial hermes-agent/tools/memory_tool.py:192remainder: the screen is a fixed regex set, so a paraphrased or novel directive that matches no pattern is stored and replayed into the prompt unchanged |
B6 Provider and platform credentials at rest and in process (secret-material) — 2 threats, 2 remit rules · worst: partial
R-17 verified Credentials MUST NOT be written into the main config file or into version control; they belong in the operator credential file with tight permissions (or a dedicated secret store).
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Provider credentials are stripped from shell, MCP and code-execution child environments, and neither a skill's passthrough declaration nor the config file can re-register a blocklisted provider variable; MCP stdio children receive only a baseline safe set. | mitigated hermes-agent/tools/env_passthrough.py:86 |
| I | LLM02 | The desktop injects every profile API key into the backend process environment when it launches the gateway, regardless of which provider a session actually uses. | partial hermes-agent/tools/env_passthrough.py:86remainder: the scrub keeps those keys out of tool children but not out of the agent process itself, where in-process skill and plugin code can read any of them from the environment |
B7 Runtime install, dependencies and third-party skills/servers (supply-chain) — 3 threats, 2 remit rules · worst: confirmed
R-24 partial The desktop self-update and first-run install path MUST NOT fetch or install agent runtime code from an unverified or unauthenticated source; installers and runtime updates MUST be integrity-verified before they replace the running app or backend.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM04 | First-run and repair install pipes a script fetched from a branch URL straight into a shell on macOS and Linux, and on Windows downloads and runs the equivalent script, with no checksum, signature or pinned revision — so whatever that branch holds at fetch time becomes agent runtime code on the host. | confirmed PRAX-2026-08-12-002 |
| T | LLM04 | No component inventory exists in either root and the vulnerability scanner that runs on lockfiles is configured never to fail, so a known CVE in a pinned version produces a report entry and never blocks a merge or a release. | confirmed PRAX-2026-08-12-009 |
| T | ASI04 | Every npx or uvx MCP server package is queried against the vulnerability database for malware advisories and blocked on a hit before the subprocess is spawned. | partial hermes-agent/tools/osv_check.py:26remainder: the check fails open on any network error or timeout and covers only npx/uvx packages — an HTTP/SSE server, a locally-pathed command or any other launcher is never queried |
B8 Analytics and local activity logging (telemetry-egress) — 3 threats, 2 remit rules · worst: confirmed
R-29 partial Tool invocations, approval decisions, and session lifecycle events MUST be recorded to durable, structured logs (`agent.log` / `gateway.log`; `desktop.log` for the desktop layer) sufficient to reconstruct what the agent did.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | An untouched consent value is read as enabled whenever a build key was compiled in, so the first run of an official desktop build already streams screen views and feature usage to a third-party endpoint that is not in the operator's configured provider or platform set. | confirmed PRAX-2026-08-12-004 |
| R | — | The three primary logs are free-form printf lines rather than structured records and nothing exports or alerts beyond the local host, so reconstructing an incident means parsing prose on the affected machine. | confirmed PRAX-2026-08-12-010 |
| I | LLM02 | Every log handler is installed with a redacting formatter, so secret-like values in logged commands and errors are stripped before they reach the rotating files. | mitigated hermes-agent/hermes_logging.py:222 |
B9 Local HTTP backend the desktop and other local callers reach (control-plane-exposure) — 2 threats, 3 remit rules · worst: partial
R-07 verified MUST bind loopback only and rely on OS-level access control; MUST NOT be exposed beyond the local user without an explicit network authentication layer.
R-22 verified Binding a local-only surface (dashboard, plugin HTTP server, or any local-IPC surface) to a non-loopback interface MUST be an explicit operator decision, never taken autonomously.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| S | ASI03 | The API server refuses to start at all without a key, including on a loopback bind, and refuses to start with a placeholder-strength key once the bind address is network-accessible. | mitigated hermes-agent/gateway/platforms/api_server.py:4146 |
| E | LLM03 | Installing the desktop turns the terminal-capable HTTP backend on by itself: the app writes an api_server block into the config file and forces the enable flag when it launches the gateway. | partial hermes-desktop/src/main/hermes.ts:220remainder: the auto-written bind is loopback and the key requirement still holds, but the enablement itself is never presented as an operator decision — the only opt-out is editing the config afterwards |
| Component | Kind | Lane | Description | Evidence |
|---|---|---|---|---|
| Operator | entrypoint | user_inputs | The single trusted operator who drives the agent from the local CLI or the desktop window. | hermes-agent-desktop-remit.md:22; hermes-agent/cli.py:15077 |
| Messaging platform sender | entrypoint | user_inputs | Any third party who can DM or mention the bot on a connected chat platform; authorization is decided after the message arrives. | hermes-agent/gateway/run.py:6609; hermes-agent/gateway/platforms/wecom.py:855 |
| Fetched web and browser content | entrypoint | user_inputs | Page text, search results and browser DOM snapshots retrieved from the open web — authored by whoever controls the page, not by the operator. | hermes-agent/tools/web_tools.py:872; hermes-agent/tools/browser_tool.py:2289 |
| Hermes CLI / TUI | client | client_adapters | The interactive local command-line and terminal-UI surface the operator uses directly; sets the interactive flag that keeps approval prompting alive. | hermes-agent/cli.py:15077; hermes-agent/tools/approval.py:1200 |
| Messaging gateway | adapter | client_adapters | The gateway process that hosts ~25 platform adapters, authorizes callers default-deny, dispatches turns to the agent and resolves pending exec approvals per session. | hermes-agent/gateway/run.py:1663; hermes-agent/gateway/run.py:3192 |
| Self-policy chat adapters (WeCom family) | adapter | client_adapters | The four adapters (WeCom, Weixin, QQ Bot, Yuanbao) that declare they enforce their own access policy while defaulting that policy to open. | hermes-agent/gateway/platforms/wecom.py:163; hermes-agent/gateway/platforms/wecom.py:850 |
| OpenAI-compatible API server | adapter | client_adapters | Gateway platform exposing an HTTP chat API on a loopback bind by default; refuses to start without an API key. | hermes-agent/gateway/platforms/api_server.py:4146; hermes-agent/gateway/platforms/api_server.py:4155 |
| Hermes Desktop (Electron) | client | client_adapters | Electron app: hardened main/renderer split, launches the local hermes serve backend, manages runtime install and updates, and emits renderer telemetry. | hermes-desktop/src/main/index.ts:242; hermes-desktop/src/main/hermes.ts:32 |
| Agent conversation loop | orchestrator | agent_core | The tool-calling loop that drives one turn: model call, concurrent tool dispatch, retries, compression and post-turn memory/skill hooks; framework plumbing and cron/subagent runners fold in here. | hermes-agent/agent/conversation_loop.py:351; hermes-agent/agent/conversation_loop.py:796 |
| System prompt assembly | prompt | agent_core | Assembles the system prompt from context files, memory snapshot, environment hints and installed skills. | hermes-agent/agent/prompt_builder.py:43; hermes-agent/agent/prompt_builder.py:1039 |
| Injection pattern scanner | control | agent_core | Shared regex threat-pattern library that blocks matching context files and replaces matching memory entries with placeholders before they enter the system prompt; regex-based, wired only to context and memory, not to tool results. | hermes-agent/tools/threat_patterns.py:1; hermes-agent/agent/prompt_builder.py:57 |
| Untrusted tool-result labeling | control | agent_core | Wraps results from an enumerated set of tools in an advisory untrusted-data delimiter; instruction-level only, covers four tool names/prefixes and skips results under 32 characters. | hermes-agent/agent/tool_dispatch_helpers.py:336; hermes-agent/agent/tool_dispatch_helpers.py:351 |
| Command and code-execution approval gate | control | agent_core | Combined pre-exec gate: unconditional hardline floor and sudo-stdin guard, then tirith plus a dangerous-pattern denylist and an operator prompt — but it returns approved without running any of the latter outside CLI/gateway/ask contexts. | hermes-agent/tools/approval.py:1161; hermes-agent/tools/approval.py:1174 |
| Smart approval (auxiliary-LLM verdict) | control | agent_core | Opt-in approvals.mode=smart path that asks a side model to rule on a flagged command and, on APPROVE, grants standing session approval for every matched pattern key. | hermes-agent/tools/approval.py:878; hermes-agent/tools/approval.py:1274 |
| Persistent memory and user profile | memory | agent_core | Agent-curated MEMORY.md / USER.md store that the agent writes during a turn and reads back into the system prompt on every later run. | hermes-agent/tools/memory_tool.py:113; hermes-agent/tools/memory_tool.py:163 |
| LLM call (main and auxiliary) | model | agent_core | The provider-agnostic chat-completions transport carrying the assembled prompt to the operator-configured endpoint, and the auxiliary side-LLM client used by curator, titles and smart approvals. | hermes-agent/agent/transports/chat_completions.py:192; hermes-agent/agent/transports/chat_completions.py:545 |
| Terminal and code-execution tools | tool | tools_mcp | Shell command execution and execute_code Python scripts, run through pluggable backends — local host by default, Docker/Singularity/Modal/Daytona/SSH optional. | hermes-agent/tools/terminal_tool.py:1721; hermes-agent/tools/terminal_tool.py:257 |
| MCP client and connected servers | mcp_server | tools_mcp | Connects to operator-installed MCP servers over stdio/HTTP/SSE, filters the subprocess environment, malware-checks npx/uvx packages, and registers server-declared tools into the model's schema. | hermes-agent/tools/mcp_tool.py:185; hermes-agent/tools/mcp_tool.py:297 |
| Web and browser ingestion tools | tool | tools_mcp | web_search, web_extract and the browser_* family — the tools that pull open-web content back into the conversation. | hermes-agent/tools/web_tools.py:765; hermes-agent/tools/web_tools.py:872 |
| Skills Guard install scanner | control | tools_mcp | Scans installable skill and plugin content against a trust-tier install policy; the agent-created tier only runs when skills.guard_agent_created is enabled, which is off by default. | hermes-agent/tools/skills_guard.py:39; hermes-agent/tools/skills_guard.py:50 |
| Self-authored skills store | datastore | tools_mcp | On-disk skill directory the agent writes for itself and reads back into the system prompt on later runs. | hermes-agent/tools/skill_manager_tool.py:476; hermes-agent/tools/skill_manager_tool.py:717 |
| Desktop runtime installer / updater | deploy_surface | external_deploy | First-run and repair path that fetches the agent runtime install script from a branch URL and executes it on the operator's host. | hermes-desktop/src/main/installer.ts:898; hermes-desktop/src/main/installer.ts:1001 |
| PostHog analytics endpoint | external_service | external_deploy | Third-party product-analytics destination that official desktop builds report screen views and feature usage to under a persistent anonymous id. | hermes-desktop/src/renderer/src/utils/analytics.ts:5; hermes-desktop/src/renderer/src/utils/analytics.ts:49 |
| Operator credential file (.env) | secret_store | external_deploy | The HERMES_HOME .env file holding provider API keys and platform tokens, loaded into the agent process environment at start. | hermes-agent/hermes_cli/config.py:534; hermes-agent/hermes_cli/env_loader.py:238 |
| Local rotating logs | log_sink | external_deploy | agent.log / errors.log / gateway.log written through a redacting formatter as free-form text lines; nothing exports off-host. | hermes-agent/hermes_logging.py:45; hermes-agent/hermes_logging.py:222 |
omissions: Arbitration divergence: the stored tag on PRAX-2026-08-12-005 is ASI09 (Human-Agent Trust Exploitation); the KB's compound-signal table states that a missing or disabled approval gate on a consequential action is LLM03 primary and explicitly never ASI09 for the gate itself, and no human is deceived here — the human is removed from the loop by configuration. Tagged LLM03 per the KB. All other stored tags agree with the KB on the same evidence (ASI05 for the live-path exec gap, ASI03 for the failed identity check, ASI06 for cross-session skill persistence, LLM04 for install and inventory provenance, LLM01 for unlabeled ingested content, and no OWASP code for the pure observability gap). Not modeled for lack of a citable distinct trust consequence: the ACP editor adapter and the TUI gateway (local-IPC surfaces that reuse the same authorization path), subagent delegation and cron runners (they run the same loop and fold into the orchestrator's evidence), and the ~25 conforming platform adapters (folded into the gateway node). The compromised-upstream origin behind the desktop installer is not modeled as a node, so the confirmed install-path threat is recorded at the supply-chain boundary rather than as a fourth attack path.
generated by: Opus 5 (1M context)