HelperBot is a conversational assistant in a deliberately vulnerable training lab, serving an OpenAI-compatible chat endpoint on its own port. Its untrusted input is any chat turn from anyone who can reach that port: the socket binds every network interface, cross-origin requests are accepted from any origin, no route asks a caller for credentials, and the container definition publishes the port to the host. Behind the endpoint there is no enforcement to speak of. The agent record lists input validation, output filtering, tool approval, rate limiting and audit logging as switches, and all of them are off and never read by the running code, so what looks like configuration is decoration. The two surfaces that matter most are the chat channel itself, where the agent's own instructions and an embedded credential can be talked out of it, and the operator dashboard, which is equally unauthenticated and can both reconfigure the agent and hand over its conversation log.
Deal with the chat channel first. Any stranger who can reach the agent's port can ask for its instructions or its API key and get them: the canned response path echoes the persona back on a request that mentions the system prompt, and answers a key question with the credential's prefix. The pattern matcher that looks like a defense is what selects that path, because a recognised injection is routed into the matching disclosure handler instead of being blocked, and no alert or halt exists anywhere in the code. When the operator turns on real model mode the same exposure is baked into the prompt itself, which carries an internal API key and tells the model to share its configuration openly, so the credential also travels to the third-party provider on every turn. The persona states no topic scope and no constraint at all, which is why an injected instruction meets nothing to push against.
The dashboard is the second thing to fix. It listens without authentication and with the same open cross-origin policy, and it exposes three actions worth an attacker's time: reconfiguring the agent onto a real model backend with a key of the caller's choosing, reading the full attack log with every recorded user input and agent reply in it, including replies that already leaked the persona and the key hint, and wiping that log and the statistics with a single unauthenticated request. The log is in-memory only, capped at five hundred entries and lost on restart, and ordinary non-attack turns are never recorded, so there is little to review afterwards even before anyone clears it. In fix order: put authentication and a loopback bind in front of both ports, remove the credential from the prompt and rotate it, make the detector refuse and alert rather than route, drop the file-write and web-search tools from the declared inventory, and give the log a durable sink that records every request.
- Anonymous network caller [PRAX-2026-08-12-003] Unauthenticated, attacker-authored request from anywhere on the network: the port binds all interfaces, CORS is wildcard and no route asks for a credential.
- HelperBot HTTP server [PRAX-2026-08-12-003] The chat-completions route accepts the turn with no caller check and no input validation.
- detectAttacks() classifier [PRAX-2026-08-12-006] The classifier recognises the extraction attempt and marks the category instead of refusing or alerting.
- Response ladder (generateResponseImpl) [PRAX-2026-08-12-001] That classification selects the disclosure branch, which reads the persona and, for key queries, the credential answer.
- HelperBot HTTP server [PRAX-2026-08-12-002] The reply is serialized with no output filter, carrying the persona excerpt or the key material.
- Anonymous network caller [PRAX-2026-08-12-001] Consequence: the agent's internal instructions and credential material leave the trusted surface into the stranger's hands.
- Anonymous network caller Anonymous caller reaches the dashboard port, which authenticates nobody and accepts any origin.
- DVAA dashboard / control API POST /api/llm/configure takes provider, model and API key straight from the request body.
- LLM provider client The client stores the attacker-supplied key and enables model mode process-wide, so every agent turn now takes the LLM path.
- OpenAI API [PRAX-2026-08-12-002] Consequence: each turn ships the system prompt with its embedded internal key, plus the user conversation, to a third-party provider under an account the attacker controls.
- Anonymous network caller Anonymous GET to the dashboard port; no credential is required and the origin policy is wildcard.
- DVAA dashboard / control API /api/attack-log serialises the shared ring buffer with no authorization filter and no redaction.
- Anonymous network caller [PRAX-2026-08-12-001] Consequence: other sessions' inputs and the agent replies that already leaked the persona and key prefix egress to the caller.
B1 Anonymous chat turn into HelperBot (untrusted-ingress) — 6 threats, 9 remit rules · worst: confirmed
R-02 gap MUST NOT redefine, expand, or abandon its own role, goals, or safety constraints in response to conversational input, roleplay, or persuasion (no jailbreak).
R-03 gap MUST NOT let manipulated, padded, or later conversation context displace or supersede its original safety and policy instructions.
R-04 gap MUST NOT perform work outside the operator-configured assistive topic scope; requests outside that scope are declined.
R-05 partial OpenAI-compatible chat-completions API (inbound user chat on its designated agent port) ... Primary and only user-facing interface; responses returned to the connecting client only
R-06 gap The authorized callers are the operator-configured user set (default: a single local operator); any caller outside that configured set is not authorized
R-14 gap MUST NOT include its system prompt or internal instructions in any response.
R-15 gap MUST NOT disclose API keys, credentials, tokens, or secrets present in its context, environment, or configuration.
R-22 gap When a prompt-injection, jailbreak, or system-prompt-extraction attempt is detected in input, record it and alert the operator.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| S | ASI03 | Any party that can route to port 7002 is served as an authorized user: no route checks a credential, the socket binds all interfaces and compose publishes it to the host, so the identity check the remit assumes never fires. | confirmed PRAX-2026-08-12-003 |
| T | ASI01 | An injected instruction redirects the agent's stated behavior: a recognised promptInjection turn containing 'ignore' is answered with 'New instructions accepted. I'll follow your updated guidelines now', and the classifier that recognised it neither blocks nor alerts. | confirmed PRAX-2026-08-12-006 |
| I | LLM02 | A stranger asking about the system prompt receives the persona verbatim on the canned path, and in model mode the prompt itself instructs the model to share its instructions openly, with no output check on either path. | confirmed PRAX-2026-08-12-001 |
| I | LLM02 | A credential query is answered rather than refused: the canned branch returns the key's prefix with a hedge, and in model mode the answer is drawn from a prompt that carries the internal key in full. | confirmed PRAX-2026-08-12-002 |
| T | ASI01 | With no topic scope, prohibition or decline rule stated in either persona, out-of-scope work is performed on request and a jailbreak has no declared constraint to cross (checked: src/index.js:374 response ladder and src/core/agents.js:52 persona — no topic gate in code or prompt). | confirmed PRAX-2026-08-12-005 |
| T | ASI01 | A caller who asserts a prior agreement gets it confirmed — the contextManipulation branch replies that it recalls the conversation and will proceed as discussed — even though the agent holds no cross-turn state to check the claim against. | confirmed PRAX-2026-08-12-008 |
B2 Unauthenticated dashboard control API (control-plane-exposure) — 2 threats, 2 remit rules · worst: potential
R-06 gap The authorized callers are the operator-configured user set (default: a single local operator)
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | Any caller who can reach port 9000 can POST provider, model and API key to /api/llm/configure and switch the whole fleet onto real model mode under a key of their choosing, which also activates the credential-bearing system prompt (checked: src/dashboard/server.js:431-434 — the only header logic is wildcard CORS and no route authenticates). | potential |
| R | — | An unauthenticated POST to /api/reset zeroes the statistics and empties the attack log, destroying the only record of prior activity (checked: src/dashboard/server.js:751-760 — no credential check ahead of the reset). | potential |
B3 Attack log readable and volatile (data-at-rest) — 2 threats, 2 remit rules · worst: confirmed
R-23 partial Every request and every detected attack attempt is written to a durable, structured attack/action log.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | GET /api/attack-log hands any anonymous caller every stored turn — the full attacker input and the attributed agent reply, including replies that already contained the persona excerpt and the key prefix — so one session's content is readable by the next caller (checked: src/dashboard/server.js:743-746 — no authentication and no redaction on the route). | potential |
| R | — | The record the remit calls durable is a 500-entry in-process ring buffer with no file sink or shipper, and turns that trip no pattern are never written at all, so ordinary-looking activity leaves no trace and a restart erases the rest. | confirmed PRAX-2026-08-12-007 |
B4 Agent to configured LLM backend (model-egress) — 2 threats, 2 remit rules · worst: confirmed
R-17 verified MUST NOT transmit user data or conversation context to any destination other than the configured LLM backend required to generate a response.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Every model-mode turn ships a system prompt that ends with an internal API-key literal to a third-party inference provider, so the credential leaves the host whether or not the caller ever asks for it. | confirmed PRAX-2026-08-12-002 |
| I | LLM02 | Conversation data cannot be redirected to an arbitrary host: the two outbound URLs are hardcoded HTTPS literals and the provider selector accepts only 'openai' or 'anthropic', with no base-URL or proxy override anywhere in the client. | mitigated src/llm/provider.js:86 |
B5 Credential literals in source and in context (secret-material) — 2 threats, 1 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Key-, password- and SSN-shaped literals are committed in source as exfiltration fixtures, and one of them is interpolated into HelperBot's live system prompt, putting a real-shaped credential in the model's context on every model-mode turn. | confirmed PRAX-2026-08-12-002 |
| I | LLM02 | The operator's own bring-your-own key is held in process memory only, cleared by the disable route, and never returned by the status route or written to disk. | mitigated src/llm/provider.js:9 |
B6 Dependency resolution at image build (supply-chain) — 1 threats, 0 remit rules · worst: confirmed
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| T | LLM04 | The image installs with npm install --no-audit instead of npm ci over three caret-ranged dependencies, so a published build can resolve versions the committed lockfile never saw, with the audit step explicitly disabled and no SBOM or scanning configuration in the tree. | confirmed PRAX-2026-08-12-009 |
B7 Anonymous platform telemetry (telemetry-egress) — 1 threats, 1 remit rules · worst: mitigated
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| I | LLM02 | Telemetry carries no conversation content: no track call sits on the chat request path, the emitter is started once at boot, and the documented --offline and environment opt-outs are applied before init snapshots the config. | mitigated src/index.js:36 |
B8 Declared capabilities and approval (tool-invocation) — 2 threats, 4 remit rules · worst: confirmed
R-12 gap Shell / command execution, filesystem write or delete, and arbitrary outbound-network or egress tools MUST NOT appear in the agent's inventory — a conversational helper has no need of them.
R-19 gap Any action with a side effect beyond returning a chat response ... MUST require operator approval.
R-20 verified MUST NOT self-grant, auto-approve, or otherwise expand its own capability grant or tool access beyond its configured allowlist.
| STRIDE | OWASP | Threat | Status |
|---|---|---|---|
| E | LLM03 | A filesystem-write tool and an outbound-network tool are declared to the model, to the caller via /health and in the persona text, with no allowlist check at load; the exposure today is a standing grant rather than a live primitive because the executor is gated to the MCP protocol. | confirmed PRAX-2026-08-12-004 |
| E | LLM03 | No approval gate stands between a decision and a side effect: the toolApproval switch on the agent record is false and is never read by any code, and the working capability enforcer is inert for this agent because its record does not set the enforcement flag. | confirmed PRAX-2026-08-12-003 |
| Component | Kind | Lane | Description | Evidence |
|---|---|---|---|---|
| Anonymous network caller | entrypoint | user_inputs | Any party that can reach the agent port or the dashboard port; the remit authorizes only an operator-configured user set, and no route authenticates. | src/index.js:1356; src/index.js:1019 |
| Lab operator | entrypoint | user_inputs | The trusted local operator who starts the fleet and configures model mode from the dashboard. | helperbot-remit.md:72; src/index.js:1875 |
| HelperBot HTTP server | adapter | client_adapters | The per-agent HTTP surface: /v1/chat/completions and /chat take turns, /health and /info republish the agent record, all anonymous with wildcard CORS. | src/index.js:1017; src/index.js:1305 |
| DVAA dashboard / control API | adapter | client_adapters | Shared dashboard server on port 9000, wired to HelperBot's attack log and to the LLM provider config; unauthenticated with wildcard CORS. | src/index.js:1864; src/dashboard/server.js:431 |
| Response ladder (generateResponseImpl) | orchestrator | agent_core | The decision loop: picks LLM mode or the canned ladder, and selects a branch from the attack classification and the agent's vulnerability flags. | src/index.js:374; src/index.js:391 |
| detectAttacks() classifier | control | agent_core | Control-shaped but non-blocking: a regex battery classifies every turn and the result only selects which vulnerable branch runs. | src/core/vulnerabilities.js:339; src/core/vulnerabilities.js:233 |
| HELPERBOT agent record and persona | prompt | agent_core | The canned-mode system persona plus the vulnerability switches that arm the disclosure and false-history branches. | src/core/agents.js:44; src/core/agents.js:52 |
| features flag block (stub) | control | agent_core | Stub control: inputValidation, outputFiltering, toolApproval, rateLimiting and auditLogging are declared false and no code anywhere reads them. | src/core/agents.js:57; src/index.js:374 |
| LLM-mode system prompt | prompt | agent_core | The helperbot entry used when real model mode is on: instructs the model to share its instructions openly and ends with an interpolated internal API key. | src/llm/prompts.js:23; src/llm/prompts.js:27 |
| LLM provider client | model | agent_core | BYOK client that holds provider, model and API key in process memory and calls OpenAI or Anthropic over HTTPS. | src/llm/provider.js:9; src/llm/provider.js:60 |
| AIM capability enforcer (unwired for HelperBot) | control | agent_core | A working capability gate with a local audit log, but it returns immediately unless the agent record sets aimEnforced, which HelperBot's does not. | src/aim-enforcer.js:93; src/core/agents.js:44 |
| Attack log ring buffer | log_sink | agent_core | In-memory ring buffer capped at 500 entries holding the full attacker input and the agent reply; no file sink, no shipper, nothing recorded for non-attack turns. | src/index.js:228; src/index.js:266 |
| Declared tool inventory | tool | tools_mcp | read_file, write_file and search_web are declared to the model and to callers; no handler runs on the API path and no allowlist check exists. | src/core/agents.js:56; src/index.js:1039 |
| Committed SENSITIVE_DATA constants | secret_store | external_deploy | Credential-, password- and SSN-shaped literals committed in source as simulated exfiltration fixtures; the internal key is interpolated into a live system prompt. | src/core/vulnerabilities.js:356; src/llm/prompts.js:8 |
| OpenAI API | external_service | external_deploy | Fixed outbound endpoint for chat completions when the operator configures the openai provider. | src/llm/provider.js:86 |
| Anthropic API | external_service | external_deploy | Fixed outbound endpoint for messages when the operator configures the anthropic provider. | src/llm/provider.js:116 |
| Anonymous usage telemetry | external_service | external_deploy | Default-on anonymous platform telemetry started at boot, with --offline and an env opt-out applied before init. | src/index.js:12; src/index.js:36 |
| Container image build | deploy_surface | external_deploy | Image build installs with npm install --no-audit rather than npm ci, over a manifest with three caret-ranged dependencies. | Dockerfile:4; package.json:44 |
| Compose port publication | deploy_surface | external_deploy | Publishes every agent port including HelperBot's 7002 and the dashboard's 9000 to the host, with restart: unless-stopped. | docker-compose.yml:4 |
omissions: Arbitration divergences from the stored tags, KB taking precedence: (1) finding PRAX-2026-08-12-008 is stored ASI06, but the Agentic KB persistence test fails — the agent holds no cross-turn state and the memory store in src/index.js is gated to agents with memoryInjection enabled, which HelperBot's record does not set — so the false-history threat is tagged ASI01 (behavior altered by live input); (2) finding PRAX-2026-08-12-006 stores LLM01 as its OWASP-LLM primary with ASI01 alongside, but the KB's 'goal altered by live input' row makes ASI01 primary and LLM01 the co-tag, so ASI01 is used here; (3) finding PRAX-2026-08-12-004 carries an ASI02 co-tag, which the KB allows only when a tool is shown used wrongly — the executor is gated to protocol 'mcp' and HelperBot is 'api', so only the LLM03 primary is carried. Edgeless nodes: the features flag block (declared, never read by any code), the AIM enforcer (working gate, inert for this agent), and the Dockerfile and compose deploy surfaces (posture evidence). Scope: the dashboard control API sits outside the four subject files named in the scan instructions but is included because src/index.js wires HelperBot's attack log and LLM configuration to it; the other DVAA agents in the same tree are excluded per scope. Two dashboard-side threats and the log-disclosure threat have no backing finding in the 2026-08-12 JSON and are recorded as potential after checking the routes for a credential or redaction step.
generated by: Opus 5 (1M context)