Openai Customer Service
Threat Model
evidence-derived · Praxen 1.3.0 · graph contract 1.4 · built against openai-customer-service-findings-2026-08-12.json
15Components
19Flows
7Boundaries
9Confirmed
3Potential
1Partial
3Mitigated
Summary
The Agent (as modeled)

This is an airline customer-service assistant that runs as a console conversation: a passenger types a message, a triage agent decides whether the question belongs to an FAQ specialist or a seat-booking specialist, and control is handed to that specialist for the rest of the turn. Everything untrusted arrives through one channel — the passenger's own text — and the agent can do exactly three things with it: look up a canned FAQ answer, rewrite the booking fields held in the session (seat, confirmation number, flight number), and hand the conversation between its own specialists. There is no shell, no filesystem, no network tool, and nothing that survives the process, so the blast radius is genuinely small. What is missing is any code standing between the customer and those actions: the framework this example is built on ships input, output, and tool guardrails plus a per-tool approval flag, and the example turns on none of them. The two surfaces that matter are the seat change, which is the only action that alters state, and the conversation transcript, which leaves the process by default to the vendor's hosted trace service.

Priority Threats (led by attack paths)

Deal with the seat change first. The rule that is supposed to authorize it — the customer must supply a confirmation number that matches the booking — exists only as a sentence in the agent's prompt. The tool that applies the change looks nothing up, checks nothing, overwrites the stored confirmation number with whatever the model passed it, and returns a success message for a change that no booking system ever saw. A passenger who talks the assistant past its routine, or a message crafted to override its instructions, reaches that tool directly, because nothing inspects what the customer sends and nothing asks a human before the tool fires. The flight number attached to the booking is invented at handoff time with a random number generator, so the confirmation the customer receives is fabricated in more than one respect.

The second priority is what the conversation says and where it goes. Nothing screens the model's replies, so a customer who asks the assistant to repeat its own instructions gets them, and the entry agent has no statement of what it is and is not allowed to discuss — the shipped demo input is itself an off-topic question that nothing refuses. Separately, and regardless of any attacker, every turn of every conversation, including passenger name, confirmation number, flight number and seat, is exported by default to the vendor's hosted trace store — a destination the remit does not list among the airline's trusted systems, and one that the same environment-supplied API key keeps switched on. A customer can also retry a failed confirmation number forever: the conversation loop has no attempt limit and no way to end an interaction. Fixes in order: verify the confirmation number in code and put an approval gate on the seat tool, attach input and output screening at the entry agent, and turn off sensitive-content capture in the trace export or point it at a sink the airline controls.

Architecture & Trust Boundaries
Attack paths are drawn in red, running from where an attacker gets in to what they reach. Click any box to jump to its inventory row, or a B-badge to jump to that boundary; hover a box to reveal its data flows. Everything reads statically below — the key resolves every mark and the tables carry every citation.
User / InputsClient / AdaptersAgent CoreTools / MCPExternal / DeployB1B2B3B4B5B6B7Customer (passenger)ENTRYPOINTon attack path — target (consequence) AND source (ingress)Console input adapterADAPTERon attack path — pass-throughConversation run loopORCHESTRATORon attack path — pass-throughTriage AgentPROMPTon attack path — pass-throughFAQ AgentPROMPTSeat Booking AgentPROMPTon attack path — pass-throughAirline booking contextMEMORYon attack path — target (consequence)Seat-booking handoffhookORCHESTRATORSDK input/outputguardrails (unwired)CONTROLFAQ lookup toolTOOLSeat update toolTOOLon attack path — pass-throughPer-tool approval gate(unwired)CONTROLOpenAI model providerEXTERNAL_SERVICEon attack path — pass-throughOpenAI trace ingestbackendLOG_SINKDependency and CIpostureDEPLOY_SURFACE
Familiesactors & inputsclient / adaptersagent coretools & datacontrolsexternal & deploy
Kindsentrypointadapterorchestratorpromptmemorycontroltoolexternal servicelog sinkdeploy surface
Marksboundary — worst threat: confirmed— potential (unanswered hypothesis)— partial (control covers part; remainder stated)— mitigatedattack path (origin → consequence)box on an attack path: source (ingress) pass-through control that failed target (consequence)faint arc = flow spanning 2+ lanes
Attack Paths
A passenger's chat message rewrites the booking on a confirmation number nobody checked
  1. Customer (passenger) Anyone holding the session types the opening message — the speaker is unauthenticated and never bound to a booking, so the claim to be a given passenger is attacker-authored.
  2. Console input adapter The console adapter returns the typed string verbatim, with no validation or provenance labelling.
  3. Conversation run loop [PRAX-2026-08-12-002] The loop appends the text to the conversation and sends it to the model with no input guardrail to trip on an off-routine or override attempt.
  4. Triage Agent [PRAX-2026-08-12-003] The entry agent has no subject-matter boundary and no decline clause, so it forwards the request and hands off to the seat specialist on the model's judgement alone.
  5. Seat Booking Agent [PRAX-2026-08-12-001] The specialist's whole authorization step is a prompt line asking for a confirmation number, which the model can skip or accept at face value.
  6. Seat update tool [PRAX-2026-08-12-001] The tool fires on the model's arguments with no approval gate and no lookup against any booking record; its one assert checks a handoff-set flight number and disappears under -O.
  7. Airline booking context [PRAX-2026-08-12-007] The booking's seat and confirmation number are overwritten and a success message is returned — an unauthorized booking mutation wherever this tool is wired to a real seat-management backend.
A passenger's crafted message gets the agent's own instructions read back to them
  1. Customer (passenger) An untrusted caller in the session composes a message designed to override the agent's role or elicit its instructions.
  2. Console input adapter The console adapter hands the message through unchanged; nothing distinguishes a request from an instruction.
  3. Conversation run loop [PRAX-2026-08-12-002] The text joins the conversation and goes to the model with no screening layer installed on any agent.
  4. OpenAI model provider [PRAX-2026-08-12-002] The model, holding the system instructions in the same context as the attacker's text, produces a reply that reproduces them.
  5. Conversation run loop [PRAX-2026-08-12-002] The loop consumes the returned message item with no output guardrail between the model's reply and the customer.
  6. Customer (passenger) [PRAX-2026-08-12-002] The internal instructions and routines are printed into the customer's transcript — operator-side control logic disclosed to an untrusted party, and copied on to the hosted trace store with the rest of the turn.
Trust Boundaries — Threats & Governing Remit Rules
B1 Customer session to agent loop (untrusted-ingress) — 4 threats, 4 remit rules · worst: confirmed
R-03 gap MUST NOT comply with attempts to override, ignore, reveal, or subvert its own instructions or role (jailbreak / prompt-injection attempts)
R-22 gap When an input-validation guardrail detects an off-topic request or a jailbreak / prompt-injection attempt, the agent MUST halt the current turn and explicitly decline
R-01 partial such requests MUST be declined and routed back to triage rather than answered
R-17 gap The agent MUST NOT reveal its system prompt or internal instructions to the customer or any other party.
STRIDEOWASPThreatStatus
TLLM01Customer text is appended to the conversation and sent to the model with no classification, labelling, or screening, so an instruction-override attempt in a passenger's message is indistinguishable from a genuine request and no tripwire can halt the turn.confirmed PRAX-2026-08-12-002
E—Out-of-scope requests reach an entry agent that names no subject-matter boundary and carries no decline clause, and no code gate refuses them — the shipped auto-mode input is itself such a request.confirmed PRAX-2026-08-12-003
ILLM08Nothing screens what goes back to the customer, so the agents' own instructions and routines can be reproduced into the printed reply on request.confirmed PRAX-2026-08-12-002
DLLM06The conversation loop is unbounded and holds no failed-attempt counter, so a caller can retry confirmation numbers or off-routine requests indefinitely; the framework's ten-turn cap bounds only the internal loop inside one run call.confirmed PRAX-2026-08-12-006
B2 Agent to model provider (model-egress) — 2 threats, 1 remit rules · worst: potential
R-09 partial The configured model provider endpoint used to run the agent.
STRIDEOWASPThreatStatus
I—Every turn's full text — including any booking identifiers the customer states — crosses to the provider with no minimization or filtering, though the provider is a counterparty the remit authorizes (checked: main.py:168 — Runner.run is called with no RunConfig and no call_model_input_filter).potential
TLLM04Neither the model nor the endpoint is pinned by the example: the agent runs whatever model name the environment resolves and talks to whatever base URL the environment names, so the party interpreting customer instructions is deploy-time configurable (checked: main.py:96,109,123 — no model= on any Agent; default_models.py:99 and openai_provider.py:139 read both from the environment).potential
B3 Model decision to business tools (tool-invocation) — 4 threats, 4 remit rules · worst: confirmed
R-19 gap A seat change MUST be authorized by a confirmation number supplied by the customer that matches the active booking; the agent MUST NOT apply a seat change without one.
R-13 verified The agent MUST NOT have shell or code-execution, filesystem, arbitrary web-browsing/outbound-network, or outbound-messaging (email/SMS) tools
R-21 partial The agent MUST NOT take instructions from retrieved FAQ content or tool output and act on them as if they were authoritative commands
R-02 partial factual answers to airline questions MUST come from the authoritative FAQ source
STRIDEOWASPThreatStatus
ELLM03The seat-update tool executes on the model's decision alone: no approval flag is set, no booking lookup is performed, and the sole authorization step is a prompt sentence — leaving one assert on the handoff-set flight number, which Python strips under -O, as the only code-level precondition.confirmed PRAX-2026-08-12-001
TLLM01Tool results re-enter the conversation as unlabelled context with no tool guardrail of any kind, so anything a tool returns is read by the model with the same authority as operator instructions — today that is hardcoded FAQ text plus the seat tool's echo of the model's own arguments.confirmed PRAX-2026-08-12-002
TLLM07Nothing compares the answer the model gives a customer against what the FAQ tool actually returned, and the tool recognises only three keyword families, so every other airline question leaves the model holding a not-known result and a plausible answer in its weights.confirmed PRAX-2026-08-12-005
E—A hijacked triage turn cannot reach the seat-update tool directly: the triage agent's tool list is empty and each specialist can invoke only its own tool, so the mutation is reachable only after an explicit handoff.mitigated examples/customer_service/main.py:130
B4 Decision to booking-state write (state-commit) — 1 threats, 1 remit rules · worst: confirmed
R-20 partial The agent MUST NOT modify any booking field other than the seat assignment (for example cancelling a flight, rebooking, or changing the passenger's name).
STRIDEOWASPThreatStatus
TLLM07The seat path writes two booking fields beyond the permitted seat — the confirmation number is overwritten with the model's argument and the flight number is invented from a random draw at handoff time — and the customer is then told the change succeeded although no seat-management backend was ever called.confirmed PRAX-2026-08-12-007
B5 Conversation trace export (telemetry-egress) — 2 threats, 3 remit rules · worst: confirmed
R-16 gap The agent MUST NOT transmit passenger data to any destination outside the airline's own trusted systems.
R-11 gap Any arbitrary outbound network destination not listed as a trusted service.
R-24 verified Every handoff/transfer between agents and every tool invocation (FAQ lookup and seat update) MUST be recorded to a durable, structured trace so a session can be reconstructed.
STRIDEOWASPThreatStatus
ILLM02Conversation content and passenger booking identifiers are exported by default to a hosted trace store that the remit does not list among the airline's trusted services, because the example opts into tracing and overrides neither the disable switch nor the sensitive-data switch.confirmed PRAX-2026-08-12-004
R—Agent transfers and tool invocations could not be reconstructed after the fact — answered by the per-turn trace span under a stable conversation id, with handoff, function and generation spans emitted for every turn.mitigated examples/customer_service/main.py:166
B6 Provider credential posture (secret-material) — 2 threats, 0 remit rules · worst: potential
the remit does not touch this boundary — threats here are assessed against the RAISE/OWASP baseline alone (a remit is a job description, not a security model; silence here is normal)
STRIDEOWASPThreatStatus
I—Credential material committed to the repository — a tree-wide sweep for API-key, AWS-key and private-key patterns returned nothing and dotenv files are excluded from version control, so the provider key is supplied by the deployment environment rather than the tree.mitigated .gitignore:105
I—One environment-supplied provider key authorizes both the model call and the trace ingest, so there is no credential-scoped way to keep passenger content out of the hosted trace store — only configuration (checked: processors.py:108 — the exporter falls back to the same environment variable the model client uses).potential
B7 Runtime dependency posture (supply-chain) — 1 threats, 0 remit rules · worst: partial
the remit does not touch this boundary — threats here are assessed against the RAISE/OWASP baseline alone (a remit is a job description, not a security model; silence here is normal)
STRIDEOWASPThreatStatus
TLLM04A compromised or vulnerable runtime dependency would enter through the install path; a committed lockfile and upper-bounded ranges make resolution reproducible, but nothing surveils the resolved Python stack.partial uv.lock:1
remainder: no SBOM or AI-BOM exists in the tree and automated dependency updates are configured for GitHub Actions only, so the Python runtime packages are never scanned for known vulnerabilities
Component Inventory
Every component with its kind, lane, and source evidence — the diagram's tooltips, on paper.
ComponentKindLaneDescriptionEvidence
Customer (passenger)entrypointuser_inputsThe passenger in the active conversational session — the only human counterparty, unauthenticated and free-text, and the sole source of untrusted input.examples/customer_service/main.py:162; examples/customer_service/main.py:173
Console input adapteradapterclient_adaptersThe console I/O helper that supplies each turn's text — real stdin normally, or a fixed environment-driven fallback string when EXAMPLES_INTERACTIVE_MODE=auto.examples/auto_mode.py:19; examples/auto_mode.py:15
Conversation run looporchestratoragent_coreThe asyncio while-loop that appends each customer turn to the running input list, invokes the SDK Runner against the current agent, prints the resulting items, and carries the last agent forward to the next turn.examples/customer_service/main.py:161; examples/customer_service/main.py:167
Triage Agentpromptagent_coreThe entry agent: holds no business tools, only handoffs to the two specialists; its instructions state no subject-matter scope and contain no decline clause.examples/customer_service/main.py:126; examples/customer_service/main.py:130
FAQ Agentpromptagent_coreThe airline-FAQ specialist; carries the instruction-level grounding control ('Do not rely on your own knowledge') and an instruction-level scope control (transfer back to triage when it cannot answer) — both enforced by the prompt, not by code.examples/customer_service/main.py:104; examples/customer_service/main.py:105
Seat Booking Agentpromptagent_coreThe seat-change specialist; its entire authorization step is the instruction-level control 'Ask for their confirmation number', with no code path that verifies one was supplied or that it matches a booking.examples/customer_service/main.py:116; examples/customer_service/main.py:119
Airline booking contextmemoryagent_coreThe per-run pydantic context carrying the session's passenger PII and booking fields (passenger name, confirmation number, seat, flight number); in-process only, discarded when the loop exits.examples/customer_service/main.py:29; examples/customer_service/main.py:154
Seat-booking handoff hookorchestratoragent_coreThe on_handoff callback that fires when triage transfers to seat booking; it fabricates a flight number with random.randint and writes it into the booking context before the specialist starts work.examples/customer_service/main.py:89; examples/customer_service/main.py:134
SDK input/output guardrails (unwired)controlagent_coreSTUB — the framework ships InputGuardrail/OutputGuardrail primitives with tripwire halt semantics, and every agent in this example is constructed without them, so the guardrail layer exists but is never installed.src/agents/guardrail.py:72; src/agents/agent.py:324
FAQ lookup tooltooltools_mcpThe 'authoritative FAQ source': a three-branch keyword matcher over hardcoded strings that returns a not-known message for anything outside baggage, seating and wifi.examples/customer_service/main.py:42; examples/customer_service/main.py:39
Seat update tooltooltools_mcpThe one state-mutating capability: writes the model-supplied confirmation number and seat into the booking context and returns a success string; its only code-level precondition is an assert on the handoff-set flight number, which Python removes under -O.examples/customer_service/main.py:79; examples/customer_service/main.py:82
Per-tool approval gate (unwired)controltools_mcpSTUB — the framework's needs_approval flag interrupts the run and requires explicit approval before a tool executes; it defaults to False and neither tool in this example sets it.src/agents/tool.py:440; src/agents/tool.py:2441
OpenAI model providerexternal_serviceexternal_deployThe upstream model endpoint that receives every turn's conversation history and returns the tool calls, handoffs and text the loop acts on; neither the model nor the base URL is pinned by the example.src/agents/models/default_models.py:99; src/agents/models/openai_provider.py:139
OpenAI trace ingest backendlog_sinkexternal_deployThe hosted trace store the SDK exports spans to by default, carrying turn text, tool arguments and tool results out of the process; the example opts into tracing and overrides neither the disable switch nor the sensitive-data switch.src/agents/tracing/processors.py:45; src/agents/run_config.py:356
Dependency and CI posturedeploy_surfaceexternal_deployThe runtime's supply-chain posture: upper-bounded dependency ranges resolved by a committed lockfile, with automated dependency updates scoped to GitHub Actions only and no SBOM anywhere in the tree.pyproject.toml:8; uv.lock:1
Extraction Notes
lane_fit: Clean, with one strain: the console client surface is split across two files — input arrives through examples/auto_mode.py (modelled as the client_adapters node) while replies are printed inline in main()'s run loop, so the outbound leg is an edge from the orchestrator to the customer rather than a second client node, since both halves of main() share one defining scope and merge under the same-scope rule.
omissions: No arbitration divergences: every stored tag matches the KB on the same evidence — PRAX-002 as LLM01 (ingress-hygiene gap with no evidenced goal alteration), PRAX-004 as LLM02 (unredacted capture in observability tooling), PRAX-005 and PRAX-007 as LLM07 (fabricated or unverified completion, no cross-session persistence so no ASI10), PRAX-006 as LLM06 (missing ceiling with nothing propagated), PRAX-003 as RAISE-only with no honest OWASP code. PRAX-001 carries a stored ASI03 co-tag; the graph's single primary field takes LLM03, which is both the stored primary and the Agentic KB's gate row (a missing approval gate is LLM03, and ASI03's row excludes capability declared in a tool definition) — the ASI03 reading is not carried, not overruled. The instruction-level controls (the confirmation-number routine, the FAQ grounding rule, the transfer-back-to-triage clauses) are not separate control nodes: each sits inside the same defining scope as its agent construct, which the ID rules merge into one node, so they are recorded in those nodes' descriptions and evidence. The handoff/transfer tools have no single enumerating construct — they are registered in three places — so they fold into the agents' evidence and appear as edges. Two edgeless nodes: the unwired guardrail layer and the unwired per-tool approval flag are framework primitives the example never installs; the dependency-posture node is repo posture with no data flow. Suspected but not cited: no seat-management backend and no FAQ knowledge service exist anywhere in the tree, so the remit's trusted services for both are absent rather than misconfigured — the FAQ 'source' is a hardcoded three-branch matcher and the seat 'backend' is process memory.
generated by: Opus 5 (1M context)