본문으로 건너뛰기
L15.5

Threat Modeling for AI Systems

Goal

Describe what an AI system must protect, who or what can influence it, how failures could occur, and which controls break those paths.

Start with a protected asset and a trust boundary​

Imagine protecting a school computer lab. Before listing fancy attack names, you would first ask: What are we protecting—student files, passwords, devices? Who or what can send input into the system? Where does information cross from a less-trusted place to a more-trusted one? What path could lead from that input to something harmful?

That is the core of threat modeling. An asset is something worth protecting. An actor is a person or system that can influence events. A trust boundary is a crossing where assumptions change and a check is needed. Once those pieces are visible, named attack techniques become examples of possible paths rather than vocabulary to memorize first.

Begin by listing the concrete assets in this system: user data, credentials, money-moving permissions, proprietary documents, system availability, model weights, audit evidence, or a user's trust in an action. If you cannot say what matters, you cannot prioritize what to defend.

Trace concrete attack paths​

Next identify the actors and trust boundaries. A normal user, administrator, external web page, retrieved document, third-party tool, model provider, and another agent may all influence the system differently. A trust boundary is where data or authority moves between components with different assumptions. Crossing it should trigger the validation or permission check that belongs there.

Consider a research assistant that reads web pages and can save notes. The web page is untrusted content. The note store is a protected asset. The model is a reasoning component, not an authorization authority. A threat path might be: malicious page text attempts to redirect the model, the model proposes a write to the note store, and the application accepts the write without checking source or scope.

That path is more useful than the label “prompt injection” by itself because it shows where a control can break the chain. The retriever can label source trust. The model-facing prompt can keep instructions and data distinct. The application can constrain which note collection may be modified. A confirmation rule can protect sensitive destinations. Logs can preserve the source, proposal, policy decision, and final action.

untrusted webpage
→ model reads instruction-like text
→ model proposes tool action
→ policy check verifies destination and authority
→ allow or block

Include accidental failures and external catalogs​

Threat models should include accidental failures too. A badly formatted tool result can be as disruptive as a malicious one. An expired credential, stale policy cache, or wrong tenant identifier can create dangerous behavior without an attacker. A tenant is one customer, organization, or isolated account scope inside a system that serves multiple groups. Security design improves when it handles both adversarial and ordinary faults at the same trust boundary.

MITRE ATLAS provides a knowledge base for adversarial techniques against AI-enabled systems. Use resources like ATLAS to expand your imagination, not to replace system-specific reasoning. A technique matters only when your architecture exposes a compatible path and the consequence matters for your assets.

Turn threats into prioritized records​

A practical threat record can use four fields: asset, precondition, attack or failure path, and mitigation. Add a test result when you can exercise the mitigation. For example: asset = private documents; precondition = assistant can retrieve external text and call a search tool; path = untrusted text tries to trigger broader retrieval; mitigation = tool policy restricts collections independently of model text; evidence = regression case proves the wider collection request is denied.

Severity and likelihood help prioritize. A low-frequency path that exposes sensitive records may deserve more urgent action than a common cosmetic failure. Do not turn the numbers into false precision. The purpose is to make assumptions and priorities visible.

Revisit the model when architecture changes​

Threat models change when architecture changes. Adding a browser tool creates new inputs and destinations. Adding memory creates persistence. Adding an agent-to-agent protocol adds remote principals and task artifacts. Every new capability should trigger a quick review of assets, trust boundaries, and permissions.

Finally, connect the threat model back to evaluation. Each important path should produce at least one testable control or detection rule. If a threat cannot be observed in any test or telemetry, the team may be relying on a design claim that has never been checked.

Predict

A model proposes reading a private collection after seeing instructions inside an untrusted web page. Which control most directly breaks the dangerous path?

Run the local Lab​

Run:

python3 labs/notebooks/level-15/l15-05-threat-model.py

The Lab checks whether each threat record names an asset, trust boundary, path, and control.

  1. Run it unchanged. Confirm the single threat has no incomplete fields and ready_for_test is true.
  2. In the first threat, find the non-empty "control" string. Before editing, predict which summary fields should change if the threat describes a path but omits the control that breaks it.
  3. Change only the control value to an empty string, then rerun.
  4. Confirm threat index 0 appears in incomplete and ready_for_test becomes false. Restore the starter afterward, then write down what executable security check the original control suggests.

Loading lab…

Quick Check

1. What should a threat model identify first?
2. Why describe a concrete path instead of only naming an attack category?
3. When should an AI threat model be revisited?

0 of 3 questions answered.

Explain it back​

Pick one tool-using AI system. Name two assets, one untrusted input, one trust boundary, one realistic failure path, and the application control that should stop it.

Key Takeaways

  • Threat models connect assets and consequences to concrete system paths.
  • Treat model output as a proposal, not trusted authorization.
  • Put controls where untrusted influence would otherwise become authority or access.
  • Include accidental faults as well as hostile behavior.
  • Convert important threat paths into regression tests with observable results.

Next Lesson

Next, L15.6 — Prompt Injection and Data Exfiltration applies threat modeling to untrusted instructions and protected information flows.

References

Lesson actions

Completion is stored locally on this device.

View progress