AI SECURITY · 09 min
AI threat modeling for modern applications
AI-enabled applications create new trust relationships between users, models, data, tools and actions. A useful AI threat model follows those relationships instead of treating the model as an isolated component.
Why AI changes threat modeling
Traditional applications often have relatively explicit control flow. AI systems can introduce probabilistic behavior, natural-language inputs, retrieval pipelines, tools and agents that decide which actions to take. Security analysis therefore needs to examine both the application architecture and the model-mediated paths through it.
Start with the system graph
Map users, models, prompts, retrieval systems, data stores, tools, APIs, identities and external services. Mark where untrusted input enters and where the system can cause consequential actions.
Key threat areas
- Prompt injection: untrusted content influences model behavior or tool use.
- Data exposure: sensitive data is retrieved, included in context or returned to an unauthorized user.
- Tool misuse: a model or agent invokes a capability beyond the intended authorization boundary.
- Memory poisoning: persistent context is manipulated so future decisions are influenced by attacker-controlled information.
- Identity and privilege compromise: an AI component inherits credentials or permissions that are too broad.
- Unsafe downstream actions: generated output becomes an input to a high-impact operation without sufficient validation.
Ask attack-path questions
Instead of asking only “What can go wrong with the model?”, ask “How could an attacker reach an objective through this system?” Trace realistic paths from an entry point to sensitive data or privileged actions, then identify the controls that break those paths.
Evidence matters
Architecture diagrams are a starting point. Effective analysis should be informed by actual permissions, network paths, identities, policies, logging, tool definitions and deployment configuration. This is particularly important for cloud-hosted AI systems whose effective attack surface can change over time.