AI security is moving beyond the question of whether a model can be tricked into producing an unsafe answer. The harder problem begins when the model is connected to tools and allowed to act.
An agent can read an email, retrieve a document, decide which function to call, query an internal API, update a ticket, invoke a cloud service or hand work to another agent. In that architecture, a successful prompt injection is not automatically a critical event. The security question is what permissions and trust relationships sit behind the manipulated model.
The model is only one trust boundary
Think of the agent as a privileged application component. It receives instructions from multiple sources, reasons over them, and can cross system boundaries. The attacker may never need to “break” the model. They may only need to influence the model’s next action.
Attack paths worth testing
1. Indirect prompt injection
Place attacker-controlled instructions in a source the agent is expected to trust: a public webpage, support ticket, PDF, email, knowledge-base record or retrieved document. The test is whether those instructions can change the agent’s objective, tool selection or output handling.
2. Tool abuse
Inventory every function exposed to the model. A function that looks harmless in isolation can become dangerous when combined with other functions. For example, a read-only database query followed by a messaging function may create an exfiltration path.
3. Identity confusion
Determine whether the agent performs actions as the requesting user, a fixed service account, or a mixture of both. A common failure mode is authenticating the user correctly but allowing the agent to call downstream systems with a much broader identity.
4. Cross-tenant or cross-role access
For multi-tenant applications, test whether retrieved context, memory, vector indexes and tool calls enforce the same tenant boundary as the primary application. Repeat the test with users who have different roles and data entitlements.
Interactive: change the attack step
What happens next?
A practical assessment methodology
- Map the agent: document prompts, memory, retrieval sources, tools, model endpoints, identities and approval gates.
- Classify actions: separate read-only, reversible and destructive capabilities.
- Test untrusted inputs: exercise direct and indirect prompt injection through every retrieval channel.
- Test authorization: verify that each tool independently checks user and tenant permissions.
- Chain capabilities: combine individually low-risk functions to identify emergent attack paths.
- Validate impact: prove data access or action execution using a controlled test object rather than relying on model output.
Controls that reduce blast radius
Design as if the model can be manipulated. Put authorization outside the model, keep tool permissions narrow, require approval for high-impact actions, validate tool arguments server-side, isolate secrets, and log every sensitive action with enough context to reconstruct the decision path.
Further reading
Google Threat Intelligence has documented the evolution of adversarial AI and the increasing use of agentic workflows. Read the Google Threat Intelligence analysis.
OWASP GenAI guidance discusses excessive agency, including excessive functionality, permissions and autonomy. Read the OWASP guidance.
