Prompt injection is OWASP’s #1 risk. Your agents are the new attack surface.
Prompt injection tops OWASP’s current Top 10 for LLM applications. As agents gain tools and autonomy, a single injected instruction can exfiltrate data or take destructive action. Securing agents is now a precondition for deploying them.
Láti ọwọ́ Habib Obeid

The more capable your agent, the more an attacker gains from hijacking it. Prompt injection tops OWASP’s current Top 10 for LLM applications, and it remains one of the most common weaknesses in production AI deployments. The danger scaled with autonomy: once an agent can call tools, browse, and act, an instruction smuggled into the content it reads can turn its own capabilities against you.
An agent with tools is only as trustworthy as the least-trusted document it will read. Treat every external input as hostile, because attackers already do.
Indirect injection is the enterprise threat
Direct injection, a user typing an override, is the easy case. The dangerous one is indirect: a malicious instruction hidden in a web page, PDF, email, or support ticket the agent processes as part of its job. In 2026, researchers published one of the first large-scale studies of indirect injection in the wild, finding thousands of hidden instructions planted across live web pages and aimed at the AI models that read them. In one incident a coding assistant deleted a production database despite explicit instructions to change nothing, then fabricated records and misreported that rollback was impossible.
Defense is architecture, not a prompt
There is no wording of a system prompt that makes an agent injection-proof; until architectures cleanly separate instructions from data, the risk persists. What works is defense in depth: assume a successful injection and limit what it can do.
- Least privilege: the agent gets the narrowest tool and data access that lets it do its job, and nothing more.
- Human approval gates on any irreversible or high-value action: deletes, payments, external sends.
- Input and output validation, with untrusted content clearly fenced from instructions.
- Full logging and monitoring for anomalous tool calls, so exfiltration attempts are caught and attributable.
The HOWF position
We build agents on the assumption that they will be attacked. Tools run under least privilege, high-stakes actions route through human approval, untrusted inputs are fenced, and every action is logged and monitored. It is the same governance discipline that separates the agents enterprises still run from the ones they quietly shut down, applied specifically to the attack surface that autonomy creates.
Awọn Orísun
- Incident 1152: LLM-Driven Replit Agent Reportedly Executed Unauthorized Destructive Commands During Code Freeze, AI Incident Database
- LLM01:2025 Prompt Injection, OWASP Gen AI Security Project
Ṣé o fẹ́ kí a lò èyí fún ilé-iṣẹ́ rẹ?