Prompt injection, explained for engineers

$ cat blog/what-is-prompt-injection.md
✓ AI security
✓ prompt injection
✓ agents
7 min read
$ 

Prompt injection is the vulnerability class that shows up the moment an LLM reads text it doesn't fully control. If your application pastes a web page, a support ticket, a file, or a pull request description into a prompt, that text can contain instructions, and the model may follow them.

Why it isn't just "SQL injection for AI"

With SQL injection, there is a clean boundary between code and data: parameterized queries fix it because the database can tell the query apart from the values. Language models have no such boundary. Everything is text, and the model decides what to act on. There is no parameterized-prompt API that makes untrusted text inert.

That is the uncomfortable core of prompt injection: you cannot fully sanitize your way out of it. You reduce the blast radius instead.

Where it bites

The risk scales with what the agent can *do*, not what it can *say*. A chatbot that summarizes text and has no tools is low-risk. An agent with a shell tool, file write access, or the ability to open pull requests is a different story: injected instructions become injected *actions*.

Common sources of untrusted text that reach a prompt:

  • Issue and PR descriptions, code comments, commit messages
  • Web pages fetched by a browsing tool
  • Files in a repository the agent reads
  • Emails, tickets, and chat messages

Patterns that actually help

  1. Least-privilege tools. Give an agent the narrowest tools it needs. An agent that only needs to read code should not have a shell.
  2. Allow-lists over deny-lists. Decide what a tool *may* do, not what it may not. Deny-lists of "dangerous" commands are routinely bypassed by an equivalent command the list didn't anticipate.
  3. Human approval for irreversible actions. Commits, deploys, sending messages, deleting data: put a person in the loop.
  4. Separate trust levels. Keep the instructions you wrote (the system prompt) distinct from content the agent ingests, and treat ingested content as data, not commands.
  5. Constrain outputs. Where you can, make the agent return structured output you validate, rather than free-form actions.

How OpenRouting helps

When OpenRouting scans a repository, it detects your agent setup — agent definitions, skills, memory, MCP servers, and the permission config that governs them — and has Claude review it specifically for prompt-injection and tool-safety risks: over-broad tool permissions, shell or file tools without allow-lists, and hooks that auto-approve dangerous actions. You get a plain inventory plus concrete findings with fixes.

Prompt injection won't be "solved" by a single setting. But most real exposure comes from a handful of over-permissive configurations, and those are findable. That's the part we automate.

Keep reading