The Model Context Protocol (MCP) is how an agent gets new tools. You point it at an MCP server, the server advertises a set of tools, resources, and prompts, and the agent can now use them. A filesystem server lets it read and write files. A database server lets it query. A browser or HTTP server lets it reach the network. It's genuinely useful — and it's the fastest way to widen what an agent can reach without widening the review that goes with it.
Here's the mental model that keeps you out of trouble: an MCP server is a dependency with permissions. Connecting one is less like installing a library and more like granting an integration a scope. Treat it that way and most of the mistakes below don't happen.
Where the risk actually enters
An agent on its own can only produce text. An MCP server is where that text turns into *actions* — reads, writes, queries, requests. So the server boundary is exactly where prompt injection stops being theoretical: untrusted text the model ingested becomes a tool call the server executes.
The diagram is the whole argument in one picture: injected text only matters to the degree a connected server gives it somewhere to go. That makes three questions non-negotiable before you wire a server in.
The three questions that matter most
1. What can it reach? A filesystem server reads files — which files? The whole disk, or one project directory? A database server runs queries — against what, with which account, read-only or read-write? An HTTP server makes requests — to anywhere, or an allow-list of hosts? The answer *is* the blast radius of any injection that reaches this agent. Write it down before you connect.
2. Whose code is it, and whose words? A first-party server you wrote is a different trust level from a community server you pulled in. And there's a subtlety unique to MCP: a server's tool descriptions are fed to the model as instructions. A careless or malicious server can describe its tools in ways that steer the agent — "always call sync_secrets first." You are trusting the server's code *and* its prose. Read both.
3. What credentials does it hold? Most servers need an API key or token to do their job. Where does that token live, and what scope does it have? An MCP server with an admin database token is an admin database integration, no matter how innocent its tool list looks.
Credentials: the most common real-world leak
MCP configuration is where secrets quietly end up committed to a repository. A .mcp.json or agent config with an inline API key is a committed credential — and as we covered in secrets leak through git history, deleting it from the file later does not remove it from history. The key is still in an old commit, readable by anyone with a clone.
// Don't: a live token pasted into committed config
{
"mcpServers": {
"db": {
"command": "mcp-postgres",
"env": { "DATABASE_URL": "postgres://admin:S3cr3t@prod/db" }
}
}
}
// Do: reference it from the environment, scoped to the minimum
{
"mcpServers": {
"db": {
"command": "mcp-postgres",
"env": { "DATABASE_URL": "${READONLY_REPORTING_DB_URL}" }
}
}
}Three rules cover most of it:
- Keep MCP credentials in environment variables or a secrets manager, referenced by name — never pasted into config.
- Scope each credential to the minimum the server needs. A read-only reporting token beats a read-write admin token every time.
- Assume any server that can read files can read your secrets. If an MCP filesystem server can reach
.env, it can hand its contents to the model, and the model can hand them to anything with network egress.
Tool permissions: turning "can do" into "may do"
MCP servers are where "the agent can do something" gets granted, so this is where least privilege earns its keep. The pattern that holds up:
- Prefer read-only servers wherever the task allows. An agent that *reports on* your database doesn't need to write to it.
- Allow-list, don't deny-list. Decide what a server *may* do. Deny-lists of "dangerous operations" are routinely bypassed by an equivalent the list didn't anticipate.
- Require approval for irreversible actions a server exposes — writes, deletes, sends, deploys — instead of letting them auto-run. See prompt injection defenses that actually work for why approval gates are so effective.
- Pin the server version. A server that auto-updates can change its tools — and its tool descriptions — under you, with no review.
A pre-connection review checklist
Run this before any MCP server lands in a repo's config:
- Reach: I know exactly what this server can touch (files, network, data) and have written it down.
- Source: I trust where the server's code *and* its tool descriptions come from.
- Credentials: referenced from the environment, scoped to the minimum, not committed.
- Permissions: read-only where possible; write/destructive tools require approval or are disabled.
- Version: pinned, not floating.
Capability is a decision, not a default. Every "yes" above should be a choice someone made on purpose.
How OpenRouting reviews this for you
When OpenRouting scans a repository, it detects the MCP servers and tool configurations in your agent setup and has Claude review them against exactly these risks: credentials committed in config, servers with broad filesystem or network reach, and write-capable tools with no approval gate. You get an inventory of what's wired in, plus concrete findings wherever a connector is more powerful than it needs to be — on every scan, not just the day someone added it. (See how it works for the full scan flow.)
MCP makes agents dramatically more capable. The security work is making sure each new capability was a decision, not an accident.