Every engineering team that turns on a security scanner meets the same wall: hundreds of findings, most of which don't matter, and no time to tell which ones do. The scanner isn't wrong to flag them — it's doing pattern matching without context. But a list nobody reads protects nobody.
The real problem is context, not detection
Tools like Semgrep are good at detection. What they lack is the context a human reviewer brings: is this user input actually reachable? Is it sanitized two lines up? Is this test code? Is the "SQL string" built from a constant?
A reviewer answers those questions by reading the surrounding code. That's exactly the kind of judgment a capable language model can apply at scale.
How the two layers fit together
OpenRouting runs proven open-source scanners first — they do the thorough, deterministic detection. Then Claude reviews each finding next to the code around it and decides whether it's a real issue or a likely false positive, with a short explanation and a suggested fix.
The scanner keeps the recall; the AI pass adds the precision. You still see everything in the findings list, but the likely-false-positives are marked as such and sorted down, so attention goes where it belongs.
Why keep the scanner at all?
Because an LLM alone is the wrong tool for exhaustive detection: it's slower, costlier, and less consistent across a whole repository than a purpose-built scanner. The division of labor matters — deterministic tools find, the model judges.
What this looks like in practice
- A finding arrives from the scanner with a rule, a file, and a line.
- Claude reads a window of code around it and returns a verdict (real / likely false positive), a confidence score, a plain-English explanation, and a fix.
- Repeat findings are matched across scans by a content fingerprint, so you never triage the same issue twice.
The outcome isn't fewer findings for their own sake. It's a list your team will actually work through, because the noise is labeled.