Coverage area · 4 checks

Input handling and parser robustness

Fuzzers find crashes by throwing input at running code. OpenRouting looks at the same weak spots from the other side: the parsers, regexes and size limits in your source that decide how much damage one malformed request can do.

api/import_users.py1EMAIL = re.compile(2 r"^([a-z0-9]+)*@corp$")3data = request.get_data()4for row in data.splitlines():Inputwhat this review answersNested quantifiers on user input?Body, file or row counts bounded?Parser depth and expansion capped?$ verdict REAL ISSUE$ fix linear regex, size and depth limits▍

Illustrative example of how a finding in this area is reviewed.

Why this area matters

Uncontrolled resource consumption (CWE-400) and inefficient regular expressions (CWE-1333) are a common cause of outages: one request with a crafted string or a deeply nested document ties up a CPU or fills memory. Several high-profile incidents have come from a single backtracking regex in a hot path.

Scanners can flag a risky regex or a missing size limit, but whether it matters depends on whether the input is attacker-controlled and whether a limit is enforced elsewhere — in a proxy, the framework config or a middleware. The review step checks those before calling it real.

What goes wrong in real applications

  • Regular-expression denial of service

    Nested or overlapping quantifiers can take exponential time on certain inputs. A single request can pin a CPU core for seconds or minutes.

  • Unbounded request bodies and uploads

    Reading a whole body or file into memory without a size cap lets one client exhaust the server's memory.

  • Unbounded nesting and expansion

    JSON, XML and YAML parsers without depth or entity limits can recurse until they crash, or expand a small document into a huge one.

  • Missing length and range checks

    Lengths, counts and offsets taken from input and used without validation cause crashes, oversized allocations or out-of-range reads.

Sample: the risky pattern and the fix

Riskyapi/import_users.py
import re

EMAIL = re.compile(r"^([a-zA-Z0-9]+)*@corp[.]example$")

@app.post("/import")
def import_users():
    data = request.get_data()             # no size limit
    for row in data.decode().splitlines():
        if EMAIL.match(row.split(",")[0]):
            save_user(row)
Saferapi/import_users.py
import re

EMAIL = re.compile(r"^[a-zA-Z0-9]+@corp[.]example$")  # no nested quantifier
app.config["MAX_CONTENT_LENGTH"] = 1_000_000           # 1 MB cap, else 413
MAX_ROWS = 10_000

@app.post("/import")
def import_users():
    rows = request.get_data().decode("utf-8", "replace").splitlines()
    if len(rows) > MAX_ROWS:
        abort(413)
    for row in rows:
        if EMAIL.fullmatch(row.split(",", 1)[0]):
            save_user(row)

The pattern ([a-zA-Z0-9]+)* can match the same text in exponentially many ways, so a long string that almost matches makes the regex engine backtrack for a very long time. The endpoint also reads an unbounded body. The fix uses an equivalent linear pattern and caps both the body size and the number of rows.

Illustrative code, simplified for clarity — not taken from a customer repository.

What OpenRouting checks in your repository

  • Regexes with nested or overlapping quantifiers applied to request data
  • Request bodies, uploads and decompression without size limits
  • XML, YAML and JSON parsing without depth or entity-expansion limits
  • Lengths and counts from input used for allocation, loops or indexing
  • Whether a limit is already enforced in framework config or middleware (the main false-positive signal)

Scope and limits

  • It doesn't fuzz or load-test your service; it reviews the code that handles input.
  • Limits enforced outside the repository — a CDN, load balancer or API gateway — aren't visible, so a finding may already be mitigated in production.
  • Resource-limit findings are reviewed by the general reviewer; parser and deserializer findings by the deserialization and injection agents.

Coverage areas reflect OpenRouting's check library. Code, secret and infrastructure scanning run today; other areas are rolling out. Scanners surface candidates and a review agent decides whether each one is real — it doesn't promise to find every issue.

Review agents for this area

Common questions

Is this a fuzzing service?

No. OpenRouting never runs your code. It reviews the input-handling code a fuzzer would exercise, which catches many of the same weaknesses earlier and explains the fix.

How does it know a regex is dangerous?

Semgrep rules flag known risky shapes, such as nested quantifiers. The reviewer then checks whether the regex is applied to untrusted input and whether input length is already capped.