Skip to main content

Security analysis without wasting tokens

· 8 min read
Vulkro
Security research

Ask an AI coding agent to "check this repository for security issues" and watch what it does. It lists the tree. It opens the route files, then the middleware, then the database layer, then a few helpers it was not sure about. Each file goes into its context window. By the time it has an opinion about one endpoint, it has paid to read forty files that had nothing to say.

That is not a flaw in the agent. It is the only way a model can look for a vulnerability on its own: by reading. And reading a codebase is the most expensive thing you can ask a model to do.

Reading code costs tokens​

Every model works inside a context window: the text it can see at once. Agents fill it with what they read, what tools return, and what they have already said. Two things follow.

It is finite, and it gets worse as it fills. Engineers who build agents describe context as "a finite resource with diminishing marginal returns" and note that recall falls as the window grows (Anthropic, 2025). Academic work found the same shape earlier: models use information at the start and end of a long input well, and "significantly" worse when it sits in the middle (Liu et al., TACL 2023). A vulnerability is, almost by definition, a detail in the middle of a lot of ordinary code.

You pay for all of it. Every file the agent opens becomes input tokens, and later turns carry it along until the context is compacted or cleared. A security review is the worst case for this, because the agent cannot know in advance which file matters. Injection in a query builder, a missing ownership check in a handler, a secret in a config file: to rule each one out, it has to read where each one could be.

So the question for anyone running agents at scale is not "can the model find the bug?" It is "how much did it read to find it, and will it find the same one tomorrow?"

Why reading code to find bugs uses so much context​

A vulnerability is rarely on one line. It is a path: a value arrives at an entry point, moves through a few functions, and reaches a call that trusts it. To see that path, a reader has to hold the entry point, every hop and the sink at once.

A model reading a repository does this the hard way:

  1. It does not know where the entry points are, so it searches for them.
  2. It does not know which calls are dangerous in this framework, so it reads the call sites to decide.
  3. It does not know whether a guard sits between the two, so it reads the middleware, the decorators and the router configuration.
  4. It does not remember the last review, so next week it does all of it again.

None of these steps produces the finding. They produce the context the finding needs. That context is what you pay for, and on a real codebase it is most of the bill.

Give the agent the finding instead​

Vulkro inverts the order. The analysis runs first, on your machine, deterministically. The agent asks for the result.

Detection uses zero model tokens. Vulkro's scan engine calls no model. It parses the code, maps every entry point the frameworks declare, follows the data from each one, and records which checks sit in the way. A finding comes back with the rule, the file, the lines, and the path that connects the entry point to the risky call.

The agent reads the finding, not the tree. Over MCP, the agent calls scan_project and gets a JSON result it can filter. It does not need the forty files. It needs the four lines where the problem is and the hops that prove it, and that is what it gets:

  • scan_project returns the findings, with a scan_id. With format: summary it returns only the counts.
  • get_findings re-filters that same scan by severity without scanning again.
  • prove returns the hop-by-hop chain for one finding: each hop's file, line and source line, plus whether the path is proven. An empty chain means no proven flow was found, never that the code is safe.
  • explain returns what a rule means and how to fix it.

Only what changed, when that is the question. Most agent work is a change, not an audit. scan_diff analyses the whole tree, so a flow that crosses files stays intact, then returns only the findings on the lines the diff added or modified. The agent sees what its own edit introduced, not the backlog it inherited. When it needs to know where to look before it edits, code_graph returns a ranked map of the most important symbols and what depends on them, instead of the agent listing directories and guessing. Both are part of Vulkro Pro.

The agent spends its context on the fix. With the finding, the lines and the path in hand, the model is doing what it is good at: changing a few lines of code correctly. It is no longer doing the part a deterministic engine does better.

Typical AI workflow

  1. Huge codebaseevery file is a candidate
  2. Massive contextfiles pulled in to be read
  3. Massive token usagereasoning over all of it
  4. Expensive analysisand a different answer next run
Model tokensBaseline

Vulkro

  1. Codeand the org metadata
  2. Targeted analysisdeterministic, no model
  3. Relevant contextthe path, the rule, the lines
  4. Actionable findingwhat to fix, and where
Model tokens98% fewer

98% fewer tokens in agentic development and security review, measured against an AI agent reading the codebase itself to find the same issues (internal measurement, vulkro 0.28.0, September 2026). Detection itself uses no model at all.

The figure, and what it is measured against​

On our own measurement, this saves 98% of the tokens in agentic development and security review, measured against an AI agent reading the codebase itself to find the same issues (internal measurement, vulkro 0.28.0, September 2026).

The baseline matters, so to be plain about it: the comparison is between an agent that asks Vulkro for the answer and an agent that reads the codebase itself looking for the same issues. It is not a claim about every task an agent does, and it is an internal measurement rather than an independent one. The part that needs no measurement is the detection itself: it uses no model, so it uses no tokens.

The same answer without paying again​

The less obvious cost of model-only review is repetition. Ask a model to review the same code twice and you may get two different lists. Research on model-based vulnerability detection has found exactly this: responses that are non-deterministic, and answers that change when only function or variable names change (Ullah et al., IEEE S&P 2024). Every rerun is paid for again, and every difference between runs has to be triaged by a person.

A deterministic engine does not have that problem. The same code gives the same findings, every run. Findings carry stable identifiers, so a finding fixed last week does not reappear under a new description, and a suppression or triage decision stays attached to the issue it was made on. You pay for the analysis once, in CPU time on your own machine, and the agent reads the result as many times as it needs.

That is also what makes the result usable as a gate. A check that might say something different on the next run cannot block a merge. A check that cannot change without the code changing can.

What this looks like in practice​

Wire Vulkro into your agent once:

# Claude Code
claude mcp add vulkro -- vulkro mcp serve

# Or the skill, for Claude Code, Cursor and Codex
curl -fsSL https://dist.vulkro.com/skill-install.sh | bash

Then the loop becomes short:

  1. The agent makes a change and commits it.
  2. It calls scan_diff and gets back only the findings its change introduced.
  3. For any finding it is unsure about, it calls prove and reads the path.
  4. It drafts a fix and calls verify_fix, which applies the diff to a temporary copy, scans again, and says whether the finding is gone and whether anything new appeared. Your working tree is never touched.

Everything the scan reads stays on your machine. Set VULKRO_OFFLINE=1 and the scans make no network call at all. The server is read-only: there is no tool that writes to your repository.

The same approach works for Salesforce: vulkro-sf mcp serve exposes the Salesforce scanner to the same agents, and in the VS Code extension the scanner is available to the editor's own agent as language-model tools.

Summary​

An agent that reads a repository to find vulnerabilities pays for every file it reads, forgets it all by the next review, and may give a different answer each time. An agent that asks a deterministic engine pays for a few findings, gets the same answer every run, and spends its context on the fix.

See how it fits your setup on the token efficiency page, read about Vulkro Core, or install it and point your agent at it.

Sources​