Deterministic analysis vs asking a model
"Paste the repo into the model and ask it to find the security bugs" has become a common first attempt at AI-assisted security review. It is easy to try, and the first run is often impressive: the model names a plausible injection, explains it well, and suggests a fix.
The trouble starts on the second run. And the third. And when somebody asks what it did not look at.
This is not an argument against models in security work. It is an argument about which job each tool should have. Finding vulnerabilities is a search problem with a right answer. Fixing them is a writing problem with many acceptable answers. Those want different tools.
Problem one: a different answer each time
A security check is useful when it is repeatable. Run it on the same code and it should say the same thing, or the result cannot gate a merge, cannot show that a fix worked, and cannot be compared with last month.
Model output is not repeatable in that sense, and the published evidence is consistent about it.
- An IEEE S&P 2024 evaluation of eight models on 228 code scenarios found that models "provide non-deterministic responses, incorrect and unfaithful reasoning, and perform poorly in real-world scenarios". Renaming functions or variables, a change that does not alter the program's behaviour at all, produced incorrect answers in 26% of cases (Ullah et al.).
- A separate study ran five models over eight tasks with settings configured to be deterministic, ten runs each, and saw accuracy vary by up to 15% between runs. Its summary: none of the models "consistently delivers repeatable accuracy across all tasks, much less identical output strings" (Atil et al.).
For a security review this has a practical cost. If run one reports five issues and run two reports four, someone has to work out whether the fifth was fixed, missed, or never real. That person's time is the most expensive line in the whole exercise.
Problem two: coverage you cannot see
When a deterministic analyser does not report a vulnerability, you can usually say why. The file was not in a supported language. No entry point reached the function. A guard sat between the input and the call. The rule did not apply to that framework. These are inspectable reasons, and a good tool will tell you which one it was.
When a model does not mention a vulnerability, there is no such record. It may have read the file and judged it fine. It may have read it and lost the detail in the middle of a long context, a pattern documented in long-context research (Liu et al.). It may never have opened it. The silence looks the same in all three cases.
Benchmarks underline how wide the gap between an impressive demo and real code can be. A study accepted at ICSE 2025 built a stricter vulnerability dataset and found a code model that scored 68.26% F1 on an older benchmark dropped to 3.09% F1 on the new one, with results "akin to random guessing" in the most stringent settings (Ding et al.). Those were fine-tuned code models classifying single functions, not today's agents reviewing a whole repository, so the figure should not be stretched further than that. The lesson it carries is narrower and still useful: a model's apparent accuracy depends heavily on how the question was set up, and in a review you do not get to see the setup.
Problem three: you pay to read the code on every run
A model can only look for a vulnerability by reading. To find a missing ownership check, it has to read the handler, the router, the middleware and the query, and to rule it out everywhere else it has to read everywhere else. Each review pays for that reading again, because nothing the model learned last time is kept.
Context is also not free in quality terms. People who build agents describe it as a finite resource whose recall drops as it fills (Anthropic). The more of the repository a model holds, the less reliably it uses any one part of it.
We have written about the token side of this in detail in Security analysis without wasting tokens.
Problem four: a person still has to triage the results
Every finding a reviewer has to check by hand costs time, and an unreliable finding costs the most, because it has to be checked from scratch. A model that explains a non-existent injection fluently is harder to dismiss than a terse rule hit, not easier.
A detector that shows its work changes this. If each finding carries the entry point, each hop the data takes, and the call it reaches, a reviewer checks a path rather than an opinion. If the path is incomplete, the finding says so.
What a detector does instead
Vulkro is built on the opposite primitive. The scan engine calls no model. It parses the code, maps every entry point the frameworks declare, follows the data from each one, and records the checks it passes on the way. Then it reports:
- The same answer every run. Same code, same findings, with stable identifiers, so a fix can be shown to have worked and a triage decision stays attached to the issue.
- Visible coverage. Each finding is marked proven, unproven or not checked, and the engine can say why nothing was reported at a given line.
- The path, not a verdict. The
provetool returns the hop-by-hop chain behind a finding. An empty chain means no proven flow was found, never that the code is safe. - Zero model tokens for detection. The analysis costs CPU time on your machine, once.
It is not perfect, and we publish where it falls short: on our benchmark the misses sit next to the hits. The difference is that a miss is a reproducible, inspectable miss, not a mood.
Where a model genuinely helps
None of this makes models useless in security work. It puts them where they are strong.
Drafting the fix. Given a finding, the exact lines and the path, a model is good at writing the change: a bound parameter instead of string concatenation, an ownership check in the right place, a narrower permission. The context is small and specific, which is the setting models handle best.
Explaining the finding. Turning a path into a paragraph a product manager can read, or into a ticket that says what will break and why, is writing. Models write.
Then the detector judges. The step that matters is what happens after the draft. In Vulkro, a fix a model proposes is checked by the same deterministic engine that found the problem: the verify_fix tool applies the diff to a temporary copy, scans it again, and returns fixed, not-fixed or regressed. The VS Code extension does the same for AI-drafted fixes and only offers a candidate after a fresh scan confirms the finding is gone and the file still parses. Vulkro's own AI features are opt-in and advisory: they never change a finding, its severity, the report or an exit code.
That division of labour is the point. The model proposes; the engine, which gives the same answer every time, decides whether the proposal worked.
The right approach
"Ask a model to find the bugs" asks a probabilistic writer to perform an exhaustive, repeatable search, and then asks a person to check its work. "Run the analysis, hand the model the finding" asks each tool to do the job it is built for.
If your team is already using agents, the change is small: give the agent the analyser as a tool, over MCP, and let it spend its context on the fix. See how that works, read about Vulkro Core, or get started.
Sources
- Ullah et al., "LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks", IEEE S&P 2024: https://arxiv.org/abs/2312.12575
- Atil et al., "Non-Determinism of 'Deterministic' LLM Settings", arXiv 2408.04667: https://arxiv.org/abs/2408.04667
- Liu et al., "Lost in the Middle: How Language Models Use Long Contexts", TACL 2023: https://arxiv.org/abs/2307.03172
- Ding et al., "Vulnerability Detection with Code Language Models: How Far Are We?", ICSE 2025: https://arxiv.org/abs/2403.18624
- Anthropic, "Effective context engineering for AI agents", September 2025: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents
