Automated SAST Remediation: Deterministic vs Probabilistic Fixes
Most teams already accept automated dependency upgrades. A bot opens a pull request, CI runs, somebody merges it. Automated SAST remediation is a harder sell, and for good reason. A dependency fix changes a version number. A code fix changes what your program does.
That difference is why the question "can AI fix our SAST findings?" is the wrong one. Almost any large language model can produce a plausible patch for a SQL injection. The real question is whether the patch is the right one, whether anyone can predict what it will be, and whether it arrives with enough evidence that a reviewer can approve it without redoing the analysis.
This post lays out how we think automated SAST remediation should work: where determinism matters, where a model genuinely helps, and what a reviewer should demand before a machine-written code fix gets anywhere near the default branch.
Why automated SAST remediation is harder than dependency upgrades
A dependency upgrade has a small, well-defined blast radius. You move from version A to version B, the lockfile changes, and the behavior change is whatever the maintainers shipped. You can reason about it from the changelog and the test suite.
A SAST finding has no equivalent. The weakness lives in code your team wrote, often in a pattern that repeats across the codebase with small variations. Fixing it means editing logic, and every edit has side effects:
- A parameterized query can change how nulls, wildcards or type coercion behave.
- Output escaping can double-encode data that was already escaped further up the stack.
- An input allowlist can reject values that production traffic actually sends.
- Path normalization can break a legitimate relative path that a feature depends on.
None of these are reasons to avoid automation. They are reasons to be precise about which parts of the process are allowed to be creative.
Deterministic vs probabilistic fixes
There are two broad ways to produce a code fix automatically.
Probabilistic generation hands the finding and some surrounding code to a model and asks for a patch. The output depends on the prompt, the context window, the model version and a sampling seed. Run it twice and you can get two different fixes. Both may compile. Both may pass tests. Only one may be correct, and nothing in the process tells you which.
Deterministic generation decides the fix from the analysis itself. The data-flow trace says a tainted value reaches a SQL sink through string concatenation, so the fix is to bind that value as a parameter. The decision is a function of the finding, not of how the model felt that afternoon. Run it twice and you get the same change.
Determinism buys you three things that matter for review:
- Predictability. The same weakness in the same shape gets the same fix. Reviewers learn what to expect and can approve a class of change instead of re-deriving each one.
- Explainability. The fix can name its strategy, because the strategy was chosen explicitly. "Parameterize" is a claim a reviewer can check against the diff.
- Auditability. When someone asks why a line changed six months from now, there is a reason attached to the change that does not depend on reconstructing a prompt.
Where a model still earns its place
Determinism decides what to change. It does not know that your repository needs a specific environment variable before the test suite will pass, that one module wraps its database calls in a helper, or that the obvious edit breaks a downstream caller. That is applied engineering work, and it is where an agent is useful: taking a known-correct change and making it fit the codebase, build and pass CI.
The split we argue for is simple. The analysis decides the fix. The agent applies it and gets it to green. The agent does not get to invent a different fix because the first one was inconvenient.
Match the fix strategy to the weakness
A useful automated fix addresses the root cause of the weakness class, not the symptom in one line. The common strategies map cleanly to the most frequent injection and traversal findings:
- Parameterize. Replace string interpolation in a query or command with safe parameter binding. This is the standard remediation for SQL injection in the OWASP SQL Injection Prevention Cheat Sheet.
- Escape. Apply the escaping that matches the output context: SQL, HTML, shell. The context is the whole point. HTML escaping in a shell command does nothing useful.
- Allowlist. Constrain input to a known-safe set. Best when the legitimate values are genuinely finite, such as a sort column or a report type.
- Path normalize. Resolve and validate file paths before use, to stop directory traversal.
Here is the kind of change a parameterize fix should produce. Before:
cursor.execute("SELECT * FROM orders WHERE customer_id = '" + customer_id + "'")
After:
cursor.execute("SELECT * FROM orders WHERE customer_id = %s", (customer_id,))
A reviewer can verify that change in seconds. The value is bound, the query text is constant, and nothing else moved. Compare that with a model-generated patch that also renames variables, reformats the function and adds a helper "for clarity". The fix might be correct, but the review just got five times longer.
Some findings should not be auto-fixed
Not every weakness has a local fix. An authorization flaw that comes from a missing policy layer, a deserialization design that trusts client input, or a crypto choice baked into a protocol all need a human decision. A good automated SAST remediation system says so. It routes those findings to written remediation guidance instead of guessing, and it tells you how confident it is in the fixes it does produce.
Validation: the step most tools skip
A fix that compiles is not a fix. A fix that passes the tests in a sandbox is closer. A fix that passes your CI, with your build, your linters and your integration tests, is the only one that should reach a reviewer as "ready".
A credible validation loop has two phases:
- Sandbox build before the pull request exists. Provision the real toolchain at the version the repository uses, apply the change, build it and run what can be run. If it does not build, do not pretend it did.
- Your CI after the pull request opens. Listen to the checks. When something fails, read the logs, fix the cause and push a follow-up commit. Stop after a bounded number of attempts and hand off with an explanation rather than looping forever.
The hand-off matters as much as the success path. An agent that silently gives up leaves a stale branch. An agent that explains what it tried, what failed and why leaves a reviewer with a head start.
What a reviewer should demand from an automated code fix
If you are evaluating any automated SAST remediation approach, including ours, hold it to this list:
- A scoped diff. Only the files the fix touches. No drive-by refactors.
- A named strategy. The pull request says which fix strategy it applied, so the reviewer checks the diff against a claim.
- A confidence level with reasons. "Lower confidence because the query spans multiple files" tells the reviewer where to look.
- A link back to the finding. Including the data-flow trace that justified the change.
- Evidence of validation. Which checks ran, which passed, and what the agent did about the ones that did not.
- Human review by default. Automation opens the pull request. A person merges it.
- A record. Every run, its outcome and its pull request, somewhere a security lead can audit later.
Anything less puts the burden back on the reviewer, and the whole point of automation was to take it off them.
How Heeler approaches SAST Auto-fix
Heeler's SAST Auto-fix follows the split described above. As part of every SAST scan, Heeler evaluates each finding and decides whether a concrete before/after fix can be generated, and stores that decision with the finding. Each fix uses one of four named strategies: Parameterize, Escape, Allowlist or Path Normalize. Where a fix is available, Heeler shows its confidence and the factors that lowered it. Findings that need a broader architectural change get written remediation guidance instead of an automated fix.
The finding carries a Suggested Fix card with the vulnerable code, the proposed change as a unified diff, the strategy, the confidence level and an effort rating. You can copy the fix into your editor, or run it with Fix Now.
When you run it, the remediation agent applies the change in an isolated sandbox with the repository's real build toolchain, then opens a pull request scoped to the files the fix touches. If the sandbox build did not pass, the pull request opens as a draft. Heeler then runs the change through your own CI and iterates with follow-up commits until the checks pass, up to five attempts, before handing off with a comment explaining what it could not resolve. The full loop is documented in Validate and Merge-Ready.
Two details are worth calling out:
- Memories shape how a fix is applied, not what it is. The agent reads repository conventions before a run and writes back what it learned afterwards, such as which suite gates a change. Those memories never change the fix strategy or the before/after code.
- Every run is recorded. SAST fix runs appear in Agent Executions with the finding, the pull request, its status and each CI iteration.
Whether a fix opens a pull request automatically or waits for review is a tenant-wide default an administrator sets. See SAST Auto-Fix in the docs for the details.
A practical rollout for automated code fixes
You do not need to trust automation with your whole backlog on day one. A sequence that works:
- Start with one weakness class. SQL injection with a parameterize fix is the easiest to review and the easiest to verify.
- Review every pull request for two weeks. Track how many you merge unchanged, how many need edits and how many you reject. That ratio is your real confidence metric.
- Write down what the agent keeps getting wrong. Usually it is a repository convention. Record it once so every later run starts with it.
- Widen by strategy, not by volume. Add escaping, then allowlisting, then path normalization, each after its own review period.
- Keep humans on the merge button. The goal is fewer hours per fix, not zero review.
FAQ
What is automated SAST remediation?
Automated SAST remediation is the generation and delivery of a code change that fixes a static analysis finding, usually as a pull request. The useful versions also validate the change in CI and explain what they changed and why.
Is an AI-generated SAST fix safe to merge?
A fix is safe to merge when a reviewer can verify it, not because a model produced it. Look for a scoped diff, a named fix strategy, a confidence level and evidence that your own CI passed.
What is the difference between deterministic and probabilistic code fixes?
A deterministic fix is decided by the analysis, so the same finding always produces the same change. A probabilistic fix is generated by a model and can differ between runs, which makes review and audit harder.
Which SAST findings can be fixed automatically?
Findings with a local, well-understood remediation, such as SQL injection, output escaping, input allowlisting and path traversal, are good candidates. Findings that need an architectural or policy decision should go to a person with written guidance.
Should automated fixes merge without review?
No. Automation should open a validated pull request. A human should decide whether it merges.
See it on your own code
If your SAST backlog is growing faster than your team can review it, the fix is not more triage. It is fixes that arrive validated, scoped and explained. See how Heeler generates, validates and delivers code fixes as reviewable pull requests on your repositories. Get a demo.


.jpg)
