Review AI-written code with evidence
Joint study with AFRL
FrozenReGrade helps AI coding agents fix 152% more bugs.
Every model we tested fixed more bugs.
A joint study with the Air Force Research Laboratory, cleared for public release. 17 language models from 6 providers, and the median model fixed 152% more bugs at a median 42% lower LLM API cost per bug fixed.
The study supplied behavioral evidence in the agent's initial prompt.
See the evidence for this →Why teams use ReGrade
Three minutes on the problem it solves: catching the changes no test asserts, knowing a release holds up under load, and proving a refactor behaves like the code it replaced.
Want to watch one run start to finish? See the runnable demos. Each one has a repo you can clone.
The code is changing. The risks are growing.
Proven in the Wild
A vulnerability hid for 7 years. Zero tests caught it. ReGrade found it on the first replay.
A widely-used open-source platform shipped a password hash disclosure bug for 7 years. Standard tests passed every day. ReGrade caught it on the first replay, without knowing the vulnerability existed.

example ReGrade merge request comments
7 Years
Undetected
First Replay
Found Immediately
Zero Prior Knowledge
Used existing tests
Powered by NCAST Technology
Compare behavior before you approve a release.
ReGrade sends identical real-world requests to your current and candidate versions, then compares the responses field by field. It runs on the traffic and the tests you already have.
1
Record
Capture traffic from any source: production, CI tests, or security scanners. ReGrade works with whatever you already have.
2
Replay
Replay the same traffic against your candidate version. The original responses were already captured in step one.
3
Compare
Review field-level and performance differences to identify unintended changes before release.
What testing and code review miss
Tests verify what you expect. ReGrade catches what you don't.
Slow before anyone complains
ReGrade compares P95 and P99 latency across versions, so a release that quietly got slower is a number you see on the merge request rather than a support ticket next week.
Zero-Day Discovery
By comparing actual responses against a known-good baseline, ReGrade surfaces vulnerabilities that no test was written to find, because no one knew they existed.
Give your agents evidence for repairs
ReGrade supplies behavioral findings to the coding agent so it can attempt repairs. Your engineers review the changes and decide which differences are acceptable.
It meets your team where they already work
MCP Server
Connect your coding agents through MCP so they can investigate behavioral differences and attempt repairs.
Merge Request Analysis
Automatically analyze candidate builds before merge. ReGrade comments appear alongside your code review with specific field-level findings.
Preview Before Deploy
Replay production traffic against your candidate version in a safe environment. Inspect differences on the requests you replay before approving a release.
Plan a ReGrade evaluation
Choose one service and compare a baseline with a candidate build. Discuss what changed, how your team will review findings, and what recording and replay will cost.
