Skip to main content
Curtail

The Research

Catch the bugs your tests can't see.

AI writes a third of new production code. The bugs it ships are the silent ones.

ReGrade pays for itself catching the first bug. ReGrade catches up to +13 extra bugs per trial, and 55 of the 58 models we could measure fixed more, across 63 AI models from 22 labs, at a marginal cost between negative (it pays for itself) and ≤16¢ per extra bug caught, versus an industry-typical $1,500–$50,000 per bug that ships to production.

-83%

on the typical model

It stops breaking what it was not asked to touch.

Unintended changes, additions or deletions to working code the agent was never asked to touch. They fell on 50 of the 58 models we could score, and stopped entirely on 13 models across 8 labs. On 8 they got worse, and every one is named below.

How this is measured
63 models · 22 labs
Alibaba
  • qwen3-coder 1.30.7
  • qwen3-coder-plus 30.3
  • qwen3.6-max-preview 30.3
  • qwen3.6-plus 1.32.3 *
  • qwen3.7-max 1.70
  • qwen3.8-max 1.50
Anthropic
  • fable-5 3.30
  • haiku-4.5 4.25.9 *
  • haiku-4.5 (Aug 2026) 4.37.2 *
  • opus-4.6 1.30
  • opus-4.7 2.10
  • opus-4.8 2.20
  • opus-5 1.80
  • sonnet-4.6 2.10
  • sonnet-5 4.11.8
Google
  • gemini-2.5-flash 117
  • gemini-3-flash-preview 1.20.3
  • gemini-3.1-pro-preview 30
  • gemini-3.5-flash 00.3
  • gemini-3.6-flash 00.2
MiniMax
  • minimax-m2.5 4.13.5
  • minimax-m3 20
Moonshot
  • kimi-k2.7-code 7.30
  • kimi-k26 2.70.5
  • kimi-k3 70
Nvidia
  • nemotron-3-ultra 3.22.2
  • nemotron-3.5-lightning 2.80
OpenAI
  • gpt-5.4 5.70.6
  • gpt-5.4-mini 3.61
  • gpt-5.5 3.80
  • gpt-5.6-luna 3.81.9
  • gpt-5.6-sol 4.90
  • gpt-5.6-terra 4.60.1
  • gpt-oss-20b 00
Sakana
  • fugu-ultra 2.40
  • sakana-namazu 4.30.2
xAI
  • grok-3-fast 5.30
  • grok-3-mini 2.50
  • grok-4-1-fast-non-reasoning 1.50
  • grok-4.1-fast 00
  • grok-4.20-0309-reasoning 7.58.1 *
  • grok-4.3 151.3
  • grok-4.5 9.50
Meta
  • muse-spark-1.1 3.30.3
  • muse-spark-1.2 110.7
DeepSeek
  • deepseek-v4-flash 20.3
  • deepseek-v4-flash-0731 2.23.1 *
  • deepseek-v4-pro 2.32
Kuaishou
  • kat-coder-pro-v2.5 30.5
Mistral
  • devstral-2 9.31.6
  • mistral-medium-3.5 2.26 *
Z.ai
  • glm-5.1 3.10.5
  • glm-5.2 2.61.6
ByteDance
  • seed-2.0-lite 2.60.8
Meituan
  • longcat-2.0 3.52.3
StepFun
  • step-3.7-flash 5.23.5
InclusionAI
  • ling-2.6-1t 4.23.3
  • ling-3.0-flash 2.23.8 *
Poolside
  • laguna-s-2.1 00.3
  • laguna-xs-2.1 4.73.7
Upstage
  • solar-pro4 2.22
ThinkingMachines
  • inkling 7.77.3
Tencent
  • hy3 3.84.2 *

improved · * got worse · – no baseline to compare against. All four measured over the same corpus: 63 models, 22 labs, 1,433 agent-mode trials. Open and reproducible.

🧹 Fixes more, and cleans up after itself.

A tool that fixes bugs but quietly breaks other things is not a safety win. So we measured the other side of the ledger: how many clean, working fields the agent damaged while making its fix. In our pre-registered head-to-head (n=30 per model), the behavioral diff cut collateral damage to near zero on every capable model.

ProviderModelClean fields brokenWith ReGrade
OpenAIOpenAIGPT-5.429.44.0
OpenAIOpenAIGPT-5.519.60.0
AnthropicAnthropicOpus 4.75.80.3
AnthropicAnthropicOpus 4.85.80.0
AnthropicAnthropicSonnet 4.62.60.3

Mean clean, v1-matching response fields the agent broke per run, lower is safer.

The behavioral safety net cuts both ways: more of the real bugs caught, far fewer new ones introduced.

The honest caveat: this is strongest on the models with the highest baseline error rates. On a few high-rescue models, more fixing comes with more editing, and collateral changes can rise, a trade-off we report in full in the paper. The clean, controlled result above is the claim we stand behind.

⚡ Productivity: 177% more bugs fixed per hour.

The ReGrade context cuts down on the back-and-forth where the agent runs shell commands and re-reads files trying to figure out what changed, the agent gets to the answer faster.

ProviderModelWithoutWithSpeed-up
QWenAlibabaQwen3.6-Max-Preview50m 7s8m 4s84% faster
Google GeminiGoogleGemini 3.1 Pro Preview10m 22s6m 46s35% faster
OpenAIOpenAIGPT-5.44m 1s2m 48s30% faster
OpenAIOpenAIGPT-5.55m 17s4m 13s20% faster
QWenAlibabaQwen3.6-Plus25m 26s21m 25s16% faster

Bonus: in four models (Sonnet 4.6, Opus 4.7, GPT-5.5, Qwen3.6-Plus) every trial produces the identical outcome, trial-to-trial randomness drops to zero. No more flaky CI retries.

💰 Tokens per bug fixed drop 51%.

A ReGrade run spends about 33% more tokens on the typical model, not fewer: the behavioral diff is extra input on every turn, and only 25 of the 60 models we measured ran leaner per run. It finds so many more bugs that each fix costs far less anyway. Tokens per bug fixed fall 51% on the typical model, and on 44 of 56. In dollars at list rates, cost per fix falls 51% on 47 of the 55 we could price. Even where it does not pay for itself outright, the marginal cost stays under 16¢ per extra bug caught, against an industry-typical $1,500–$50,000 per bug that ships to production.

ProviderModelWithout $/bugWith $/bugChange
QWenAlibabaQwen3.6-Max-Preview$0.333$0.036-89%
XxAIGrok 4.20 Reasoning$0.235$0.077-67%
OpenAIOpenAIGPT-5.5$0.156$0.091-42%
AnthropicAnthropicSonnet 4.6$0.126$0.092-27%
OpenAIOpenAIGPT-5.4$0.020$0.015-25%
AnthropicAnthropicOpus 4.7$0.257$0.199-23%

47 of the 55 models we could price get cheaper per bug fixed with ReGrade. We report tokens on the cards because a token count does not go stale when a vendor reprices.

📊 The test, by the numbers.

ReGrade is a behavioral-diff context block your AI coding agent reads alongside its existing prompt, it shows the agent how the new code's runtime behavior differs from the old, the way a human checks for regressions at code review. No model retraining. No changes to your build or CI pipeline.

What we tested. Each trial = one AI coding agent (running autonomously, the way Claude Code or Codex CLI operate: shell tools, file access, code edits) attempts to fix 18 known bugs planted in a 300,000-line Python codebase. 1,433 agent-mode trials across 63 AI models from 22 labs, US and non-US, from frontier flagships down to free tiers. Every model gets a control run and a ReGrade run on the same bugs, so each model is its own comparison.

What counts as a hallucination. Unintended changes, additions or deletions to working code the agent was never asked to touch. We measure it by replaying the agent's patch against the original behavior and counting what changed that nobody requested. A tool that fixes bugs while quietly breaking other things is not a safety win, so we report both sides of that ledger.

63

AI models

22

labs

1,433

agent-mode trials

18

known bugs

300,000-line Python

codebase

-83%

hallucinations

🛠️ Works with every major coding agent and model API.

3 agent CLIs tested · 17 LLMs across 6 providers.

Coding-agent CLIs

Anthropic

Claude Code

  • Haiku 4.5
  • Sonnet 4.6
  • Opus 4.6
  • Opus 4.7
OpenAI

Codex CLI

  • GPT-5.4
  • GPT-5.5
QWen

qwen-code

  • Qwen3-Coder
  • Qwen3-Coder-Plus
  • Qwen3.6-Plus
  • Qwen3.6-Max-Preview

Models tested via API

Google Gemini

Gemini

  • Gemini 3 Flash Preview
  • Gemini 3.1 Pro Preview
X

Grok

  • Grok 3 Mini
  • Grok 3 Fast
  • Grok 4.20 Reasoning
DeepSeek

DeepSeek

  • V4 Pro
  • V4 Flash

One ReGrade context block in the agent's prompt. No model retraining. No changes to your build or CI pipeline. Tested working across 63 models from 22 labs, US and non-US.

🎯 Nearly every model improves.

+6% to +437% bug-fix improvement, on 55 of the 58 models we could measure. Without ReGrade, the best model fixes only 13.9 / 18. With ReGrade, four frontier models hit the ceiling and almost every other one moves up.

ProviderModelWithoutWithΔ
AnthropicAnthropic ★Fable 5 (Mythos-class)11.5 / 1818.0 / 18+57%
AnthropicAnthropicOpus 4.710.6 / 1817.8 / 18+67%
AnthropicAnthropicSonnet 4.67.3 / 1817.6 / 18+141%
OpenAIOpenAIGPT-5.513.3 / 1818.0 / 18+36%
OpenAIOpenAIGPT-5.413.9 / 1817.9 / 18+29%
DeepSeekDeepSeekDeepSeek V4 Pro5.3 / 1816.6 / 18+211%
XxAIGrok 4.20 Reasoning3.2 / 1812.3 / 18+284%
Google GeminiGoogleGemini 3.1 Pro Preview8.3 / 1814.5 / 18+76%
QWenAlibabaQwen3.6-Max-Preview5.7 / 1813.8 / 18+144%
MoonshotMoonshot ★Kimi K2.6 (native CLI)4.4 / 1816.0 / 18+261%

Bug-fix counts from blind agent-mode evaluation, regenerated from the corpus 2026-08-11. Depth varies by cell: the frontier models run deepest (Opus 4.7 n=84, Sonnet 4.6 n=87, GPT-5.4 n=79, GPT-5.5 n=67 across both arms), while the thinnest cells sit at n=8. Treat a small cell as provisional, means move as they fill out. ★ Fable 5 (Mythos-class, n=32) is one of the most advanced models we tested, even it left a third of the bugs unfixed on its own, and ReGrade took it to a perfect 18/18. ★ Kimi K2.6 is a preliminary n=13 cell, engaged-subset means.

🌍 The newest labs, too.

The behavioral-diff effect isn’t a quirk of the big US providers. We ran the same test on the newest releases from five more labs (each at 30 trials per model) and the rescue held every time, statistically decisive in each case.

LabModelWithoutWithΔ
ByteDanceSeed-2.0-Lite2.9 / 1814.7 / 18+407%
MiniMaxM2.52.1 / 1810.9 / 18+419%
ZhipuGLM-5.14.5 / 1817.6 / 18+293%
StepFunStep 3.7-flash3.4 / 1811.0 / 18+221%
MistralDevstral 25.4 / 1813.0 / 18+141%

63 models across 22 labs, US and non-US, from frontier flagships down to free tiers. 55 of the 58 we could measure fixed more silent regressions with ReGrade.

✅ The gains are real: proven, not assumed.

In a 600-trial head-to-head test whose design we specified in advance (so the results couldn't be cherry-picked after the fact), ReGrade beat all three comparison conditions, across every model tested.

Control conditionBugs fixedvs ReGrade
No extra context (just the test suite)8.6 / 1816.9 / 18
Random gibberish, same length as ReGrade8.7 / 1816.9 / 18
Unrelated source code, same length as ReGrade8.4 / 1816.9 / 18

What moves the needle is the behavioral signal in ReGrade, not the volume of extra context. Less than 1-in-a-billion-trillion chance the difference is a fluke.

🕵️ It catches what your tests can't.

Functional tests check status codes and required fields. They routinely miss the behavioral drift that actually breaks production: a CORS header that quietly changed, a list that's now an object, a timestamp format that broke a downstream consumer. ReGrade compares runtime behavior directly, so the agent sees the drift the test suite waved through.

35%

of AI patches that passed every functional test still silently broke behavior

191 of 548 test-passing patches, across 14 of 26 tasks

0 → 9

multi-file bug sites fixed, weaker models fixed 0 of 9 on their own

with ReGrade, agents clear all 9 sites of a cross-file regression

18

classes of silent drift planted, headers, encoding, ordering, response shape

the behavioral defects that pass type-checks, lint, and existing tests

Your test suite says “green.” ReGrade says “but the API’s behavior drifted.” That gap is exactly where production incidents come from.

🔎 Better answers, and better detection.

Finding a bug is only half the job, a reviewer still has to understand and trust the fix. Even when a model already spots a regression, ReGrade’s behavioral evidence helps it explain exactly what changed and where.

more precise bug explanations on a production codebase (Ghost CMS)

explanation precision rose from 0.40 to 0.91

9 of 10

models explained the bug better on a 291,000-line codebase (NetBox)

+3.2 points on a 20-point rubric, even where the model already found it

The difference between “something looks off” and a root-cause a human can act on with confidence.

Want ReGrade for your AI coding agent?

Drop the context block into your agent's prompt. No retraining. No CI changes.