More bugs fixed
Curtail research · 63 models, 22 labs, updated 2026-08-29
The AFRL study measured this across 17 models in April to May 2026. These figures cover 63 models. Read the study as published.
Every model tested, on bugs fixed
| Verdict | |||||||
|---|---|---|---|---|---|---|---|
| laguna-xs-2.1 | 0.70 | 9.70 | +1350% | — | Poolside | 6 | ✓ |
| qwen3-coder-plus | 2.00 | 11.8 | +488% | — | Alibaba | 7 | ✓ |
| haiku-4.5 (Aug 2026) | 2.70 | 14.5 | +437% | 44 | Anthropic | 20 | ✓ |
| minimax-m2.5 | 2.10 | 10.9 | +427% | — | MiniMax | 60 | ✓ |
| haiku-4.5 | 2.50 | 12.8 | +413% | 44 | Anthropic | 60 | ✓ |
| seed-2.0-lite | 2.90 | 14.7 | +408% | — | ByteDance | 60 | ✓ |
| kat-coder-pro-v2.5 | 3.70 | 17 | +364% | — | Kuaishou | 5 | ✓ |
| nemotron-3-ultra | 3.00 | 13.8 | +360% | 49 | Nvidia | 10 | ✓ |
| ling-2.6-1t | 1.60 | 7.00 | +338% | — | InclusionAI | 10 | ✓ |
| qwen3.6-plus | 4.00 | 16 | +300% | 55 | Alibaba | 7 | ✓ |
| ling-3.0-flash | 2.90 | 11.3 | +294% | 51 | InclusionAI | 29 | ✓ |
| glm-5.1 | 4.50 | 17.6 | +292% | 56 | Z.ai | 60 | ✓ |
| grok-4.20-0309-reasoning | 3.20 | 12.3 | +284% | — | xAI | 21 | ✓ |
| nemotron-3.5-lightning | 1.40 | 5.20 | +275% | 27 | Nvidia | 9 | ✓ |
| gpt-5.4-mini | 3.60 | 13.2 | +267% | 56 | OpenAI | 10 | ✓ |
| kimi-k26 | 4.40 | 16 | +261% | 62 | Moonshot | 13 | ✓ |
| step-3.7-flash | 3.40 | 11 | +221% | 40 | StepFun | 60 | ✓ |
| grok-4.3 | 2.00 | 6.30 | +217% | 42 | xAI | 6 | ✓ |
| sonnet-5 | 5.50 | 17.4 | +216% | 72 | Anthropic | 60 | ✓ |
| kimi-k2.7-code | 5.00 | 15.7 | +213% | 61 | Moonshot | 6 | ✓ |
| deepseek-v4-pro | 5.30 | 16.6 | +211% | 59 | DeepSeek | 8 | ✓ |
| qwen3-coder | 1.70 | 5.00 | +192% | — | Alibaba | 15 | ✓ |
| solar-pro4 | 3.80 | 10.6 | +179% | 53 | Upstage | 10 | ✓ |
| opus-4.6 | 7.00 | 17.7 | +152% | — | Anthropic | 6 | ✓ |
| glm-5.2 | 6.70 | 16.6 | +150% | 69 | Z.ai | 39 | ✓ |
| qwen3.6-max-preview | 5.70 | 13.8 | +144% | — | Alibaba | 8 | ✓ |
| devstral-2 | 5.40 | 13 | +142% | 31 | Mistral | 60 | ✓ |
| sonnet-4.6 | 7.30 | 17.6 | +141% | 63 | Anthropic | 87 | ✓ |
| minimax-m3 | 7.30 | 17.7 | +141% | 59 | MiniMax | 6 | ✓ |
| sakana-namazu | 5.00 | 12 | +140% | — | Sakana | 7 | ✓ |
| hy3 | 6.00 | 14 | +133% | 59 | Tencent | 9 | ✓ |
| qwen3.7-max | 7.30 | 16.7 | +127% | 66 | Alibaba | 6 | ✓ |
| mistral-medium-3.5 | 6.80 | 15.4 | +127% | 47 | Mistral | 10 | ✓ |
| gpt-5.6-luna | 7.50 | 16.8 | +123% | 71 | OpenAI | 60 | ✓ |
| deepseek-v4-flash-0731 | 7.80 | 16.7 | +116% | 69 | DeepSeek | 22 | ✓ |
| opus-4.8 | 9.10 | 17.9 | +96% | 74 | Anthropic | 60 | ✓ |
| deepseek-v4-flash | 3.80 | 7.00 | +84% | 56 | DeepSeek | 8 | ✓ |
| gemini-3-flash-preview | 9.60 | 17.6 | +83% | — | 47 | ✓ | |
| fugu-ultra | 10.2 | 18 | +77% | — | Sakana | 10 | ✓ |
| gemini-3.1-pro-preview | 8.20 | 14.5 | +76% | 69 | 8 | ✓ | |
| opus-4.7 | 10.6 | 17.8 | +67% | 74 | Anthropic | 84 | ✓ |
| longcat-2.0 | 6.70 | 11 | +65% | 45 | Meituan | 6 | ✓ |
| inkling | 9.30 | 15.3 | +64% | 52 | ThinkingMachines | 6 | ✓ |
| grok-4.5 | 11.2 | 18 | +60% | 72 | xAI | 8 | ✓ |
| fable-5 | 11.5 | 18 | +57% | 77 | Anthropic | 32 | ✓ |
| opus-5 | 11.9 | 18 | +52% | 78 | Anthropic | 60 | ✓ |
| gpt-5.6-terra | 12 | 17.9 | +49% | 77 | OpenAI | 60 | ✓ |
| muse-spark-1.2 | 12 | 17 | +42% | 72 | Meta | 6 | ✓ |
| kimi-k3 | 12.3 | 17 | +38% | 76 | Moonshot | 6 | ✓ |
| gpt-5.5 | 13.3 | 18 | +36% | 75 | OpenAI | 67 | ✓ |
| muse-spark-1.1 | 13.2 | 17.7 | +33% | 71 | Meta | 7 | ✓ |
| gpt-5.6-sol | 13.9 | 18 | +30% | 78 | OpenAI | 60 | ✓ |
| gpt-5.4 | 13.9 | 17.9 | +29% | 71 | OpenAI | 79 | ✓ |
| qwen3.8-max | 15 | 17 | +13% | 72 | Alibaba | 6 | ✓ |
| gemini-2.5-flash | 5.30 | 5.60 | +6% | — | 15 | ✓ | |
| grok-3-fast | 6.30 | 5.50 | −13% | — | xAI | 21 | * |
| grok-3-mini | 2.00 | 1.70 | −14% | — | xAI | 14 | * |
| grok-4-1-fast-non-reasoning | 1.30 | 0.30 | −75% | — | xAI | 6 | * |
| gemini-3.5-flash | did not edit | 16.7 | N/A | 70 | 6 | no baseline | |
| gemini-3.6-flash | did not edit | 6.00 | N/A | 69 | 13 | no baseline | |
| gpt-oss-20b | did not edit | – | N/A | 21 | OpenAI | 9 | no baseline |
| grok-4.1-fast | did not edit | – | N/A | — | xAI | 8 | no baseline |
| laguna-s-2.1 | did not edit | 10 | N/A | — | Poolside | 6 | no baseline |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.
How we tested.
ReGrade is a behavioral-diff context block your AI coding agent reads alongside its existing prompt. The block shows the agent how the new code's runtime behavior differs from the old, the way a human checks for regressions at code review. No model retraining. No changes to your build or CI pipeline.
How we measured it. At Curtail® we tested 63 AI models on one task: fix 18 real bugs in a working 300,000 line Python service. Every model ran that task twice. On the first run the model worked blind, the way a coding agent normally works today. On the second run we gave the model ReGrade's patented behavioral comparison, which runs the same live traffic against both versions of the service and shows the model exactly what its last changes did to the running system. 1,431 agent-mode trials across 63 AI models from 22 labs, US and non-US, from frontier flagships down to free tiers.
What counts as an AI coding hallucination. An unintended change, addition or deletion to working code the agent was never asked to modify. The code is not what hallucinates; the AI writing it is. We measure it by replaying the agent's patch against the original behavior and counting what changed that nobody requested, deduplicated to root causes so one mistake surfacing at nine call sites counts once.
63
AI models
22
labs
1,431
agent-mode trials
18
known bugs
300,000-line Python
codebase
-83%
AI coding hallucinations
🕵️ It catches what your tests can't.
Functional tests check status codes and required fields. They routinely miss the behavioral drift that actually breaks production: a CORS header that quietly changed, a list that's now an object, a timestamp format that broke a downstream consumer. ReGrade compares runtime behavior directly, so the agent sees the drift the test suite waved through.
35%
of AI patches that passed every functional test still silently broke behavior
191 of 548 test-passing patches, across 14 of 26 tasks
0 → 9
multi-file bug sites fixed, weaker models fixed 0 of 9 on their own
with ReGrade, agents clear all 9 sites of a cross-file regression
18
classes of silent drift planted, headers, encoding, ordering, response shape
the behavioral defects that pass type-checks, lint, and existing tests
Your test suite says “green.” ReGrade says “but the API’s behavior drifted.” That gap is exactly where production incidents come from.
Measured on BaxBench and the multi-site experiment in July 2026. These three figures come from that experiment rather than from the running corpus, so they do not move as more models are tested.
🔎 Better answers, and better detection.
Finding a bug is only half the job, a reviewer still has to understand and trust the fix. Even when a model already spots a regression, ReGrade’s behavioral evidence helps it explain exactly what changed and where.
2×
more precise bug explanations on a production codebase (Ghost CMS)
explanation precision rose from 0.40 to 0.91
9 of 10
models explained the bug better on a 291,000-line codebase (NetBox)
+3.2 points on a 20-point rubric, even where the model already found it
The difference between “something looks off” and a root-cause a human can act on with confidence.
Want ReGrade for your AI coding agent?
Drop the context block into your agent's prompt. No retraining. No CI changes.