Cost reduction
Curtail research · 63 models, 22 labs, updated 2026-08-29
Across 55 model cells in Curtail’s current corpus, the typical cell cut its cost per bug fixed by 51.6% with ReGrade. 47 of the 55 got cheaper per bug fixed. Cost per bug fixed is what a run costs, divided by the bugs the agent actually eliminated, at each provider’s list price.
Adding the behavioral report raises what a run costs, and the agent finds enough more that the cost per bug fixed falls.
Every model tested, on cost per bug fixed
| Verdict | |||||||
|---|---|---|---|---|---|---|---|
| laguna-xs-2.1 | $0.15 | $0.009 | −94% | — | Poolside | 6 | ✓ |
| ling-2.6-1t | $0.09 | $0.006 | −93% | — | InclusionAI | 10 | ✓ |
| solar-pro4 | $0.02 | $0.002 | −91% | 53 | Upstage | 10 | ✓ |
| qwen3.6-max-preview | $0.34 | $0.05 | −87% | — | Alibaba | 8 | ✓ |
| minimax-m2.5 | $0.02 | $0.003 | −83% | — | MiniMax | 60 | ✓ |
| deepseek-v4-flash-0731 | $0.02 | $0.004 | −83% | 69 | DeepSeek | 22 | ✓ |
| kimi-k2.7-code | $0.05 | $0.010 | −81% | 61 | Moonshot | 6 | ✓ |
| minimax-m3 | $0.06 | $0.01 | −80% | 59 | MiniMax | 6 | ✓ |
| sakana-namazu | $0.16 | $0.04 | −77% | — | Sakana | 7 | ✓ |
| seed-2.0-lite | $0.01 | $0.003 | −75% | — | ByteDance | 60 | ✓ |
| deepseek-v4-flash | $0.02 | $0.005 | −71% | 56 | DeepSeek | 8 | ✓ |
| haiku-4.5 | $0.18 | $0.05 | −71% | 44 | Anthropic | 60 | ✓ |
| hy3 | $0.02 | $0.004 | −71% | 59 | Tencent | 9 | ✓ |
| glm-5.2 | $0.03 | $0.008 | −69% | 69 | Z.ai | 39 | ✓ |
| haiku-4.5 (Aug 2026) | $0.15 | $0.05 | −68% | 44 | Anthropic | 20 | ✓ |
| opus-4.6 | $0.29 | $0.10 | −67% | — | Anthropic | 6 | ✓ |
| step-3.7-flash | $0.04 | $0.01 | −64% | 40 | StepFun | 60 | ✓ |
| qwen3.6-plus | $0.06 | $0.02 | −64% | 55 | Alibaba | 7 | ✓ |
| gpt-5.4-mini | $0.02 | $0.009 | −63% | 56 | OpenAI | 10 | ✓ |
| grok-4.20-0309-reasoning | $0.27 | $0.10 | −62% | — | xAI | 21 | ✓ |
| fable-5 | $0.59 | $0.25 | −57% | 77 | Anthropic | 32 | ✓ |
| gpt-5.6-luna | $0.01 | $0.005 | −57% | 71 | OpenAI | 60 | ✓ |
| glm-5.1 | $0.05 | $0.02 | −56% | 56 | Z.ai | 60 | ✓ |
| fugu-ultra | $0.49 | $0.21 | −56% | — | Sakana | 10 | ✓ |
| qwen3-coder-plus | $0.07 | $0.03 | −55% | — | Alibaba | 7 | ✓ |
| muse-spark-1.2 | $0.07 | $0.03 | −54% | 72 | Meta | 6 | ✓ |
| grok-4.5 | $0.03 | $0.01 | −53% | 72 | xAI | 8 | ✓ |
| sonnet-5 | $0.26 | $0.13 | −52% | 72 | Anthropic | 60 | ✓ |
| gpt-5.5 | $0.17 | $0.09 | −50% | 75 | OpenAI | 67 | ✓ |
| muse-spark-1.1 | $0.04 | $0.02 | −49% | 71 | Meta | 7 | ✓ |
| opus-4.8 | $0.40 | $0.21 | −49% | 74 | Anthropic | 60 | ✓ |
| qwen3.7-max | $0.05 | $0.03 | −48% | 66 | Alibaba | 6 | ✓ |
| kat-coder-pro-v2.5 | $0.05 | $0.03 | −46% | — | Kuaishou | 5 | ✓ |
| ling-3.0-flash | $0.004 | $0.002 | −45% | 51 | InclusionAI | 29 | ✓ |
| qwen3.8-max | $0.09 | $0.05 | −45% | 72 | Alibaba | 6 | ✓ |
| gemini-3-flash-preview | $0.02 | $0.01 | −45% | — | 47 | ✓ | |
| gpt-5.6-terra | $0.02 | $0.01 | −43% | 77 | OpenAI | 60 | ✓ |
| kimi-k3 | $0.06 | $0.04 | −37% | 76 | Moonshot | 6 | ✓ |
| grok-4.3 | $0.02 | $0.01 | −35% | 42 | xAI | 6 | ✓ |
| opus-5 | $0.25 | $0.17 | −31% | 78 | Anthropic | 60 | ✓ |
| longcat-2.0 | $0.01 | $0.008 | −30% | 45 | Meituan | 6 | ✓ |
| gpt-5.6-sol | $0.07 | $0.05 | −29% | 78 | OpenAI | 60 | ✓ |
| gpt-5.4 | $0.03 | $0.02 | −28% | 71 | OpenAI | 79 | ✓ |
| inkling | $0.03 | $0.02 | −21% | 52 | ThinkingMachines | 6 | ✓ |
| sonnet-4.6 | $0.11 | $0.09 | −20% | 63 | Anthropic | 87 | ✓ |
| nemotron-3-ultra | $0.07 | $0.06 | −20% | 49 | Nvidia | 10 | ✓ |
| opus-4.7 | $0.20 | $0.17 | −18% | 74 | Anthropic | 84 | ✓ |
| gemini-3.1-pro-preview | $0.12 | $0.12 | +2% | 69 | 8 | * | |
| qwen3-coder | $0.10 | $0.12 | +21% | — | Alibaba | 15 | * |
| gemini-2.5-flash | $0.008 | $0.010 | +30% | — | 15 | * | |
| grok-3-fast | $0.05 | $0.07 | +40% | — | xAI | 21 | * |
| deepseek-v4-pro | $0.01 | $0.02 | +44% | 59 | DeepSeek | 8 | * |
| grok-3-mini | $0.005 | $0.010 | +106% | — | xAI | 14 | * |
| mistral-medium-3.5 | $0.02 | $0.05 | +107% | 47 | Mistral | 10 | * |
| grok-4-1-fast-non-reasoning | $0.02 | $0.06 | +200% | — | xAI | 6 | * |
✓ improved · * got worse · N/A: the model did not edit the file in the control run, so there is no baseline to compare against. Click a column to sort. Measured over 63 models, 22 labs, 1,431 agent-mode trials. Capability is the Artificial Analysis Coding Index, retrieved 2026-08-18, best listed configuration per model. It is not measured by us; it is here so the table can be sorted by how capable a model is. — means the model is not listed there.
🏆 Three recommended models, ranked by cost with ReGrade.
For teams choosing their AI coding agent today: best results-per-dollar with ReGrade.
Ranked from Curtail's current corpus of 63 models, by cost per bug fixed with ReGrade among those fixing at least 16 of 18. This list moves as later testing continues.
🥇
deepseek-v4-flash-0731
16.7 / 18
bugs fixed
$0.004
per bug
🥈
gpt-5.6-luna
16.8 / 18
bugs fixed
$0.005
per bug
🥉
glm-5.2
16.6 / 18
bugs fixed
$0.008
per bug
24 more model and ReGrade combinations clear the same threshold, at least 16 of 18 bugs fixed. Works equally well across providers.
ReGrade is a single addition to the agent's prompt, you pay only the marginal LLM tokens for that context, with no separate ReGrade fee per bug.
Want ReGrade for your AI coding agent?
Drop the context block into your agent's prompt. No retraining. No CI changes.