AI

Which AI model should write your code? Price per task

By · Fri Oct 09 2026 · 11 min read · 0 views

View as a Web Story

AI#gemini#claude#AI coding models#model comparison#GPT#SWE-bench

Balance scale with a price tag on one side and a checkmark on the other

Pick your coding model by the cost per passing task rather than by leaderboard rank, because the leaderboards no longer separate the leading models in any way that predicts your results. On October 9, 2026, two public leaderboards named different models first on SWE-bench Verified, and the top five scores sit within about 10 points of each other, while the price of a typical coding task differs by about 40 times.

This post gives you the official prices for ten models, a per-task cost table, and a script that turns your own test runs into a cost per passing task, and it also explains why the leaderboards disagree with each other and what practical measurement you should use in their place.

Why can't you trust one leaderboard?

Because the same benchmark gives different winners depending on who publishes it. SWE-bench Verified is a benchmark of real GitHub issues that a model must fix, and it is the most quoted coding score in 2026. The original SWE-bench paper, first submitted on October 10, 2023, collected 2,294 problems from 12 Python repositories, and the best model then, Claude 2, solved only 1.96% of them, per the SWE-bench paper on arXiv. Two sites that list it, both last updated October 9, 2026, rank the top models differently.

Rank llm-stats BenchLM
1 Claude Fable 5 (95.0%) Claude Opus 5 (96%)
2 Claude Mythos Preview (93.9%) Claude Mythos 5 (95.5%)
3 Ember-1 (92.2%) Claude Fable 5 (95%)
4 Claude Opus 4.8 (88.6%) Ember-1 (92.2%)
5 Claude Opus 4.7 (87.6%) Claude Opus 4.8 (88.6%)

Two ranked lists of the top five SWE-bench Verified models from llm-stats and BenchLM on October 9, 2026, with lines showing different orderings and different models in first place

The llm-stats leaderboard states that all 117 results are self-reported, with none verified. The BenchLM page goes further. It says the benchmark's tests are contaminated and saturated, which is the reason it excludes SWE-bench Verified from its weighted scoring and shows the results for reference only.

Read that as a warning about interpretation. When the top scores cluster between 85% and 96%, the benchmark no longer separates the models well, and because the results are also unverified, a 1-point gap in that range tells you almost nothing about how a model will behave on your own codebase.

How much does each model cost per coding task?

A typical agentic coding task costs between $0.02 and $0.93 on current models, which is a spread of about 40 times, and the table below uses the official prices from each vendor together with one fixed task size so that the comparison stays consistent.

The task size is an assumption that you should replace with your own measurements. Take a coding agent that reads files, edits code, and runs tests over many turns, and assume that it consumes 150,000 input tokens in total, with 120,000 of them served from the prompt cache, plus 12,000 output tokens.

Model Input / cached / output per 1M Cost per task
Claude Fable 5.1 $10 / $0.25 / $50 $0.930
GPT-5.6 Sol $4 / $0.40 / $20 $0.408
Claude Opus 5.5 $4 / $0.20 / $20 $0.384
GPT-5.3 Codex $1.75 / $0.175 / $14 $0.242
GPT-5.6 Terra $2 / $0.20 / $12 $0.228
Gemini 3.1 Pro $2 / $0.20 / $12 $0.228
Claude Sonnet 5.5 $2 / $0.10 / $10 $0.192
Gemini 3.8 Flash $0.75 / $0.075 / $3.75 $0.077
Claude Haiku 5.5 $0.50 / $0.05 / $2.50 $0.051
GPT-5.6 Luna $0.20 / $0.02 / $1.20 $0.023

Horizontal bar chart of estimated cost per agentic coding task for ten models, from GPT-5.6 Luna at 2 cents to Claude Fable 5.1 at 93 cents

The prices come from three official pages. Claude prices are from Anthropic's pricing page; GPT prices are from OpenAI's API pricing page. Gemini prices are from Google's Gemini API pricing page.

Advertisement

Four details change the numbers.

  1. Claude Haiku 5.5 uses its higher tier here, because the context is above 100,000 tokens.
  2. OpenAI says GPT-5.6 Sol's promotional pricing runs "at least through November 21, 2026."
  3. OpenAI charges more for prompts above 272,000 tokens. Sol rises to $8 input and $30 output.
  4. Cache writes cost extra on Claude, and this table ignores them.

The biggest caveat is token use, because different models consume different amounts of tokens for exactly the same job, so you should treat this table as a rough floor for each model and replace it with your logged usage as soon as you have it.

Does the cheaper model save money?

Not when developer time is the real cost; a token bill of $0.19 is small beside a developer's hour. If a failed attempt costs 10 minutes of review at $60 an hour, that failure costs about $10. That is 50 times the token cost of the Sonnet run.

So for interactive coding the pass rate decides, since a model that finishes one more task in ten saves more than any price gap shown above. For unattended work the arithmetic flips, because an agent that opens 1,000 pull requests a day at $0.93 each spends $930 a day, while the same load on GPT-5.6 Luna costs about $23.

Use this rule of thumb. For a close look at one frontier price list, see our notes on claude opus 5.5 pricing.

Workload Optimize for Typical pick
Interactive coding in an editor Pass rate A frontier model
Code review or triage in CI Cost, then pass rate A mid-tier or flash model
Bulk refactors, 1,000+ tasks a day Cost per passing task A cheap model plus a retry
Hard, rare problems Pass rate The strongest model, used sparingly

Many teams route work by difficulty, so that a cheap model drafts the change and a stronger model reviews it or retries whenever the tests fail, which is the routing pattern from agent design and which reduces cost significantly without capping the quality of the final result. Open-weight options such as DeepSeek change the math again, and our deepseek vs claude comparison covers that case.

How much does prompt caching change the bill?

Caching cuts the cost of a long agent session by more than half, and the task above assumes that 80% of the input tokens come from the cache. If you remove the cache, the same task costs far more, because every turn has to read the whole context again at the full input price.

Model With cache (80% hits) No cache Saving
Claude Sonnet 5.5 $0.192 $0.420 54%
Claude Opus 5.5 $0.384 $0.840 54%

For Sonnet 5.5, the no-cache price is 150,000 tokens at $2 per million, which is $0.30, plus $0.12 for output. For Opus 5.5 it is $0.60 plus $0.24. Cache reads on Sonnet 5.5 cost $0.10 per million, one twentieth of the input price, per Claude pricing.

Two habits protect the discount, and a careless cache configuration can erase it entirely, as the bedrock prompt caching cost audit found when cache write pricing consumed most of one organization's bill. Teams that cut claude code token usage by trimming context report similar savings. Keep the beginning of your prompt stable by placing system instructions and repository context first, and put the changing part, such as the latest user request, last, because any edit near the top of the prompt can invalidate the cache for everything that follows it.

Why does price per token mislead?

Models use different numbers of tokens for the same job, so a low rate can still produce a higher bill when a model rambles, retries failed steps, or reads extra files. Anthropic's own analysis of its research agent found that token usage explained 80% of performance variance on BrowseComp, which shows how tightly tokens and results are linked.

Here is a worked example with assumed numbers. Suppose Model A costs $2 per million output tokens and uses 20,000 output tokens per task, which produces a bill of $0.040, while Model B costs $3 per million but finishes in 10,000 tokens, which produces a bill of only $0.030, so the model with the higher rate is actually cheaper per task.

Your usage log settles the question, because the script above already reads real token counts and therefore reflects the habits of each model. Look at output tokens per task as a column of its own, and investigate any model that uses twice the median.

How do you find your own cost per passing task?

Run 10 real tasks from your backlog through each candidate model and log the tokens. Then divide total cost by the number of tasks that passed your tests. Cost per passing task is the total spend on a model divided by the number of its runs that pass your acceptance tests. It captures both price and quality in one number.

Follow these steps.

  1. Pick 10 closed issues or tickets with tests. Use a mix of easy and hard.
  2. Reset the repository to the commit before each fix.
  3. Run each model on each task with the same prompt and tools.
  4. Log input, cached, and output tokens from the API usage field.
  5. Mark a run as passed only when your tests pass without edits.
  6. Feed the log to the script below.

Flow diagram from ten backlog tasks through reset, run on each model, log tokens, run tests, to a cost per passing task table

Save the prices as prices.json, with one entry per model in dollars per million tokens. Save your runs as runs.json. Then run the script with Node.js 20 or newer.

// cost-per-pass.mjs — usage: node cost-per-pass.mjs prices.json runs.json
import { readFileSync } from "node:fs";
const prices = JSON.parse(readFileSync(process.argv[2], "utf8"));
const runs = JSON.parse(readFileSync(process.argv[3], "utf8"));
const by = {};
for (const r of runs) {
  const p = prices[r.model];
  const cost =
    (r.inputTokens * p.in + r.cachedTokens * p.cached + r.outputTokens * p.out) / 1e6;
  const m = (by[r.model] ??= { runs: 0, passed: 0, cost: 0 });
  m.runs++; m.cost += cost; if (r.passed) m.passed++;
}
const rows = Object.entries(by).map(([model, m]) => ({
  model, runs: m.runs, passRate: +(m.passed / m.runs).toFixed(2),
  costPerRun: +(m.cost / m.runs).toFixed(3),
  costPerPass: m.passed ? +(m.cost / m.passed).toFixed(3) : null,
}));
console.table(rows.sort((a, b) => (a.costPerPass ?? 1e9) - (b.costPerPass ?? 1e9)));

The script was checked on an invented log of 10 runs per model. With the fixture's 70% pass rate for Claude Sonnet 5.5 and 90% for Claude Opus 5.5, the script printed $0.274 and $0.427 per passing task. Those pass rates are made up to check the math; they are not a measurement of either model.

The fixture still shows a useful pattern: the pricier model cost about twice as much per run, but only about 56% more per passing task, and the gap would shrink further if the stronger model passed an even larger share of its runs.

Which model fits which job?

Match the model tier to the cost of a mistake; these are starting points, and your own test log should override them.

  • Autocomplete and small edits. Use the cheapest fast model, such as GPT-5.6 Luna or Claude Haiku 5.5. Mistakes are cheap to spot.
  • Everyday feature work. Use a mid-tier model, such as Claude Sonnet 5.5, GPT-5.6 Terra, or Gemini 3.1 Pro. These three cost between $0.19 and $0.23 per task here.
  • Large refactors and ambiguous bugs. Use a frontier model, such as Claude Opus 5.5 or GPT-5.6 Sol. They cost about $0.38 to $0.41 per task here.
  • Rare, high-stakes problems. Use the top tier, such as Claude Fable 5.1, and only when a cheaper model has failed.

GPT-5.3 Codex sits between tiers. Its output price of $14 per million is high for its input price of $1.75, so it favors tasks with long inputs and short answers.

How should Laravel teams choose?

Laravel's AI SDK lets you switch providers with an enumeration value, so changing the model becomes a configuration change rather than a rewrite of your application. The Laravel AI SDK documentation lists text support for OpenAI, Anthropic, Gemini, Azure, Bedrock, Groq, xAI, DeepSeek, Mistral, Ollama, and OpenRouter. That makes an A/B test cheap.

Run your 10-task test once per quarter and keep the provider in an environment variable, so that whenever the published prices move you can repeat the comparison and reach a decision within an hour.

What should you re-check when a new model ships?

Re-run your 10-task test and read the pricing page again, because a new model changes both quality and cost. Check four things before you switch.

  1. Tokens per task. A new model may think longer or call more tools, which raises output tokens even at the same rate.
  2. Context tiers. Claude Haiku 5.5 changes price above 100,000 tokens, Gemini 3.1 Pro above 200,000, and GPT-5.6 above 272,000, per the three pricing pages. Long agent sessions cross those lines quickly.
  3. Cache terms. Cache reads and writes are priced separately, and a model with cheap reads but costly writes can lose its advantage on short sessions.
  4. Promotions. OpenAI lists GPT-5.6 Sol's promotional pricing as available at least through November 21, 2026, so budget with the post-promotion price in mind.

What are the limits of this comparison?

The cost table assumes one task size and equal token use across models, although real agents differ and some models take more turns. The benchmark scores are self-reported and unverified by the sites that list them, and prices change, so the OpenAI promotion may end after November 21, 2026.

The method holds up better than the numbers, so measure on your own code, count the cost per passing task, and repeat the test whenever a new model ships, because a one-hour experiment is more informative than a month of reading leaderboards.

Advertisement

FAQ

What is the best AI model for coding in 2026?

No single model wins on every measure, because leaderboards dated October 9, 2026 name Claude Fable 5 and Claude Opus 5 first and every score on them is self-reported. For your team, the best model is the one with the lowest cost per passing task on your own repository.

How much does one AI coding task cost?

About $0.02 to $0.93 under the assumptions here: 150,000 input tokens with 120,000 cached, and 12,000 output tokens. GPT-5.6 Luna is the cheapest at $0.023, and Claude Fable 5.1 is the most expensive at $0.93.

Is SWE-bench Verified still a useful benchmark?

It is a rough guide at best, since BenchLM calls its tests contaminated and saturated and excludes the benchmark from its scoring, and because the top scores now cluster between 85% and 96%, small differences between models are not meaningful.

Should I use the cheapest model to save money?

Only for unattended, high-volume work. For interactive coding, a failed attempt costs considerably more in developer time than the token savings could ever recover, so you should compare the cost per passing task rather than the cost per run.

How can I test models on my own code?

Pick 10 closed tickets with tests, run each model on them, and log the token usage together with the result of each run, then divide the total cost by the number of passes, which the script in this post calculates with a single command.

Comments

Loading…

Sign in to join the conversation.

Related posts

Claude Sonnet 5 to Sonnet 5.5 migration showing which API changes return errors

Claude Sonnet 5.5 migration: which changes return a 400

Claude Sonnet 5.5 launched on September 28, 2026 at the same price as Claude Sonnet 5, $2 per million input tokens and $10 per million output tokens. Moving to it is not a model-ID swap. Several

Tue Sep 29 2026 · 6 min read · 2 views

AI