Two leaderboards updated on October 9, 2026 name different models first on SWE-bench Verified, and every score on them is self-reported.
One agentic coding task costs about $0.02 on GPT-5.6 Luna and $0.93 on Claude Fable 5.1 under a fixed 150,000-token assumption, a spread of about 40 times.
Developer time usually outweighs token cost, so pick by cost per passing task and measure it on your own repository with 10 tasks and a short script.