The cheapest AI API is not the cheapest to run

DeepSeek's own pricing documentation carries a notice that it plans to raise API prices, "with a significant increase expected," which moves the floor every budget comparison is anchored to.

Reasoning models spend at least 3.11 times and up to 14.78 times more tokens than non-reasoning models on the same benchmarks, so a lower per-token price often produces a higher per-task bill.

Prompt caching moves DeepSeek V4 Flash input from $0.14 to $0.0028 per million tokens, a 50x gap that usually matters more than which model you picked.

Read the full post

Read the full post