Spotify Cut Claude Code Token Use 90% With Shunt
By Nihar Ranjan Das · Thu Oct 08 2026 · 6 min read · 1 views
View as a Web StoryAI Tools#claude code#ai-costs#Spotify#token optimization#Portal

Spotify Portal cuts Claude Code costs 90%
Most of what Claude Code spends tokens on is not thinking. It is reading big files, copying test patterns, and typing out boilerplate. Spotify engineer Dimitri Mazmanov noticed this and sent that work to a cheaper model. The result, published by Spotify Engineering in September 2026, was an average 90% drop in Claude token use on bulk reads. The pieces are Portal by Spotify's AiKA Modes and an open Claude Code plugin called shunt. This guide explains how it works, what the 90% really measures, how to set it up, and how to copy the idea without Portal.
Why Claude Code bills get big
Every file Claude reads becomes input tokens, and those tokens ride along in context for the rest of the session. Every line it writes costs output tokens, which are priced higher than input. Spotify's write-up notes that engineering leaders reportedly spend $200 to $500 per developer per month on tokens, and some spend more than $2,000.
Mazmanov's finding was that a large share of those tokens went to work that needs no deep reasoning:
- Reading large files to answer a narrow question.
- Copying the pattern of an existing test to write a new one.
- Generating configs, stubs, and other predictable code.
Reasoning about a bug or an architecture needs a top model. Summarizing a 900-line class does not.
What Portal and AiKA Modes are
Portal is Spotify's developer platform. AiKA Modes are declarative agents that run on temporary runtimes, a bit like AWS Lambda for AI agents. For each mode you define:
- the instructions,
- the model and parameters such as temperature,
- the MCP tools it can use.
Portal handles the infrastructure, API keys, and server management. A mode can be public or private, and you call it from the Portal CLI or API. That means Claude Code can hand a job to a mode with a single shell command and get back a short answer.
The two modes Spotify built
| Mode | What it does | Model in Spotify's tests | Settings |
|---|---|---|---|
bulk-reader |
Reads several files and answers a specific question about them | Gemini 2.5 Flash | Temperature 0.2 |
code-writer |
Writes predictable code (tests, configs, stubs) matching a reference file | Gemini 2.5 Flash | Needs a reference file |
Both are configurable, so you can swap the model without touching the plugin. Spotify's design rule sums it up: "The plugin decides when to delegate. The mode decides how to respond."
Usage looks like this:
bulk-read --question "Where is retry logic configured?" --paths src/a.java src/b.java
code-write --spec "Unit test for the cache eviction path" --reference ExistingTest.java --target NewTest.java
How the shunt plugin enforces delegation
Asking Claude nicely to delegate does not work reliably, so shunt uses PreToolUse hooks to make it happen.
check-file-sizeblocks a full-file read when the file is above a threshold (default 350 lines). It tells Claude to usebulk-readerinstead. Reads with an offset and limit pass through, because Claude already knows which section it needs.check-bash-readcatches shell reads such ascatandgreppiped over large files, so Claude cannot dodge the first hook.
You can change the threshold with the SHUNT_MIN_LINES environment variable or in .claude/settings.json.
Advertisement
For code generation, shunt sends the spec and a reference file to code-writer, strips any markdown fences from the reply, and can write the result straight to disk. Claude never sees the generated code, so you save the expensive output tokens as well as the input tokens.
What the 90% number really means
Read the claim carefully. Spotify reports a mean saving of about 90% on the bulk-read work that was delegated, measured in Claude tokens on a Java monorepo. It is not a promise that your whole monthly bill falls by 90%.
Other coverage makes the same point: shunt is worth testing when Claude repeatedly opens large files to answer narrow questions, but it is not proof that your entire AI coding bill will drop by that much. Your real saving depends on how much of your usage is reading and boilerplate.
| Your usual work | Delegation fit | Realistic outcome |
|---|---|---|
| Exploring a large unfamiliar codebase | High | Big savings on reads |
| Writing tests and config from existing patterns | High | Good savings on output tokens |
| Fixing a subtle bug | Low | Small savings, higher risk |
| Architecture and design decisions | Low | Keep it on the main model |
| Editing code by line number | Low | Summaries give unreliable line numbers |
Spotify did not publish per-scenario percentages, so measure your own sessions before trusting any figure.
Where delegation breaks down
Spotify names the failure modes directly:
- Editing tasks. A summary from a worker model can give wrong line numbers, and Claude then edits the wrong place.
- Reasoning work. In testing, the cheaper model missed subtle bugs that Claude would have caught.
- Small files. Below the threshold, network overhead costs more than the tokens you save.
- Latency. Each delegation takes 10 to 30 seconds, and Portal enforces a 30-second limit per call. A long agent run with many delegations feels slower.
There is also a hidden cost: verification. If the summary is wrong, you spend tokens and time finding out. The useful test is whether the cheap reader returns enough accurate context to finish the task without extra checking.
How to set up shunt
You need Claude Code and access to a Portal instance. Install the plugins from Spotify's marketplace:
claude plugin marketplace add spotify/portal-ai-plugins
claude plugin install portal@portal
claude plugin install shunt@portal
Then run /portal:setup inside Claude Code to authenticate against your Portal instance.
The shunt plugin registers two scripts that wrap the Portal CLI. Claude calls them with named arguments, and the scripts build the request, call the API, unwrap errors, and print token usage so you can see the savings.
The bulk-reader and code-writer modes are public, so you can reuse them or copy them as a starting point for your own.
No Portal? Copy the idea with plain Claude Code
If you cannot use Portal, the pattern still works with tools you already have.
- Use a subagent for reading. Create a Claude Code subagent that runs on a smaller model, such as Haiku, with a prompt like "answer only the question, quote line numbers, return under 200 words". Let it open the big files and return a summary. Only the summary enters your main context.
- Add a PreToolUse hook. Write a hook that blocks
Readon files over a line limit and tells Claude to use the subagent instead. Allow reads that include an offset and limit. - Keep generated boilerplate off the main model. Have a script call a cheaper model API for tests and configs, then write the file directly.
- Measure first. Log tokens for a week. If reading is under 20% of your usage, this will not move your bill much.
This version has no 30-second cap and no extra platform, but you maintain the hooks yourself.
A short checklist before you roll this out to a team
- Pick a repo with large files and lots of exploration, not a small service.
- Set
SHUNT_MIN_LINEShigh enough to skip files where overhead beats savings. - Keep editing and debugging sessions on the main model.
- Compare token and time totals for the same task with and without delegation.
- Review a sample of delegated summaries for accuracy before trusting them.
Sources
Advertisement
FAQ
Do I need to work at Spotify to use shunt?
No. The plugins are published in the `spotify/portal-ai-plugins` marketplace, and the example modes are public. You do need access to a Portal instance to run the modes.
What does "90% savings" mean?
It is Spotify's average reduction in Claude token use on delegated bulk-read work in a Java monorepo. It does not mean your whole bill drops by 90%.
Can I use a model other than Gemini 2.5 Flash?
Yes. The model is set per mode, so you can change it without touching the plugin.
Does shunt work with tools other than Claude Code?
The plugin is built for Claude Code. The AiKA Modes themselves are callable from the Portal CLI and API, so other tools can use them.
What happens when a delegated call fails or times out?
The scripts return the error to Claude, which can retry or read the file directly. Portal stops any single call at 30 seconds.
Should I delegate debugging?
No. Spotify found that the worker model missed subtle bugs and gave unreliable line numbers for edits. Keep debugging and design on the main model.
Comments
Loading…
Sign in to join the conversation.
Related posts

What Is a Proxy on Janitor AI? How It Works
A proxy on Janitor AI is a connection that lets the site use an outside language model, such as DeepSeek, instead of its built-in one. It is not a VPN and it does not hide your traffic. Janitor AI's
Wed Oct 07 2026 · 8 min read · 0 views

Janitor AI Suspended From Gemini? What to Check
Short answer: if Gemini stopped working in Janitor AI, the cause is usually one of three things: a rate-limit or quota error (429), a content filter block, or an API key that stopped working. Only the
Wed Oct 07 2026 · 7 min read · 0 views

Is Janitor AI Down? How to Check and Fix Errors
Short answer: to check whether Janitor AI is down, open its official status page at status.janitorai.com, which Janitor AI's own help centre points to. On October 7, 2026 at 16:58 UTC it showed All
Wed Oct 07 2026 · 5 min read · 0 views