Agentic AI architecture: five patterns and what each costs
By Nihar Ranjan Das · Fri Oct 09 2026 · 10 min read · 0 views
View as a Web StoryAI#ai agents#agentic ai#multi-agent#claude#Architecture#Laravel AI SDK

Most agent projects should use a fixed workflow, not a free-roaming agent. Anthropic's own guidance says to find "the simplest solution possible, and only increase complexity when needed." The reason is cost. Anthropic reports that agents use about 4 times the tokens of a chat, and multi-agent systems use about 15 times.
This post covers the five patterns that matter, a decision tree for choosing between them, and a cost model built on current Claude prices. It also shows how a Laravel app wires one agent to another, using the documented AI SDK.
What is agentic AI architecture?
Agentic AI architecture is the way you connect language model calls, tools, and control logic so a system can finish a multi-step task. It answers three questions. Who decides the next step, code or the model? How do model calls pass work to each other? What stops the system from looping forever?
Anthropic draws one useful line in its Building Effective Agents guide, published December 19, 2024. Workflows are systems where code paths decide how models and tools interact. Agents are systems where the model directs its own process and tool use.
That line decides your cost, your debugging effort, and your risk. A workflow is predictable. An agent is flexible and can compound its own errors, as Anthropic warns.
What are the five patterns?
Anthropic names five workflow patterns, and each fits a different task shape. The same guide pairs each with a "when to use" rule.

| Pattern | Use it when | Model calls per task |
|---|---|---|
| Prompt chaining | The task splits into fixed steps | n, one after another |
| Routing | Inputs fall into distinct categories | 1 classifier plus 1 handler |
| Parallelization | Subtasks run at once, or you want several attempts | k workers plus 1 merge |
| Orchestrator-workers | Subtasks cannot be predicted in advance | 1 plan, k workers, 1 synthesis |
| Evaluator-optimizer | Clear criteria exist and revision helps | 2 per round, times the rounds |
Prompt chaining is a pattern where each model call handles one step and passes its output to the next. It trades latency for accuracy. Think of an outline, then a draft, then a polish pass.
Routing is a pattern where a first call classifies the input and sends it to a specialized handler. A support inbox that splits billing, bugs, and sales questions is the classic case.
Parallelization runs several calls at once. You can split a task into independent parts, or run the same task several times and vote. Code review across many files fits well.
Advertisement
Orchestrator-workers is a pattern where a central model breaks a task into subtasks at runtime and hands them to worker models. Use it when you cannot list the subtasks ahead of time, such as a research question or a change that touches an unknown number of files.
Evaluator-optimizer is a loop where one model drafts and another critiques until the output passes a bar. It only pays off when you can state clear evaluation criteria.
How do you choose a pattern?
Start from the task, not from the framework. Ask these questions in order and stop at the first yes.

- Can one well-written prompt with the right context do the job? Use one call.
- Are the steps always the same? Use a chain.
- Does the input type decide the handler? Use a router.
- Are the parts independent? Use parallel calls.
- Are the subtasks unknown until runtime? Use orchestrator-workers.
- Does output improve with critique against a rubric? Add an evaluator loop.
Most teams stop at step 2 or 3. Anthropic's guidance recommends starting with simple prompts and optimizing them with evaluation before adding agentic structure.
What does each architecture cost?
Cost grows with the number of calls and the context each call carries. Anthropic gave two ratios in its write-up of its multi-agent research system, published June 13, 2025. Agents typically use about 4 times the tokens of chat interactions. Multi-agent systems use about 15 times the tokens, per the Anthropic engineering post.
Apply those ratios to current prices. Anthropic's pricing page lists Claude Sonnet 5.5 at $2 input and $10 output per million tokens, and Claude Opus 5.5 at $4 input and $20 output, according to Claude pricing. Assume one chat answer uses 3,000 input tokens and 700 output tokens.
| Model | One chat | Agent (4x) | Multi-agent (15x) |
|---|---|---|---|
| Sonnet 5.5 | $0.013 | $0.052 | $0.195 |
| Opus 5.5 | $0.026 | $0.104 | $0.390 |
At 10,000 tasks a month, the Sonnet figures become $130 for chat, $520 for an agent, and $1,950 for multi-agent. The Opus figures become $260, $1,040, and $3,900.

Treat this as a model, not a quote. The 4x and 15x ratios come from one company's research system, and your workload will differ. Cached input tokens also cost far less. Claude Sonnet 5.5 cache reads are listed at $0.10 per million tokens, so a long shared prompt changes the math.
The point is the ratio. A multi-agent design must deliver 15 times the value of a chat to break even. Anthropic's own multi-agent setup beat a single agent by 90.2% on its internal research evaluation (Anthropic, June 2025), so the gain can be real. It is also a self-reported internal result.
What goes wrong in production?
Agents fail in predictable ways. Researchers who annotated more than 1,600 traces from 7 multi-agent frameworks grouped 14 failure modes into three categories: system design issues, inter-agent misalignment, and task verification, per the MAST study on arXiv. Anthropic documented several of the same patterns in its multi-agent write-up. Plan for each one before launch.
- Over-spawning. Early agents created excessive subagents for simple queries, including as many as 50.
- Endless searching. Agents searched for sources that did not exist.
- Over-continuing. Agents kept working after they already had enough results.
- Duplicated work. Vague task instructions made subagents repeat the same searches.
- Poor source choice. Agents favored SEO-optimized content farms over authoritative sources.
The fixes are mostly prompt and design work. Embed explicit effort-scaling rules in the orchestrator's instructions. Give each worker a clear objective, an output format, and boundaries. Anthropic also reported that rewriting tool descriptions cut task completion time by 40% for future agents (June 2025).
Add hard limits on top. Long runs also suffer from context rot, where answers degrade as the window fills, so keep each worker's context small. A shared agent token budget per run makes the limit concrete. Cap the number of workers per task, the number of tool calls, and the dollars per run. A budget guard in code is more reliable than a request in a prompt.
When does multi-agent earn its 15 times cost?
Multi-agent pays off when the work is wide, parallel, and valuable. Anthropic's research system is the model case. It read many sources at once, and the post reports that parallelization cut research time by up to 90% for complex queries (June 2025).
Anthropic also analyzed its BrowseComp results. Three factors explained 95% of the performance variance, and token usage alone explained 80% (Anthropic, 2025). That finding cuts two ways. More tokens often buy better answers. A multi-agent design is partly a way to spend more tokens in parallel.
Use that to judge your own case. If you are still picking tools, our ai agent framework comparison ranks them by churn. Multi-agent fits tasks with these traits.
- The work splits into independent reads, such as many documents or many files.
- A wrong answer costs more than the extra tokens.
- Latency matters, so parallel workers beat one long chain.
It fits poorly when steps depend on each other. Most coding changes are sequential, because each edit depends on the last. A single agent with good tools usually suits those jobs better.
What does a routed support inbox cost?
A router in front of specialist handlers costs almost nothing compared with the handlers themselves. Consider a support inbox that sorts messages into billing, bugs, and sales. The classifier needs roughly 500 input tokens and 20 output tokens per message.
On Claude Haiku 5.5, priced at $0.10 per million input tokens and $0.50 per million output tokens for contexts up to 100,000 tokens, that call costs about $0.00006, per Haiku 5.5 price list. A thousand messages cost about six cents to classify. The expensive part is whichever handler the router picks, so keep the router small and let only hard cases escalate to a larger model.
What should you log in production?
Log every model call with its parent, its input size, its output size, and its cost. Without that trace you cannot tell a healthy run from a runaway one. Four fields do most of the work.
- Run ID and parent ID. They rebuild the tree of calls for any task.
- Tokens in and out per call. They show which worker eats the budget.
- Tool name and arguments. They expose loops, such as the same search repeated ten times.
- Stop reason. It separates a finished task from a hit limit.
Alert on two numbers. Set one alert for calls per run, and one for dollars per run. A run that makes 10 times the usual calls is almost always a loop or an over-spawn, and you can kill it before it finishes.
How do you build one in Laravel?
Laravel's AI SDK models a sub-agent as a tool. The Laravel AI SDK documentation describes an agent as a PHP class that implements the Agent contract and uses the Promptable trait. You expose tools by returning them from tools().
Delegation needs no separate orchestration API. You return one agent from another agent's tools() method. This example comes from the documented pattern.
class CustomerSupportAgent implements Agent, HasTools
{
use Promptable;
public function instructions(): string
{
return 'You help customers with account, order, and billing questions. Delegate refund policy questions to the refunds specialist.';
}
public function tools(): iterable
{
return [
new RefundsAgent,
];
}
}
The RefundsAgent can implement CanActAsTool to set a clear name and description. Name it refunds_specialist and describe when to use it. The model picks tools by reading those descriptions, so write them with care.
The docs state that each sub-agent invocation runs in isolation and does not receive the parent's conversation history. That is a feature for cost, because the worker carries only the task you hand it. It is also a trap, because the worker knows nothing you did not pass along. Put the order number, the policy text, and the customer's question into the task itself.
This design is the orchestrator-workers pattern with one worker. Multi agent security deserves a look before launch, because one agent can pass a poisoned instruction to the next. Cached prompts also change the bill, as the bedrock prompt caching cost audit showed. It stays cheap while the parent routes most turns itself. Watch the cost when a parent starts calling several sub-agents per turn.
How should you evaluate an agent before shipping?
Build an evaluation set before you build the architecture. Twenty real tasks with known good outcomes is enough to start. Run each pattern against the set and record success rate, tokens, and wall-clock time.
Use this scoring rule. Move up one pattern only if it raises the success rate by more than the cost increase justifies. A chain that solves 90 percent of tasks at $0.05 beats an agent that solves 93 percent at $0.20, unless the last 3 percent are expensive failures.
Track these four numbers per run:
- Success against your rubric.
- Total input and output tokens.
- Number of model calls and tool calls.
- Time to the final answer.
Watch the call count. A jump from 3 calls to 30 on the same task usually signals over-spawning or looping, and it shows up before the bill does.
What are the limits of this advice?
The cost ratios are Anthropic's, measured on a research workload that reads many web pages. A coding agent, a data pipeline, and a support bot will each have a different multiplier. The prices were current on the date of writing and will change. The 3,000 and 700 token chat size is an assumption you should replace with your own average.
Even so, the order of the decisions holds. Try one prompt first, then a workflow, and only then an agent. Each step up multiplies cost and shrinks predictability, so make each one earn its place with data from your evaluation set.
Advertisement
FAQ
What is the difference between a workflow and an agent?
A workflow runs a path that your code defines, with models filling in steps. An agent lets the model decide its own steps and tools. Workflows are cheaper and more predictable. Agents handle open-ended tasks that fixed paths cannot.
How many tokens do multi-agent systems use?
Anthropic reported that multi-agent systems use about 15 times the tokens of a chat, and single agents about 4 times. These figures describe its research system. Measure your own workload before you budget.
When should I use orchestrator-workers?
Use it when you cannot predict the subtasks in advance, such as research questions or changes that touch an unknown number of files. If you can list the steps ahead of time, a prompt chain is cheaper and easier to debug.
How do I stop an agent from looping or over-spawning?
Set hard caps on workers, tool calls, and spend per run, and add effort-scaling rules to the orchestrator prompt. Anthropic saw early agents spawn up to 50 subagents for simple queries. Code-level limits are more reliable than prompt instructions.
Can a Laravel app run multi-agent systems?
Yes. The Laravel AI SDK lets one agent call another by returning it from the `tools()` method. Each sub-agent runs in isolation without the parent's conversation history, so pass the full task context in each call.
Comments
Loading…
Sign in to join the conversation.
Related posts

Which AI code review tool should you trust with your PRs?
Every AI code review vendor ranks first on its own benchmark. Compare the published numbers, then run a one-hour seeded-bug test on your own repo.
Fri Oct 09 2026 · 11 min read · 0 views

Which AI model should write your code? Price per task
Coding benchmarks now agree within a few points while prices differ 40 times. See cost per task for ten models and a script to find your own cost per passing task.
Fri Oct 09 2026 · 11 min read · 0 views

Claude Sonnet 5.5 migration: which changes return a 400
Claude Sonnet 5.5 launched on September 28, 2026 at the same price as Claude Sonnet 5, $2 per million input tokens and $10 per million output tokens. Moving to it is not a model-ID swap. Several
Tue Sep 29 2026 · 6 min read · 2 views