AI Agent Token Budget: Prevent Context Overflow
Three chained agents. The first extracts data from a contract, the second analyses clauses, the third drafts an executive summary. In local development, against four-page test documents, the pipeline runs flawlessly.
In production, a fifty-page contract arrives on Tuesday at 9am. The third agent throws an error. Or worse: it throws nothing and returns a truncated response that nobody catches until the client calls. The context exploded and nobody saw it coming.
The problem that doesn't show up in logs
Context overflow in multi-agent pipelines has a particularly treacherous trait: it doesn't always throw an exception.
When you exceed a model's context limit, the API can respond with a context_length_exceeded error. Fine so far. But in architectures where each agent receives the previous agent's output as part of its context, the problem accumulates silently:
- Agent 1 processes 40,000 input tokens. OK.
- Agent 2 receives Agent 1's output (3,000 tokens) plus the original document: 43,000 tokens. OK.
- Agent 3 receives both previous outputs plus accumulated context: 195,000 tokens. Limit exceeded.
If your code doesn't check context size before calling the API, you find out about the problem in Agent 3's response — after you've already spent the credits for the first two calls and the work is half done.
The solution isn't handling the error when it happens. It's never reaching the error.
Count tokens before sending
The Anthropic API exposes a token counting endpoint you can call without consuming generation tokens. It counts exactly how many tokens your message will occupy before you send it.
import Anthropic from "@anthropic-ai/sdk";
const anthropic = new Anthropic();
async function countTokens(
model: string,
messages: Anthropic.MessageParam[],
system?: string
): Promise<number> {
const response = await anthropic.messages.countTokens({
model,
messages,
...(system ? { system } : {}),
});
return response.input_tokens;
}
This call is cheap — it generates nothing, just counts. Use it as a guard before every agent call:
const MODEL = "claude-sonnet-4-6";
const CONTEXT_LIMIT = 200_000;
const SAFETY_MARGIN = 0.85; // Act at 85% of the limit
async function safeAgentCall(
messages: Anthropic.MessageParam[],
systemPrompt: string
): Promise<Anthropic.Message> {
const tokenCount = await countTokens(MODEL, messages, systemPrompt);
if (tokenCount > CONTEXT_LIMIT * SAFETY_MARGIN) {
throw new Error(
`Token budget exceeded: ${tokenCount} / ${CONTEXT_LIMIT} tokens`
);
}
return anthropic.messages.create({
model: MODEL,
max_tokens: 8192,
system: systemPrompt,
messages,
});
}
Per-agent budgets in a pipeline
The pattern above throws when an agent exceeds its budget. But in a real pipeline, you want something smarter: trim the context in a controlled way before the agent rejects it.
Define an explicit budget for each agent in your pipeline:
interface AgentBudget {
name: string;
maxInputTokens: number;
onBudgetExceeded: "trim" | "summarize" | "error";
}
const PIPELINE_BUDGETS: AgentBudget[] = [
{ name: "extract", maxInputTokens: 80_000, onBudgetExceeded: "trim" },
{ name: "analyse", maxInputTokens: 100_000, onBudgetExceeded: "summarize" },
{ name: "summarise", maxInputTokens: 60_000, onBudgetExceeded: "trim" },
];
With this structure, before running each agent you count tokens and, if over budget, apply the configured strategy:
async function runWithBudget(
budget: AgentBudget,
messages: Anthropic.MessageParam[],
system: string
): Promise<Anthropic.Message> {
let workingMessages = [...messages];
const tokenCount = await countTokens(MODEL, workingMessages, system);
if (tokenCount > budget.maxInputTokens) {
if (budget.onBudgetExceeded === "trim") {
workingMessages = trimToFit(workingMessages, budget.maxInputTokens, system);
} else if (budget.onBudgetExceeded === "error") {
throw new Error(`${budget.name} exceeded budget: ${tokenCount} tokens`);
}
// "summarize" → delegate to another agent; see context summarization article
}
return anthropic.messages.create({
model: MODEL,
max_tokens: 8192,
system,
messages: workingMessages,
});
}
function trimToFit(
messages: Anthropic.MessageParam[],
maxTokens: number,
system: string,
target = 0.8
): Anthropic.MessageParam[] {
// Drop oldest messages until under 80% of the limit
let trimmed = [...messages];
while (trimmed.length > 1) {
trimmed = [trimmed[0], ...trimmed.slice(2)];
if (trimmed.length <= Math.ceil(messages.length * target)) break;
}
return trimmed;
}
Tokens as a system health signal
Once you have per-agent token counts, you get a free monitoring signal. An agent's context size reflects the complexity of the document it's processing.
If the analysis agent starts consuming 150,000 tokens regularly for documents that previously cost 40,000, something changed: the documents are longer, the system prompt grew, or a bug is passing redundant data into the context.
Emit a metric with every call:
// In your telemetry system (Datadog, Prometheus, or plain logs)
console.log(JSON.stringify({
event: "agent_call",
agent: budget.name,
input_tokens: tokenCount,
budget_used_pct: Math.round((tokenCount / budget.maxInputTokens) * 100),
timestamp: new Date().toISOString(),
}));
An alert when budget_used_pct exceeds 80% at the P95 gives you advance warning before the problem reaches production. Same principle as monitoring disk usage before the drive fills up.
The cost of skipping this
Without token budgets, your pipeline has two failure modes:
- Explicit error: the API rejects the call. You find out, but you've already spent credits on the previous agents.
- Silent failure: the model truncates output or returns a coherent but incomplete response. You don't find out until someone reviews the result manually.
In AI integration projects, the second case is the most common and the most expensive. A system that fails silently can spend weeks producing incorrect results before anyone detects it.
If you're building agent pipelines for your business, I also recommend reviewing how to implement AI-Driven Development with resilience patterns from the initial design — it's far cheaper than bolting them on later.
Conclusion
Context overflow in AI agents isn't an error handling problem: it's an architecture problem. Counting tokens before sending, assigning explicit per-agent budgets, and emitting usage metrics is what separates a demo pipeline from one that runs reliably in production for months.
The token counting API costs nothing extra. Implementation time is two hours. The cost of skipping it can be weeks of debugging and unhappy clients.
Want to review how your production agents are managing context? Message me on WhatsApp and let's take a look together.