Designing a tool-calling agent loop with deadlines and budgets
— AI Agents, TypeScript, Reliability — 9 min read
Strip away the framework and a tool-calling agent is a while loop. You send messages to a model, it asks for a tool, you run the tool, you send back the result, and you repeat until it stops asking.
The loop is easy to write. What takes more work is making it stop. A loop with no limits will happily spend money for ten minutes on a task that should have taken ten seconds, and nobody finds out until the invoice arrives or a user gives up.
This post covers the limits I put on every agent loop now, and the TypeScript I use for them. The code uses the Anthropic TypeScript SDK, but the same ideas work with any provider that supports tool calling.
The four limits
Every run gets a budget object, and the loop checks it before each model call:
export type Budget = {
maxTurns: number // hard cap on model calls
deadlineMs: number // wall-clock time for the whole run
maxCostUsd: number // spend ceiling, from token usage
}Each limit catches a different problem:
| Limit | What it catches |
|---|---|
| Turn cap | The model keeps calling tools and never converges |
| Deadline | A slow tool or a slow model response holds the user hostage |
| Cost ceiling | Long contexts make each turn more expensive than the last |
| Finish tool | "Done" is a decision the model makes on purpose, not a guess |
You need all of them. A turn cap of 20 does nothing for you if one of those turns is a tool call that hangs for five minutes. A deadline does nothing if the model is fast but each turn resends a 150k-token context.
The loop
The whole loop is about 80 lines, and it type-checks against the current SDK.
import Anthropic from '@anthropic-ai/sdk'
const client = new Anthropic()
const MODEL = process.env.AGENT_MODEL ?? ''
// Example rates in USD per million tokens. Copy the real ones from your
// provider's pricing page; they change.
const USD_PER_MTOK = { input: 3, output: 15 }
const TOOL_TIMEOUT_MS = 20_000
const MAX_TOOL_CHARS = 20_000
const MAX_REPEATS = 2
type StopReason = 'max_turns' | 'deadline' | 'budget' | 'truncated' | 'looping'
export type Outcome =
| { status: 'done'; answer: string; costUsd: number; turns: number }
| { status: 'stopped'; reason: StopReason; costUsd: number; turns: number }
export type ToolFn = (input: unknown, signal: AbortSignal) => Promise<string>
const finishTool: Anthropic.Tool = {
name: 'finish',
description: 'Call this exactly once, when you have the final answer.',
input_schema: {
type: 'object',
properties: { answer: { type: 'string' } },
required: ['answer'],
},
}
function cost(u: Anthropic.Usage): number {
return (u.input_tokens * USD_PER_MTOK.input + u.output_tokens * USD_PER_MTOK.output) / 1e6
}
export async function runAgent(
task: string,
tools: Anthropic.Tool[],
impl: Record<string, ToolFn>,
budget: Budget,
): Promise<Outcome> {
const signal = AbortSignal.timeout(budget.deadlineMs)
const messages: Anthropic.MessageParam[] = [{ role: 'user', content: task }]
const allTools = [...tools, finishTool]
const seen = new Map<string, number>()
let costUsd = 0
let turns = 0
const stop = (reason: StopReason): Outcome => ({ status: 'stopped', reason, costUsd, turns })
while (turns < budget.maxTurns) {
if (signal.aborted) return stop('deadline')
if (costUsd >= budget.maxCostUsd) return stop('budget')
turns++
let res: Anthropic.Message
try {
res = await client.messages.create(
{ model: MODEL, max_tokens: 2048, tools: allTools, messages },
{ signal },
)
} catch (err) {
if (signal.aborted) return stop('deadline')
throw err
}
costUsd += cost(res.usage)
// A tool call cut off by max_tokens has incomplete JSON. Never run it.
if (res.stop_reason === 'max_tokens') return stop('truncated')
messages.push({ role: 'assistant', content: res.content })
const calls = res.content.filter((b): b is Anthropic.ToolUseBlock => b.type === 'tool_use')
const finish = calls.find(c => c.name === 'finish')
if (finish) {
return { status: 'done', answer: (finish.input as { answer: string }).answer, costUsd, turns }
}
if (calls.length === 0) {
messages.push({ role: 'user', content: 'Use a tool, or call finish with your answer.' })
continue
}
const results: Anthropic.ToolResultBlockParam[] = []
for (const call of calls) {
const key = `${call.name}:${JSON.stringify(call.input)}`
const n = (seen.get(key) ?? 0) + 1
seen.set(key, n)
if (n > MAX_REPEATS) return stop('looping')
results.push(await runTool(call, impl, signal))
}
messages.push({ role: 'user', content: results })
}
return stop('max_turns')
}And the tool runner:
async function runTool(
call: Anthropic.ToolUseBlock,
impl: Record<string, ToolFn>,
runSignal: AbortSignal,
): Promise<Anthropic.ToolResultBlockParam> {
const base = { type: 'tool_result' as const, tool_use_id: call.id }
const fn = impl[call.name]
if (!fn) return { ...base, is_error: true, content: `Unknown tool: ${call.name}` }
try {
const signal = AbortSignal.any([runSignal, AbortSignal.timeout(TOOL_TIMEOUT_MS)])
const out = await fn(call.input, signal)
return { ...base, content: out.slice(0, MAX_TOOL_CHARS) }
} catch (err) {
return { ...base, is_error: true, content: `Tool failed: ${String(err)}` }
}
}One signal for the whole run
The deadline is a single AbortSignal made with AbortSignal.timeout(). It goes into the SDK call as a request option, and into every tool.
Each tool also gets its own shorter timeout, merged with the run signal through AbortSignal.any(). The tool stops at whichever comes first: its own limit or the end of the run.
Checking Date.now() at the top of each loop is not enough on its own. That check only runs between turns, so a tool that hangs never gives the loop a chance to look at the clock. The signal cancels the work that is in flight.
This only helps if your tools respect the signal. Pass it to fetch, to your database driver, and to child processes. A tool that ignores its signal will still run to the end in the background, even after the loop has moved on without it.
Cost accounting per call
Every response includes a usage object with input and output token counts. Add them up after each call, not at the end, so the guard at the top of the loop sees the real number.
Two things to know:
- Input cost grows every turn. The whole conversation is resent on each call. Turn 15 costs much more than turn 1, so a turn cap alone is not a budget.
- Cached tokens are priced differently. If you use prompt caching,
usagealso reports cache reads and writes. Add those fields tocost()or your numbers will be off in one direction or the other.
The ceiling is a soft limit. One call can push you over it, because you only find out what a call cost after it returns. If you need a hard limit, estimate the next call's input tokens before making it and stop early.
Make "done" explicit
Without a finish tool, the usual rule is "stop when the model replies with no tool calls." That breaks in two common ways. Sometimes the model asks the user a question when nobody is there to answer it. Sometimes it writes a progress update ("I'll now check the second file") and then stops, and your loop treats that as the final answer.
A finish tool with a schema turns "I am done" into a structured call. The answer comes back in a known field, you can validate it, and a plain-text reply becomes a signal to nudge the model rather than a silent exit.
Failure modes worth naming
Looping. The model calls the same tool with the same arguments over and over, usually because the result did not contain what it expected. The seen map counts identical calls and stops the run after the second repeat. That is crude, but it catches the common case cheaply. You could also return an error result that says "you already called this" and give the model one more chance.
Truncated tool calls. If the model hits max_tokens while writing tool arguments, you get a tool_use block whose JSON is incomplete. The stop reason docs cover this case. Never execute a cut-off call. Either stop, as the sketch does, or retry the turn with a higher max_tokens.
Huge tool results. A tool that returns a whole web page fills the context, and every turn after that costs more. Truncating to a fixed character count is the blunt version. A better version summarizes or paginates, and tells the model the result was truncated so it doesn't treat a partial result as complete.
Tool errors. Return them to the model as is_error: true results instead of throwing. Models are often good at recovering from "file not found" when they can see it.
Return why it stopped
The return type is a union with a reason on every stopped run. If I had to drop everything else in this design, I'd keep that field. Log it, chart it, and alert on it. If budget stops climb after a prompt change, you know where to look. If looping shows up for one tool, the tool's output format is probably confusing the model.
Takeaways
- Put four limits on every loop: turns, wall-clock deadline, spend, and an explicit finish.
- Use one
AbortSignalfor the run and combine it with per-tool timeouts. Pass it all the way down. - Count cost after every call. Input cost grows each turn.
- Treat
max_tokensas a stop, not as a tool call to execute. - Return a typed stop reason and track it like any other production metric.
Anthropic's Building effective agents makes the same point from a different angle: start with the simplest loop that works, and add structure only where it pays off. In my experience, limits are the first bit of structure that earns its keep.