MCP Tool-Result Token Budget

The token budget each MCP tool result gets so one unbounded response cannot blow the context window.

Budget for all tool results combined-
Per tool call-
Roughly, in characters-
Utilization of the context window-
-

How the math works

The tool budget is the context window minus the system prompt, conversation reserve, and answer reserve, less a safety margin, divided across the tool calls. Four characters per token converts that into a truncation length for MCP server responses.

Frequently asked

Why budget tool results at all?

Tool and MCP results are the fastest way to blow a context window. One unbounded search or file dump can consume more tokens than the rest of the conversation combined, and cost scales with every token in the window on each subsequent turn.

Is 4 characters per token accurate?

It is a good planning average for English and code. JSON with heavy punctuation runs closer to 3 characters per token, prose closer to 4.5. Use 3.5 if your tool returns dense JSON.

Should truncation happen in the MCP server or the agent?

In the server when you know the shape of the data, for example keep the first N rows of a query result. In the agent only when the relevance depends on the task. Server-side truncation is cheaper because the tokens never reach the model.

About the author

I build and ship the ClaudFlow workflow builder, including its eight block types and JSON export format. The agent loop patterns these calculators price are patterns implemented in shipped code, not invented for a blog post. Everything runs client side by design: nothing you type leaves this page. More on the about page.

ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip. Token counts are estimates, and actual usage depends on encoding, model version, and prompt structure.

Prompt Caching Savings Estimator Agent Loop Cost Calculator All Tools