The token budget each MCP tool result gets so one unbounded response cannot blow the context window.
The tool budget is the context window minus the system prompt, conversation reserve, and answer reserve, less a safety margin, divided across the tool calls. Four characters per token converts that into a truncation length for MCP server responses.
Tool and MCP results are the fastest way to blow a context window. One unbounded search or file dump can consume more tokens than the rest of the conversation combined, and cost scales with every token in the window on each subsequent turn.
It is a good planning average for English and code. JSON with heavy punctuation runs closer to 3 characters per token, prose closer to 4.5. Use 3.5 if your tool returns dense JSON.
In the server when you know the shape of the data, for example keep the first N rows of a query result. In the agent only when the relevance depends on the task. Server-side truncation is cheaper because the tokens never reach the model.
I build and ship the ClaudFlow workflow builder, including its eight block types and JSON export format. The agent loop patterns these calculators price are patterns implemented in shipped code, not invented for a blog post. Everything runs client side by design: nothing you type leaves this page. More on the about page.
ClaudFlow is a free, privacy-first suite of AI workflow tools. No accounts, no tracking, no data leaves your browser. Part of the Zovo Tools network by Michael Lip. Token counts are estimates, and actual usage depends on encoding, model version, and prompt structure.