Your AI assistant can burn through thousands of dollars a year before it answers a single question. That is the uncomfortable truth behind AI agent token costs: every tool you connect — your practice management system, your ticketing platform, your CRM — loads a definition into the model’s context window at the start of every conversation, whether the agent uses it or not. Connect enough tools and you are paying a standing tax on work that has not happened yet.
Where AI agent token costs actually come from
Anthropic published concrete numbers on this. In a five-server setup — GitHub, Slack, Sentry, Grafana and Splunk — 58 tools consume roughly 55,000 tokens before the conversation even starts, and adding a system like Jira pushes that another ~17,000 tokens. Internally, Anthropic saw tool definitions reach 134,000 tokens before optimization.
The second leak is intermediate data. When an agent pulls a full document or a 10,000-row report just to answer one question, every row passes through the model — often twice. That is the part most businesses in NJ, NY and PA never see on an invoice until the bill arrives.
The Context Tax, Measured
Source: Anthropic engineering, “Introducing advanced tool use” and “Code execution with MCP.”
Three proven ways to cut the bill
Load tools on demand. Instead of dumping every schema into context, the agent searches for the three to five tools it actually needs. Anthropic measured an 85% reduction in token usage — and better accuracy, with one model improving from 79.5% to 88.1% on tool-selection evaluations.
Let the agent write code, not copy data. Presenting connected systems as code APIs lets the agent filter a 10,000-row result down to five before anything reaches the model. Anthropic’s code execution with MCP approach cut one workflow from 150,000 tokens to 2,000 — a 98.7% saving.
Keep intermediate results out of context. Orchestrating multiple tool calls in a single execution step dropped average usage from 43,588 to 27,297 tokens — a 37% reduction — on complex research tasks. It also keeps sensitive records out of the model entirely, which matters enormously for healthcare and dental practices.
Why this is a governance problem, not just a billing one
Every token that flows through a model is also data that flows through a model. The same architecture that trims AI agent token costs — on-demand tool loading, filtering before the model sees anything, execution inside a controlled environment — is the architecture that keeps patient records and client files out of a vendor’s context window. Cost efficiency and data governance are the same design decision.
How Aufsite builds cost-efficient AI agents in NJ, NY and PA
Aufsite’s Secure MCP Framework connects AI assistants to your business systems with governed access, scoped authentication and guardrails — and with on-demand tool loading and in-environment filtering built in from day one, so you are not paying a five-figure context tax for tools your agent rarely touches. It applies this same discipline to all business use case workflows.
As an AWS Select Partner based in Princeton, we design, deploy and manage these integrations for businesses and healthcare practices across New Jersey, New York and Pennsylvania — with your AI spend, security posture and audit trail treated as one problem, not three.
Paying more than you expected for AI that has not earned it yet? Talk to Aufsite about managed AI and cloud support — we will show you exactly where the tokens are going and what a governed, efficient agent architecture looks like for your practice or business.
