AI Agents: Ongoing Costs Scale with Tasks, Not Users

In short
AI agents don't scale per user—they scale per completed task. Inference costs rise with every agent task, while maintenance, monitoring, and compliance often make up two to four times the build cost over 24 months. The operating curve is the real lever, and it can be made predictable.
The New Cost Logic: Agents Bill per Task
Classic software costs scale with the number of users. AI agents work differently: once an agent independently runs research, checks documents, or takes over a process, compute is consumed for every single execution. Altimeter Capital projects installed AI compute capacity to rise from roughly 18 gigawatts in 2025 to around 115 gigawatts in 2028—with the largest jump of about 43 gigawatts in 2027 alone. For your budget, that means the operating cost curve follows completed tasks, not the number of logged-in employees.
Inference Is the Engine—and the Cost Driver
Inference—the actual execution of a model—already accounts for about two-thirds of all AI compute, according to Bloomberg Intelligence via ClearML, and continues to grow. At the same time, the unit price per inference has dropped by 99 percent over the past two and a half years. That sounds like relief, but it is a trap: inference-time reasoning has driven token consumption parabolic. You buy the unit cheaper, but your agents consume far more of it. Total cost becomes a function of workload.
≈ 2/3
of AI compute is already inference (Bloomberg Intelligence via ClearML, 2026)
Volume is exploding while unit prices fall: Gerstner revised the realistic 2027 capacity addition forecast to around 25 gigawatts instead of 43 gigawatts, citing grid, interconnection, and power bottlenecks. The decisive lever is not the daily price, but the cumulative task volume over months.
The Hidden Operating Costs Over 24 Months
Token costs are only a small part of the truth. Our own RAG analyses show that, depending on the setup, tokens account for just 10 to 30 percent of real costs. Monitoring, human-in-the-loop escalations, maintenance, model drift, and compliance dominate the bill—and can reach two to four times the build cost over 24 months. Add a silent share of 15 to 30 percent engineering time per month. This is the difference between a demo budget and an operating budget. Comparing API prices alone means planning past reality—as the Swiss AI podcast explains, most 2026 AI pilot projects die for exactly this reason.
For Swiss SMBs, Discretion Is an Operating Cost
revDSG compliance, data residency, and confidentiality are not one-time costs. They run through every agent task and must be built into the operating curve.
Why In-House Agent Development Gets Expensive
Building an agent is only the beginning. MIT NANDA shows that buying from specialized providers succeeds in about 67 percent of cases, while pure in-house development succeeds in only about 33 percent. With agents, this is amplified because ongoing costs scale with every task taken over. Building internally means carrying development, operations, escalation, and model maintenance yourself.
Predictable Operating Costs Through AI Outsourcing
Outsourcing changes the cost logic: you bring in an agent as an external division, with agreed deliverables, escalation paths, and defined operating costs—instead of financing the self-dynamics of growing agent tasks on your own. The Swiss AI podcast highlights how to make AI a management topic and anchor scaling at the executive level. That makes the inference and maintenance curve calculable.
The Decisive Lever: Steering Task Volume
Not every process is ready for autonomous agents. What matters is how often an agent actually runs, how many escalations it triggers, and how stable the data situation is. A feasibility matrix helps you assess task volume and thus realistic ongoing costs before scaling.
Frequently asked questions
- Why do AI agents scale with tasks instead of users?
- Because every executed task consumes its own compute. With classic software, the number of licenses drives cost; with agents, the volume of completed tasks determines the operating cost curve.
- What are the biggest ongoing costs after go-live?
- Beyond token costs, monitoring, human-in-the-loop escalations, maintenance, model drift, and compliance are the largest items. Together they can reach two to four times the build cost over 24 months.
- Is in-house development worth it for AI agents?
- Rarely. MIT NANDA shows buying from specialized providers succeeds in about 67 percent of cases, while pure in-house development succeeds in only about 33 percent—and in-house teams also carry the full operating burden.
- Why does this matter for Swiss SMBs?
- Because agents with growing autonomy can quickly become an unplanned operating expense. revDSG compliance and discretion add ongoing requirements that further burden the budget.
- How can AI outsourcing make costs predictable?
- Outsourcing defines deliverables, escalation paths, and operating costs contractually. Instead of financing the self-dynamics of growing agent tasks, the cost curve becomes calculable and tied to tasks rather than unclear user numbers.
Sources
Would you like to explore this topic for your company?
Check Availability