Insights

How to Forecast Agent Workloads for AI Budgets

By Brian Diamond

Published October 5, 2026

A monthly AI vendor invoice can tell you what was spent. It cannot tell you whether next quarter's spend will rise because sales is deploying a new research agent, support volume is seasonal, or an existing workflow is silently retrying failed calls. To forecast agent workloads, finance and technology teams need to model the operational drivers behind AI usage, then translate them into attributable cost.

This is different from forecasting a conventional software subscription. Agent costs are usually variable, usage is distributed across teams and vendors, and one business process can trigger multiple model calls, retrieval steps, tool calls, and retries. A credible forecast starts with the workload, not the invoice total.

Start with the unit of work

The most useful forecast is built around a business event an operator can observe and a finance leader can understand. Depending on the agent, that unit may be a support ticket resolved, sales account researched, claims document reviewed, marketing asset generated, or code change analyzed.

Avoid beginning with total tokens. Tokens are necessary for calculating cost, but they are not a planning unit most business owners can forecast. A support leader can estimate ticket volume from staffing plans, backlog, seasonality, and product launches. A platform team can then convert that expected ticket volume into model usage.

For each production agent, define four relationships:

  1. The forecasted business demand, such as 80,000 support tickets per month.
  2. The adoption or automation rate, such as 35% of tickets routed to the agent.
  3. The technical consumption per agent run, including input tokens, output tokens, model calls, and non-model services.
  4. The fully loaded cost per run, assigned to a team, project, and cost center.

This structure makes assumptions visible. If the forecast misses, teams can identify whether demand changed, adoption changed, or the agent became more expensive to operate.

How to forecast agent workloads from demand signals

Begin with a baseline from recent production data. Three to six months is often enough for an initial model, provided the agent and its underlying workflow have not changed materially. If usage is newer, use weekly data and document that the forecast has lower confidence.

Separate the demand forecast from the consumption forecast. Demand answers how many business events will occur. Consumption answers what the agent does for each event. Combining them too early can hide the source of a variance.

Consider a customer support agent. Historical data shows 50,000 tickets per month, with 30% eligible for agent handling. The agent is used on 60% of eligible tickets, creating 9,000 runs. Next quarter, the support organization expects ticket volume to increase 20% after a product release and plans to raise agent adoption to 75%.

The workload forecast is 50,000 × 1.20 × 30% × 75%, or 13,500 agent runs per month. That is the operational workload finance should review with the support owner. It is also a far more useful planning number than a blanket assumption that the AI bill will grow by 50%.

Demand signals vary by function. Revenue agents may follow pipeline volume, account coverage, or outbound campaign plans. Finance agents may follow invoice count, close-calendar activity, or expense report volume. Engineering agents may track active developers, pull requests, or deployment frequency. The right driver is the one with a plausible causal connection to agent runs.

New programs require a different approach. Use a rollout curve rather than assuming full adoption on day one. For example, forecast 10% of eligible users in month one, 30% in month two, and 50% in month three, with a separate downside scenario for slower enablement. That is more honest than presenting an annualized usage number as a committed operating plan.

Convert runs into AI consumption and cost

Once expected runs are clear, calculate consumption per run from telemetry. At a minimum, capture input tokens, output tokens, model or deployment, and request count. For agents that use retrieval, tools, or orchestration, include embedding activity, vector storage or queries, gateway fees, compute, and third-party tool charges where they are material.

A simple cost formula is:

`Monthly agent cost = forecasted runs × average cost per run + fixed platform costs`

The average cost per run should reflect the model mix actually used. Many agents route simple requests to a lower-cost model and escalation cases to a higher-cost model. A forecast that uses one blended rate is acceptable for an early budget, but it becomes unreliable when routing logic or task complexity changes.

For example, suppose an account research agent is expected to run 20,000 times per month. Its average model and infrastructure cost is $0.18 per run, and it has $1,400 in monthly fixed costs for shared services. The forecast is $5,000 per month: $3,600 in variable cost plus $1,400 fixed.

That number should not be treated as final until the team examines outliers. Averages can conceal costly behavior such as unusually long prompts, repeated agent loops, failed tool calls, or fallback to a premium model. Use median and percentile consumption alongside the average. If the 95th percentile run costs ten times the median, investigate whether that behavior is intentional and budgetable.

Account for the drivers that make forecasts miss

Agent workload forecasts are most often wrong for understandable reasons, not because the math is difficult. The model needs explicit assumptions for changes that affect either volume or unit cost.

Seasonality affects many workflows. A retailer may see more support contacts during promotions; a finance team may process more documents at month-end or quarter-end. Product launches, acquisitions, geographic expansion, and policy changes can alter demand quickly.

Technical changes matter just as much. A revised system prompt can increase input tokens. Adding retrieval context can increase both token volume and database activity. A model routing change can lower unit cost while increasing request counts. Better caching can reduce spend without reducing business activity. Track these changes as forecast assumptions, not as unexplained variance.

Retries deserve special attention. An agent that averages 1.2 model calls per successful task is fundamentally different from one that averages 2.5. Forecast both attempted runs and completed business outcomes. Otherwise, a reliability problem can appear as higher demand and be incorrectly charged to the business team.

Forecast at the ownership level

A forecast becomes financially useful when it has an owner. Aggregate vendor-level projections are useful for procurement, but they do not support accountability. Each forecast line should identify the consuming team, agent, project, cost center, and business owner, along with the allocation method for shared services.

Direct costs should follow measured usage whenever possible. If a marketing agent uses 12% of a shared gateway's attributable requests, it should receive 12% of the variable gateway cost. Fixed shared costs can be allocated by request volume, active users, reserved capacity, or another documented driver. The best method depends on what actually causes the cost and whether stakeholders view it as fair.

Consistency matters more than theoretical perfection. Finance needs an allocation policy that can be explained, repeated, and reconciled to the vendor invoice. Platform teams need enough granularity to troubleshoot. A monthly process should compare forecasted cost, metered actuals, vendor-billed amounts, and the amount allocated to each owner.

This also supports accounting discipline. If a shared AI platform cost is initially recorded in a central technology cost center, an allocation entry can move the appropriate share to consuming departments. For example, debit Sales AI Expense and credit Central AI Platform Expense for the sales allocation. The exact accounts vary by chart of accounts, but the ownership trail should remain intact.

Use scenarios, not one budget number

AI usage is still evolving in most organizations, so a single-point forecast can create false confidence. Build at least a base case, a low case, and a high case. The cases should differ based on operational assumptions, such as ticket volume, adoption, routing mix, or cost per run, rather than arbitrary percentages applied to total spend.

A useful planning view might show the base case at 13,500 monthly support-agent runs, the low case at 10,000 because adoption stalls, and the high case at 18,000 because ticket volume rises and the agent is expanded to another queue. Finance can then reserve contingency where it is warranted, while technology can identify the capacity and controls needed for the high case.

Review the forecast monthly, with a faster cadence for rapidly changing deployments. Explain material variance in plain language: run volume was 18% above plan due to a product release; unit cost was 12% below plan after routing changes; retries added $4,200 after a tool integration failure. Those explanations create a management process, not just a reporting exercise.

The practical goal is not to predict every token perfectly. It is to give every meaningful agent workload a business driver, a cost owner, and a documented path from operational demand to the general ledger. That is how AI spending becomes a forecastable operating expense rather than a recurring surprise.

Not sure which of your AI costs are being booked? Run the free Agent Spend Assessment.

Onaro Meridian is FinOps for agentic AI: the system of record that attributes, controls and books what AI agents spend.

Brian Diamond

Brian Diamond

Brian Diamond is a fractional Chief AI Officer and founder of Onaro. He has spent 30 years running infrastructure operations and founded LANStatus, a Connecticut managed services provider and Microsoft partner, in 2001. He holds a Chief AI Officer certification and writes the CAIO Brief on AI leadership for finance and operations.

LinkedIn · CAIO Brief · Author page

Markdown version