How to Monitor and Control LLM Token Spend: A Controller-Led Playbook

Untitled design (1)
  • Most AI-spend advice assumes a CFO who isn’t there. At the small and midsized business level, watching software and AI spend is a controller’s job. The controller builds the coding, sets the budget, and runs the review directly with the founder.
  • LLM token spend breaks the old software budget model. Seat-based tools scale predictably with headcount. Large language model (LLM) APIs bill by consumption instead, so the same task can cost dramatically more or less depending on the model and how it’s used.
  • Dormant SaaS seats are the same disease with a slower clock. Industry data puts license utilization near half, and the reconciliation that catches orphaned seats after turnover also catches duplicate tools.
  • The fix already lives in your accounting system. Coding every software and AI dollar to a vendor and a department, then reviewing budget versus actual monthly, catches both problems before they compound.
  • Classification questions belong with your tax preparer before year-end. How AI and software costs get characterized can affect deduction timing and, where a company builds AI features internally, eligibility for the research and development (R&D) tax credit.

 

Ask a CFO in 2026 what line item worries them, and the answer keeps shifting. Not payroll. Not rent. The AI bill.

The FinOps Foundation’s 2026 State of FinOps report found that 98% of teams now manage some form of AI spend. That’s up from just 31% two years earlier. The capability practitioners say that what tey most lack is granular monitoring: tokens, requests from large language models, GPU utilization. 

Zylo’s 2026 SaaS Management Index found AI-native application spend up 108% year over year. And 78% of IT leaders reported unexpected charges tied to consumption-based or AI pricing. Deloitte published a CFO guide to AI token economics in April 2026. Two years ago, that topic barely registered on the finance radar.

Most companies adopted AI tools faster than they built the controls to watch them. The good news is that the controls aren’t new. They’re the same budget disciplines you should already be running on the rest of your software stack. The AI bill just made the cost of skipping them impossible to ignore. 

One thing the enterprise coverage gets wrong for most of the market: it assumes a CFO is in the seat

Fusion CPA serves small businesses (roughly $0 to $5 million in revenue) and midsized ones ($5 to $50 million). Many of them run without a CFO, full-time or fractional. Instead, the structure is a controller and a founder. Whether you run a small or midsize business, without a CFO on staff, a good controller ends up doing a scaled-down version of the CFO’s job: watching cash, setting budgets, and flagging what needs a decision. When it comes to spend monitoring specifically, the controller doesn’t just support the effort; the controller leads it, building the coding structure, running the reviews, and flagging the trends worth the founder’s attention. Everything in this article assumes this structure is in place. 

Why Is LLM Token Spend So Hard to Budget?

For twenty years, software followed one pricing model: pay per seat, scale predictably. A CFO could budget software the way they budget headcount. That’s because software spend moved with headcount. Buy 40 seats in January, and you pay for 40 seats in December.

Token pricing broke that model. Most LLM APIs, and a growing share of AI products, bill by consumption instead. You pay for what the model reads and writes, measured in tokens: the small chunks of text a model processes. Rates vary enormously by model. In 2026, lightweight production models run pennies per million tokens. Frontier reasoning models can run $100 or more per million. Route the same task to a different model, and it can cost many times as much.

Usage is just as variable as pricing. A simple query might cost a fraction of a cent. A long, multi-step workflow with document analysis and code generation can cost dollars per run. Agentic workflows are different still: rather than answering once, the model works through a task in a sequence of steps on its own, searching, drafting, checking its own work, and moving to the next step, before returning a result. According to Gartner’s 2026 analysis, these consume meaningfully more tokens than a standard chat interaction. A pilot that costs a few hundred dollars a month can turn into a five-figure monthly line item once it ships to the whole team. There’s often no contract amendment and no purchase order to flag it. The invoice is frequently the first place finance finds out. 

Seat-Based vs. Token-Based Spend

Seat-based spend (Microsoft, Adobe, HubSpot) Token-based spend (LLM APIs, AI tools)
How it’s billed Fixed price per user per month Variable, by consumption
How it grows Linearly, with headcount Nonlinearly, with usage and model choice
Where it leaks Dormant seats after turnover; duplicate tools; auto-renewals Runaway workflows; expensive model defaults; no per-team attribution
When you find out At renewal, if anyone looks On the next invoice
The fix Seat audits tied to offboarding; renewal calendar Department-level budgets; usage caps and alerts; monthly budget versus actual

The Dormant Seat Problem Was the Warm-Up Act

Before tokens, there was shelfware: software nobody opens but still pays for. Specifically, Zylo’s 2026 SaaS Management Index puts license utilization at 54%, up from 47% the year before. In other words, the average organization still leaves roughly 46% of its licenses unused. According to Zylo’s dataset, that’s worth close to $19.8 million a year in waste. Your business is probably not wasting millions. Even so, the percentage travels down-market just fine. In our experience, the mechanism is always the same.

First, someone leaves the company. Because it’s a security checklist item, IT cuts their email and laptop access on day one. However, nobody cuts their Adobe Creative Cloud seat or their HubSpot seat. Likewise, nobody touches their Google Cloud Platform project access, or the project management tool marketing signed up for two years ago. That’s because those live on a corporate card and renew automatically. As a result, six months later, you’re paying for a department that’s 20% smaller at 100% of the old software cost.

To see this in practice, run the math on a modest version. Take a 40-person company spending $250 to $400 per employee per month on software. Over a year, it loses six people and backfills four. Say five or six orphaned seats survive across four or five tools at $30 to $80 per seat. As a result, that’s roughly $2,500 to $7,000 a year gone, from turnover alone. And that’s before counting duplicate tools or unused premium tiers. Naturally, scale the headcount, and the leak scales with it.

In short, token spend is the same disease with a faster clock. A dormant seat costs you the same wrong number every month, while a runaway workflow costs you a different, bigger wrong number every month. That’s why the fix has to live in your financial system, where both kinds of spend land. It can’t, in other words, live in a spreadsheet someone updates when they remember.

Which Businesses Feel This First?

For AI startups, this isn’t an overhead conversation at all. Token spend is typically cost of goods sold (COGS). Every customer interaction burns inference, the computing cost of running a trained model. That means model choice and usage controls flow straight into gross margin, and gross margin is the number every investor reads first. Industry benchmarks report meaningful gross margin erosion tied to AI workloads at a significant share of enterprises surveyed. A startup that budgets tokens by department and reviews the trend biweekly in QuickBooks Online or NetSuite isn’t just controlling costs. It’s protecting the margin story it will tell at the next raise.

The rest of our client base feels the same pressure a step behind. Technology and eCommerce companies run the largest tool stacks. They’re usually first to hit consumption-billed AI features inside software they already own. Law firms, staffing agencies, and marketing agencies buy seats in proportion to billable headcount. Staffing in particular churns people fast enough that orphaned licenses are almost guaranteed without an offboarding checklist. Healthcare groups and management services organizations (MSOs) multiply seats per provider per location, so one unwatched tool becomes twelve. Private equity firms see the duplication problem at portfolio scale: the same platform, bought separately by three companies that share an owner. Real estate, entertainment, family businesses, and field service companies tend to run lean back offices. The honest answer to “who reviews software spend” is often nobody, which is exactly the gap a fractional controller closes.

Different industries, same formula. The dimensions change (names, matters, providers, properties, crews), but the structure and the cadence don’t.

What Does the Monitoring Playbook Actually Look Like?

There’s a complicated formula Fusion CPA sets up for clients on overhead and marketing spend, now pointed at software and AI. 

1. Structure the chart of accounts and dimensions so spend has an address.

Every software and AI dollar gets coded to a vendor and a department. That’s class or location in QuickBooks, department or class segments in NetSuite

An entry like “Software subscriptions, $18,400” tells you nothing. “OpenAI API, $2,100, Engineering” and “HubSpot, $1,650, Marketing” tell you where to look. Most books we take over have one undifferentiated software account. That single design choice is why nobody caught a drift in the first place.

2. Set a budget at the level you coded.

QuickBooks Online Plus and Advanced support budgets built by class or location. That gives each department’s software line its own number. NetSuite goes further, with budgets across department, class, and location segments. It also offers NetSuite Planning and Budgeting for companies that want driver-based forecasts. For consumption-based AI spend, many teams budget a range rather than a point number, and treat the top of the range as an alert threshold.

3. Put a gate in front of new spend.

Approval workflows catch new vendors and plan upgrades before they hit the card. That means bill approvals in QuickBooks Online Advanced, or SuiteFlow approvals in NetSuite. Spend management platforms such as Ramp add virtual cards with per-vendor limits, a clean way to hard-cap a department’s AI tooling at a dollar figure. On the API side, the major LLM providers offer usage limits and alerts at the workspace level. Finance should ask engineering to set those to match the budget, not leave them at the default.

4. Review budget versus actual monthly, by department, by vendor.

This is the step that actually catches things. The report takes minutes to run in either system once the coding is in place. Look for drift: a vendor line that grew 40% with no headcount change, a department 30% over its software budget three months running, a new vendor nobody approved. Every one of those is a conversation. Most conversations end in money recovered.

None of this requires new software. It requires the accounting system you already pay for to be configured like a monitoring system instead of a filing cabinet.

Who Should Be Watching This, and How Often?

This only works if someone owns the cadence, which is where the controller earns their keep. Most companies review the general ledger biweekly or monthly, with the controller walking spend by department and vendor and flagging drift. At most of our clients’ size, that’s a direct review with the founder; where a CFO is in the seat, the controller runs it with them instead. Either way, the controller watches; the review is where findings become decisions. Twenty minutes with the right report beats two hours with the wrong one.

The cadence is the point. Reviewed once a year, at tax time, the damage is already twelve months deep: a dormant seat caught in month two costs two months, caught at year-end it costs the full year. A vendor line up 15% two periods running is a conversation; the same line 90% higher a year later is just a loss.

An outsourced controller engagement is built for exactly this. Recovered spend, reclaimed seats, cancelled duplicates, capped AI usage, often offsets a meaningful share of the monthly fee, and spend monitoring is just one piece of what gets reviewed alongside close quality and cash trends.

Schedule a Discovery Call

A Pattern We See

A recent controller client had software spend that grew significantly over a year or so. Headcount, meanwhile, grew only modestly. At first, everything was coded to a single “Dues and Subscriptions” account. Once we split it by vendor and department, though, a few findings covered most of the gap. Specifically, seats of departed employees were still renewing across several tools, and more than one department was paying for duplicate tools. A single AI tool was also involved. It started as one seat, but quietly grew into a consumption-billed plan with a charge that swung sharply and no one watching it. None of this was fraud or carelessness. Instead, it was invisible, because the books weren’t built to show it. So the fix was the four-step structure above, and now the review takes about twenty minutes. (Note: to protect our clients’ privacy, this is a composite taken from recent client engagements.)

How Does Software and AI Spend Show Up on the Tax Return?

One more reason to get the coding right: classification isn’t just a management question. How AI and software costs are characterized can affect the timing of deductions. Where a company builds its own AI features on top of LLM APIs, a portion of that spend may also relate to research and experimental activity. That activity carries its own federal treatment. In many cases, it also carries research and development (R&D) tax credit implications. State treatment of these provisions varies, and multi-state businesses often see different answers in different filings. Put this question in front of your tax preparer before year-end, not after. It’s a standing part of the year-round planning conversation Fusion CPA runs with clients. For Georgia-based businesses, our Georgia state tax guide covers how the state approaches conformity questions generally.

Where to Start This Week

If you’re doing this for the first time, work through it in order. Start by pulling twelve months of software and AI spend and splitting it by vendor. Most teams find vendors they forgot they had on the first pass. Next, assign each vendor to a department and recode go-forward transactions with class, location, or department dimensions. Then reconcile seats to the current roster for your five to ten largest per-seat vendors, and add license review to the offboarding checklist so the leak stays closed. Finally, set department-level software budgets, with a separate line or range for consumption-billed AI spend, and put budget versus actual on the monthly close calendar with a named owner.

The companies that struggle here are rarely missing a tool. They’re missing the structure and the cadence.

Why Work With Fusion CPA on This

Software and AI spend rarely announces itself; it just shows up bigger on next month’s invoice. Fusion CPA works with founders and controllers whose software budget has outgrown a single “Dues and Subscriptions” line, including remote offices in Atlanta, GA, Tampa, FL, San Juan, Puerto Rico, and Park City, Utah. We code every software and AI dollar to a vendor and department, build budgets by class or location in QuickBooks or NetSuite, and put a monthly budget-versus-actual review on the calendar. We also flag classification questions for your tax preparer, since how AI and software costs are characterized can affect deduction timing and R&D credit eligibility.

Schedule a Discovery Call to get your software and AI spend under a real monthly review before the next invoice surprises you.

Frequently Asked Questions

How do CFOs track LLM token spend by department? 

Generally by combining two views. One is provider-side: usage dashboards and API keys separated by team or project. The other is accounting-side: coding that books each AI vendor’s invoice to a department dimension in the financial system. The accounting view is the one that supports budget versus actual reporting. Fusion CPA sets this structure up as part of outsourced controller engagements. 

How much do companies typically waste on unused SaaS seats? 

Industry studies consistently put license utilization well below full use. Zylo’s 2026 SaaS Management Index found utilization at 54%, meaning roughly 46% of provisioned licenses go unused. That works out to an average of $19.8 million a year in waste across organizations in its dataset. Actual waste varies with company size and turnover. That’s why a seat-to-roster reconciliation on your largest vendors is usually worth the hour it takes.

How often should a controller review the general ledger for spending trends? 

For most businesses, biweekly or monthly, depending on transaction volume. The controller reviews spend by department and vendor against budget. Therefore, the controller walks the findings with the CFO, or with the founder directly in companies without one. An annual review, typically at tax time, is generally too late. By then, a dormant seat or a runaway consumption charge has been billing for a year. Fusion CPA’s outsourced controller service runs this cadence for clients as a standard part of the monthly close.

Does Fusion CPA help set up budget monitoring in NetSuite or QuickBooks? 

Yes. Fusion CPA provides tax preparation, tax planning, outsourced accounting, and CFO advisory for businesses across 40+ states. Configuring department and vendor-level budget monitoring in QuickBooks Online and NetSuite is a core part of its bookkeeping, controller, and CFO work.

Schedule a Discovery Call

 


About the author

Trevor McCandless, CPA, MTax is the founder and CEO of Fusion CPA, a tax, outsourced accounting, and advisory firm serving business owners and high-achieving individuals across 40+ states from offices in Atlanta GA, Tampa FL, San Juan Puerto Rico, and Park City Utah. He works with owners on entity structure, owner compensation, and the multi-state exposure that arrives quietly with a growing team.

This article is reviewed by Steven Sumners, CPA, MAcc, Senior Tax and Accounting Manager at Fusion CPA.

About this article and how we use AI

This article is provided for general informational and educational purposes only and does not constitute tax, legal, accounting, or financial advice. Tax laws change and apply differently depending on your specific circumstances. Nothing here creates a client relationship with Fusion CPA, and it should not be relied upon or acted on without consulting a qualified professional about your own situation. To discuss how these rules apply to you, contact Fusion CPA at info@fusiontaxes.com.

Fusion CPA articles are grounded in the professional experience of our CPAs and the situations we encounter in practice. Scenarios described are illustrative composites, not any individual client’s facts. We use AI tools to assist with drafting and research. Before publication, every article is verified against primary sources such as the IRS and state departments of revenue and is reviewed for technical accuracy by a licensed CPA, who is named on the piece