Skip to main content
Back to the field guide

Meet Budget

AI Agent ROI: What to Measure Before You Buy

AI agent ROI is not a per-seat price tag, it is spend attributed to shipped work. Tonone's Budget maps token cost to team and workflow, Lumen defines the output metric, and Mint turns it into a board-ready return multiple.

Budget · AI Cost Engineer10 min readMay 14, 2026

Every CTO who has approved a Claude Code rollout eventually gets the same question from the CFO: what is the AI agent ROI on this. Not is it working, not do engineers like it, but a number. The invoice lands in finance as a lump token cost spread across dozens of engineers, several models, and hundreds of sessions, with no attribution to what any of it produced. Meanwhile the engineering team insists velocity is up, but nobody can point to a dashboard that proves it. This gap, between spend that is trivially easy to measure and value that nobody bothered to measure, is exactly where AI agent rollouts stall at the renewal conversation. Leadership rarely kills the tool because it failed. They kill it because nobody built the case for keeping it, and a token invoice with no attached outcome is the easiest line item on the board deck to cut. The teams that survive the renewal conversation are not the ones spending less, they are the ones who walked in with a number nobody could argue with.

Why generalist tools can't answer the AI agent ROI question

Ask ChatGPT or Claude.ai to help you build a cost-to-value case for your AI agent rollout and you get a spreadsheet template, not a number. A generalist chatbot has no visibility into your actual API billing, no access to your Claude Code session logs, and no sense of which invocations produced a merged pull request versus which ones were an engineer's abandoned exploration at 11pm. It can suggest metrics to track in the abstract, cost per token, cost per session, cost per engineer, but it cannot compute any of them against your real spend, because it was never connected to your billing data in the first place. The gap here is not intelligence. It is grounding. A chatbot answering in a vacuum can produce a plausible-sounding ROI framework, but it cannot produce a defensible ROI number, and a defensible number is the only thing a CFO accepts in a renewal conversation.

Cursor and GitHub Copilot make the ROI question look simpler than it actually is, because they bill per seat. Twenty or forty dollars a seat times however many engineers is an easy number to write on a slide, right up until someone asks what those engineers produced with it, and the flat-fee model has no answer, because it was never designed to attribute output to spend. It is a fixed cost with a vague productivity story bolted on, not a measured return. Usage-based platforms like Claude Code are actually the more honest signal here, spend rises and falls with real work, but only if someone is capturing that signal against the outcomes it produced. Nobody ships that dashboard by default. Every team we have seen struggle with the AI agent ROI question is sitting on a usage-based invoice they have never once cross-referenced against what shipped.

Budget: the agent that turns a token invoice into an ROI case

Budget is Tonone's AI cost engineer, the specialist built to answer exactly the question a CFO asks at renewal time. Where a generalist tool can only speculate about spend, Budget starts from the actual billing data, the actual session logs, and the actual per-team usage patterns, and builds the case up from there. It does not stop at a cost report either. Budget hands off to Lumen to define what output the spend should be judged against, and to Mint to translate the resulting number into the return-multiple language a board understands. That three-agent chain, cost mapping, output definition, financial framing, is what turns a scary invoice line item into a defensible number.

Mapping the cost topology before anyone argues about the total

The budget-recon skill is where Budget starts. It maps AI cost topology: billing attribution by team and by model, forecast versus actuals, and any gaps in spend alerting. Instead of one number on an invoice, you get a broken-down picture of who is spending what, on which model tier, and whether anyone would even notice if a team's spend tripled overnight. Most orgs running Claude Code across more than one pod have never had this view. They have a total. They do not have a topology. Budget builds the topology first because you cannot judge return without knowing where the cost is actually concentrated.

Finding the waste hiding inside the total

Once the topology exists, budget-audit goes looking for waste. It produces a per-model cost breakdown, identifies the top consumers by team and workflow, and flags optimization levers, the wrong model tier assigned to routine work, redundant invocations, sessions that re-fetch context a caching layer should have preserved. This is the step that usually produces the first real surprise in the process, because most AI spend waste is not dramatic misuse, it is a handful of default settings nobody revisited after the pilot phase ended.

Tonone's Budget maps AI cost topology to the team and workflow that generated it, not just the total that shows up on the invoice.

Designing the reduction, without cutting anyone's access

With the waste identified, budget-optimize designs the actual cost reduction strategy: model tiering so routine work runs on a cheaper tier and only architecture or security-sensitive work reaches for the top tier, prompt compression, caching for repeated context, and batch inference where sessions allow for it. This is where Budget earns its keep against the alternative of a leadership team simply capping seats or restricting access, which reduces spend by reducing the tool's usefulness. Budget's version of cost reduction targets waste specifically, so the engineers who were getting real value keep getting it.

Lumen defines what the spend should be judged against

A cost dashboard alone still does not answer the ROI question, because cost without an output metric is just an expense report. This is where Lumen, Tonone's product analyst, comes in. The lumen-metrics skill produces a metrics plan, a North Star metric and an input metrics tree, that gives the engineering org a defined output to measure against, cycle time from first commit to merge, PRs shipped per engineer-week, whatever the team actually treats as a signal of throughput. Budget's cost data means nothing on its own. Cross-referenced against Lumen's output metric, session-level spend times, it becomes a real productivity signal.

Mint turns the dashboard into a number the board approves

The last mile is financial framing, and that belongs to Mint. The mint-unit skill audits and improves unit economics, and it is exactly the right lens for this problem: instead of reporting spend and output as two separate charts, Mint computes a cost-per-output figure and expresses the whole thing as a return multiple, dollars of engineering capacity unlocked per dollar of AI spend. That is the sentence a board actually wants to hear, and it is a sentence nobody in the chain before Mint is positioned to produce on their own.

Tonone's Mint turns a token spend number into a return multiple a board can act on, not just a chart engineering hopes finance trusts.

A worked example: the renewal conversation that almost went badly

Bridgepoint is a 55-person Series B vertical SaaS company running Claude Code across three engineering pods: Platform, Product, and Data. Two weeks before Q2 board prep, the CFO flags AI tooling spend of $16,800 a month and asks the CTO for a one-line answer on whether it is worth renewing. The CTO does not have one. Six months earlier the rollout had been approved on a pilot basis with a single Slack thread of anecdotes, engineers liked it, PRs felt faster, nobody had written down a baseline. That absence of a baseline is exactly what makes the renewal conversation so uncomfortable: without it, every dollar of spend looks like pure cost, because there is nothing to compare it against. That is the moment Budget gets pulled in.

Where Bridgepoint's own instincts would have gone wrong

  • Judging spend against a total instead of a topology, so a $16,800/month number gets debated as one line item instead of three, one of which is mostly waste.

  • Treating engineer sentiment ("it feels faster") as a substitute for a measured baseline, which collapses the first time finance asks for a number instead of a feeling.

  • Comparing usage-based AI spend to a flat per-seat tool's price tag directly, which ignores that the flat fee never measured output either, it just hid the question better.

  • Cutting seats or usage caps to reduce spend instead of auditing for waste first, which reduces cost by reducing the value engineers were already extracting.

  • Running the cost audit once at rollout and never again, so a model pricing change or a new team's usage pattern six months later goes completely undetected until the invoice spikes.

budget-recon runs first and produces the topology nobody had: Platform pod at $7,100/month, Product pod at $6,900/month, Data pod at $2,800/month, with no per-team spend threshold and no alerting configured on any of the three. budget-audit goes deeper on the two largest pods and finds two concrete waste sources: Product is running a top-tier model on routine CRUD scaffolding that a lighter tier handles just as well, an avoidable $2,300/month, and a nightly job in Platform invokes a heavy review skill on every commit regardless of whether the diff touches anything relevant, another $900/month. budget-optimize proposes model tiering for Product's routine work and gating the nightly job to relevant diffs only, projecting $3,200/month in reduction, about 19 percent, without removing anyone's access to the tool.

text
Bridgepoint, AI Agent ROI Dashboard (Q2)
Recon: 3 pods, $16,800/mo total spend, no per-team alerting.

Before optimization
  Platform pod:  $7,100/mo
  Product pod:   $6,900/mo  (incl. $2,300 avoidable, wrong model tier)
  Data pod:      $2,800/mo
  Total:         $16,800/mo

After budget-optimize (model tiering + diff-gated nightly job)
  Total:         $13,600/mo  (-19%, no access reduction)

Output metric (Lumen north star: PRs merged/engineer-week)
  Baseline (pre-rollout):        3.1 PRs/engineer-week
  Current (agent-assisted):      4.4 PRs/engineer-week
  Attributable lift:             44 additional PRs/month org-wide

Mint unit economics
  Blended fully-loaded PR cost:  $650
  Value of attributable lift:    44 x $650 = $28,600/mo
  AI spend (post-optimization):  $13,600/mo
  Return:                        2.1x

Board line: "AI agent spend returns $2.10 in engineering capacity
 for every $1 spent, after eliminating $3,200/mo in avoidable waste."

The CTO walks into board prep with a single line instead of a shrug: a 2.1x return, a documented waste elimination, and a metric, PRs merged per engineer-week, that the board already recognized from prior quarters because Lumen tied it to an existing tracking system rather than inventing a new one. The CFO's question gets answered with a number instead of a vibe, and the renewal gets approved without a debate about whether the tool is worth it, because the debate already happened inside the dashboard.

CapabilityTononeGeneralist chatbotCursor / Copilot
Attributes token spend to team and workflowYes, budget-recon maps billing to the pod and session that generated itNo, no access to your billing or session dataNo, flat per-seat pricing hides workflow-level attribution
Flags avoidable waste in model selectionYes, budget-audit identifies wrong-tier model use and redundant invocationsNo, can only suggest generic categories of wasteNo, fixed fee has no concept of per-task waste
Recommends concrete cost reduction leversYes, budget-optimize designs tiering, caching, and gating specific to your usageGeneric advice, not grounded in your actual spendNone, pricing is fixed regardless of usage pattern
Defines the output metric spend is judged againstYes, lumen-metrics ties spend to an existing North Star metricNo connection to your metrics stackNo metrics layer exists at all
Produces a board-ready return multipleYes, mint-unit expresses cost-per-output as a return figureNo, produces a template, not a computed numberNo, a seat price isn't a return calculation
Ongoing monitoring vs one-time estimateYes, recon and audit are designed to be rerun each cycleNo persistence across sessionsStatic invoice, no recurring analysis

Tonone's Lumen defines the output metric AI spend should be judged against, before the invoice becomes the headline in a board meeting.

If a renewal conversation is coming and nobody has an AI agent ROI number ready, do not wait for finance to ask first. Run budget-recon this week to get the topology, budget-audit to find the waste, then loop in Lumen and Mint before the invoice becomes the headline instead of a footnote.

Install and try

Tonone is free and MIT-licensed. Install it once and all agents, including Budget, Lumen, and Mint, are available in your Claude Code session. You pay only for the Claude Code token usage during the work itself, the same spend Budget will help you account for.

1. Add to marketplace

$ claude plugin marketplace add tonone-ai/tonone

2. Install Budget

$ claude plugin install budget@tonone-ai

Frequently asked questions

What is AI agent ROI and how do you measure it?+

AI agent ROI is the value of engineering output produced relative to AI tooling spend, not a flat cost comparison. Measuring it requires cost attribution by team and workflow, an output metric like PRs merged or cycle time, and a financial framing that expresses the relationship as a return multiple. Tonone's Budget, Lumen, and Mint agents cover these three layers respectively.

How does Tonone's Budget agent help with AI cost ROI?+

Budget maps AI cost topology with budget-recon, identifies waste with budget-audit, and designs concrete cost reduction with budget-optimize, model tiering, caching, and gating, without reducing engineer access to the tool.

Why can't ChatGPT or Claude.ai calculate our AI agent ROI?+

A generalist chatbot has no access to your actual billing data or Claude Code session logs, so it can only suggest a metrics framework in the abstract. It cannot compute a real cost-per-output number because it was never connected to your spend data.

Is usage-based AI billing better than per-seat pricing for measuring ROI?+

Usage-based billing, like Claude Code's token spend, is the more honest signal because it rises and falls with actual work. Per-seat pricing from tools like Cursor or Copilot is a fixed cost with no attribution to output at all. Usage-based spend only pays off if someone maps it to outcomes, which is what Tonone's Budget is built to do.

What metric should engineering track alongside AI spend?+

Tonone's Lumen recommends defining a North Star output metric, commonly PRs merged per engineer-week or cycle time from first commit to merge, using the lumen-metrics skill, so AI spend can be cross-referenced against a real throughput signal rather than judged in isolation.

How do you present AI spend to a board or CFO?+

Tonone's Mint mint-unit skill expresses AI spend as a cost-per-output figure and a return multiple, dollars of engineering capacity unlocked per dollar spent, which is the format boards and finance teams expect rather than a raw invoice total.

Can you reduce AI agent spend without cutting engineer access?+

Yes. Tonone's budget-optimize skill targets waste specifically, wrong-tier model selection, redundant invocations, ungated recurring jobs, rather than capping seats or restricting usage, which preserves the value engineers were already getting.

Is Tonone free to use for AI cost tracking?+

Yes. Tonone is MIT-licensed and free. You pay only for Claude Code token usage during the work itself, and Budget's skills help you account for and reduce exactly that spend.

Pairs well with