← Back to blog

Guides

How to write the business case for AI support your CFO will sign

Priya Nair, Co-founder · August 11, 2026 · 8 min read

flowtux|Blog · Guides

Most AI support proposals die because they lead with the tool. The ones that get signed lead with a measured baseline and a named rollback. Here is the one-page structure.

flowtux.com/blogGuides

Most internal proposals for AI support fail in the same way. They open with a product, quote a vendor deflection figure, project a saving from it, and land on a finance desk that has seen four of these this quarter. The number is unfalsifiable, so it is discounted to zero, and the request is deferred pending more information that nobody will gather.

The version that gets signed is duller and shorter. It leads with a baseline you measured, prices the work honestly, expresses the return in hours rather than heads, and names the conditions under which you would turn the thing off. One page. Here is how to build it.

Start with a baseline you measured

Pull 90 days of tickets and break them down by category, by volume, by median time to resolution, by reopen rate, and by who actually resolved them. The last dimension is the one people skip and the one that carries the argument, because in most companies a meaningful share of tickets are resolved by people who are not in the support budget at all — engineers pulled out of focused work to answer a question they have answered before.

Do this before you talk to a vendor, for two reasons. It stops your baseline from being shaped by whatever the vendor measures well, and it frequently changes what you buy. Teams that run this exercise honestly often discover the top category by volume is not the one they were planning to automate.

Price the work defensibly

Compute cost per ticket the fully-loaded way: compensation with benefits, the pro-rated time of partial responders, tooling, and a defensible slice of overhead, divided by tickets resolved net of reopens over a quarter. Segment it by category and handling path, because a blended average cannot support a decision — it hides exactly the routine, expensive band that makes the case.

A finance reader will test two things: whether you used resolved rather than received, and whether you picked a flattering period. Address both in the document before they ask. If your quarter contained an unusual incident, say so and show the number with and without it. Volunteering the inconvenient version is what makes the rest of the page credible.

Baseline

90 days by category, resolution time, reopens, and who really resolved

Unit cost

fully-loaded cost per ticket, segmented, net of reopens

Hours returned

the return framed in labour hours, not headcount

Rollback

the named trigger and the person who can pull it

Four blocks. A case missing any of them is the kind finance defers rather than rejects.

Frame the return as hours, not heads

The instinct is to convert saved tickets into saved salaries, because that is the number that looks biggest. Resist it for three reasons, all practical. It is the least defensible claim in the document, because headcount reduction depends on organisational decisions you do not control. It hardens into a commitment — once finance models a role removed, you own that outcome whether or not the technology delivers it. And it destroys adoption, because the people whose cooperation you need to make the rollout work now understand the project as aimed at them.

Labour-hours returned is the better frame and is verifiable. If a category consumes a known number of responder hours a month and automation removes most of the handling, the returned hours are measurable in the same system that produced the baseline. Then say what those hours go to — a backlog, faster response on the tickets that do need humans, engineering time back on the roadmap. Finance can decide for itself whether returned hours eventually become a hiring decision; that is their job, not a claim in your document.

Name the risk and the rollback

Every AI proposal has the same risks, and a document that omits them reads as naive rather than confident. The agent answers wrongly with confidence. It closes tickets requesters reopen, which converts a saving into a loss and annoys people twice. Data goes somewhere it should not. Quality degrades quietly after a model or prompt change. The team learns to route around it.

Write each of those down with the control beside it. Suggest mode before autonomous mode. Allow-lists per category rather than a global switch. Reopen rate as the primary quality metric, with a threshold that triggers review. An audit trail that lets you replay any decision. A named person who can disable autonomous action without a meeting, and a stated trigger — for example, reopen rate on automated categories exceeding the human baseline for two consecutive weeks. A named rollback is what turns an irreversible-looking bet into a reversible experiment, and reversible experiments are what get approved.

What not to promise

Do not promise a deflection percentage before the pilot. You do not know it, the vendor number was measured on someone else queue, and stating it converts your credibility into a hostage. Do not promise savings in month one — implementation, connector setup, and knowledge cleanup all land before any return does, and a case that ignores the J-curve looks wrong by week six.

Do not promise quality improvements you cannot measure. Do not promise the AI will handle the hard categories, because the hard categories are hard for reasons the model does not remove. And do not promise headcount reduction, for the three reasons above. A case with a modest, defended range beats one with an ambitious number that collapses under one question.

The one page itself

Structure it in six blocks. Baseline: volume, cost per ticket, and the top three categories, with method stated. Scope: which categories you are automating first and why those. Expected effect: a range, not a point, with the reasoning visible. Cost: subscription plus implementation hours plus ongoing review time, because the hidden cost of AI support is the human who reads what it did. Review date: when you will report actuals against this page. Rollback: the trigger and the owner.

Run your numbers through the ROI calculator to sanity-check the arithmetic and, more usefully, to see which inputs the result is most sensitive to — that tells you which assumption to defend hardest in the meeting. Then keep the page to one page. The proposals that get signed are the ones a CFO can check in four minutes, and every extra page reduces the share that gets read.

Frequently asked questions

How do you justify AI support spending to a CFO?

Lead with a measured 90-day baseline, a fully-loaded cost per ticket segmented by category, a return expressed as labour hours returned rather than headcount removed, the full cost including implementation and ongoing review time, a review date, and a named rollback trigger with an owner. Keep it to one page.

Why not promise headcount reduction?

Because it is the least defensible claim in the document, it hardens into a commitment you may not control, and it destroys adoption among the people whose cooperation the rollout needs. Labour-hours returned is measurable in the same system that produced your baseline, and finance can draw its own staffing conclusions.

What should the business case say about risk?

Name the specific failure modes — confident wrong answers, closes that get reopened, data exposure, silent quality drift after a model change — and put a control beside each: suggest mode first, per-category allow-lists, reopen rate as the primary quality metric with a review threshold, an audit trail, and a named person who can disable autonomous action without a meeting.

Ready to let Tux AI run your queue?

Flat pricing from $49/month. Every team, no per-agent fees.

Start free trial →