Agents on the P&L: the GTM finance playbook for underwriting AI spend
On this page
Agent spend confuses budget review because it behaves like three familiar categories at once: a fixed platform fee that acts like software, managed deployment work that acts like labor, and usage credits that scale like media. This playbook is the underwriting manual for GTM finance: how to classify the spend, what evidence should release budget, how decision traces make the spend auditable, how to compute CAC payback period with credit costs included, and the oversight cadence that keeps an agent fleet honest.
Finance has two public poles to steer between. Salesforce's April 2026 profile of Asymbl describes about 200 digital workers running beside 170 humans, accounted for as labor on the P&L and reviewed line by line every week, on the stated logic that "thumbs-up and thumbs-down buttons won't get the job done." Five weeks later, Deloitte's DART team published accounting guidance for outcome-priced agentic software. Same spend category, filed in different worlds. ICONIQ's 2026 study of more than 150 B2B software GTM organizations supplies the yardstick that explains why finance cannot sit this one out: high-AI-adoption companies generate about $640K of net-new ARR per GTM FTE against $370K at low adopters.
The numbers GTM finance is underwriting against
Sources: ICONIQ, Leaner, Smarter, Flatter, 2026 · Gartner CMO Spend Survey, May 2025 · Gartner agentic AI forecast, June 2025
The stakes on the other side are just as documented. Gartner forecasts over 40 percent of agentic AI projects canceled by end of 2027, citing cost, unclear value, and weak risk controls, which is a list of underwriting failures. Marketing budgets are flat at 7.7 percent of company revenue per Gartner's May 2025 spend survey, with 39 percent of CMOs cutting labor and agency costs, so every agent dollar displaces a dollar somebody already defends. RevSure's stated thesis puts the backdrop bluntly: GTM efficiency dipped more than 40 percent post-ZIRP. The CFO has become the quiet approver of every GTM platform decision. This playbook is written for that approver.
Classify the spend before you approve it
The first finance question is what kind of money this is, because the classification decides the governance. Spend that behaves like software gets a shelfware review once a year. Spend that behaves like labor gets performance management. Spend that behaves like media gets efficiency scrutiny every month. Agent contracts contain all three behaviors, and the errors come from picking one lens for the whole invoice, which is easy to do because the invoice arrives as one number with a SaaS vendor's logo on it.
RevSure's disclosed pricing makes a useful worked example precisely because it is public. The published package for 200,000 contacts and 10 managed agents runs $74,880 a year all-in, and it decomposes into the three behaviors cleanly. The platform license is $32,000 at that contact tier, fixed for the term, scaling in published steps with database size: software. Managed deployment is $20,000 for the package of 10 agents (or $2,500 per agentic workflow), which buys configuration, testing, and ongoing management by RevSure's GTM engineering: labor, and priced well under the human alternative, with GTM engineer salaries running $99K to $310K (Revnu, 2026) and an in-house build taking 2 to 3 engineers 3 to 4 months per RevSure's deck. The remaining $22,880 is net usage credits at $0.02 per credit, consumed per action and scaling with volume: media. Add-ons follow the same media logic with published tiers of their own; visitor deanonymization, for example, runs $24K a year for 2,000 identified visitors a month up to $70K for 20,000, with additional volume at $500 per 1,000 visitors.
Three cost behaviors on one agent invoice
Source: RevSure published worked example: 200,000 contacts, 10 managed agents, $74,880 all-in per year, including 80,000 testing credits.
Treat each slice by its behavior. The software slice gets a usage review: are the agents live, or is this the shelf. The labor slice gets an output review: what did managed deployment ship this quarter. The media slice gets unit economics: cost per action against value per action, monthly, the way a paid channel would be read.
The two public poles both hold a piece of the answer. Asymbl's labor framing buys the right instinct, which is recurring performance review of individual agents. Deloitte's software framing buys the right rigor on revenue and cost recognition for outcome-priced contracts. The caution flag on over-literal labor framing comes from HBR: a May 2026 study by BCG and Boston University researchers found that treating AI agents like employees "reduced individual accountability, increased unnecessary escalation, lowered review quality" without improving adoption. Classify by cost behavior. Manage by evidence.
Less effective: booking the entire agent contract to the software line because the vendor is a SaaS company, then discovering at renewal that credit consumption doubled and nobody owned the variance.
Recommended: three budget lines from day one (platform, managed work, usage), each with an owner, a review cadence, and a threshold that triggers a conversation.
Underwriting: the evidence that should release budget
Underwriting means releasing money against evidence rather than against a demo, and demand is strong enough that the demo often wins by default. In RevSure's 2026 study of 306 senior marketing, sales, and RevOps leaders, 76 percent are already deploying or implementing agentic AI, while 47 percent name lead quality and data reliability as their top barriers: the same organizations racing to deploy are the ones reporting that the inputs are shaky. The discipline gap is measurable too. LangChain's June 2026 survey of 1,340 agent builders found 89 percent of teams run observability while only 52.4 percent run offline evals, and quality is the number one blocker to production. Trust, meanwhile, is running ahead of verification: Gong's December 2025 study of 3,048 revenue leaders found 70 percent already trust AI for regular business decisions. Finance's job is to be the institution that still asks for the file. Deepinder Singh Dhingra's read on enterprise buying applies inside the building as well: "We think buyers want speed. But what they're really paying for is certainty."
Three artifacts should gate an agent budget, in tranches.
First, a backtest. Before an agent touches live pipeline, its decision logic should run against your own history: last year's leads scored, last year's budget reallocations proposed, then compared with what actually happened. A vendor that resists backtesting on your data is asking you to underwrite on their data.
Second, pilot guard metrics, defined before the pilot and tracked past it. The metric that releases money must sit downstream of the activity, because activity is what agents inflate first. The founder's warning from a working session with a customer team draws the line exactly:
you might get 100 meetings but none of those meetings are converting... you reduce the cost but that agent is still generating the same quality meetings, the conversion of the meetings is not going to improve just because you deployed an agent.
from a working session with a customer team
Meetings are an activity metric. Meeting-to-opportunity conversion is a guard metric. Underwrite on the second.
Third, eval discipline as a standing condition. The agent's owner should show a versioned evaluation suite that reruns on every model change, with regressions blocking release. This is the same logic as requiring audited statements from a borrower: the point is a repeatable process, and a process that only 52.4 percent of the industry runs is a genuine screen. Vendor-side process discipline has a checkable marker as well. ISO/IEC 42001:2023 is the first international standard for AI management systems; certification attests that a management process exists and is audited, and it does not by itself prove outcomes, so treat it as a floor in the memo rather than a substitute for the three artifacts above.
Less effective: releasing the full annual budget at signature because the pilot demo impressed the sponsor, then reviewing outcomes at renewal, eleven months after the evidence went stale.
Recommended: tranche release. A capped pilot budget opens on the backtest. The production budget opens when guard metrics clear their floor for two consecutive months. The expansion budget opens on a portfolio review with finance in the room.
The audit trail: retiring "grades its own homework"
There is a reason marketing budgets get underwritten more skeptically than sales headcount, and a CFO once said it to RevSure's founder directly:
marketing is the only department that grades its own homework.
a CFO, from the founder's published field notes
He could not laugh at it, because the operating reality behind the joke is real. Marketing reports its own attribution, from systems marketing configured, on models marketing chose. The historical answer was to assemble evidence by hand: in one software testing company's record, a weekly dashboard update consumed about 90 minutes every Monday evening and board reporting once took roughly 100 person-hours a quarter. The current answer at many companies is worse. "We do everything in spreadsheets. Run the whole business outta a spreadsheet," a revenue leader at a cybersecurity company told us, which is why Deepinder Singh Dhingra's line lands: "Our biggest competitor is Excel plus gut feeling."
Agentic systems change the auditability question structurally, because agents generate their own paper trail as a byproduct of acting. Every decision RevSure's agents take records a Decision Trace: what the agent saw, the evidence it weighed, the action it took, what the action cost. Decision Attribution is computed on the Full Funnel Data Graph, so a credited outcome decomposes to records rather than to a model's say-so. For finance the practical consequence is sampling. An auditor does not read every transaction; an auditor samples against the ledger. Decision traces give GTM its ledger, and finance can pull twenty traces a quarter and check that the evidence supports the action, at a marginal cost of an afternoon. Headcount productivity was always argued from averages and anecdotes. Agent spend can be audited action by action, which makes it, on this one dimension, the most governable money in GTM. Marketing attribution reporting hours reclaimed at three companies in our field record: roughly 100, 50, and 40 hours per quarter. The homework is still marketing's. The grading no longer has to be.
Efficiency accounting: yardsticks, CAC payback, and the 15x question
Three calculations belong in the finance pack once agents join the cost base.
ARR per GTM FTE. ICONIQ's benchmark gives the spread: about $640K of net-new ARR per GTM FTE at high-AI-adoption software companies against $370K at low adopters, with top performers running 20 to 30 percent leaner and one AI CSM covering the workload of roughly 20 human CSMs. Keep the denominator human. Adding agents to the FTE count creates an unauditable conversion rate between agents and people; the cleaner pattern is a human denominator with agent costs carried in the numerator as a cost line, plus a companion ratio of net-new ARR per fully loaded GTM dollar. The paired view shows whether agents are raising output per person or merely re-labeling cost.
CAC payback period with credit costs included. The classic CAC payback period divides acquisition cost per customer by monthly gross profit per customer. The discipline agents demand is on the numerator: platform fees, managed deployment, and usage credits all belong in acquisition cost, allocated the way agency fees and media already are. The worked package above meters out at the equivalent of $6,240 a month all-in, and its published run volumes make the allocation concrete: 4,000 personalized emails a month is $1,600, 6,000 lead scores is $720, 500 enrichments is $250. Leaving credits out of CAC because they feel like software subscriptions understates acquisition cost and flatters the payback trend exactly when the board is watching it most closely. There is an upside to the recomposition worth stating in the same breath: metered acquisition cost is correctable at monthly speed. A team that overshoots on credits can turn the volume down next month; a team that overshoots on headcount corrects over quarters. Finance should expect agent-heavy CAC to be more volatile month to month and more governable over the year, and read the payback trend accordingly.
What one agent action costs
Source: RevSure published pricing model, 2026. Credit rate $0.02; a personalized email also consumes about 31,000 LLM tokens.
Unit costs also settle a common anxiety honestly: the expensive part of agentic GTM is context, and the model is comparatively cheap. At the lightest published model rate ($2.50 per million tokens), the roughly 31,000 tokens inside a personalized email cost about 8 cents, while the email is priced at about 40 cents. The spread pays for the work around the model: identity resolution, enrichment, orchestration, and the trace.
Then the 15x question. RevSure's own deck presents a 15x year-1 ROI panel, and the honest instruction to finance is to treat it the way you should treat any vendor's multiple, ours included: as a claim with levers. The published levers are a 20 percent lift in MQL-to-opportunity conversion, a 10 percent lift in win rate, a 40 percent lift in booking value, a 30 percent reduction in cost per deal, and 50 percent productivity gains. A multiple is an output. Underwrite the levers one at a time: demand the baseline each lift is measured against, the guard metric that would confirm it, and the trace evidence behind the claimed lift. Any lever that cannot name its baseline is a hope, and hopes do not clear underwriting, at 15x or at 2x.
The oversight cadence: revenue operations runs it, finance signs it
Underwriting without oversight is a loan with no covenants. The public reference point for intensity is Asymbl's weekly line-by-line review of every digital worker; the HBR finding cautions against dressing that review up as employee ritual. The workable middle for a GTM fleet is an operational review run weekly by RevOps inside the control surface (in RevSure's case, the GTM Harness and its Propose, Approve, Commit, Roll back loop) and a quarterly agent portfolio review with finance in the room. Anthropic's Project Vend research makes the case that oversight structures genuinely work and genuinely have limits: an agent run on a real P&L under a quota-issuing supervisor agent largely eliminated negative-margin weeks, while helpfulness-over-profit failure modes persisted. Supervision improves agents. It does not finish the job, which is why the kill criteria get written at underwriting time, while everyone is still calm.
| Review | Cadence | Who runs it | Who signs |
|---|---|---|---|
| Run-level operations: volumes, overrides, cost per action | Weekly | RevOps | RevOps leader |
| Agent portfolio review: guard metrics vs underwriting memo, kill list | Quarterly | RevOps with the owning executive | CFO or GTM finance |
| Credit and usage true-up against budget lines | Quarterly | Finance | CFO |
| Kill or retire decision on a failing agent | Event-driven | Owning executive | CRO or CMO, finance countersigns |
The signature question deserves one paragraph of its own, because ambiguity here is what turns agent data political. The owning executive (CMO or CRO) answers for outcomes. Revenue operations operates the fleet day to day: configurations, approvals, override handling, the run-level review. Finance underwrites and countersigns kills. Nobody else changes what an agent may do. The founder's warning about the human version of this fight applies verbatim to the agent version: "No company dies from lack of data. They die when data becomes political instead of analytical."
Kill criteria worth writing down: a guard-metric floor (downstream conversion below the underwriting threshold for two consecutive months), an override ceiling (humans rejecting a majority of the agent's proposals after the calibration period), and a unit-cost ceiling (cost per accepted action above the human alternative). A kill also needs a what-happens-next clause: configurations documented, traces archived, credits reassigned to the survivors.
One concession belongs in the memo, because the field record demands it. Deployments in this category can stall in configuration, and the stall consumes contract months:
we are almost more than halfway through our contract for the year and we're still in setup mode.
a marketing leader at a healthcare benefits company
RevSure's stated onboarding is live in under 4 weeks, and the quote above is the failure mode every buyer, ours included, should underwrite against. Put time-to-first-committed-action into the underwriting memo as its own covenant, with a named owner on both sides. A platform that has not committed an action is a software line with no labor and no media, and the review above will surface it in one quarter instead of at renewal.
What to do next quarter
- Reclassify current agent spend into the three lines (platform, managed work, usage) and name an owner for each. This is an afternoon with the invoices.
- Write the one-page underwriting memo template: the backtest required, the guard metrics with floors, the eval condition, the tranche schedule, the kill criteria, the signatures.
- Recompute CAC payback period with agent platform fees, managed deployment, and credits in the numerator, and publish both versions once so the discontinuity is explained rather than discovered.
- Add net-new ARR per GTM FTE and per fully loaded GTM dollar to the board pack, benchmarked against ICONIQ's $640K and $370K poles.
- Sample 20 decision traces from one live agent and have finance, not marketing, walk the evidence. The exercise costs an afternoon and settles the homework question for good.
- Book the first quarterly agent portfolio review now, with the CFO's signature block already on the template.
Where this comes from
Built from our work inside enterprise GTM teams: indexed, verbatim working sessions with marketing, RevOps, and revenue leaders, quoted here at descriptor level with permission discipline. Pricing figures are RevSure's published 2026 pricing model, and the worked example is one vendor's economics, ours, offered because it is the disclosed one. This playbook is operating guidance for budget governance; formal accounting treatment of agent contracts belongs to your auditors, and Deloitte's DART guidance is where that conversation starts.
Frequently asked questions
What is CAC payback period?
CAC payback period is the number of months it takes for the gross profit from a new customer to repay the cost of acquiring that customer. When AI agents join the GTM motion, the calculation should include agent platform fees, usage credits, and managed deployment costs alongside people and media, or the metric flatters the new cost base.
Should AI agents be accounted for as software or labor?
Both patterns exist in public. Salesforce profiled Asymbl accounting for about 200 digital workers as labor with weekly performance reviews, while Deloitte's DART guidance treats outcome-priced agentic products as software accounting questions. In practice agent contracts mix a fixed platform fee, managed work, and volume-priced credits, so finance should classify each component by its cost behavior.
What evidence should release budget for AI agents?
Three artifacts: a backtest of the agent's decisions against the company's own historical data, pilot guard metrics that track downstream conversion rather than activity volume, and a versioned evaluation suite that reruns on every model change. Budget releases in tranches as each artifact clears its threshold, the way credit releases against collateral.
How do decision traces make AI agent spend auditable?
A decision trace records what an agent saw, the evidence it weighed, the action it took, and what the action cost. Finance can sample traces the way auditors sample transactions, inspecting GTM claims without relying on marketing's own reporting. RevSure records a Decision Trace for every agent decision on the Full Funnel Data Graph.