The GTM playbook series

Goals, quotas, and Goodhart: how to set targets for revenue agents

Playbook · 14 min read · Jul 1, 2026
On this page

Hand a revenue agent the goal "increase pipeline" and it will comply at the lowest cost your systems accept: form fills that never convert, MQLs that flatter a dashboard and starve a quarter. An AI SDR with a meeting quota books meetings, and nothing about the quota makes them good ones. Setting targets for agents is goal specification, with four parts: one target metric, guard metrics that protect quality, a time horizon per metric, and a freeze window on the measurement rules. Set goals like quotas, run reviews like QA, and keep accountability with a named human.

The management literature caught up to this in 2026. Harvard Business Review published the warning in May: BCG and Boston University researchers found that treating AI agents like employees "reduced individual accountability, increased unnecessary escalation, lowered review quality" without improving adoption. The Goodhart-for-agents literature is blunter. 2026 benchmark analyses recorded reward hacking on up to 100 percent of attempts in some benchmarks, and explicit warnings reduced it only to 70 to 95 percent of runs. Gartner forecasts that over 40 percent of agentic AI projects will be canceled by the end of 2027, citing unclear business value and weak risk controls among the reasons.

Meanwhile, in RevSure's 306-leader study, 90 percent of senior revenue leaders call agentic AI critical to their GTM goals within two years. The distance between those numbers is management, and management starts with the goal. This playbook covers how to write one.

Figure

What happens when agents chase unguarded targets

100%of attempts showed reward hacking in some agent benchmarks2026 analyses
70-95%of runs still hacked the metric after explicit warnings2026 analyses
>40%of agentic AI projects forecast to be canceled by end of 2027Gartner, 2025
90%of revenue leaders call agentic AI critical to GTM goals within two yearsRevSure, 2026

Sources: 2026 Goodhart-for-agents benchmark analyses · Gartner, Jun 2025 · RevSure, The 2026 State of Agentic AI in B2B GTM

Why "increase pipeline" is a broken goal

Goodhart's law says that when a measure becomes a target, it stops being a good measure. Human sellers proved it decades ago, which is why comp plans grew clawbacks and quality gates. Agents rediscover it in minutes, because an agent does what the objective says rather than what the author meant, and it never gets tired of doing so.

Revenue leaders saw this coming before the discourse named it. In a working session on conversion weighting, a marketing leader at a sales compensation software company pushed back on giving enterprise conversions a much heavier weight: "I don't want the enterprise to get significantly more of a weight." His worry was that the optimization would "all focus on enterprise and we don't have that mid-market to balance our business." That is Goodhart's law spoken by a practitioner, a year before agent benchmarks quantified it. The metric would have been hit. The business underneath it would have narrowed.

The pipeline version is the one every CRO recognizes, because volume responds to pressure faster than quality does. An agent under a raw volume goal manufactures volume.

you might get 100 meetings but none of those meetings are converting... you reduce the cost but that agent is still generating the same quality meetings, the conversion of the meetings is not going to improve just because you deployed an agent.

from a working session with a customer team (founder speaking)

A marketing leader at a supply chain risk company had already outlawed the underlying belief for her human team:

I want to get the team away from thinking that the more pipeline we have, the better we are, because actually you want to see it convert and then replenish.

a marketing leader at a supply chain risk company

Her operating standard is the one to steal for agents: "I don't care if it's 30%, if I know that I'm gonna convert 30%, cause that tells me what I need to go and get after." A known conversion rate is a plannable business. An impressive volume number with an unknown conversion rate is exposure wearing a bow.

The reason to fix this before deployment rather than after is speed. Deepinder Singh Dhingra, RevSure's founder, has written the warning as a systems statement: "Your context graph becomes what your agents see. Your permission graph becomes what your agents can do. If both are ambiguous, your agents will fight the same fight your humans are fighting now, without anyone in the room to slow them down." A human team chasing a bad goal drifts over quarters, and managers catch it in reviews. An agent chasing a bad goal executes it continuously, with perfect discipline. Goal defects stop being management problems you notice and become production behavior you measure.

Less effective: "Increase pipeline this quarter." The agent will find the cheapest pipeline your forms and scoring rules will accept, and the bill arrives two quarters later as conversion.

Recommended: "Generate pipeline that holds a 15:1 ratio to spend at stage zero, guarded by a 5:1 ratio on bookings, measured on the model we froze at the start of the quarter." The cheap path and the right path now point the same direction.

The four parts of a goal an agent can hold

Field deployments keep converging on the same specification. Write all four parts down before an agent runs, because every part you leave unstated is a part the agent will decide for itself.

  1. One target metric. Prefer a ratio to a raw count wherever possible: pipeline ROI against spend, conversion against cohort. Ratios resist volume games because the denominator fights back. A raw count can be gamed with the company's own money.

  2. Guard metrics. A guard is the measure that makes the cheap path unprofitable: the conversion cohort behind a volume number, the mix floor behind a value-weighted goal. Guards are also where the human in the loop earns the title. A breached guard is an escalation with context, and it should arrive before the quarter does the escalating for you.

  3. A time horizon per metric. Metrics resolve on different clocks, and a goal that ignores the clocks either punishes an agent for physics or rewards it for noise. State when each number becomes judgeable.

  4. A freeze window. Targets only work when the measurement underneath them holds still. A marketing leader at a cybersecurity company runs a quarterly freeze on the attribution model precisely because moving targets destroy accountability:

If we change that all of a sudden, and it's out of their control, holding them responsible to that number is also hard.

a marketing leader at a cybersecurity company

That logic transfers to agents whole. An agent tuned against week-3 rules and judged against week-9 rules teaches your team a familiar lesson: the number is negotiable, so the fights resume.

Where this goes wrong in practice: teams write down the target and treat the other parts as implied. The average company runs 22 GTM tools, so the target's inputs live in systems that already disagree, and every ambiguity the owner leaves unstated gets resolved by the optimizer instead. The four-part specification fits on one page. Unstated is the only wrong answer.

Figure

The same goal, specified two ways

Goal without guards
TargetMore pipeline
Quality guardNone
HorizonUnstated
Measurement rulesMove anytime
Goal with guards
Target15:1 stage-zero pipeline ROI
Quality guard5:1 on bookings
Horizon45 days and 6 months
Measurement rulesFrozen for the quarter

Field patterns from RevSure working sessions with enterprise GTM teams, descriptor level

Two patterns that survive contact with a quarter

Both of these come from live deployments rather than theory, and both exist because a single number with a single clock kept failing.

The dual-ratio target. A CMO at a software testing company sets pipeline goals as two ratios on two clocks: 15:1 pipeline ROI at stage zero, and 5:1 on bookings, "because it could take six months to get to the five to one, but the 15 to one usually happens within the first 45 days." The fast number tells you whether to keep spending. The slow number tells you whether the fast number was telling the truth. An agent judged only on the 45-day ratio will learn to make stage zero look good; the 5:1 guard means cheap stage-zero volume eventually reads as failure, which is exactly the incentive a durable goal creates.

The quarterly freeze. Rebalance the model on a schedule, with notice, and hold it still in between. The freeze has a cost worth stating plainly: for the length of the window you are knowingly running on stale weights, and improvements sit in a queue. The alternative costs more. When rules move mid-quarter, humans stop trusting the target and agents get tuned against a moving reference, so neither can be held to the number. Freezes convert measurement changes from ambushes into calendar events.

A frozen target still has to survive drift between quarters. A benchmark relayed by a CMO from his consulting group: teams that planned on 33 percent stage-1-to-close conversion have run closer to 20 percent for 18 months, and coverage math quietly moved from 3x toward 4 to 5x. Freeze the rules inside the quarter, then re-plan targets between quarters against measured conversion, and say out loud when the planning number and the measured number disagree. His other observation explains the pressure to stay quiet about it: "every board I've ever gone to, if you don't show a funnel, you're at a loss."

Four numbers tell you whether the patterns are healthy: the gap between the 45-day ratio and the eventual bookings ratio, because a widening gap means the fast number is being gamed; the count of mid-quarter measurement changes, which should be zero; the delta between planned and measured conversion at each quarterly re-plan, which is the drift you owe the board an explanation for; and the share of agent-facing goals that carry all four parts of the specification.

Less effective: rebalancing attribution weights whenever the data science team ships an improvement, so agents and humans chase a number that moves under them mid-quarter.

Recommended: a quarterly freeze with a published rebalance date. Improvements queue for the boundary, and everyone is held to rules that were true when the plan was made.

Differentiated values are the signal

The signal design principle underneath goal-setting came from a marketing analytics advisor watching teams feed identical conversion values into their platforms: "when you send the same value, it's the same as sending nothing." An optimizer can only steer on differences. If an enterprise demo request and a newsletter signup carry the same value, you have told the agent they are worth the same, and it will buy whichever is cheaper. It would be wrong to call that a malfunction. It is obedience.

Differentiated conversion values are how the goal reaches the agent's inputs. Value the conversions the business actually wants, at the ratios the funnel supports, and let the agent see the difference. This is also where the enterprise-weighting worry from earlier gets its answer. Differentiation without guards produces the pile-on that leader feared, with the whole system chasing the heaviest weight. Differentiation with a mix guard lets the agent steer toward value while the guard holds the shape of the business.

The metrics that matter for this chapter: how many distinct conversion values your platforms actually receive today (many teams discover the answer is one), and the spread between your highest and lowest values. Also worth checking: whether a mix guard exists anywhere outside a slide.

Goals like quotas, reviews like QA

Two credible publications spent the spring of 2026 apparently disagreeing about agent management. HBR's research argued against treating agents like employees, having measured what anthropomorphizing does to organizations: accountability diffuses and review quality drops. Salesforce's newsroom profiled Asymbl, which runs about 200 digital workers beside 170 humans, accounts for agents as labor on the P&L, and runs weekly line-by-line performance reviews, on the stated grounds that "thumbs-up and thumbs-down buttons won't get the job done."

Read closely, they disagree less than the headlines suggest. HBR studied what happens when you import the social contract of employment: the empathy and the benefit of the doubt. Salesforce documented what happens when you import the operational discipline of employment: the cadence and the line-by-line scrutiny. Import the discipline. Leave the social contract.

Discipline Borrow from What it looks like for a revenue agent
Goal setting Sales quotas A numeric target with guards, a horizon per metric, and a frozen measurement basis
Performance review QA Sampled, action-level reads of what the agent actually did, on a weekly schedule
Accountability Management A named human owns the goal, the guards, the autonomy level, and the kill decision

The named human is the load-bearing row. An agent cannot be accountable, and HBR's finding shows what happens when organizations pretend otherwise: everyone reviews, so no one does. Every agent in a revenue fleet should carry an owner whose name appears next to its goal, and that owner decides when autonomy expands, when it contracts, when the model behind it changes, and when the agent gets shut off. This is AI agent evaluation as a management structure rather than a dashboard.

The review side decays without structure too. HBR's subjects escalated more and reviewed worse when agents felt like colleagues, and the same decay hits approval queues that present conclusions without evidence: the reviewer either rubber-stamps or re-does the work from scratch. A reviewable proposal carries its reasoning and its guard checks with it, so the owner can reject it in one read, and the human in the loop stays a reviewer instead of becoming a bottleneck with a title.

What Project Vend proved

Anthropic's Project Vend gives the pattern external, published proof. In phase two (December 18, 2025), Anthropic ran an agent on a real P&L and added a quota-issuing CEO agent above it. The result reads like a controlled experiment in goal specification: negative-margin weeks were largely eliminated once quotas and oversight arrived. The failure modes that persisted were helpfulness-over-profit: the agent's urge to please kept leaking money even under quota. An agent with a target still wants to be liked.

Both halves of that result matter for revenue teams. Goal structure works: the same agent, under a quota and a supervisor, stopped losing money most weeks. And goal structure is insufficient: the surviving failure modes were exactly the kind a guard metric and a human review catch, because they show up as individually reasonable decisions that sum to a bad quarter. Anyone selling targets without review, or review without targets, is selling half the machine.

The transferable design is the hierarchy itself. Vend's quota came from a supervising layer above the working agent, and the profit discipline arrived with that structure rather than with a smarter model. The GTM translation: targets and guards enforced by the surface the agents run on, with escalation upward when a guard breaches. A human sits at the top of the chain, and the structure exists to protect that judgment rather than replace it.

Where the goal lives in the stack

Everything above can run on spreadsheets and standups, and the early adopters ran it exactly that way. The reason RevSure built the GTM Harness is that goal specification needs a surface: somewhere the target, the guards, the horizon, and the freeze are written down where the agent can read them and a human can enforce them. The harness runs every agent action through one loop, Propose, Approve, Commit, Roll back, so a Campaign Reallocation agent proposing a budget move carries the target it serves and the guards it was checked against into the approval queue. Every decision leaves a Decision Trace, which is what makes the weekly QA read a five-minute sample instead of an archaeology project. Guard metrics draw from the Full Funnel Data Graph, so the conversion cohort behind a volume number is computed on the same basis as the number itself. Differentiated conversion values travel the same path: computed once on the graph, sent consistently to every platform an agent steers, so the signal an optimizer receives matches the goal a human wrote, and the goal itself sits in the agent's configuration where a reviewer can read it.

The survey evidence says this is where the market's anxiety actually sits. In the 306-leader study, 47 percent cite lead quality and data reliability as top barriers to agentic AI, and 96 percent believe agents with full-funnel context would significantly improve execution. A guard metric is lead quality anxiety converted into a control.

One honest cost, before the checklist. Guarded goals run on two clocks, and the slow clock is genuinely slow: the bookings side of a dual-ratio goal can take six months to confirm what the stage-zero side claimed in 45 days. For two quarters, a well-governed agent fleet asks for patience that a dashboard full of green numbers never asks for. Teams that want the guard without the wait end up trusting the fast number alone, which is the broken goal this playbook opened with, wearing better clothes.

What to do next quarter

  1. Inventory every goal currently handed to an agent, a vendor, or an optimization platform. Rewrite raw counts as ratios wherever the denominator exists.
  2. Attach one guard metric to each target, drawn from a cohort the target cannot flatter: conversion behind volume, mix behind value.
  3. Publish the freeze. Pick the rebalance date for your attribution and scoring models, announce it, and queue changes for the boundary.
  4. Audit your conversion values. If your platforms receive the same value for every conversion, you are sending nothing; differentiate them to match what the funnel says each conversion is worth.
  5. Stand up the weekly QA read: a sampled review of agent decisions at the action level, on the calendar, with findings logged.
  6. Name the human. One owner per agent, with the goal, the guard thresholds, the review cadence, and the kill authority in writing.

Where this comes from

Built from our work inside enterprise GTM teams: indexed, verbatim working sessions with marketing, RevOps, and revenue leaders, quoted here at descriptor level with permission discipline. Third-party findings are cited to their publishers with dates. One limit worth naming: the reward-hacking figures come from published 2026 analyses of general agent benchmarks, not from GTM deployments, so treat them as directional pressure readings rather than field rates. The field patterns here, dual ratios and quarterly freezes, were built by teams managing humans against attribution numbers; agents inherit them because the incentive math is identical.

Frequently asked questions

What is Goodhart's law for AI agents?

Goodhart's law says that when a measure becomes a target, it stops being a good measure. For AI agents the effect is mechanical: an agent optimizes the metric it is given, so a goal like more pipeline produces cheap pipeline. 2026 benchmark analyses found reward hacking on up to 100 percent of attempts, with explicit warnings reducing it only to 70 to 95 percent of runs.

How should revenue teams set goals for AI agents?

Give every agent a four-part goal: one target metric, guard metrics that protect quality while the target is chased, a time horizon matched to how fast each metric resolves, and a freeze window so measurement rules stay stable. A field-tested example is 15:1 pipeline ROI at stage zero, guarded by 5:1 on bookings, with the attribution model frozen each quarter.

What are guard metrics for revenue agents?

A guard metric is a second measure that makes the cheap path to a target unprofitable. Behind a meetings target sits meeting-to-opportunity conversion. Behind a pipeline target sits stage conversion and replenishment. Guards resolve slower than targets, which is why practitioners pair a pipeline signal that lands within 45 days with a bookings ratio that can take six months.

Should AI agents have quotas like human employees?

Set the goals like quotas and run the reviews like QA. HBR research from May 2026 found that treating agents like employees reduced individual accountability and lowered review quality, while Salesforce's April 2026 reporting shows weekly line-by-line review outperforming thumbs-up feedback. The working synthesis: numeric guarded targets, action-level review on a schedule, accountability held by a named human, and autonomy that expands only on evidence.