What agentic monetization means in due diligence
Due diligence on software companies has long treated pricing as a per-seat subscription question. Agentic monetization breaks that habit. A copilot assists a human inside an existing workflow and is priced per seat, much like classic SaaS. An agent performs tasks or entire workflows autonomously, so the unit of value stops being the seat and becomes the action, the output, or the outcome. Bessemer Venture Partners separates the market into copilots, agents, and AI-enabled services precisely because each carries a different charge metric: copilots are typically priced per seat or on consumption, while agents are increasingly tied to workflow or outcome pricing linked to tangible ROI.
| Business model | What it does | How it is priced |
|---|---|---|
| Copilot | Assists a human user inside an existing workflow | Per seat or consumption, like classic SaaS |
| Agent | Executes tasks or entire workflows autonomously | Workflow or outcome based, tied to tangible ROI |
| AI-enabled service | Blends automation with human oversight to deliver a service | Per output or outcome, often anchored to legacy service pricing |
Per-seat economics fail when one license triggers thousands of variable-cost actions. Coding copilots became the canonical illustration: heavy users consumed more inference compute than a flat seat fee covered, because every query carries a real marginal cost rather than the near-zero marginal cost of classic SaaS. When usage scales with engagement, the seat becomes a margin trap rather than a margin engine. This is why the shift belongs in revenue-quality diligence rather than in a technology annex: the charge metric now determines whether reported growth converts into durable, margin-positive revenue, a theme we develop further in our piece on SaaS monetization.
Zuora identifies four agentic pricing models, per agent, per activity, per output, and per outcome, and reports that pure per-outcome pricing represents less than 10% of the roughly 60 agentic AI services its analysts studied, because defining and tracking outcomes is materially harder than metering activity. That gap between outcome-based ambition and per-activity practice is exactly where an investor's attention should sit, and it is where the rest of this article operates.
The three-part test: defined, measurable, attributable
The framework this article develops is deliberately simple, and it is an investor's test set rather than any vendor's methodology. Outcome-based pricing is only credible where the outcome can be (1) defined in contract language, (2) measured by an instrument both parties accept, and (3) attributed sufficiently to the product rather than to the customer's own process, staff, or market conditions. Each gate is a diligence question. A vendor that fails any one of them is selling something other than a pure outcome, whatever its marketing says, and its revenue should be read as activity or consumption revenue instead.
Measurable versus subjective outcomes
The first practical split is between measurable and subjective outcomes. A resolved support ticket is countable, timestamped, and auditable. Customer satisfaction, improved productivity, or a better hiring decision is not, at least not without an instrument both parties trust. Where the underlying outcome is subjective, expect the vendor to substitute a stand-in metric, and expect the stand-in, not the outcome, to be what the invoice actually references. Diligence should always ask which one the contract prices.
Who operates the agent, and who picks the stand-in
Vendor documentation shows how elastic a stand-in can be. Zendesk's help center explains that on email a conversation ends 72 hours after the last email, and an automated resolution can be counted after that silence if no human agent intervened and an LLM verification step confirms the reply was relevant. Intercom counts an assumed resolution when a customer simply exits the conversation without requesting further help, and bills one outcome per conversation. In both cases the vendor defines the resolution event, verifies it with its own model, and sends the invoice against it. Whether the agent is customer-operated or vendor-operated matters for the same reason: the party controlling the measurement is never a neutral counterparty.
The third gate is attribution. Zuora's agentic pricing work grades outcomes on a spectrum from diffuse (an assistant whose contribution runs quietly in the background) through medium to strong (a sales rep whose closed deals are directly tracked and credited): where the AI contributes to an outcome influenced by many factors, per-agent or per-activity pricing fits, and pure per-outcome pricing fits only where attribution is direct and the measurement is unambiguous. Diffuse attribution should push vendors toward hybrid structures, and it should push a diligence team toward the same conclusion when it reads the revenue build.
Pricing metrics and structures to know before reading a contract
Before reading a single contract, an investor should know the market's standard meters. The table below collects publicly listed or documented pricing from vendors an investor is likely to meet in agentic AI data rooms. Prices change, so treat them as reference points for the metric, not as a valuation input.
| Vendor | Pricing metric | Published or documented price |
|---|---|---|
| Intercom Fin | Per outcome (resolution, procedure handoff, disqualification) | $0.99 per outcome; qualifications at $9.99 |
| Zendesk AI Agents | Per verified automated resolution | $1.50 per automated resolution on a committed pack, $2.00 pay-as-you-go, above a small per-agent allowance |
| Salesforce Agentforce | Per conversation | $2 per customer service interaction as originally launched |
| Salesforce Agentforce | Consumption credits | $0.10 per action; $500 per 100,000 Flex Credits |
| Oracle Fusion AI | Consumption (AI Units) | Roughly one US cent per AI Unit; premium LLM actions consume about 5 units, with 20,000 units per month included in Fusion subscriptions |
| Vantage | Shared savings | 5% of actual AWS storage cost savings delivered |
Hybrid subscription-plus-usage structures are becoming the default compromise. ServiceNow explicitly combines license and usage components, with management arguing the mix gives customers predictability while tying usage to the outcomes customers see. Oracle includes a monthly AI Unit allocation in every Fusion subscription and sells additional units in $1,000 increments pooled across the contract. Salesforce itself moved from its launch pricing of $2 per conversation toward Flex Credits so that cost scales with discrete actions rather than with sessions.
Two reading habits matter when these structures appear in a data room. First, the same vendor can run several meters at once, so the headline per-outcome price may sit on top of seats, minimum commitments, or included allowances that expire. Second, the spread between committed and overage pricing is itself a signal: a vendor that discounts committed volume heavily is telling you it fears the cost variance of its own model.
What investors should test: predictability, margins and gaming
The three-part test converts into a concrete set of diligence questions. Each of the following tests maps back to one of the three gates, and each has already produced real-world evidence of where outcome-based models strain. In practice the work splits across the commercial diligence workstream, which owns the metric definitions and customer behaviour, and the financial diligence workstream, which owns the margin and recognition consequences.
- Usage variability and procurement certainty. Uber burned through its entire annual AI budget in four months and then capped employee spending at $1,500 per month per agentic coding tool. Where consumption is volatile, buyers push cost control back into the contract: procurement demands caps, and vendor revenue gets capped with it.
- Revenue predictability. When billing depends on customer-defined outcomes, forecasting degrades into guesswork. IDC's Ritu Jyoti notes that outcome-based pricing can lead to disputes over whether the desired effect was achieved, and that enterprises ask for subscription tiers because they want predictability.
- Margin alignment. AI vendors run gross margins around 50-60% against 80-90% for classic SaaS, because every query carries real inference cost. A fixed outcome fee loses money whenever the compute behind a resolution exceeds the resolution price, and the seller absorbs that gap.
- Customer gaming and metric design. The party that chooses the resolution stand-in is the party sending the invoice. Silence timers, assumed resolutions, and LLM self-verification all bias the meter toward the vendor, and a sophisticated customer will learn to work the timer.
- Post-contract dispute exposure. Contested attribution is not hypothetical. It surfaces as credits, refunds, and renegotiations, and it belongs in the diligence file as a recurring revenue adjustment, not a one-off.
Margin alignment deserves the most emphasis because it is the least visible. Zuora's guide is explicit that per-outcome pricing transfers cost variance from buyer to seller, and that it is only commercially viable when the seller can predict cost per outcome within a tight band, the outcome is auditable so disputes can be resolved, or the seller has the cash flow tolerance to absorb high-variance contracts. A target that cannot show which of those three conditions it satisfies is booking revenue it may not be able to keep.
Evidence checklist and red flags
The document checklist
- Master agreements and order forms with the outcome definition in writing, not in marketing decks
- Metric definitions, including how silence, inactivity timers, and human handoffs are treated
- Billing and usage logs, reconciled to issued invoices
- Revenue recognition memos under ASC 606, distinguishing a stand-ready obligation to provide access from an obligation to deliver a specified quantity of successful outcomes
- Dispute and credit records, with the contractual clause each dispute turned on
- Gross-margin bridges by pricing model, isolating inference cost per outcome
The ASC 606 distinction is not academic. Deloitte's accounting guidance on agentic AI notes that whether the promise is stand-ready access to the agent or delivery of a specified quantity of successful outcomes changes when revenue is recognized, and that outcome-based variable fees qualify for the variable consideration allocation exception only under narrow conditions. A target that has not documented which promise it is making has a revenue-recognition question, not just a pricing question. For the broader document set, our due diligence checklist for AI-native targets covers the adjacent ground.
Red flags
| Red flag | What it looks like in the data room | Why it matters |
|---|---|---|
| Vague outcome definition | Resolution or success is not defined in the contract, only in sales material | The vendor, not the customer, decides what was delivered |
| Vendor-scored outcomes | The vendor's own LLM verifies whether the outcome occurred | No independent measurement instrument exists; gaming risk sits entirely with the invoicing party |
| Minimum fees under per-outcome pricing | A base fee or minimum commit sits beneath the headline per-outcome price | The pay-only-for-outcomes story is softened by a stand-ready floor that bills regardless of results |
| Outcome fees below inference cost | The price per outcome looks aggressive against plausible compute cost per outcome | Margin erosion grows with adoption; the vendor is buying revenue |
| Hybrid mixes masking seat decline | Consumption line items grow while seat revenue shrinks | Net revenue quality may be deteriorating beneath a growing topline |
Implications for PE, M&A and how Plausity supports the work
For valuation, the mix matters more than the label. Consumption and outcome revenue behaves differently from committed ARR: it is more elastic in downturns, more sensitive to customer-side cost controls, and harder to forecast cohort by cohort. ServiceNow illustrates the upside case and the hedge at once: management reports that around half of net new business is already priced on a non-seat basis, including token and other usage meters, while keeping seat-based options because customers want a predictable entry point. For investment professionals sizing that trade-off, the question is not whether outcome pricing lifts the multiple, but which portion of revenue survives the three-part test, and our piece on valuation risk extends the same logic to premium-multiple AI targets.
IC and lender narratives carry their own risk when pricing models are in flux. A revenue build that assumes per-outcome billing can be invalidated when a target converts to hybrid, or when dispute-driven credits compress net revenue while gross bookings grow. Stress-test the model under both readings, and treat the pricing-metric mix as a cohort variable rather than a constant. Advisory teams that need to standardize this analysis across mandates will find the M&A advisory use case directly relevant.
- Data Room Ingestion scans contracts, order forms, and billing records within minutes of upload
- The AI-Analysis Engine extracts and cross-references pricing-metric definitions across documents, surfacing where the outcome definition differs between the master agreement and the order form
- Risk Radar grades findings by materiality, so contested outcome definitions and below-cost outcome fees reach the top of the risk register
- Report Builder structures the memo with source traceability back to the underlying clause
- Collaboration Hub keeps commercial, financial, and legal workstreams aligned on one set of findings
The platform supports the analysis; it does not make the pricing decision. Judgment about whether an outcome is genuinely defined, measured, and attributable stays with the deal team, which is where it belongs. What an AI-native due diligence platform changes is the cost of running the test at all: the definitions, the billing logs, and the margin bridges are extracted and connected rather than reconstructed by hand.
How to use this in your next diligence workflow
Run the framework as a sequence, not a brainstorm. The steps below assume a live deal with a data room open and a workstream split across commercial, financial, and legal reviewers. Teams that want to see how the steps fit a broader diligence workflow automation effort can map them directly onto their existing risk register.
- Request the pricing-metric definitions and a sample of billed outcomes in the first document request, before the data room is curated around you.
- Run the three-part test per product line and classify each revenue stream as defined, measurable, and attributable, or not.
- Reconcile billed outcomes to recognized revenue, and test margin per outcome against inference cost using the gross-margin bridge.
- Log red flags in the risk register with evidence links to the underlying clauses and billing records.
- Brief the IC on which revenue is genuinely outcome-backed versus subscription-backed, and on what happens to the model if the hybrid mix shifts or disputes crystallize.
The evidence checklist above is the reusable artifact. It converts this framework into a document request that can be dropped into any deal involving agentic AI revenue, and it pairs naturally with the custom workflows a team already runs for IC preparation. Plausity was built for exactly this kind of evidence-first diligence: it ingests the data room, extracts the definitions that matter, and keeps every finding traceable to its source, so the three-part test can be run on every AI-native target rather than only on the ones with clean data rooms. It is the same three steps every time: ingest, analyze, and report with the evidence attached.
How Plausity accelerates this workflow
Plausity is an AI-native due diligence and deal intelligence platform that helps M&A advisory firms, VC and PE funds, corporate development teams and investment-banking teams structure evidence, findings and questions across a data room. Plausity supports evidence extraction, source grounding, findings management and IC preparation — it does not replace human analysts, advisers or investment professionals, does not provide legal, tax, audit, regulatory or investment advice, and does not make autonomous investment decisions. All findings require human review.
To explore the underlying capabilities, see the Plausity AI analysis engine and the findings and risk intelligence product page. For team-level workflows, see how VC and PE funds and M&A advisory firms use Plausity across live deals.



