What pricing-model due diligence means for AI software
Pricing-model due diligence is the discipline of testing whether a target's monetization model, whether seat-based, usage or consumption-based, outcome or value-based, or a hybrid of these, actually fits the underlying product economics and the buying context it sells into. It is not a review of the pricing page. It asks what unit the company bills, who triggers that unit, whether the company can predict it, and whether the model will survive contact with enterprise procurement, AI-driven automation and the target's own cost structure. Getting this wrong misprices growth quality, retention and margin in the valuation.
Why the pricing model matters in diligence
The reason this has become a live diligence question is that AI breaks the assumption traditional SaaS pricing was built on. Classic SaaS assumed zero marginal cost: once the software was built, the next customer cost roughly the same to serve as the last, so flat per-seat subscriptions worked. AI products carry variable and meaningful inference cost, where every query, agent action or generated artifact triggers compute spend the vendor pays. When the pricing model does not track that variable cost, margin shrinks with adoption rather than expanding.
Diligence teams should calibrate against market reality rather than hype. AlixPartners analyzed 65 major SaaS companies and AI-native competitors and found that only 4 have fully adopted outcome-based pricing, while 72% run a hybrid approach incorporating consumption through AI credits or usage metering, and more than half of the remaining companies still rely primarily on per-seat models. Migration is real but slower and more uneven than the commentary suggests, which is why the correct model is a matter of product economics, not ideology.
- Per Seat: a fixed fee per user with access to the AI, familiar to buyers but fragile when variable inference cost rises with usage.
- Per Token: metered model input and output, the canonical infrastructure-layer model, transparent to engineers but hard for business buyers to budget.
- Per Activity: a charge per discrete action such as a conversation, API call or workflow run.
- Per Output: a charge per generated artifact such as an image, document or drafted report.
- Per Outcome: a charge only when the AI delivers a defined business result, such as a resolved ticket.
- Hybrid: a subscription base combined with consumption overage, prepaid credits or outcome-based components, the pattern most production AI businesses run.
Each of these models carries a distinct diligence profile. For a deeper treatment of how seat compression, inference COGS and hybrid structures flow through transaction analysis, see our companion piece on AI pricing due diligence. The sections below give buyers a matching framework, model-by-model tests, an evidence checklist and the red flags that connect pricing findings to valuation.
A framework for matching pricing models to product economics
Before judging whether a target's pricing model is good or bad, buyers should apply a diagnostic that separates fit from fashion. A model that looks aggressive in one context can be exactly right in another. Six questions structure that judgment:
- What is being monetized: access to the software, a countable activity, a generated output, or a delivered business outcome?
- Who triggers consumption: a human user, an automated workflow, or the AI agent itself? A model billed per user behaves very differently when agents do the work.
- How predictable is usage: is consumption stable and forecastable, or does it swing with seasonality, customer behavior and query complexity?
- Can value or outcome be measured and attributed: is success discrete and verifiable, or diffuse across many contributing factors?
- Does pricing align with enterprise procurement: can a CIO budget it, get it through purchasing, and defend it internally?
- How much pricing power does the company really have: is the pricing model defensible because the product is embedded in workflows and hard to replace, or is it exposed to competitive repricing?
Two published frameworks help answer these questions. Zuora's COMPASS framework sorts the six common AI pricing models, per seat, per token, per activity, per output, per outcome and hybrid, by the unit of value the AI delivers and how clearly that value can be measured, and notes that most production AI businesses end up running hybrid models combining a subscription base with consumption overage or prepaid credits. The framework does not tell a company what to charge, but it shows when an outcome claim is credible and when it is a marketing overlay: where success cannot be measured in a way the customer will accept, an outcome meter is a negotiating position rather than a pricing model.
AlixPartners adds three structural exposure conditions that tell buyers how much pressure a pricing model is under: a high concentration of seats in automatable workflows such as customer service or IT ticketing, selling into functions where AI ROI is easily measurable, and reliance on seats with limited secondary pricing layers. Companies at the intersection of all three face the most urgent need to act. Pricing power itself should be tested separately: where a product is deeply embedded and switching is costly, a vendor has more room to defend seat value, a dynamic we examine in our analysis of workflow switching costs.
What investors should test in seat, usage, outcome and hybrid models
Each pricing model fails in its own way, so each demands its own diligence questions. The table below summarizes the core tests; the paragraphs that follow expand on the two areas where deals most often go wrong.
| Model | Core diligence questions | Illustrative market terms |
|---|---|---|
| Seat | Where are seats exposed to automation of repetitive workflows? How much shelfware sits in the installed base? How deep is discounting, and what is real price realization? | Familiar per-user subscriptions; margin risk when heavy users consume more inference than the seat fee covers. |
| Usage / consumption | How is the meter defined? Who or what triggers consumption? How seasonal and forecastable is volume? Does price track inference cost? | A single complex query can consume 50 to 100 times the tokens of a simple one, making bills hard to predict. |
| Outcome / value-based | How are outcomes defined and scored? What are the attribution rules? What resolution rate makes the economics work? Who bears performance risk? | Intercom charges $0.99 for a solved conversation; Zendesk sells resolutions at $1.50 committed or $2.00 as you go, and HubSpot cut its price from $1.00 to $0.50. |
| Hybrid | How are allowances and overages designed? What does packaging include versus meter? How are features and modules monetized on top? | Subscription base plus consumption overage or prepaid credits; the dominant pattern among mature production AI businesses. |
Seat models: compression is the silent churn
For seat-based targets, the central test is seat-compression risk: where AI agents can handle the workload of users who would otherwise hold licenses, the value-per-seat argument weakens with every deployment cycle. Diligence should quantify how much revenue sits in workflows that are repetitive and rules-based, and separately examine shelfware, the gap between licensed and active seats, because it flatters reported adoption and understates renewal risk. Discounting depth and price realization complete the picture: a company growing seats by cutting price is not demonstrating pricing power.
Outcome models: the definition is the product
For outcome-based targets, the pricing definition is effectively part of the product. There is no agreed industry standard for scoring outcomes, so published per-resolution prices are not comparable across vendors. Buyers should establish how success is verified, how customer silence is treated, what happens when the AI hands off to a human mid-task, and who bears the cost when resolution rates disappoint. AlixPartners found that the four companies in its cohort with confirmed outcome-based pricing all operate in customer service workflows where outcomes are discrete and verifiable, which is a useful benchmark for how narrow the viable space currently is.
Evidence and document checklist for pricing diligence
Pricing findings are only as strong as the documents behind them. The data room should be mined for the following, and gaps in this list are themselves a finding:
- Customer contracts and order forms with pricing schedules, including renewal terms and any usage or outcome commitments.
- Price books and discount approval records, to test realized price against list price.
- Usage and consumption data by customer, including credit burn-down and true-up mechanics for prepaid or allowance-based models.
- Revenue recognition memos covering how usage, credits and outcome fees are treated under ASC 606.
- Churn, downgrade and expansion evidence around past pricing or packaging changes, at cohort level where available.
Contract terms determine the accounting treatment, which in turn shapes reported revenue quality. Deloitte's DART guidance highlights the key judgment for agentic AI arrangements: whether the vendor's promise is a stand-ready obligation to provide continuous access to the agent, typically recognized over time, or an obligation to deliver a specified quantity of successful outcomes, typically recognized as those outcomes occur. Success criteria and roll-over rights matter here, and prepaid outcome credits may be nonrefundable on a use-it-or-lose-it basis or refundable, which changes how deferred revenue and variable consideration behave.
Usage evidence also needs to be tied to the meter actually in force, because meters themselves are moving. GitHub Copilot, for example, meters interactions by input, output and cached tokens at rates that differ by model, converts them into AI credits where one credit equals one US cent, includes an allowance per plan and bills usage above that allowance per token, while code completions stay outside the credit meter. A usage dataset built on an old meter can be worthless for forecasting the current one. Tools for financial document analysis help deal teams extract these terms from dense contracts and cross-reference them against the usage files.
Red flags in AI software pricing: a diligence table
The following patterns should trigger deeper testing before they reach the investment committee.
| Red flag | Why it matters | What to test |
|---|---|---|
| Revenue concentrated in seats exposed to automation of repetitive workflows | The seat base that drives current revenue is the seat base AI can compress first. | Map revenue by workflow type; quantify the share in ticketing, routine analysis and basic content generation. |
| Outcome definitions controlled by the vendor with no independent audit trail | The vendor both defines success and bills on it, so reported outcomes may not survive scrutiny. | Request the scoring methodology, dispute history and any third-party validation of resolution counts. |
| Heavy discounting needed to sustain usage growth | Growth bought with price is not pricing power and will not hold at renewal. | Compare list price to realized price over time; review discount approval records. |
| Unexplained revenue movements after pricing or packaging changes | Mix shifts can mask underlying churn or downgrades behind headline growth. | Rebuild revenue bridges by cohort around each pricing change. |
| Revenue recognition treatment inconsistent with contract terms | Misclassified stand-ready versus outcome obligations distort both revenue timing and quality. | Read revenue memos against the actual contract language on success criteria and roll-over rights. |
| Pricing migrations mid-flight with no cohort-level evidence of retention | A model change resets the baseline for NRR and forecasting precisely when buyers need it most. | Request cohort retention and usage data for customers migrated to the new model. |
The pattern behind several of these flags is risk allocation. Of the 65 companies AlixPartners analyzed, 53 put performance risk on the customer, 8 share it, and only 4 bear it as the vendor, meaning that for 82% of companies the customer takes the punishment if the software does not deliver. And in outcome pricing specifically, the party choosing the stand-in for success is the party sending the invoice, which is why outcome definitions deserve the same scrutiny as the price itself.
Implications for valuation, revenue quality and post-close value creation
For PE and VC investment professionals, pricing findings should feed directly into the revenue quality assessment. Under variable consideration, Deloitte notes that where an arrangement qualifies as a series of distinct services, outcome-based fees may be recognized as successful outcomes occur rather than over time, which changes revenue timing and makes reported growth more sensitive to resolution volumes. Forecasting a business whose revenue depends on outcomes the customer must achieve is a different exercise from forecasting seats, and the ARR quality analysis should reflect that.
Gross margin is the second exposure. When pricing does not track variable cost, adoption expands inference spend faster than revenue, the margin trap Zuora describes. Diligence should model cost to serve per billing unit, not just blended gross margin, and stress-test it against usage growth. This connects directly to the financial model due diligence workstream, where unit economics assumptions often hide pricing-model risk.
Third, NRR mechanics and multiple risk concentrate wherever a model is mid-migration. AlixPartners warns that moving to outcome-based pricing too early risks cannibalizing per-seat revenue, while moving too late risks ceding market position. A target caught between models can show compressed NRR on the old base and unproven retention on the new one, which is a multiple question, not a footnote.
- Repricing: correcting realized price toward list where discounting has eroded realization.
- Repackaging: restructuring allowances, overages and module entitlements to match how customers actually consume.
- Migration sequencing: moving customer cohorts to new models in an order that protects seat revenue while the new model's retention is proven.
These levers are where post-close value is created or destroyed, and teams advising M&A advisory mandates should treat the sequencing question as a value-creation plan item rather than a commercial afterthought.
How Plausity supports the pricing diligence workflow
The framework above only matters if it reaches the deal team's workstreams in a structured, evidence-backed form. The platform supports that workflow end to end:
- Use Data Room Ingestion to scan customer contracts, order forms and price books within minutes of the room opening, so pricing terms enter the analysis from day one.
- Use the AI-Analysis Engine to extract pricing terms, meter definitions and success criteria from dense contracts, and to cross-reference usage data against reported revenue.
- Use Risk Radar to surface pricing risks, score them by materiality and deal relevance, and rank them alongside the rest of the risk register.
- Use the Collaboration Hub to align the commercial, financial and tech DD workstreams on a single set of pricing questions and evidence.
- Use Report Builder to draft evidence-backed findings with full source traceability, so every pricing claim in the report links back to a document in the room.
One boundary should be stated plainly: Plausity helps deal teams structure evidence and organize diligence questions, it does not automatically determine the correct pricing model. That judgment stays with the team, and it is best made the way TechTarget's experts advise buyers to make it, by walking real cases through the pricing before signing rather than arguing about headline rates. Deal teams that want to see how this fits into a broader evidence-backed deal team setup, or how AI supports diligence workflow automation across data-room triage and risk registration, will find the pricing questions above slot directly into those workflows.
How to use this in your next diligence workflow
In practice, run the pricing workstream in four steps. First, classify the target's billing unit and packaging from the order forms rather than the pricing page, and note where seats, meters and outcome fees overlap. Second, apply the six diagnostic questions above to test fit against product economics and procurement reality. Third, tie every claim to a document: contracts, price books, usage extracts and revenue-recognition memos, flagging gaps as findings in their own right. Fourth, translate the findings into the valuation case, showing what they mean for growth quality, gross margin, revenue predictability and the post-close repricing, repackaging and migration-sequencing plan. Commercial diligence teams can carry that structure straight into the investment committee pack.
How Plausity accelerates this workflow
Plausity is an AI-native due diligence and deal intelligence platform that helps M&A advisory firms, VC and PE funds, corporate development teams and investment-banking teams structure evidence, findings and questions across a data room. Plausity supports evidence extraction, source grounding, findings management and IC preparation — it does not replace human analysts, advisers or investment professionals, does not provide legal, tax, audit, regulatory or investment advice, and does not make autonomous investment decisions. All findings require human review.
To explore the underlying capabilities, see the Plausity AI analysis engine and the findings and risk intelligence product page. For team-level workflows, see how VC and PE funds and M&A advisory firms use Plausity across live deals.



