What AI unit economics due diligence means
AI unit economics due diligence is the economic layer of technology diligence: the structured test of whether a target's revenue scales faster than the costs of delivering AI, as usage grows. It sits below the product and technology review and asks a narrower question than 'is the AI any good?'. It asks what one incremental unit of delivered work, a processed document, an answered query, a completed workflow, actually costs to serve, and whether that cost falls or rises as the customer base expands. The three core variables are cost-to-serve, gross-margin sensitivity and pricing power: what the AI workflow costs to run, how exposed the margin is to changes in that cost, and whether the company can hold or raise prices while costs move.
This layer deserves its own workstream because AI cost of goods behaves differently from traditional software cost of goods. In classic SaaS, serving an additional customer is close to free, so gross margins cluster at high levels and diligence can treat them as stable. In AI products, a large share of delivery cost is variable and metered: every workflow consumes tokens, and every token is billed by a model provider or paid for as cloud and GPU capacity. Whether a target's margins look like traditional software or something materially thinner is an empirical question about its own architecture, contracts and customer mix, not a category law.
The discipline for deal teams is to avoid blanket claims in either direction. Commentary that declares AI software structurally lower-margin is as untested as the assumption that AI products carry SaaS margins. Published venture benchmarks for 2025 show a wide spread within the same category: some of the fastest-scaling AI companies operate at gross margins far below the software norm, sometimes negative at the unit level, while slower-growing peers sit much closer to conventional software levels. Both profiles exist in the same category, at the same time. The task in diligence is to locate the target on that spectrum using its own numbers, and then to test how durable its position is. This is the economic complement to the technology view set out in our AI-native software moat diligence framework.
- Definition: testing whether revenue scales faster than inference, third-party API and hosting costs as customer usage grows.
- Core variables: cost-to-serve per workflow and per customer, gross-margin sensitivity to cost and usage shifts, and pricing power to hold or pass through cost changes.
- Why now: margin outcomes are company- and product-specific, so sector averages and category narratives are a poor substitute for the target's own cost stack.
- Diligence discipline: no blanket margin claims; the framework below treats gross-margin durability as a hypothesis to be evidenced, not assumed.
The AI cost stack: from tokens to cost-to-serve
The first analytical step is mapping where AI cost of goods actually sits. The visible line item is per-token inference spend with model providers, billed per million tokens at list rates that vary by model tier. Around it sit third-party model and API fees for embeddings, transcription or other specialised services, cloud and GPU hosting for any self-hosted models, vector storage and retrieval infrastructure, and the orchestration layer that routes requests between them. Wrapped around all of this are the human costs that AI products still carry: engineering time on prompts and evaluation, customer support, and in many cases review or QA work on AI output.
From this stack, build a variable cost per workflow and per customer, the point at which this analysis meets financial due diligence. Take a representative workflow, measure the tokens consumed per run, apply the blended token price, add the share of hosting, storage and orchestration costs that scale with usage, and include the human cost per unit where output is reviewed. Repeating this across customer segments produces a cost-to-serve view that can be set against revenue per customer to yield contribution margin by segment. The output is rarely what management presents: most management materials show an aggregate gross margin, not a customer-level cost picture.
A consistent finding is that the true cost of an AI workflow exceeds the raw API bill. The token invoice is the most legible number, so it anchors most internal analysis, but it understates delivery cost in predictable ways. Retrieval and storage costs scale with the document corpus, not with revenue. Orchestration often calls cheaper models multiple times per visible output. Failed runs, retries and evaluation traffic consume tokens that never reach a customer. And the human review layer, where it exists, is a genuine cost of goods that no token log captures. A diligence-grade cost-to-serve model therefore starts from the API bill but must be reconciled upward to fully loaded delivery cost.
- Per-token inference spend with model providers, at rates that differ sharply by model tier and volume commitment.
- Third-party model and API fees for embeddings, transcription, safety filtering and other specialised services.
- Cloud and GPU hosting for self-hosted or fine-tuned models, including capacity that is reserved rather than consumed on demand.
- Vector storage, retrieval infrastructure and orchestration, which scale with corpus size and workflow complexity rather than with revenue.
- Human and support costs: prompt engineering, evaluation, output review and customer success, which are part of cost to serve even though no token log shows them.
Falling inference prices and margin durability
The deflation curve is the best-evidenced trend in AI economics, and it is genuinely steep. a16z's analysis of historical pricing since GPT-3 found that for an LLM of equivalent performance, inference cost falls roughly 10x per year, with the cost of a fixed MMLU score dropping by a factor of 1,000 in three years. The Stanford AI Index 2025 records the cost of querying a model at GPT-3.5-level performance on MMLU falling from $20.00 per million tokens in November 2022 to $0.07 by October 2024, a more than 280-fold reduction in approximately 18 months, and notes that depending on the task, LLM inference prices have fallen anywhere from 9 to 900 times per year. That range, rather than any single headline multiple, is the right input to a margin model, which should treat deflation as a scenario variable rather than a constant.
For diligence, the curve cuts both ways, and the question is not whether prices fall but who captures the fall. If a target's gross margins are improving primarily because its model provider keeps cutting list prices, the improvement is a pass-through of an industry trend, not evidence of company-specific economics. In competitive markets, providers tend to compete those savings away: customers expect prices to fall as the provider's costs fall, and rivals price aggressively to win share. A margin bridge that depends on continued input-price deflation should be stress-tested against a scenario in which the deflation slows, or in which competition forces the target to pass the savings to customers. Equally, a target whose costs are falling but whose prices are fixed in contract may be one of the genuine winners of the trend; the point is to identify which side of the pass-through the target sits on, with evidence.
Management also has efficiency levers it can actually pull, and the diligence questions are which of them the target uses and what each is worth. a16z attributes the price decline to several independent forces: better GPU cost-performance, model quantization, software optimisation that reduces compute and memory requirements, smaller models trained on more data, better instruction tuning, and open-source competition compressing margins across the value chain. A target can route routine requests to smaller models and reserve frontier models for hard cases, quantise models it hosts itself, cache repeated queries, and cut wasted tokens through better retrieval. Each lever has an engineering cost and a quality risk, so the test is whether the target has a measured roadmap for these levers or treats falling prices as its only margin plan.
- The deflation evidence: roughly 10x per year at constant performance, a more than 280-fold drop for GPT-3.5-level quality in about 18 months, and declines of anywhere from 9 to 900 times per year depending on the task.
- Who captures the savings: does the target retain falling input costs as margin, or does price competition pass them to customers? The margin bridge should state which, with evidence.
- Efficiency levers: model routing, quantization, caching, prompt and retrieval optimisation, and smaller models for routine work. Which are deployed, which are planned, and what is each worth in basis points of gross margin?
- The stress test: re-underwrite the margin profile assuming input-price deflation slows materially, and see how much of the plan survives.
Customer-level economics: usage intensity and contribution margin
Aggregate gross margins hide the most important distribution in an AI business: the relationship between usage intensity and profitability. In a usage-metered cost structure, the heaviest users consume the most tokens, so the customers who value the product most can be the most expensive to serve. The diligence test is to rank customers into usage deciles and compute contribution margin per customer, revenue minus fully loaded variable delivery cost, for each decile. If the top deciles are the most profitable, usage-based growth is compounding margin. If the top deciles are at or below zero contribution, growth is destroying value at the margin and the pricing model is misaligned with the cost structure.
Pricing model alignment is the second half of the analysis, and it belongs in commercial due diligence as much as in the financial workstream. Seat-based contracts decouple revenue from consumption: a customer that doubles its usage under a flat seat price converts falling unit costs into a margin gain, but a customer that quadruples usage can push its contribution negative. Usage-based or hybrid contracts re-couple revenue to consumption and pass cost shocks through, but expose the target to price competition on a metered input. Neither model is inherently superior; the question is whether the contract structure matches the cost structure, and whether pricing was set with a view to the cost-to-serve analysis above.
The market data shows how wide the company-specific spread is. Bessemer's 2025 benchmarks put fast-scaling AI Supernovas at roughly 25% gross margins, often negative at the unit level, against about 60% for Shooting Stars that grow somewhat more slowly. At the infrastructure layer, TechCrunch reports that Anthropic expects a 50% gross profit margin this year and 77% by 2028, up from negative 94% the prior year. A frontier lab and an application company are different businesses, but the point for diligence is the spread itself: businesses in the same category span the full range, and the target's position is determined by its own usage mix, contract structure and cost discipline.
Pass-through mechanics complete the picture. When AI input costs rise, whether through price changes at a model provider, capacity constraints or heavier consumption, can the target reprice? Contracts with fixed pricing and no usage caps absorb the shock entirely. Contracts with usage-based pricing pass it through automatically. Annual commitments with renegotiation windows sit in between. The diligence questions are concrete: what share of revenue renews in the next twelve months, what do the pricing clauses actually say, and has management tested customer tolerance for price increases that reflect higher delivery costs?
What investors should test: evidence checklist and red flags
The framework reduces to five core tests, each of which should be answerable from documents rather than from management narrative. Together they establish whether revenue scales faster than cost-to-serve, what assumptions underpin management's future margin profile, and whether the company can pass increased AI costs to customers.
- Revenue versus cost-to-serve: does revenue per customer grow faster than fully loaded variable cost per customer, at each usage decile, over the last eight quarters?
- Customer-level contribution: are the highest-usage customers actually the most profitable, or is growth concentrated in customers with thin or negative contribution?
- Margin bridge assumptions: what exactly drives management's projected margin improvement, input-price deflation, efficiency levers, pricing actions, or mix, and is each driver evidenced?
- Pass-through capacity: can the company pass increased AI costs to customers through contract terms, usage pricing or repricing at renewal, and what share of revenue is protected?
- Vendor and pricing dependency: how exposed is the cost base to a single model provider's pricing, capacity and product decisions, and what alternatives exist at what switching cost?
The evidence request list should go into the data room early, because these documents take time to produce and are the ones most often missing. Cloud and API invoices across all providers, token usage logs by customer and by workflow, a customer-level profit and loss or at minimum revenue and cost attribution per account, model-provider contracts including committed-spend and rate terms, and the pricing policy documents that govern repricing and pass-through. Where a document does not exist, that fact is itself a finding: a target that cannot attribute cost per customer cannot defend its margin story.
| Red flag | What it looks like in the data | Why it matters |
|---|---|---|
| Flat pricing with rising usage | Revenue per account flat while token consumption per account climbs across several quarters | The target is absorbing consumption growth; contribution margin per heavy user is compressing even if aggregate margin looks stable |
| Single-model dependency | One model provider accounts for the substantial majority of inference spend, with no committed rates and no tested alternative | Pricing, capacity and product decisions at one vendor flow directly into the target's cost of goods |
| Margin gains driven only by input price cuts | Margin bridge shows improvement almost entirely from provider price deflation, with no efficiency or pricing actions | The margin profile depends on an industry trend the target does not control, and competition can pass the savings to customers |
| Absent cost attribution per customer | No customer-level P&L, no usage logs by account, cost-to-serve estimated only at company level | Management cannot see which customers are unprofitable, so the margin story cannot be evidenced or defended |
Vendor concentration, contracts and margin compression risk
Concentration risk in the AI value chain is best understood through a worked example from a public filing. CoreWeave's S-1, filed ahead of its 2025 IPO, disclosed that Microsoft was its largest customer in both 2023 and 2024, accounting for 35 percent and then 62 percent of revenue. Its two largest customers together accounted for 77 percent of 2024 revenue, and committed take-or-pay contracts, typically two to five years in length, accounted for 96 percent of revenue that year. The filing itself flagged that any negative change in demand from Microsoft would adversely affect the business. The lesson for target-level diligence is structural: at every layer of the AI value chain, a small number of counterparties control critical inputs, and the terms they set, pricing, capacity, contract length, flow directly into a dependent company's cost of goods.
Translated to a target, the upstream questions are direct. Does the target depend on a single model provider or a single cloud for the majority of its inference spend, and what share of that spend is on committed rates versus on-demand list prices? Committed-spend contracts buy price stability and capacity priority but create floor costs that persist if usage disappoints; on-demand pricing flexes with usage but exposes the target to list-price movements it cannot control. The concentration analysis should also run downstream: a target whose own revenue depends heavily on one or two customers inherits the same fragility CoreWeave disclosed, with the added exposure that its largest customers may be negotiating equivalent terms with its competitors.
Margin compression scenarios should then be modelled explicitly. The most instructive recent case is Anthropic, which per TechCrunch's reporting expects a 50% gross margin this year against a prior-year figure of negative 94%, with a 77% target by 2028. Even a frontier lab with scale advantages and its own efficiency programmes treats margin as a trajectory to be managed, not a stable number. For a target, the scenarios to stress are: input prices that stop falling or rise, usage patterns that shift toward heavier consumption, a major customer renegotiating at renewal, and a model provider changing its pricing structure or deprecating a model the target's product depends on.
For PE and M&A, these findings translate into deal structure. Where margin durability cannot be fully evidenced, the valuation should reflect it, either as a haircut or as a price that assumes the stressed rather than the management case. Earnout structures can tie consideration to realised gross margin, aligning sellers with the margin profile they have represented. Representation and warranty coverage should address cost assumptions explicitly: the accuracy of cost attribution, the terms of model-provider contracts, and any committed-spend obligations that become the buyer's liability at closing. None of this is adversarial; it is the standard response to a cost structure that is more contract-dependent than traditional software.
How to use this in your next diligence workflow
The framework converts into a sequenced workflow that runs from data-room access to investment committee. Each step maps to the tests and benchmarks set out above, and each produces a specific artefact the deal team can defend.
- Ingest the data room: connect to the VDR and process the invoices, contracts, financial models and management materials that carry the cost evidence, including cloud and API invoices, model-provider agreements and the pricing policy documents from the checklist above.
- Pull the cost evidence: extract token usage logs and invoice line items across providers, and reconcile the raw API bill upward to a fully loaded variable cost per workflow, covering hosting, storage, orchestration and human review.
- Build the cost-to-serve view: construct contribution margin per customer and per usage decile, and set it against revenue per customer to establish whether the heaviest users are the most profitable.
- Stress the margin bridge: decompose management's projected margin profile into input-price deflation, efficiency levers and pricing actions, and re-underwrite it under a slowed-deflation scenario, benchmarked against the BVP State of AI 2025 margin data.
- Package the findings: assemble the margin and concentration findings, the red flags and the pass-through analysis into an evidence-backed memo for the investment committee.
Plausity supports each step of this workflow. Data Room Ingestion connects to the VDR and processes invoices, contracts and financial models within minutes, so the cost evidence enters analysis immediately rather than waiting on manual indexing. The AI-Analysis Engine reads and cross-references those documents, connecting invoice line items and usage logs to the management model and surfacing where the numbers do not reconcile. Risk Radar evaluates the findings by materiality, financial impact and deal relevance, so margin compression and vendor concentration risks are ranked alongside the rest of the risk register. Report Builder drafts the evidence-backed memo with full source traceability back to the underlying documents, and Collaboration Hub keeps the workstreams aligned as the cost-to-serve view, the margin bridge and the IC materials come together. The output is a margin analysis in which every figure traces to a document in the data room, which is the standard the five tests above demand.
How Plausity accelerates this workflow
Plausity is an AI-native due diligence and deal intelligence platform that helps M&A advisory firms, VC and PE funds, corporate development teams and investment-banking teams structure evidence, findings and questions across a data room. Plausity supports evidence extraction, source grounding, findings management and IC preparation — it does not replace human analysts, advisers or investment professionals, does not provide legal, tax, audit, regulatory or investment advice, and does not make autonomous investment decisions. All findings require human review.
To explore the underlying capabilities, see the Plausity AI analysis engine and the findings and risk intelligence product page. For team-level workflows, see how VC and PE funds and M&A advisory firms use Plausity across live deals.



