Software M&A Due Diligence Playbook: ARR to Defensibility

Software M&A Due Diligence Playbook: ARR to Defensibility

Image: Plausity

Key Takeaways

  • Median B2B SaaS NRR is 102%, with usage-based targets at 108% versus 98% for seat-based models (Aleph and Benchmarkit, 2025 data).
  • Median software gross margin is 80%, but usage-only pricing models run at 62%, a structural trade-off against their stronger retention.
  • Median SaaS CAC payback was 18 months in 2024, improving to 16 months in 2025, and must be read alongside NRR.
  • Q4 2025 enterprise SaaS M&A deal value reached $83.7 billion, capping the strongest year for enterprise SaaS deals since 2021 (deal databases).
  • PwC warns that NRR can mask AI-driven seat contraction, making GRR and cohort-level analysis essential in current diligence playbooks.

Why software M&A diligence is a category of its own

What should a software or SaaS buyer actually diligence before committing to a transaction? The answer spans six evidence streams: the quality and durability of recurring revenue (ARR definitions, cohort retention, churn composition), the customer base (concentration, contract terms, change-of-control exposure), the pricing architecture and unit economics behind it, the product's embeddedness and technical condition, the target's exposure to AI-driven substitution, and the credibility of management's plan to create value post-close. None of these streams can be read in isolation: a 102% net revenue retention means something very different for a seat-based vendor than for a usage-based one, and a clean ARR bridge is worthless if the contracts underneath can be terminated on change of control. The playbook below moves through each stream in order and closes with the evidence pack an investment committee actually needs.

The stakes are high because the market is. Global enterprise SaaS M&A deal value climbed 23.9% quarter on quarter to $83.7 billion in Q4 2025 across 245 deals, making 2025 the strongest year for enterprise SaaS dealmaking since 2021, according to deal databases. With that much capital chasing scaled software assets, the cost of a diligence miss (an overvalued cohort, a terminable contract, a replaceable product) compounds across an entire fund cycle.

AI has also changed what buyers are buying. Sponsors now rank software targets into tiers of AI disruption exposure, from insulated system-of-record owners to replaceable point solutions, and AI has become a standalone diligence workstream in almost every deal, probing not only technical capability but competitive positioning against AI-native entrants. The framework below reflects that reality: AI exposure sits alongside, not after, the classic commercial and technical workstreams.

Diligence stageCore questionKey evidence
Quality of revenueIs the ARR real, recurring and durable?ARR bridge, cohort retention matrices, churn composition
CustomersIs the base concentrated and contractually secure?Revenue distribution, contract terms, change-of-control clauses
ProductIs the product embedded or discretionary?Usage telemetry, integration count, system-of-record status
TechnologyWhat does the codebase and roadmap really cost?Tech debt inventory, architecture review, capitalized development
EconomicsDo unit economics support the underwriting case?Gross margin by revenue line, CAC payback, pricing model
AI exposureWill AI substitute or amplify the target?Tier assessment, AI roadmap proof points, seat-contraction risk
Value creationWhat does the buyer do with the asset?Stress-tested forecast, workstream evidence pack, value-creation plan

Each stage feeds the next. Revenue quality questions determine which customers to examine, customer contracts constrain pricing action, pricing architecture shapes unit economics, and the product and technology findings determine whether AI is a threat or a lever. The sections that follow take them in turn.

ARR definition and recurring revenue quality

Everything in software diligence starts with a clean ARR definition, because ARR is the denominator for every multiple and every retention metric that follows. The first request to management is not a number but a bridge: a monthly reconciliation of opening ARR to closing ARR that separates the four movements any buyer needs to see distinctly.

  • New logo ARR: revenue from customers acquired in the period, which must be separated cleanly from expansion, especially under usage-based pricing where the boundary is easy to blur.
  • Expansion ARR: upsell, cross-sell and usage growth within existing accounts, the engine behind net revenue retention.
  • Contraction ARR: downgrades and reduced commitments, which signal a value-realization or pricing problem rather than a lost customer.
  • Churned ARR: revenue from cancelled contracts, best read alongside the logo count to distinguish logo churn from revenue churn.

NRR and GRR: cohort performance and cohort discipline

On top of the bridge, buyers should read net revenue retention (NRR) and gross revenue retention (GRR) together, because the gap between them is the diagnostic. Full-year 2025 benchmark data from the Aleph and Benchmarkit study of 342 SaaS and AI-native companies puts median NRR at 102%, with usage-based companies at 108% versus 98% for seat-based models, while median GRR sits at 84%, down four points year on year. An 18-point spread means expansion is doing heavy lifting to cover real churn underneath; a healthy NRR built on a weak GRR is a warning sign, not a victory. The premium is real: industry research's analysis of more than 100 B2B SaaS companies finds that top-quartile NRR performers reach 113% retention.

Churn deserves its own read, separate from the blended retention metrics, because its three components point to different problems. Logo churn (customers lost outright) signals product-market fit failure or competitive displacement; revenue churn (ARR lost from the base) weights those losses by account size and can be far worse than the logo count suggests when the losses sit in the largest accounts; and downsell or contraction (customers paying less for the same relationship) signals a value-realization gap, an over-scoped original contract, or seat erosion from AI-driven headcount efficiency. The diligence artefact is a churn decomposition by cohort, contract size and reason code: a base losing logos but almost no revenue is a very different asset from one losing few logos but a large share of ARR, and only the decomposition tells you which one you are buying. Reason codes matter most: churn attributed to budget cuts or M&A on the customer side is exogenous, while churn attributed to missing features, poor support or a cheaper AI-native alternative is a product verdict.

Churn: logo, revenue and downsell

The most revealing artefact in this workstream is a cohort matrix by vintage: every annual customer cohort tracked for retention in each year since acquisition. Newer cohorts decaying faster than legacy ones is one of the clearest signals that growth was bought through discounting, weak fit, or a shifting market rather than earned through product value. For a deeper treatment of ARR durability in take-private contexts, see the firm's analysis of ARR quality in mission-critical SaaS.

Customer concentration, contract duration and change-of-control risk

Revenue quality is a distribution question as much as a retention question. Diligence should map ARR across the customer base and test the standard thresholds: any single account above roughly 10% of ARR, the top five customers as a share of the total, and concentration within specific verticals or geographies that a macro shock could hit simultaneously; industry best practice holds that no single customer should exceed 10% of ARR, with the top ten customers ideally below half of recurring revenue. Concentration interacts with retention: a top-five account that is also a legacy cohort member with deep integration is a different risk than one on a recent discount, and the cohort matrix from the previous section tells you which kind you are looking at.

Contract review is where sponsor-grade diligence earns its name. Private equity buyers were involved in nearly 58% of all SaaS transactions in 2025, one of the most sponsor-heavy years on record, so the market norm is a full population review of customer contracts rather than a sample. The clauses that matter most are the ones that constrain what the buyer can do post-close.

  • Change-of-control termination triggers: customer rights to exit or renegotiate on a sale, which can turn headline ARR into a fraction of itself the day the deal closes.
  • Contract duration and renewal mechanics: remaining term, auto-renewal, and whether renewals require re-procurement or competitive bids.
  • MFN (most favoured nation) pricing clauses: commitments that cap the buyer's ability to reprice the base or introduce new pricing tiers.
  • SLA penalties: service-level credits that scale with usage or incidents and directly affect post-close margin.

Each of these should be quantified into an exposure figure (ARR subject to change-of-control rights, ARR under MFN constraints) rather than left as a legal observation, because that figure flows directly into the revenue adjustments the deal model underwrites.

Pricing architecture: seat, usage and outcome models

Pricing architecture is the strongest structural determinant of retention, and it is where AI cuts both ways. Seat-based pricing, the foundational unit economics of SaaS, is under structural pressure: PwC's deal advisory team warns that traditional metrics like NRR can mask AI-driven seat contraction, which makes GRR and cohort-level analysis essential in current diligence, and that seat-based growth models face sharper pressure as AI lowers barriers to competition. Usage-based models, by contrast, convert customer consumption into organic expansion, which is why they post materially higher retention.

Gross margin and CAC payback

The unit economics behind each model are measurable. The 2025 benchmark medians are 80% gross margin on software revenue and 76% on total revenue, with usage-only pricing models running materially lower at 62% because compute and infrastructure costs weigh on every consumption dollar. CAC payback tells the go-to-market side of the story: the median B2B SaaS company recovered acquisition cost in 18 months in 2024 before improving to 16 months in 2025. Neither metric should be read in isolation. A usage-based model trading margin for a 108% NRR can be the better asset; a seat-based model with a long payback and a 98% NRR is exposed on both ends. For a full treatment of model-selection risk, see the firm's guide to AI software pricing due diligence.

Product usage, switching costs and system-of-record status

The central product question is whether the target is a system of record or a discretionary tool. A system of record owns data and workflows tied to a financial or regulatory outcome; a discretionary tool competes on user experience and can be swapped in a budget cycle. PwC frames the durable version of this as workflow gravity: products that own the system of record and are tied to financial or regulatory outcomes tend to have AI agents layered on top of them rather than routed around them, while standalone tools competing primarily on UX face severe pricing pressure.

Embeddedness is testable, and the evidence should be requested in the data room rather than inferred from the sales deck.

  • Usage telemetry: DAU/MAU ratios, session depth and feature-level adoption that distinguish daily operational use from periodic log-ins.
  • Integration count and depth: number of production integrations, bidirectional data flows, and whether the product writes back into customer systems of record.
  • Feature depth: the share of contracted modules actually deployed, which reveals whether the contract outruns the footprint.
  • Roadmap credibility: shipped-versus-promised history over the last 24 months, and whether the roadmap addresses the AI question or defers it.

Product roadmap credibility

Roadmap credibility is the forward-looking half of the product workstream, and it is where the shipped-versus-promised history earns its place in the evidence pack. The request is simple: the last 24 months of roadmap commitments set against what actually shipped, with dates. A team that promised an AI capability three quarters running and shipped none of it is telling you something about engineering capacity, prioritization or both, and the gap flows directly into the value-creation plan, because whatever was deferred becomes the buyer's capex. The credibility test also extends to the roadmap's direction: PwC's deal advisory view is that the companies commanding premium exits are those using AI to accelerate their own product development, shipping faster and deepening customer stickiness, while a roadmap that defers the AI question entirely is a roadmap the market has already started to discount. Diligence should therefore score the roadmap on shipped velocity, capacity honesty and whether the AI items on it are funded engineering commitments or positioning.

Technical debt

Technical diligence then prices what the product review surfaces. Monolithic codebases, deprecated dependencies and thin engineering documentation all translate into post-close investment requirements, and capitalized development costs deserve particular scrutiny because they can inflate adjusted EBITDA with spending that is, economically, maintenance. A structured approach to this workstream is set out in the firm's tech due diligence checklist for software M&A.

AI substitution risk

AI exposure is now a workstream of its own, and this playbook deliberately does not repeat the framework for testing workflow defensibility in the AI reset; that analysis, including the diagnostics for system-of-record status, proprietary data loops and replication risk, is covered in the firm's companion piece on the AI reset in software M&A. What the broader playbook needs is the tier logic that sponsors now apply and the upside case that credible AI proof points can support.

On the risk side, the market has formalised the sorting. Specialist investors have built internal frameworks ranking SaaS companies from most insulated (owners of proprietary data and core system records adopted company-wide) to most vulnerable (narrow point solutions), with the middle tiers, department-level and productivity tools, under the greatest threat of AI replacement. Every target in a process should be placed in a tier, and the placement should be evidenced rather than asserted.

AI opportunity: where AI can add value inside the acquired company

On the opportunity side, AI proof points have moved from nice-to-have to expected in the equity story. The clearest recent example comes from treasury software: GTreasury's agentic AI tool reached 100% weekly active usage across its existing install base within two months of launch, and 40% of new customer inbounds came from GSmart AI-related content, evidence its backer cited in the exit process. Diligence should therefore ask not only whether AI threatens the target but whether the target has shipped anything that changes retention or expansion, and whether the usage data supports the claim. A structured approach to both sides of the question is set out in the firm's AI impact due diligence playbook.

Management forecasts and stress tests

The final stage turns findings into an investment committee decision. Management forecasts should be stress-tested against the specific failure modes the diligence surfaced: cohort decay rates from the vintage matrices rather than management's blended retention assumption, seat contraction under AI-driven headcount efficiency for seat-based models, and pricing-model migration costs where the underwriting case assumes a shift toward usage or outcome pricing. A forecast that only works at the diligence-stage NRR, with the GRR underneath it unexamined, is not a forecast; it is a hope with a spreadsheet.

Evidence checklist and IC-ready synthesis

Each workstream should close with a defined evidence artefact, so the IC reads conclusions that trace back to documents rather than to summaries of summaries.

  • ARR bridge and cohort matrices by vintage, reconciled to the general ledger.
  • Retention analysis separating logo churn, revenue churn and contraction, benchmarked by pricing model.
  • Contract review quantifying change-of-control, MFN and SLA exposure in ARR terms.
  • Usage telemetry and integration evidence supporting the system-of-record assessment.
  • Tech debt inventory with remediation cost estimates and capitalized development analysis.
  • AI exposure memo placing the target in a disruption tier with supporting proof points.

Running these workstreams at this depth is where Plausity's platform earns its place in the process. Plausity supports the full software playbook described here: Commercial DD for ARR bridges, cohort retention and concentration analysis, Financial DD for unit economics and quality-of-earnings work, Tech DD for the technical debt and roadmap assessment, and AI Impact DD for placing the target in a disruption tier and testing its AI proof points. Every conclusion traces back to the underlying document, so the investment committee reads evidence rather than summaries of summaries. Deal teams preparing a live software transaction can arrange a platform demonstration with the Plausity team to see the playbook applied to a real data room.

How Plausity accelerates this workflow

Plausity is an AI-native due diligence and deal intelligence workspace that helps M&A advisory firms, VC and PE funds, corporate development teams and investment-banking teams structure evidence, findings and questions across a data room. Plausity supports evidence extraction, source grounding, findings management and IC preparation — it does not replace human analysts, advisers or investment professionals, does not provide legal, tax, audit, regulatory or investment advice, and does not make autonomous investment decisions. All findings require human review. Built for today's investment and deal teams. Trusted by >200 firms.

To explore the underlying capabilities, see the Plausity AI analysis engine, findings and risk intelligence and evidence gap detection product pages, plus the IC memo and AI Q&A Assistant product pages. For team-level workflows, see how VC and PE funds and M&A advisory firms use Plausity across live deals, and how AI Impact due diligence, value creation, Tech DD and Commercial DD workstreams support the analysis.

Sources

Frequently Asked Questions

PLAUSITY

AI Summary

Ask an AI assistant to summarise Plausity.