Software M&A in the AI Reset: How to Test Workflow Defensibility Before Paying a SaaS Multiple

Software M&A in the AI Reset: How to Test Workflow Defensibility Before Paying a SaaS Multiple

Image: Plausity

Key Takeaways

  • 78% of dealmakers report that AI is fundamentally altering their M&A assessment criteria.
  • industry analysts predicts 40% of enterprise SaaS spend will shift toward usage or outcome-based pricing by 2030.
  • Diligence must assess multi-party workflow orchestration and system-of-record status to determine true AI defensibility.

The AI Reset: How Generative AI Changes Software M&A

How does AI change software M&A? It invalidates the single question that anchored a decade of software diligence: how sticky is the software? Usage logs, renewal rates and low churn still matter, but they measure the past. Generative AI can now reproduce interfaces, draft the same outputs and execute parts of the same workflow without the vendor's product in the loop, so a target's historical retention is no longer proof that its revenue will survive the next contract cycle. The question that replaces it is harder and more useful: what does this software uniquely control when AI can reproduce parts of the interface or the workflow around it?

This is not a fringe concern. In KPMG's M&A Pulse survey of senior dealmakers, 78% say AI is changing their M&A assessment criteria, yet there is still no established approach for incorporating AI's impact on target defensibility into diligence, valuation and thesis development. The market has accepted that the risk is real faster than it has built the instruments to measure it. That gap is where deals are mispriced in both directions: buyers overpay for workflow software that an AI-native entrant can hollow out, and walk away from genuinely defensible platforms they could not distinguish from wrappers.

KPMG's survey also shows where dealmakers believe the durable defenses actually sit. Only two characteristics command majority support as indicators of AI defensibility: regulatory, compliance or security barriers (55%) and workflow integration (52%). Beyond those two, opinions fragment across proprietary data, domain expertise, switching costs and network effects. The practical implication for a deal team is that defensibility can no longer be asserted in a management presentation. It has to be tested, evidence by evidence, across a structured set of diagnostics.

  • Shift the diligence question from 'does the target use AI?' to 'does AI strengthen or erode the drivers of this business?'
  • Treat historical retention as a necessary but insufficient data point; test what the platform controls that an AI agent cannot bypass
  • Score every defensibility claim against documentary evidence in the data room, not management narrative
  • Price the target on the durability of its future cash flows under AI substitution, not on category benchmarks alone

The sections that follow set out those diagnostics in the order a software-focused deal team should work through them: system-of-record status and proprietary data, workflow ownership and integrations, seat compression and pricing durability, a consolidated scorecard, and the substitution risks that determine whether a SaaS multiple is supportable at all. For advisers building this into a repeatable diligence workstream, our earlier analysis of AI-native software moats and due diligence covers the underlying framework in depth.

System-of-Record Status and Proprietary Data Moats

The first diagnostic is the oldest and still the most discriminating: is the target the system of record for its domain, or merely a tool that touches the domain? A system of record owns the ledger of truth. It is where the customer's contracts, claims, filings, patient records, shipments or financial positions formally live, and where other systems and people must go to establish what is actually the case. Tools that sit beside the ledger, however elegant, are replaceable by whatever interface is cheapest this quarter. Platforms that are the ledger are not, because replacing them means migrating the customer's canonical data and re-certifying every process that depends on it.

Foundational models are trained on public text and code. They have no native access to a target's accumulated transaction history, its customer-specific schemas, or the network effects that form when counterparties, auditors and regulators all read from the same system. That exclusivity, not model quality, is what an AI-native competitor cannot casually replicate. When evaluating the data moat, deal teams should distinguish carefully between data the target merely processes and data it uniquely accumulates and can lawfully productise.

  • Ownership: does the vendor hold write authority over the customer's system of record, or does it sync from someone else's?
  • Data exclusivity: does the platform accumulate data no third party can access, such as cross-counterparty activity, multi-year outcome histories or regulated records?
  • Schema depth: are the data structures and taxonomies embedded in customer processes, so that migration is a re-engineering project rather than an export?
  • Network effects: do all parties to the workflow converge on the platform, so that its value grows with each additional participant?
  • Rights to productise: does the contract and privacy posture actually permit the vendor to train or differentiate on this data?

The failure mode here is data theatre: a large volume of customer data that is either generic (public-equivalent information any model can approximate), borrowed (ingested from a customer's other systems under restrictive terms) or legally encumbered. Diligence should therefore test the moat with the same rigour applied to revenue: sample the contracts, inspect the data flows, and establish which data assets would survive a hypothetical AI-native competitor with unlimited model access. Our analysis of workflow, data and replication risk sets out the evidence base for this test.

Workflow Ownership and Embedded Integrations

The second diagnostic asks whether the target orchestrates a multi-party workflow that no LLM can casually replicate. Interface depth is now cheap; orchestration depth is not. A single user asking an AI assistant to draft a document is exactly the interaction layer AI commoditises. But a workflow that routes a claim from an adjuster to an external assessor to a compliance officer, enforces approval hierarchies, timestamps each decision for audit and reconciles outcomes against a ledger is a different asset. Reproducing it requires not just model capability but the customer's process re-engineering, change management and counterparty onboarding, which is why 52% of dealmakers in KPMG's survey cite workflow integration as a primary defense and 55% cite regulatory or compliance barriers.

Decision rights are the most under-examined component. Where the software encodes who may approve, reject or escalate, and where those rules are configured into the customer's own control environment, the product has become part of the customer's governance. Displacing it means re-papering internal control documents, retraining approvers and, in regulated functions, re-validating the control with audit and compliance functions. That is a multi-quarter programme with internal political cost, which is precisely the friction that defends revenue.

Integrations deserve equal scrutiny, and the depth spectrum matters more than the count. A read-only REST connection is decorative; write access into the customer's ERP, event-driven flows that trigger downstream processes, and certified bi-directional syncs constitute structural embeddedness. Ask which integrations were commercially negotiated, which carry certification or partner-programme status, and how long a replacement vendor would need to rebuild them. Swap breadth for depth and a seemingly integrated product can be unbundled in a quarter; genuine write-level orchestration typically cannot.

  • Map every integration by depth: read-only, write, event-driven, certified partner programme
  • Identify which workflows involve external counterparties, since multi-party adoption is the hardest thing for an AI-native entrant to replicate
  • Verify where the product's approval and escalation logic is documented in the customer's own control framework
  • Check regulated workflows specifically: compliance-owned processes resist substitution far longer than productivity tools

Seat Compression Risk and Pricing Durability

The third diagnostic tests whether the revenue model survives AI-driven cost compression, and it is where the reset is most quantified. industry analysts estimates that up to $234 billion of enterprise application spending is exposed to what it calls agentic arbitrage between now and 2030, roughly 20% of enterprise application SaaS spending, as AI agents complete tasks across systems and reduce the need for users to interact with traditional interfaces. industry analysts is explicit that this breaks the link between user growth and revenue growth for many enterprise software vendors. For a target priced on per-seat contracts, that is a direct challenge to the ARR base the multiple is applied to.

The mechanism is seat compression: when AI agents do the work of three analysts, the customer needs one licence, not four. industry research's analysis of generative AI's economic potential finds that about 75% of the value generative AI use cases could deliver falls across four areas: customer operations, marketing and sales, software engineering, and R&D, which are exactly the enterprise functions that anchor most SaaS workflows. A buyer should therefore model seat-count sensitivity the way it models churn: which customers would need fewer seats if they deployed agents against this workflow, and over what horizon.

Pricing durability is the other half of the same test. industry analysts predicts that by 2030 at least 40% of enterprise SaaS spend will shift to usage-, agent- or outcome-based models, with seat-based revenue share declining. Vendors are already layering AI consumption charges on top of existing seat pricing: industry analysis's analysis of more than 30 major SaaS vendors found roughly 65% have done so. Diligence should establish whether the target's contract structure can migrate toward usage or outcome pricing without margin erosion, and whether its AI features are generating new monetisation or merely defensible discount leverage. A deeper treatment of seat, usage and outcome models in diligence is in our pricing-focused analysis AI software pricing due diligence.

  • Map every revenue line to its pricing unit: seat, account, transaction, usage or outcome
  • Identify which workflows an agent could execute end-to-end, and estimate the seat exposure per customer cohort
  • Test whether the target can raise effective price per unit of work as headcount falls, or whether AI pricing pressure compresses it instead
  • Review AI feature bundling and uplift clauses in recent renewals for signs of forced migration

Customer Concentration, Roadmap Credibility and the Remaining Dimensions

The last two diagnostics frame the others. Customer concentration determines how much any single defect matters: a defensible platform whose revenue sits in a handful of large accounts carries a very different risk profile from the same scorecard spread across a diversified base, because AI-era repricing tends to arrive first where the largest customers hold the most negotiating leverage. Roadmap credibility then asks whether the target can deliver an AI-native version of its own product before a competitor delivers it for them. Examine engineering capacity against the AI initiatives in the plan, the delivery track record against prior roadmaps, and whether shipped AI capabilities exist in the product today or only in the investor deck. A target that scores well on every structural dimension but cannot execute its roadmap is holding a defensible position on a depreciating clock.

Two further dimensions separate durable friction from assumed friction. Switching costs are frequently asserted but rarely evidenced: the test is what verifiably happens when a customer leaves, which means migration project records, actual exit timelines and win-back data rather than management's estimate of re-training pain. Usage depth and workflow penetration are the companion test: a platform licensed broadly but used shallowly, with dormant seats and single-team adoption, has penetration risk that mirrors seat compression risk, while daily use across roles, modules and regions signals a workflow the customer has genuinely built its operation around. Both dimensions appear as rows in the scorecard below, but they deserve standalone diligence because they are the two places where a management presentation most often diverges from the data room.

The SaaS Defensibility Scorecard for AI-Era Diligence

The diagnostics above consolidate into a working scorecard. Its purpose is not to produce a valuation, it is to force every defensibility claim onto a single scale so that deal teams, investment committees and sellers are arguing from the same evidence. Score each dimension from 5 (strongly defensible, evidence-backed) to 1 (claim does not survive contact with the data room), then let the aggregate pattern inform how much of a SaaS multiple the cash flows can carry. A consistent framework matters: KPMG finds that organisations currently tailor AI defensibility assessments to individual transactions and deal teams rather than applying a consistent evaluation approach, which is precisely where diligence quality diverges. The same discipline applies to commercial diligence on ARR durability.

DimensionCore questionEvidence to requestScore anchorsImplication for valuation ceiling
System-of-record statusIs the platform the ledger of truth for the domain?System architecture documentation, data flow diagrams, integration inventory5: canonical ledger with write authority. 3: syncs from another system. 1: parallel toolLedger status supports the multiple; parallel-tool status caps it at an application-tool multiple
Proprietary data moatDoes the platform hold data no foundation model or competitor can access?Data inventory, training rights in customer contracts, privacy and DPA terms5: exclusive cross-counterparty data with productisation rights. 1: generic or legally encumbered dataExclusive data justifies a premium; encumbered data removes the moat from the thesis
Workflow ownershipDoes the target orchestrate a multi-party workflow an LLM cannot casually replicate?Process maps, counterparty adoption data, approval-flow configuration exports5: multi-party, externally onboarding workflow. 3: single-enterprise workflow. 1: single-user utilityMulti-party orchestration supports durability; single-user tools face compression
Embedded integrationsAre integrations write-level, event-driven and certified, or decorative?API documentation, partner-programme certifications, integration usage telemetry5: write access with event-driven flows. 3: scheduled syncs. 1: read-only connectionsWrite-level embeddedness raises switching costs; shallow APIs leave the product unbundlable
Compliance and regulated statusIs the workflow owned by compliance or audit functions?Regulatory certifications, control documentation, audit findings, customer approval workflows5: product is the control itself. 3: supports a regulated process. 1: unregulated convenience layerRegulated workflows defend revenue longest and support the multiple accordingly
Switching costs (real vs mythical)What verifiably happens if the customer leaves?Migration project records, contract terms, renewal and win-back data, churn interviews5: documented multi-quarter migrations. 3: moderate friction. 1: export-and-leaveReal switching costs defend the base; mythical ones invite AI-era repricing at renewal
Usage depth and penetrationHow much of the customer's workflow runs through the product?Seat utilisation, feature adoption telemetry, module-level usage by role and region5: daily use across roles and modules. 3: single-team usage. 1: dormant licencesDeep penetration supports expansion revenue; shallow use invites displacement
AI-native substitution riskWhich parts could a competitor rebuild in 12-24 months?Build-versus-buy reconstruction estimate, dependency list, code and architecture review5: core is data and orchestration, not interface. 1: core is a UI over a public modelHigh substitution risk caps the multiple regardless of growth
Pricing durabilityCan the model migrate from seats to usage or outcomes without margin erosion?Contract structure, renewal uplift history, AI feature monetisation data5: outcome-aligned pricing in market. 3: hybrid transition planned. 1: pure per-seat exposurePricing flexibility protects unit economics; seat-only exposure invites compression
Customer concentration and roadmap credibilityDo the top customers depend on the roadmap, and can the team deliver it AI-natively?Revenue concentration analysis, delivery track record, AI engineering capacity5: diversified base with shipped AI capabilities. 1: concentrated base and slid roadmapConcentration plus weak velocity compounds every other risk above

Scored honestly, the pattern matters more than any single number. A target that scores 5 on system-of-record, data moat and regulated status but 1 on pricing durability is a defensible asset with a revenue-model problem, which is a negotiable risk. A target scoring 3 across the board is the more dangerous profile, because it will present well in management meetings while every layer of its moat is simultaneously erodible. The scorecard is also where switching-cost myths get exposed: claims of high switching costs should be reconciled against actual migration timelines and churn-and-win-back data, a test our scorecard-and-evidence analysis develops in detail moat diligence scorecard.

AI-Native Substitution and Service-to-Software Risks

The final diagnostic stress-tests the product against a capable AI-native entrant with a 12-to-24-month build horizon. The exercise is deliberately concrete: decompose the product into interface, orchestration logic, data assets, integrations and compliance surface, then estimate for each layer how long a well-funded team using foundation models would need to reach feature parity. Interfaces and generic content generation now measure in weeks. Orchestration logic calibrated to a specific industry's decision rules measures in quarters. What remains stubbornly slow is accumulated proprietary data, certified integrations and regulatory validation, which is exactly why the earlier dimensions carry the most weight in the scorecard.

Buyers should run this substitution test against evidence rather than instinct. Build a minimal working replica of the target's core workflow using current AI tooling and time it. If a deal team can reproduce the demo in a sprint, an AI-native competitor can reproduce the product roadmap, and the multiple being discussed is pricing in durability that does not exist. The same logic applies to the target's services revenue, where service-to-software substitution is now a two-sided risk: AI is automating the implementation, customisation and managed-service work that justified blended gross margins, while simultaneously making it easier for competitors to eat those services with agents. Diligence should separate recurring software revenue from AI-exposed services revenue and margin-test each on its own economics hybrid business models.

The market is already repricing for these risks, meaning these diagnostics are closely tied to the final deal structure. Premium multiples are flowing to businesses with durable growth, strong cash flow and defensible AI capabilities, while basic AI integrations drive no premium at all. In practice, the scorecard does not change whether a SaaS premium applies; rather, it changes how much of it survives diligence. A structured review of architecture, security and technical debt remains the backbone of that test tech due diligence checklist.

  • Decompose the product into interface, orchestration, data, integrations and compliance layers
  • Time-box a working replica of the core workflow with current AI tooling and record what resisted reproduction
  • Separate software revenue from AI-exposed services revenue and stress-test each margin structure
  • Benchmark the implied multiple against the disclosed market range and let the scorecard justify the premium or the discount

Modernizing Tech and Commercial DD with Plausity

Testing fifteen defensibility dimensions across a compressed deal timeline is an evidence problem before it is a judgement problem. The evidence lives in thousands of VDR documents: customer contracts that reveal switching terms and data rights, architecture documentation that shows which integrations write and which merely read, usage reports that expose seat compression risk before the model does. Plausity was built for exactly this workload. Its Data Room Ingestion connects to virtual data rooms and processes PDFs, spreadsheets, contracts and financial models within minutes, and the AI-Analysis Engine reads, cross-references and reasons over that material to produce source-grounded findings, every claim traceable back to the document it came from rather than asserted from management narrative.

In practice, deal teams use the platform across AI Impact DD, Commercial DD and Tech DD as one connected workstream. The Risk Radar evaluates findings by materiality, financial impact and deal relevance, which is how seat-compression exposure, mythical switching costs or an unbundlable integration surface get surfaced as scored risks rather than buried in a data room annex. The Collaboration Hub aligns the deal team, advisers and workstreams around those findings in real time, and the Report Builder drafts investor-ready deliverables with full source traceability, so the defensibility scorecard a committee sees is the same one the analysis produced. The consistent framework matters: KPMG's research finds that organisations applying a disciplined, repeatable approach to AI defensibility are better positioned to improve diligence quality and valuation confidence. Software-focused deal teams can arrange a walkthrough of the platform against a live data room via the demo page on the Plausity site.

How Plausity accelerates this workflow

Plausity is an AI-native due diligence and deal intelligence workspace that helps M&A advisory firms, VC and PE funds, corporate development teams and investment-banking teams structure evidence, findings and questions across a data room. Plausity supports evidence extraction, source grounding, findings management and IC preparation — it does not replace human analysts, advisers or investment professionals, does not provide legal, tax, audit, regulatory or investment advice, and does not make autonomous investment decisions. All findings require human review. Built for today's investment and deal teams. Trusted by >200 firms.

To explore the underlying capabilities, see the Plausity AI analysis engine, the findings and risk intelligence and evidence gap detection product pages, and the IC memo product page. For team-level workflows, see how VC and PE funds and M&A advisory firms use Plausity across live deals, and how AI Impact due diligence and value creation workstreams support the analysis.

Sources

Frequently Asked Questions

PLAUSITY

AI Summary

Ask an AI assistant to summarise Plausity.