The debt was already there
For fifteen years, B2B software sold outcomes and delivered systems of record. The pitch was pipeline predictability, faster collections, lower churn. The product was a place to type things and a place to read them back.
Assembling the outcome — connecting the record to the work, the work to the decision, the decision to the result — stayed the customer's problem. Ops teams exist because of that gap. We have made a living inside it.
Nobody minded much while software was priced per seat and judged on adoption. A system of record everyone logs into is a successful product by the only metric anyone was measuring.
Then the measurement changed.
Then everyone bolted an assistant onto it
An AI assistant is a very fast reader. Point it at a clean, connected, complete picture of a business function and it produces something close to judgment. Point it at the average GTM stack — three sources of truth for account, two definitions of a qualified lead, a close date field that is a negotiation rather than a fact — and it produces confident, fluent, wrong answers at speed.
That describes most of what shipped in the last two years. Menlo Ventures' read on incumbent AI is blunt: bolt-on features on legacy platforms. The bolt-on wasn't laziness. It was the only thing the underlying data model could support. You cannot ship reasoning on top of records that never agreed with each other.
The assistant didn't create the problem. It published it — for the first time, in a demo, in front of a buyer.
So the industry shipped more
Faced with a visible gap, software companies did the thing quarterly earnings reward. Feature velocity became the proof of life. Roadmaps that used to move in halves started moving in weeks. Release notes got longer than the product.
By early 2026, AI-assisted development was the default workflow, with roughly 85% of professional developers using AI coding tools weekly. The bill is arriving in a form ops teams recognize immediately.
- ~45%of AI-generated code carries an exploitable OWASP Top 10 vulnerability — a figure that has not improved in two years, despite dramatically better models.
- 6 → 35CVEs attributable to AI-generated code, January to March 2026. Near sixfold in a quarter.
- The 20%that gets skipped: error handling, permissions, observability, the states nobody demoed. The code mostly works. The missing part compounds.
Look at the shape of that failure. Something that works in the happy path, sold as something that works. It is the original SaaS promise again, one layer down.
Starting clean has its own failure mode
The obvious counter-move is to build without the legacy: no fifteen-year-old schema, AI in the foundation. A meaningful number of companies did exactly that, and built the product almost entirely with AI assistance.
They have the opposite problem, and it isn't obviously better. When a product is a prompt, a retrieval step, and a frontier model, every model release is an existential event rather than an upgrade. Sometimes the release absorbs the feature outright and the thing you charged for last quarter becomes a checkbox in the platform underneath you. Sometimes it merely shifts behaviour enough to break the scaffolding — and nobody notices for a week, because no test existed that would have caught it.
From a defensibility standpoint, no moat and a moat with a ninety-day expiry behave identically. Both make you re-earn the customer every quarter.
Three categories, one lesson
No names needed. The pattern is structural.
Outbound personalization
In 2023, a system that read a prospect's context and wrote a relevant first line was a product with a valuation. In 2026 it is a prompt, and the enrichment beneath it is a commodity API. The category didn't lose to a competitor. It lost to the model getting good enough that the feature stopped being a feature.
Meeting capture
Transcription accuracy was the differentiator, until it went to roughly free. Everything now competes on what happens after the transcript: which system it updates, which decision it changes, which risk it surfaces to which person. The moat left the recording and moved to the connections around it.
Support deflection
Every product in this category demos at high accuracy. Deployed against real ticket distributions, most fall over on the long tail. The survivors did not win on model quality. They won on retrieval scoped to one specific business, escalation rules that fire before a wrong answer reaches a customer, and an eval suite that catches regressions before customers do.
Where the capability was the product, the moat lasted about a quarter. Where the capability was wrapped in something specific to the customer, it held.
Enterprise buys insurable. Mid-market buys fast.
There's a widespread belief that large buyers prefer large brands because they think large brands have better AI. They don't. They prefer them because those brands absorb liability.
An enterprise buying committee is not optimizing to be right. It is optimizing not to be individually blamed for a breach, an audit finding, or a hallucinated number in a board deck. An incumbent arrives with certifications, indemnity language, data residency options, an audit trail, an existing MSA, and a name that makes the decision defensible in a post-mortem. That isn't conservatism. It is a correct read of the incentive.
Mid-market runs the other way and is equally rational. AI capability has become a primary filter rather than a nice-to-have, and these buyers name faster innovation as the reason they choose AI-native vendors. With no procurement apparatus to defend and no committee to satisfy, speed is the sane optimization.
Two consequences worth planning around. Sell to mid-market, and capability is your entry while liability absorption is your ceiling — you cross into enterprise the quarter you can indemnify, not the quarter you get better. Sell to enterprise, and your competitor is not a better model; it is a procurement process you either survive or don't.
Neither segment is buying intelligence. Intelligence is available to everyone at list price.
Three layers, in order of how often they get skipped
Function-wide context. Not app-wide. Function-wide. Revenue does not live in the CRM — it lives across CRM, product usage, billing, support tickets and call recordings, and an agent with access to one of those is structurally incapable of a complete answer no matter how good the model gets. The unit of work is the function, not the tool. This is unglamorous: entity resolution, one definition per metric, an explicit decision about which system is authoritative for which field. It is also the only layer that gets more valuable as models get cheaper.
Guardrails at the action boundary. Most AI governance conversation is about prompts. The risk isn't in the prompt. Reading is cheap; writing is where money, reputation and compliance exposure move. Every write an agent attempts should pass a fixed sequence: permission, evidence, authority at this value, human approval above a threshold, and a log an auditor can read without help. Build it once at the boundary and you can add agents without re-litigating risk each time.
A harness. Prompts, tool definitions, retrieval config, evals and fallbacks — versioned and tested like application code. Not for tidiness. Because it converts a model change from an incident into a config change. Almost nobody has this layer, and it is the difference between a stack that improves with each frontier release and one that is destabilized by it.
One measurement, hard to fake
The metric
Time to model swap
How long does it take to move your primary workflows onto a newly released model — and how do you know nothing regressed?
A team with function-wide context, action-level guardrails and a real eval harness answers in days, with evidence. A team that shipped features answers in months, or answers confidently and is wrong, which is worse. The metric is useful because it can't be gamed with a demo, it isn't a matter of opinion, and it correlates with everything else you care about.
Five questions in the same spirit — for a diligence conversation, or your own planning:
- Which system is authoritative for each field an agent can write to?
- What is the approval threshold above which an agent stops and asks a human?
- Where is the eval suite, and when did it last fail?
- What breaks if your model provider changes pricing or deprecates a version next month?
- Who can read the audit log without engineering help?
Two ways this argument fails
If integration stops being hard
The context argument assumes reconciling messy, contradictory enterprise data stays difficult. If models get good enough to do it reliably on the fly, careful context engineering becomes an expensive solution to a problem that solved itself, and the fast shippers were right. Watch model performance on ambiguous, conflicting internal data specifically — not on benchmarks.
If liability becomes purchasable
The insurability argument assumes procurement stays slow. If AI-specific certification and indemnity products commoditize the way SOC 2 did, the enterprise moat thins considerably and advantage swings back to raw capability. Watch whether coverage for AI outputs becomes something you buy rather than something you earn.
Both are live. Neither has resolved. We would rather say so than sell certainty.
The short version
Systems of record are a solved, commoditized problem. Model access is a commodity with a price list. Features built directly on model capability have a shelf life of about a quarter.
What is left is the context a model operates in, the rules governing what it is allowed to do, and the scaffolding that keeps it working when the model underneath it changes.
That's the moat. Less exciting than a launch, and the only part that is still yours next year.