Software categories usually evolve slowly. A feature war here, a pricing shakeup there, the occasional acquisition. It is rare to watch an entire enterprise category tear up its value proposition and rebuild it in under two years. That is what has just happened in one of the most consequential software markets most technologists have never examined, and the mechanics of the reinvention are a case study in how regulation reshapes product.
The category, before
The market is software for healthcare risk adjustment: the systems American insurers use to extract, validate, and submit the diagnosis data that determines their government payments. Insurers covering older adults are paid according to how ill their members’ records show them to be, so for fifteen years the category’s pitch was uncomplicated: our platform reads charts faster and finds more codable diagnoses than the competition. Dashboards celebrated codes captured and revenue identified. Natural language processing was the engine; payment uplift was the promise.
Product roadmaps followed the pitch. Extraction accuracy improved relentlessly. The direction of extraction did not: platforms were architected to find diagnoses that could be added, not to flag recorded diagnoses that lacked support, because customers were not asking to find fewer billable conditions.
The eighteen months
Then the environment flipped. Federal auditors scaled to roughly two thousand certified coders running quarterly review cycles, with sample error rates extrapolated across entire contracts. Reviews published in March 2026 found 81 to 91 percent of certain sampled high-risk codes unsupported at three plans. The Department of Justice extracted a 117.7 million dollar settlement from a major insurer whose chart-review programmes, prosecutors argued, functioned as one-way code-addition machines. Regulators separately moved to discount diagnoses that could not be traced to genuine patient encounters.
Every one of those developments converted a selling point into a liability. “Finds more codes” became “manufactures audit exposure at scale.” The category’s buyers, suddenly answerable to auditors, rewrote their evaluation criteria almost overnight, and the software had to follow.
The category, after
Survey the leading AI risk adjustment platforms today and the product architecture reads like a different industry. The headline capability is bidirectional review: the same pass that surfaces missed diagnoses flags recorded ones lacking evidence, with removals treated as first-class outputs rather than reluctant afterthoughts. Every suggestion ships with its evidence, the source text in the clinical note, the documentation rule satisfied, the confidence, and the recorded human decision, so any output can be reconstructed under audit years later.
Explainability moved from marketing slide to load-bearing architecture; opaque models that emit conclusions without traceable reasoning are effectively unsellable in the category now, whatever their benchmark scores. Encounter linkage became a core data model concept, because diagnoses unanchored to real clinical visits are losing payment validity. And audit workflow, once an afterthought, is now a product surface of its own: evidence assembly, mock-audit tooling, response tracking against the government’s five-month record windows.
Even the metrics inverted. Sales decks that once led with revenue identified now lead with validated accuracy rates and audit outcomes. One widely cited benchmark in the category is out-of-box AI accuracy in the low nineties, with human-in-the-loop review pushing final accuracy higher, numbers that matter because buyers now sample vendor output adversarially during procurement, exactly the way an auditor would.
The pattern for every category
The reinvention holds three lessons that travel well beyond healthcare.
Regulation does not just constrain categories; it re-founds them. The vendors thriving today are largely those whose architecture happened to align with the new rules before they arrived, evidence trails, bidirectional correction, human checkpoints. Their moat is not a feature but a multi-year head start on structure competitors must now retrofit.
Buyer criteria can flip faster than roadmaps. Eighteen months is shorter than most enterprise release cycles. Categories exposed to regulatory risk should treat the compliance scenario as a product requirement while it is still hypothetical, because when it lands, the market does not wait for the next major version.
Anti-features become features. The capability nobody demoed five years ago, telling customers their data contains errors that reduce their revenue, is now the first thing sophisticated buyers ask to see. In any category where the customer’s incentives and the truth can diverge, productised honesty eventually commands a premium, usually right after the first headline settlement.
Enterprise software rarely offers natural experiments this clean. An entire category, forced mid-flight to swap its objective function from maximisation to defensibility, and completing the swap inside two years. The vendors that made it are shipping the blueprint every regulated AI category will eventually need. The ones that did not are the cautionary slide in someone else’s pitch deck.



