Chapter 2.2 — Applications & Monetization
Adoption is broad; durable return on investment is not. The useful divide is not “AI winners versus AI laggards,” but trials versus production workflows whose output can be measured, trusted, and funded again. This chapter follows the budget from first experiment to recurring spend and asks which products become part of the work rather than another optional tool.
Applications are where AI becomes a product that a consumer or enterprise pays for. Durable application revenue ultimately supports spending on models, chips, data centers, and power. Enterprise returns vary sharply by workflow, so adoption must be evaluated task by task rather than treated as a single market.
Why high adoption can coexist with weak returns
The bearish case has a number attached. MIT's Project NANDA, in an August 2025 study, found that roughly 95% of organizations saw no measurable profit-and-loss return on their generative-AI pilots despite $30–40B of spending, with only about 5% of integrated projects extracting real value. The study's own explanation matters: the failures came from tools that did not fit workflows or retain context, not from weak models. That distinction is the key to the whole layer.
Set against that failure rate is an equally real explosion in spending. Enterprise AI spend reached $37B in 2025, up 3.2 times in a single year from $11.5B, the fastest growth in the history of enterprise software.
Of that $37B, applications received $19B, slightly more than half. Spending is concentrated: coding accounts for about $4B of the departmental total, healthcare $1.5B, and legal $650M. AI deals convert at roughly 47% versus 25% for traditional software. The data therefore supports strong returns in a limited set of deep workflows and weaker results in horizontal deployments. Startups now capture about 63% of application-layer revenue, up from 36% a year earlier, showing that existing distribution has not prevented new entrants from gaining share.
Adoption is not a yes-or-no event. A large insurer can give every employee an AI assistant and record heavy use for document summaries and email drafting, while its finance team remains unable to identify incremental profit. Ten minutes saved does not automatically become ten minutes of additional output; the time may be absorbed by other work.
Now imagine that the same insurer uses AI to examine a defined class of claims. The system gathers the relevant documents, checks them against policy rules, flags exceptions, and sends only uncertain cases to a human reviewer. Management can compare the old and new process: cost per claim, time to resolution, error rate, fraud loss, customer retention, and the number of claims each employee can handle. A narrower deployment can therefore have more economic value than a company-wide assistant, because its result has a denominator and appears in an operating metric.
This is the first reading rule for the application layer: count the business process, not the number of people who have opened the tool. A product becomes investable when repeated use changes a measurable outcome and the customer funds it from an operating budget, not merely an innovation budget.
A trial budget has to pass five gates before it becomes durable revenue
An enterprise application normally travels through five stages. Vendors often report them as though they were equivalent, which is why application announcements can sound more mature than the underlying business.
- Individual use. Employees try a tool on their own. This proves attraction, but the employer may not be paying and may not even know the tool is being used.
- Central purchase. A department buys seats or API credits. Revenue begins, but the contract may still be a reversible experiment.
- System integration. The product connects to identity, permissions, internal data, and existing software. Deployment becomes slower and more expensive, but switching also becomes harder.
- Workflow redesign. The customer changes approvals, staffing, service rules, or the sequence of work around the product. The application is no longer an optional assistant; it has become part of how the organization operates.
- Financial realization. Revenue, cost, error rates, cycle time, or working capital improve in a way that finance and operations can audit. The contract can now compete for recurring operating budget.
The first two gates produce user counts and press releases. The final three produce retention, pricing power, and eventually cash flow. A vendor with one thousand pilots and twenty production deployments may have a weaker business than one with one hundred customers and ninety production deployments. The second has fewer logos but a much larger share of each customer's real work.
The same path explains where an application moat comes from. Connecting permissions requires security review. Using internal records requires data governance. Taking an action requires an audit trail, an exception path, and someone who owns the error. Changing a workflow requires training and organizational agreement. A model provider can copy a feature quickly; it cannot instantly copy all the institutional work already completed inside a customer.
Where trials have become recurring production budgets
Commercial traction clusters by vertical:
| Company | Vertical | Scale (approx.) | Access |
|---|---|---|---|
| GitHub Copilot | coding | ~20M users, 4.7M paid | MSFT |
| Cursor (Anysphere) | coding | ~$1B ARR; $29B val | PVT |
| Claude Code (Anthropic) | coding | ~$2.5B run-rate | PVT |
| Cognition (Devin) / Windsurf | coding | ~$490M run-rate | PVT |
| Sierra | customer support | ~$200M ARR; >$15B val | PVT |
| Decagon / Intercom Fin | customer support | Decagon $4.5B val | PVT |
| Harvey / Legora | legal | Harvey ~$300M ARR / $11B val | PVT |
| Abridge | healthcare (scribing) | $117M contracted ARR | PVT |
| OpenEvidence | healthcare (clinical Q&A) | ~40% of US physicians; $12B val | PVT |
| Glean | enterprise search | ~$300M ARR; $7.2B val | PVT |
| Clay | sales / GTM | ~$100M ARR | PVT |
| Salesforce Agentforce | enterprise agents | ~$1.2B ARR | CRM |
| ServiceNow Now Assist | enterprise agents | ~$750M ACV | NOW |
| Palantir | applied-AI platform | AI-product revenue not separately disclosed | PLTR |
| Doubao / DeepSeek / Ernie / Yuanbao | China consumer | Doubao 382M MAU | PVT / BIDU / 0700.HK |
The values in this table do not all mean the same thing. Annual recurring revenue is contracted or recurring revenue expressed on a one-year basis. Annual contract value describes the value of contracts, not necessarily revenue recognized in the current period. A run-rate may simply annualize a recent month or quarter, and private-company estimates are not audited public filings. The table is useful for locating demand, not for ranking business quality.
The more important pattern is concentration by workflow. Coding, customer support, legal research, clinical documentation, and enterprise search all have a defined user, a high cost of labor, and an output that can be checked. Horizontal assistants reach more employees, but they often save fragments of time that are difficult to convert into a budget owner’s profit-and-loss account. The application market is therefore broad in usage and narrow in proven economics.
Coding shows how capability becomes revenue
Most fast-growing application companies remain private. The common feature among those with meaningful revenue is control of a specific, high-value workflow from input to measurable result.
In coding, the standouts are Anthropic's Claude Code, at roughly a $2.5B annual run-rate by early 2026, and Cursor (Anysphere), which crossed $1B in annualized revenue and raised at a $29.3B valuation, with GitHub Copilot's 20M-plus users and Cognition's Devin (which absorbed Windsurf) rounding out the field. In customer support, Sierra reached about $200M in revenue at a valuation above $15B, and Decagon tripled to a $4.5B valuation, both on outcome-based pricing. In legal, Harvey grew to about $300M of run-rate at an $11B valuation, with Legora close behind. In healthcare, Abridge's ambient clinical documentation reached $117M of contracted revenue across Kaiser, Mayo, and Johns Hopkins, and OpenEvidence, used by roughly 40% of US physicians, doubled to a $12B valuation. In enterprise search and sales, Glean tripled to over $300M and Clay reached around $100M with net revenue retention above 200%. The common thread is defensibility through workflow depth and proprietary data, not model access.1234
Coding is the cleanest example because four conditions line up. Developers are expensive, so an hour saved is valuable. Code can be tested, which makes quality less subjective than the quality of a marketing paragraph. Mistakes can usually be caught before deployment and rolled back. And the product lives inside the editor, repository, terminal, testing system, and deployment process, so repeated use creates context and switching friction.
That does not make every coding assistant a durable company. It shows what a commercially mature AI workflow looks like. The product does not sell an impressive answer; it occupies a recurring decision point inside work, produces output that can be verified, and gives the customer a way to calculate return. Legal, healthcare, and customer support can achieve the same depth, but their error costs, regulation, and implementation cycles are heavier. The economic prize may be larger, while the route to revenue is slower.
Outcome pricing expands revenue and transfers risk to the vendor
Pricing is shifting as agents perform tasks previously assigned to people. Seat-based pricing fell from 21% to 15% of vendors in a year, while hybrid and usage-based models rose to 41%. Outcome-based pricing goes further: Intercom's Fin charges $0.99 per resolved ticket, and Decagon prices per resolution. Vendors that can verify completed work may address labor budgets, while incumbents dependent on seat count face pressure to change packaging and prove incremental value.
The larger addressable market comes with a less forgiving income statement. Under seat pricing, the vendor gets paid even when the customer uses the product lightly. Under outcome pricing, the vendor may absorb model inference, data retrieval, third-party APIs, implementation, human review, and the cost of failed attempts before it earns one billable result. The relevant unit economics become:
Contribution profit per outcome = revenue per successful outcome − model and infrastructure cost − human review and exception handling − implementation and support
Lower model prices help only if the vendor keeps the savings. A competitive market may pass them straight to the customer. Higher automation helps only if exception handling does not rise at the same time. Investors should therefore ask for contribution profit, human-escalation rates, and cost per completed task—not merely the number of tasks processed.
Agents must clear a reliability gate before they clear a budget gate
The next phase of this layer is agents, software that takes actions rather than just answering, and the early revenue is genuine. Salesforce's Agentforce reached a $1.2B run-rate, up more than 200% year on year, and ServiceNow's Now Assist passed $750M in annual contract value and restructured its entire product line around AI tiers. Real deployments show the potential: Klarna's OpenAI-powered assistant handled 2.3M conversations, the equivalent of 700 human agents, and cut resolution time from eleven minutes to under two, an estimated $40M profit improvement, though Klarna later rebalanced toward a human-plus-AI hybrid after over-automating.5
The constraint, again, is reliability rather than capability, and it is the same gate described in Chapter 2.1 from the model side. Gartner projects that more than 40% of agentic-AI projects will be canceled by the end of 2027 on cost, unclear value, and inadequate controls, even as it forecasts that a third of enterprise applications will embed agents and 15% of day-to-day work decisions will be autonomous by 2028. Only about 16% of enterprises run "true" plan-and-adapt agents today; the rest are fixed workflows. The signal to watch is the conversion rate from pilot to production, which remains in the single digits for genuine autonomous agents.
Distribution wins the first user; workflow control retains the profit
The national contrast in this layer is about business model as much as capability. The United States monetizes applications directly, through enterprise API contracts and consumer subscriptions; ChatGPT alone has around 50M paying subscribers, and the $19B of enterprise application spend flows to a dense field of $100M-to-$1B revenue startups. China largely does not charge. Its leading assistant, ByteDance's Doubao, reached 382M monthly users, and DeepSeek's app around 130M, but the dominant model is free access funded by advertising and the broader super-app ecosystem. Baidu made its Ernie assistant fully free in 2025 after subscriptions failed, and ByteDance only began piloting paid Doubao tiers in mid-2026, priced from around $10 a month.
China's application layer generates large usage but comparatively little direct subscription revenue. Monetization occurs mainly through platform ecosystems and advertising rather than the standalone application revenue more common in the United States.
A wrapper fails when the model takes over its control point
A recurring concern is that many AI applications are thin interfaces over third-party models and have little protection from competition. Application gross margins are estimated at 50–60% versus 70–90% for traditional software, and new model features can quickly absorb standalone products. Counter-evidence includes $19B of enterprise application spending, startups' 63% share of application-layer revenue, and more than ten applications above $1B in annual recurring revenue. Model access itself is widely available; distribution, proprietary data, workflow integration, and switching cost determine whether an application retains value. Products that own a specific workflow are better protected than features a model provider can reproduce.
The evidence investors should wait for
Enterprise returns need to broaden beyond today's strongest workflows. Relevant evidence includes pilot-to-production conversion, any follow-up to MIT's 5% success finding, and whether coding and vertical applications sustain growth and margin as model prices fall. Agent reliability also matters: dependable production use would support longer contracts and more recurring revenue.
Where application economics can endure
Many fast-growing application companies remain private and are available to public investors only indirectly or after an IPO. Listed exposure comes mainly from incumbents with existing distribution: Salesforce (CRM) and ServiceNow (NOW) in enterprise agents, Microsoft (MSFT) through Copilot, and Palantir (PLTR) as a more concentrated, higher-valuation applied-AI company. Anthropic's coding-led profitability, indirectly represented through Amazon and Google, supports the case that deep workflows can monetize. Companies that control a valuable workflow, proprietary data, and customer distribution should retain more margin than products whose main asset is access to a third-party model.
Public-market exposure also needs to be sized honestly. A large software company can report strong AI bookings while the new product merely replaces revenue that would previously have arrived through ordinary seat expansion. A cloud or platform company can benefit from application usage while simultaneously funding the data centers that make the usage possible. A strategic investment in a private model or application company does not give public shareholders a dollar-for-dollar claim on that company’s valuation.
For each listed company, the useful sequence is therefore: Is the customer paying separately for the AI product? Does the product increase the total contract rather than relabel an existing one? Who bears inference cost? Does AI improve or dilute gross margin? And is the new profit large enough to matter to the parent company’s valuation? These questions prevent a real application success from becoming an exaggerated security thesis.
What would settle—or break—the ROI case
The constructive case weakens if enterprise returns remain confined to coding and a few verticals, if MIT's 95% pilot-failure finding does not improve, or if Gartner's forecast that 40% of agent projects will be canceled proves accurate. It strengthens if reliable agents reach broad production use and outcome-based pricing captures part of labor budgets. Those outcomes can be tested through production conversion, renewal, gross margin, and measured customer profit.
The chapter’s conclusion is deliberately narrower than “AI software will win.” Application value becomes durable when a product clears four tests at once: it enters production rather than remaining a pilot; it controls enough context and workflow that replacing it is costly; it accepts responsibility for an outcome the customer can measure; and the customer renews at equal or greater spend after seeing the result.
That gives the investor a practical quarterly sequence. First, look for pilot-to-production conversion. Next, look for deeper data and permission integration. Then examine retention, contract expansion, and the share of tasks completed without expensive human rescue. Finally, compare customer ROI with the vendor’s own contribution margin. If only usage rises, the story is still adoption. If all four improve, adoption has begun to become a business.
Sources
Linked evidence for this chapter's figures and load-bearing claims: 1 2 3 4 5
Footnotes
-
Anthropic raises $30 billion in Series G funding at $380 billion post-money valuation. Anthropic, 2026-02-12; accessed 2026-07-25. ↩ ↩2
-
Cursor Recurring Revenue Doubles in Three Months to $2 Billion. Bloomberg, 2026-03-02; accessed 2026-07-25. ↩ ↩2
-
Glean Surpasses $300M ARR. Glean, 2026-05-28; accessed 2026-07-25. ↩ ↩2
-
Big Ideas 2026. ARK Invest, 2026-02-01; accessed 2026-07-25. ↩ ↩2
-
2025: The State of Generative AI in the Enterprise. Menlo Ventures, 2025-12-19; accessed 2026-07-25. ↩ ↩2