The model is commoditizing and value is migrating to the application layer. Most firms are buying the wrong one.
Where enterprise AI actually stands
Walk into a large enterprise in mid-2026 and the picture is remarkably consistent. Copilots are deployed across every function, adoption dashboards are green, and spend is climbing. Ask the people doing the work what has changed, and the honest answer is usually: not much. AI is everywhere and has changed very little, because most of what firms bought is pass-through. The same model, accessed through a vendor’s interface, produces the same artifacts the team was already producing, somewhat faster. The productivity gain is real, but it accrues identically to every buyer of the same product, and it leaves when the contract does.
This piece is about why that gap exists and what to do about it. The argument, stated up front: the model layer is commoditizing, and value is migrating to the application layer. The application layer most firms have today was built to bill for usage, not to capture anything. Firms that build a capture-first application layer will compound an asset no competitor can buy. Firms that keep buying pass-through will not. And the path to the thing everyone now wants, a model of the firm’s own, runs through a sequence most firms are executing in the wrong order.
The state of the market
Three forces are moving value away from the model layer, where most executive attention has been spent.
First, public benchmarks are saturating faster, while leading models are clustering more tightly on head-to-head evaluations. The 2026 Stanford AI Index reports four developers within 25 Arena Elo points and notes that benchmarks intended to last for years now saturate in months. A leaderboard position confirms that a model is in the capability band; it does not decide which one is right for a specific workflow. [1]
Second, the model itself is commoditizing. Published API prices span a wide range even among highly capable models, while the 2025 AI Index measured a more than 280-fold drop in the cost of querying a model at GPT-3.5-level MMLU performance between November 2022 and October 2024. A capable model is becoming a utility input, and its price is still moving. [2] [3]
Third, adoption has outrun workflow integration. Stanford reports that 88 percent of surveyed organizations used AI in 2025 and 70 percent used generative AI in at least one function, while agent deployment remained in the single digits across nearly every function. An MIT field report similarly found that access to tools was widespread but production transformation remained difficult. The gap is no longer access to a model. It is the work required to connect that model to a measurable process. [4] [5]
Usage-based pricing also makes the buyer's economics depend on routing and workload design, not just a seat price. Satya Nadella has described the strategic risk directly: firms need to own a learning loop that compounds their human and “token” capital, instead of letting institutional knowledge become a commodity inside a small number of general models. Microsoft has a commercial position in that argument, as every model provider does. The useful test is operational: if two firms can rent the same model, durable differentiation has to come from what surrounds it. [6]
Where value accrues, and where it leaks
It goes to the application layer: the surface where the model meets the firm’s actual work, context, and judgment. That is the only place differentiation is possible, because that is where proprietary information lives and where processes actually change. Value accrues to whoever owns the application layer that captures it.
The problem is that the application layer most firms buy does not capture anything. Copilots, coding agents, assistant seats, and vertical tools are all pass-through. Every company runs the same model through a vendor-owned interface and produces the same artifacts its teams already produced. The gain is identical for every buyer. There is genuine demand for AI in the workflow, but nothing about the business changes: the model passes through, the artifact gets produced, and the work goes back to how it was. Value accrues to the application layer in theory and leaks out in practice, because a pass-through interface is not a capture surface.
The subtler version of this catches firms two years in. Fifteen copilots deployed across functions, dashboards showing healthy usage, a procurement narrative about AI everywhere. Underneath it, fifteen rented surfaces feeding fifteen vendor-owned learning loops, none of which the firm controls or can carry forward. Tokenmaxxing is not just overpaying for tokens. It is paying to build other people’s application layers while your own never gets built.
The frontier labs make this worse when they sell the application layer themselves. A router owned by the seller of the most expensive tokens cannot credibly optimize your cost down. The compounding asset accumulates in the lab’s environment rather than yours. And a lab racing to build the best general models cannot simultaneously act as the neutral custodian of the firms it learns from. Use the labs as an input. Do not let them be the custodian.
The right application layer
The answer is not another tool. Every tool on the market shares the same structural failure: it asks the firm to adapt its workflow to the tool, and it produces the same artifacts slightly faster instead of enabling the net-new processes that would actually change how the team works. The face of enterprise AI today is old artifacts with AI inside. The opportunity is new processes the team owns, processes that did not exist before and that improve with use. No tool that requires you to adapt your workflow to it can deliver that, because the workflow is the thing you are trying to reinvent.
What is needed is a workspace: an application layer where every team, not just engineers, can build use cases in an environment the firm owns, on context the firm controls, with proprietary signal captured by default. The capture layer is not a separate engineering project. It is the substrate the team builds on, which means every use case shipped is also an installment of the asset. The team describes the work; the platform handles routing, evaluation, governance, cost telemetry, and memory.
This answers the objection that matters most. Building a capture layer sounds like a lot to ask of a business that adopted AI to accelerate, not to become an AI infrastructure company. On the right platform, capture, evaluation, and governance are properties of the workspace rather than engineering projects. The firm does not need to become an infrastructure company. It needs an application layer that already is one.
What compounds
The durable asset is not the firm’s data. External data commoditizes and proprietary data ages. The durable asset is the tacit judgment in senior practitioners’ heads and in the flow of work: the exception handled a particular way, the draft corrected for a particular reason, the grounds on which an output was accepted or rejected. This signal is created continuously and captured almost nowhere. The capabilities that compound are the ones that capture it at the moment of creation and convert it into assets the firm owns.
Nadella frames this as two kinds of capital: human capital, the knowledge and pattern recognition of the firm’s people, and token capital, the AI capability the firm builds and owns. “You can offload a task, or even a job,” he wrote, “but you can never offload your learning.” The test is simple. Swap the model. Does the capability survive? If yes, the firm built genuine token capital. If no, it built a wrapper. [6]
Qualcomm ran a version of this test in production. The same foundation model went from a 23 percent task pass rate to 98 percent in four months, with no fine-tuning and no model change. Every point of the gain came from the application layer: retrieval over indexed documentation, learned memories, and an eval-gated improvement loop. The capability was never in the weights. It was in what the weights could see. [7]
Five capabilities compound, and together they form one lifecycle: routing to the cheapest sufficient model (one sponsor cut token cost 57 percent month over month while usage doubled); private evaluation scored against the firm’s own rubrics; observability into where AI value accrues, firm-wide and per team; ring-fenced sovereignty with an ejection path that keeps everything; and firm-specific models trained on the firm’s captured work. That last capability is the one everyone now talks about. It is also the one most firms get wrong.
Every company wants its own model
A thesis is forming in the more ambitious corners of the market: every company should eventually have its own model, trained on its captured work, reasoning from its first principles, owned outright. It is the logical terminus of the capture strategy. If the application layer is the firm’s proprietary judgment accumulated over time, then at sufficient volume that judgment can be encoded in weights the firm owns. The model becomes the asset.
Getting there requires four things, in order. Capture: the traces, rubrics, and preference data generated in the flow of work. Dataset: that captured signal turned into something a model can learn from, through curation, labeling, and structured annotation. Synthesis: much of the training data will be synthetic, generated from captured patterns, and this is a specialized discipline where model quality is largely determined. Training infrastructure: compute, orchestration, evaluation loops, and validation, a cloud-scale build that no financial services firm, manufacturer, or healthcare system should attempt in-house.
Most firms get the sequence backwards. They stand up infrastructure first and discover they have nothing to train on, because the capture layer was never running. It is building the factory before securing the materials. The right sequence is to start capture now, on a workspace that captures by default; let the dataset accumulate from work the team is already doing; and stand up training infrastructure, with a partner, once the capture is deep enough to justify it. The firm owns the capture and the dataset. The synthesis pipeline and the training infrastructure come from a partner whose business is building them, and the partner is present throughout: the workspace captures from day one, the platform turns capture into training-ready data as it grows, and the training infrastructure is there when the firm is ready. The firm never has to become an AI infrastructure company. It has to become a company that captures its own work.
One caution from the field before treating the model as the finish line. When Qualcomm distilled its captured work into small models it owns, the distilled models scored within two points of the frontier teacher at a fraction of the cost. Cut off from the live context graph, the same distilled model lost eight points. The weights are an artifact of the loop, not a replacement for it. Even when the firm owns the model, the loop is the asset.
Reinforcement learning as a service
The capabilities above are properties of a workspace built to have them. What makes them compound into genuine token capital is reinforcement learning as a service: the loop in which captured traces and expert feedback are fed back into the system so that every run improves the next one. This is not a feature the firm builds. It is a service the platform provides on top of the capture layer. Traces show what happened. Evals measure whether it was good. The reinforcement loop trains the system to produce more of the good and less of the bad, against the firm’s own definition of good. The firm owns the capture and the judgment; the platform runs the loop. This is why the partner matters throughout the lifecycle: the loop runs continuously on the firm’s captured work, and the infrastructure that runs it belongs to the partner.
Where forward-deployed engineers fit
Every vendor now sells forward-deployed engineers, and most firms end up using them as a permanent dependency. The FDEs land, build the workflows, and stay, and the capability runs only as long as the vendor’s engineers are in the room. That is a managed service dressed up as a partnership.
The right model is on, off, on. In phase one, FDEs embed alongside the first team for enablement and discovery, explicitly time-boxed, with a defined deliverable: a team that can build without them. In phase two, the FDEs step back, the team self-serves on the workspace, and it finds the use cases nobody scoped. In phase three, FDEs return only for the high-leverage moments, the work that compounds. The tell is in the contract. If the FDEs never leave, the firm is renting engineers. If the phases are scoped and named, the firm is buying a capability that compounds.
Deployment takeaways
The question is not which model to standardize on. The model is the cheapest and most fungible input, and the answer matters less every quarter. The questions that matter concern the application layer: who owns it, where the capture accrues, and whether the firm is building an asset or renting one. Five steps, in order.
1. Understand and consolidate your AI use cases. Today there is no governed way to know where the needle is moving inside most firms, only a scatter of invisible use cases and vertical tools with no cross-team visibility. Fix that first: see what is happening across the firm and where AI is creating or leaking value. Then identify the processes most ready to be documented and transformed, the workflows that are high-toil, well understood, and ripe for capture. You cannot prioritize what you cannot see.
2. Decide what to build versus what to externalize. If you were going to build a cloud, would you build your own, or build on someone else’s platform? The same decision applies here. Own your IT team and your data model. Externalize the data context layer, the training infrastructure, and the synthesis pipeline. The infrastructure is not the asset; building it in-house is building the factory before you have the materials. The asset is the capture. Own that, and externalize the rest.
3. Build on a platform that connects data, people, and context. If the platform separates data, people, and context across silo platforms or vertical tools, you are back to the fifteen-copilot problem. Once the platform is in place, move fast on capture: get the flow of work into it, and capture the workload by default. Then build the second layer, augmented workflows that create net-new ways of doing the work with AI. Not just automating the low-hanging fruit, but building new processes that leverage intelligence and that the team owns. The companies pulling ahead are not plugging into a tool handed to them. They are building net-new systems that turn the business itself into an asset.
4. Capture traces, private evals, and human expert feedback. This is the mechanism that compounds. Every workflow generates execution traces, the full path of tool calls, steps, and decisions; capture those by default. Then build private evals that score every run against the firm’s own rubrics, so you know whether the work is getting better, not whether the model is getting better on public benchmarks. Then capture the human expert feedback: the edits, corrections, overrides, and the reasons an output was accepted or rejected. That feedback is the firm’s judgment encoded as data, and it is the thing no competitor can copy. Traces show what happened, evals measure whether it was good, and feedback teaches the system what good means in your shop. Skip this step and you have a workspace. Do it and you have a capture layer that appreciates over time and eventually feeds the firm’s own model. One diligence question separates the two: ask for the failure half-life on your tasks, the number of eval-gated iterations that removes half of the remaining failures. A vendor running a real loop knows this number. A pass-through vendor has a benchmark slide.
5. Build toward continual enterprise intelligence. This is the end state, and it is an operating model rather than a project. Once the capture layer is running, the evals are scoring, the feedback is flowing, and the reinforcement loop is compounding, the firm has a system that gets better at its work every day without anyone retraining it manually, because the work itself is the training signal. The model is rented. The application layer is owned. The capture compounds. The firms that reach this state first will hold an asset that cannot be bought or copied, because it is built from their own work. The firms that do not will keep renting intelligence and paying for tokens while the learning walks out the door.
Sources and measurement notes
- [1]Stanford HAI, 2026 AI Index: Technical Performance — Frontier-model convergence, benchmark saturation, and benchmark reliability.
- [2]Stanford HAI, 2025 AI Index: Research and Development — Measured decline in inference cost at a fixed capability threshold.
- [3]Published provider pricing: OpenAI, Anthropic, and Google — Prices are a July 2026 snapshot and can change.
- [4]Stanford HAI, 2026 AI Index: Economy — Organizational adoption and agent deployment by business function.
- [5]MIT NANDA, State of AI in Business 2025 — Field study on adoption, production deployment, and organizational learning gaps.
- [6]Satya Nadella, “A frontier without an ecosystem is not stable” — Primary statement on human capital, token capital, and owned learning loops.
- [7]Context and Qualcomm deployment case study — Held-out evaluation methodology, deployment timeline, ablations, and measured production results.
- [8]Context deployment telemetry — The 57% routing-cost reduction is measured from an anonymized sponsor deployment; usage roughly doubled over the same period.