The Four Layers of the AI Industry
Compute, models, platforms, and applications. See what each layer sells, pays for, and builds its moat on.
The industry splits into four layers. Once you see what each layer sells, who pays whom, and where the barriers sit, you can answer one practical question: does the product I'm building actually have a moat?
What you'll run into:
- A model vendor cuts prices once, and your carefully calculated cost model is obsolete
- A feature you built becomes a built-in model capability three months later
- Your boss asks "what's our moat?" and all you can say is "we integrated early"
Who sits on each layer
Compute sits at the bottom: chip designers and clouds that rent out GPUs. This layer is marked by tight supply and hard substitution, which gives it the strongest pricing power and the thickest margins. The industry proverb "the best money is made selling shovels" refers to exactly this layer.
Models are the vendors that train and sell models. This layer burns the most cash: training buys compute, and inference keeps buying compute—the money passes through and lands on the layer below. With capability gaps narrowing, price wars rage hardest here. A single page on OpenRouter lists hundreds of models for sale at publicly comparable prices. openrouter-models
Middleware is everything that makes models usable: model platforms in the cloud, vector databases, agent frameworks, evaluation and observability tools. This layer trains no models; it earns by lowering the cost of integration and providing engineering capability.
Applications sit on top, facing end users directly. This layer is closest to the money—user payments enter here—but its technical barrier is the lowest.
Money flows downward
Follow the money and the diagram comes alive.
Users pay applications; applications pass part of it to model vendors for API access; model vendors pass a large share to cloud and chip vendors for compute. Every layer is the customer of the layer below. That's why revenue gets more concentrated and more stable the lower you go, while the higher you climb, the more players and the more fragmentation. claude-pricing
This structure yields a few conclusions you can act on.
First, in applications, your cost structure isn't in your hands. A single upstream price change, rate limit, or model retirement forces you to rework margins and experience. So when choosing a provider, don't look only at today's price—look at their ability to keep supplying, and be ready to switch from day one. Stanford HAI's AI Index Report is a useful baseline for overall industry investment and pricing trends. hai-index
Second, the boundary of model capability keeps eating application-layer features. The prompt engineering, document parsing, and format fixing you carefully built today may be built into the next model generation. The test is simple: if your product's value comes mostly from "patching the model's weaknesses," its shelf life is the next model update.
Third, what actually survives is almost never inside the model. Proprietary data, embedding into real workflows, industry trust and compliance credentials, and users' accumulated habits—none of these get carried away by a model upgrade. The moat has always lived in these places; AI didn't change that.
Where you stand
Most product people sit in applications, a few in middleware. Once you place yourself, the metrics to watch differ completely:
| Your layer | What you're really selling | The risk to watch |
|---|---|---|
| Applications | Finishing a specific job; users pay for the outcome | Features absorbed by the model; upstream price hikes eating margins; commoditized competition |
| Middleware | Helping others integrate faster, run more reliably, see more clearly | A cloud vendor turns your feature into a built-in toggle on their platform |
| Models | Capability itself, sold per token | Training cost and price wars squeezing from both sides; open-source models dragging the price floor down |
| Compute | Compute itself, sold per GPU-hour or instance | Cyclical supply-demand swings; customers designing their own chips |
One last note: the same company can sit on several layers at once. Large cloud vendors sell GPUs, train their own models, and offer model platforms. So when you negotiate, first figure out which layer's identity the other party is wearing today—the price and terms they give you are often serving a business on a different layer.
Data baseline: August 2026. This layering exists to clarify business relationships, not as a strict statistical taxonomy. Specific players and market shares shift fast, so verify current public information before making decisions.
