Models are plateauing, Capabilities are becoming a commodity - What should you do about it?

Part 1: Models are plateauing, Capabilities are becoming a commodity - What should you do about it?

TMLS Newsletter article May 21, 2026

Graham Toppin and Josh Goldstein

May 21, 2026

16

Written by Graham Toppin Co-chair, TMLS, Co-founder and Analyst at Peerlabs.ai.

Part 1: The Pattern

Apologies for not getting an AI in production out this past Monday. We had put together an extensive essay on logprobs and model / harness customization. A number of things happened to change our stance toward the end of the week.

Today, we’re going to discuss:

  1. Current trends in Generative AI and ML in general
  2. Possible implications to how we’re building systems
  3. What you can do today to plan for these events

While we’re going to speculate a bit, we’ll try to ground this in provenance you can reason about.

There’s a lot here, so this will be a two part essay.

We’re going to publish the articles on customization later this week, focused on open-weight and local models given the changes we’re seeing in Frontier models.

Let’s dive into it!

BLUF (Bottom Line Up Front) What you should do

To sum up:

  1. Generative AI is useful.
  2. The technology is improving, though at a reduced rate.
  3. The technology is also commoditizing, the financial structure supporting it is under stress, and the vendors are responding by restricting practitioner control.
  4. These things are all true simultaneously, and they all point in the same direction: build your AI practice on foundations you control.

If you read nothing else, read this section.

Optionality is the strategy

Frontier models’ value propositions are shifting from “irreplaceable capability” to “convenient, well-integrated, professionally supported.”

This is for a few reasons:

More concretely, this means:

If you’re interested in how we came to these conclusions, read on!

What we’re seeing

The TMLS Steering Committee tracks and actively discusses trends in AI and Machine learning. In the last few weeks, an unanticipated structural pattern has emerged, surfaced through the accumulation of evidence over eight weeks of tracking.

This pattern has three layers:

  1. Technology
  2. Finance
  3. Control

The order is important and the connection between them matters (Technology-->Finance-->Control). The interaction between these three explains a lot of what we are experiencing right now. We think you may be seeing the same thing.

At the technology layer, Frontier Model capabilities are plateauing and the gap between proprietary and open-weight models is narrowing. This is commoditization, and is the same trajectory every maturing technology follows.

Commoditization creates pressure at the financial layer. Frontier Labs and their investors have committed hundreds of billions in capex on the assumption of sustained pricing power and exponential growth. If the technology is commoditizing, the margins to service those commitments compress; and the quantum of spend compressed into this narrow a window makes the payback math unforgiving.

This financial pressure drives behaviour at the control layer. When the technology itself is no longer a durable moat, the rational response is to build non-technical moats: to restrict access, remove practitioner control levers, and create switching costs. This is what we are observing across the major Frontier Labs right now.

The practitioner implication flows from the same chain in reverse: if control is being restricted, if the financial structure is fragile, and if the technology is commoditizing, then building deep dependencies on a single frontier provider carries increasing risk. When we discuss optionality, we are referring to the ability to move between providers and between proprietary and open-weight.

Given all of the above, we believe optionality is the appropriate response.

In this post (Part 1), we’ll cover the technology and financial layers. In Part 2, we’ll cover the control layer, the open-weight counter-narrative, and our confidence assessments.

Terminology note: throughout this document, we’ll be referring to the “open-closed gap” to refer to the gap between closed (frontier) models and open models.

The technology layer: capability is a commodity

Frontier model capabilities have been showing signs of plateauing, arguably for about 12-18 months. The improvements are real but incremental. The gap between “frontier” and “good enough” is narrowing faster than the frontier is advancing.

Widely cited is Claude Opus 4.5, representing a “watershed” moment for the usability of Generative AI in coding. This is correct, however when you look more closely at benchmarks (and yes, benchmarks are flawed) the picture becomes clearer: We have been approaching asymptotic improvement for a while, but Opus 4.5 represented a threshold being crossed, not a fundamental change in the trajectory of improvement.

Consider:

So, what does all of this mean? The most probable short-to-medium term outcome is commoditization of the model layer.

We need to be clear: this is not meant to be a prediction of collapse or irrelevance. It is more than likely the normal trajectory of a maturing technology. Commoditization will mean better prices for consumers and practitioners, but a margin-constrained (or margin-compressed) business for providers.

And margin-constrained business models matter because of what it does to the financial layer.

The financial layer: the spending is real, the revenue is not (yet)

Many AI skeptics and apologists are debating the merits of the technology. We would argue the more meaningful challenge to Generative AI is making the economics make sense.

What makes our current moment unusual is the large amount of upfront and planned investment in Generative AI and the implications in the likely scenario of it not reaching expectations.

Microsoft, Meta, Alphabet, Amazon, and Oracle collectively plan to spend $630-700B+ in AI infrastructure in 2026, an eye-watering figure rivalling Sweden’s GDP.

Morgan Stanley estimates ~$2.9 trillion in global data centre construction through 2028, with 80%+ of spending still ahead.

Why is this important? Front-loaded investment with lagging revenue is normal for infrastructure CapEx. What is not normal is the quantum of spend compressed into this narrow a window. $630-700B in a single year makes the emphasis on cash flow and payback acute, and a credible analysis of payback timelines at current revenue trajectories is noticeably absent from the public discourse, including from the labs themselves.

At the same time, physical infrastructure is under stress from outside the technology sector:

This has created or exacerbated a cost inversion already visible at the operational level. Bryan Catanzaro, NVIDIA’s VP of Applied Deep Learning, told Axios in May 2026: “For my team, the cost of compute is far beyond the costs of the employees.”

This follows an emerging picture in the data:

An analysis of SEC filings across 32 companies that publicly linked layoffs to AI between 2023 and Q1 2026 found operating margins declined or held flat at every company that buys AI. The study found only companies selling AI infrastructure have improving margins. Further, only one company out of thirty-two, Salesforce, showed genuine, measurable improvement where margin gains, headcount reductions, and a named AI product all aligned. The authors’ framing: the payroll savings at AI-buying companies are becoming revenue at AI-selling companies. Heads up: the framing of the data in this study is more polemic than we would prefer; however the data are strong evidence of the stress the frontier labs are under.

The thesis “cheap compute replaces expensive humans”, is generating a lot of uncertainty, and it is currently running in reverse at the organizations purchasing and building out Generative AI product and operations.

This does not mean the economics will never work. (E.g. inference costs are falling (DeepSeek V4-Flash at $0.14/M input; Qwen3 Coder at $0.07/M on Novita.ai)), and the open-weight cost structure is dramatically cheaper than frontier proprietary.

But it does mean at the moment the cost structure has not caught up to the adoption curve; and the organizations most aggressively adopting AI tools are the ones feeling this margin pressure the most.

More specifically:

An important dissenting opinion is Anthropic’s recent [$30B revenue run rate]. We treat this number with care, however, given it reflects growth based on the same pricing model and usage changes we have seen disrupt Anthropic’s user base; however, this is a narrative worth considering when reading this essay.

This means there are perverse incentives at play, which leads to our discussion of the control layer.

In Part 2, we’ll examine how frontier labs are responding to these pressures, why open-weight models are gaining a structural advantage beyond capability, and what we can say with varying degrees of confidence about where this is heading.