How "tokenomics" might save AI PCs

Bob O'Donnell

Posts: 141   +2
Staff member

There are some interesting things happening in the PC market these days and, more importantly, the potential for even more impactful changes over the next year or so. At a high level, overall PC shipments were predicted to decline and are indeed starting to do so on a unit basis.

Driven primarily by huge increases in memory and storage costs, average PC selling prices have risen significantly over the last year. That has dampened overall demand, particularly among cost-sensitive consumers and other low-end PC segments.

Despite this, the major PC makers, notably Dell, HP, and Lenovo, have all reported higher overall PC revenues and expect that growth to continue through 2026. Demand for higher-end, more capable PCs remains strong, and the additional revenue generated by these much more expensive machines is more than offsetting the losses from lower-priced models. That, in itself, is both interesting and surprising.

However, I believe we could be on the precipice of an even larger shift: a significant increase in demand for AI-capable PCs. The explosive growth in enterprise AI usage is creating an entirely new corporate expense category, tokens, and the need to control that spending could finally provide the economic rationale for AI PCs that the industry has struggled to articulate.

Several structural changes have already occurred in how businesses think about their computing needs, and a few more still need to happen, that point toward this shift.

First is the fact that companies have suddenly had to deal with a new multimillion-dollar expense category: AI inference spending, increasingly measured and managed through token consumption. Between trends like tokenmaxxing and the now generally accepted belief that companies risk falling behind their competitors unless they leverage new GenAI and agentic AI capabilities as aggressively as possible, organizations are having to redirect enormous sums of money toward something they had barely considered before.

Initially, the general excitement and sense of urgency around leveraging GenAI and agentic AI meant that little oversight or analysis of these efforts was taking place. Now, however, companies are realizing that these tokenomics issues will represent a large and long-term part of their corporate budgets, so significantly more attention is being paid to how those costs can be managed.

At first, virtually all token requests were sent to cloud-based services. Early on, that wasn't a major concern because most tokens were being generated for free, a situation that is still largely the case in other parts of the world, notably China.

That changed dramatically about a year ago when major model providers started charging for token usage. Already, we have major model suppliers like Anthropic supposedly reaching staggering annualized revenue rates of $65 billion.

At the same time, several technical advances are opening new options for generating tokens. Improvements in the performance of smaller models, the dramatic rise in the use of customizable open-weight models, and the growing availability of AI infrastructure designed specifically for enterprise data centers are all driving new ways of thinking about how AI-focused computing demands can be met.

Notably, the rise of hybrid AI architectures that combine cloud, on-premises, and on-device computing is giving organizations more choices about where tokens are generated. There's also growing recognition that not all AI requests need to be handled by frontier-level models. Smaller and more specialized models can handle many of these requests and, in some situations, provide better and more accurate responses.

Suitably-equipped AI PCs, "deskside" workstations like those powered by Nvidia's GB10 chip and AMD's Ryzen AI Max/Max+ 400x, fit perfectly into this new scenario. They can run many of the more powerful, more compact models that are appearing on a daily basis and provide an intriguing new economic alternative to token generation.

In fact, some have argued that these systems are capable of running the equivalent of frontier-level models from about a year ago. Plus, because of the growing number of AI applications and agentic platforms that leverage open-weight models, these machines can be put to work on more specialized applications and solutions.

Beyond all these technical reasons, however, the simplest and most compelling argument for AI PCs is economics, or rather, tokenomics. For sufficiently heavy AI users, diverting even 20% of token consumption from expensive cloud models to local inference could materially shorten the payback period on a $4,000 AI PC. In some high-usage scenarios, it could potentially reduce that period to months rather than years.

Toss in the potential to send, say, another 30% of token requests to an on-premises, GPU-equipped server, or an enterprise AI factory, as Nvidia's Jensen Huang has labeled them, and the savings could be even greater.

Of course, there would be an initial capital outlay for these purchases, but building an ROI model to justify them is getting easier every day. Tapping into the newly created token funds or budget lines that organizations have been forced to establish could make the argument easier still.

As logical as this all sounds, there are certainly obstacles in the way. First, the existing PC procurement process at most organizations is structured in a way that could never justify these kinds of purchases.

Companies typically have three-to-five-year lifecycle plans for their PCs and look to keep purchase prices within the historical range of $1,000 to $2,000 for corporate systems. Shifting to a fleetwide deployment of significantly more expensive machines would require C-suite-level changes to well-established, and likely highly protected, processes.

The critical change is that enterprises need to stop thinking about AI PCs strictly as endpoint purchases and start thinking about them as distributed AI infrastructure.

Second, the challenge of how best to split and orchestrate the different elements of a single AI prompt or workflow across the various computing resources in a hybrid AI environment has yet to be solved.

Many organizations are clearly working on this challenge, and industry standards such as MCP and A2A help create the interoperability foundation needed for more heterogeneous AI computing. What remains immature is the intelligence layer capable of dynamically determining which model and compute resource should handle each part of a workflow based on capability, cost, latency, privacy, and availability.

Thankfully, we are also seeing efforts by some of the largest model providers to make versions of their models available in different environments. Several months ago, Google announced a deal with Dell that allows versions of Gemini to run on Dell AI infrastructure within enterprises.

In addition, there are clear economic incentives for companies like Anthropic and OpenAI to create tools that allow them to maintain some degree of usage within a hybrid AI environment. Some portion of most AI prompts will likely continue to require a response from the latest frontier models. By creating tools that make that process seamless, these companies can guarantee a level of usage even as organizations work to reduce their dependence on the latest frontier models, as they inevitably will.

To put it succinctly, it is better to orchestrate 100% of an enterprise's AI activity and directly monetize 30% of it than to insist on processing 100% and risk being bypassed.

To be clear, AI PCs aren't the only solution to the growing economic challenge created by massive token consumption. In the same way that early, unfettered cloud computing usage led to FinOps platforms for managing cloud consumption, we're bound to see a dramatic rise in both the number and variety of solutions designed to help companies manage their tokenomics challenges as AI usage continues to grow.

From a practical and economic perspective, however, powerful AI PCs and workstations can and should prove to be formidable tools in helping organizations achieve those goals. Ironically, the killer app for AI PCs may not turn out to be a particular AI application at all. It may simply be economics.

If enterprises discover that putting more AI compute on employees' desks can meaningfully reduce the rapidly growing cost of cloud-based inference, the tokenomics of AI could finally provide AI PCs with the compelling ROI story they've been missing.

Bob O'Donnell is the founder and chief analyst of TECHnalysis Research, LLC a technology consulting firm that provides strategic consulting and market research services to the technology industry and professional financial community. You can follow him on Twitter @bobodtech

Permalink to story:

 
Companies are justifying the insane AI spend by wasting tokens generating slop, completing pointless tasks and creating work that still has to be human checked and rewritten... and calling it progress. Looking busy and producing masses of trivial data is not the same as increased productivity.

When the MBA (execs) realise this the equation will change again.

 
The future of AI is local AI, especially when you consider the privacy issues. Once the model is trained it really doesn't take that many resources to run. I see AI going from a cloud model back to a "well sell you a software licsense you run on your hardware" type thing. It might not even get that far. There are open source models you can download RIGHT NOW that compete with the big guys. if you have a $1million in hardware setting around, you can download and run these open models that can often outperform ChatGPT and Claude.
 
They've invented a shortage and several "crises" with intent to sustain, a "tokenization crisis", a "memory crisis", a "gpu crisis", an "electricity crisis", a "storage crisis", a second "privacy crisis", and several others depending on which articles you read and yet in the midst of all this... No legitimate artificial intelligence. They can't reach it so they bring the hurdle down, by questioning the definition of "intelligence", and then breaking it off into separate terms so they can say their systems are some form of intelligence. Generative artificial intelligence. Not to be confused with AGI, which is the real thing. Complete clown show.

Ultimately a few big corporations will survive to run their datacenters where viable, and the rest will be open sourced, which will drive development long after suits have lost the stomach for the math. Open source can't be defunded, bankrupted, or stopped when driven by the passion of people who do not care what the salary is. Just look at what the internet is actually built on.
 
They've invented a shortage and several "crises" with intent to sustain, a "tokenization crisis", a "memory crisis", a "gpu crisis", an "electricity crisis", a "storage crisis", a second "privacy crisis", and several others depending on which articles you read and yet in the midst of all this... No legitimate artificial intelligence. They can't reach it so they bring the hurdle down, by questioning the definition of "intelligence", and then breaking it off into separate terms so they can say their systems are some form of intelligence. Generative artificial intelligence. Not to be confused with AGI, which is the real thing. Complete clown show.

Ultimately a few big corporations will survive to run their datacenters where viable, and the rest will be open sourced, which will drive development long after suits have lost the stomach for the math. Open source can't be defunded, bankrupted, or stopped when driven by the passion of people who do not care what the salary is. Just look at what the internet is actually built on.
In the same way that servers run Linux for basically the cost of labor and maintenance, AI is going to be run for basically the cost of labor and maintenance. We (the US) wanted to horde AI so badly that we actually accelerated open source development. I would love to pick up a rack if like 8, H100s when they start flooding the market in a few years. Be fun for the homelab if I could pick one up for used car money.
 
And not one of these AI PCs can actually run an AI model locally. Every single request gets sent to the same metascalers. What an absolute scam.
 
My perspective is from the coding tools side which maybe is heavier requirements than some others. I'm told that to replicate the work I currently do using the cloud services I would need the equivalent of multiple 5090s and even then it would probably be slower. Those requirements are not going to be a fit for pushing down to individual workstations because it would be wasteful every second it was not in use (and make the hardware and even the power supply to each cube/office a lot trickier to build and maintain correctly.) I'm not sure it's even a fit for local servers in even good-sized offices given the non-work hours when it would be more idle than during work hours.

There may be other tasks which already fit better at least for 'managed at the office level' but I think we're a fair ways away from pushing any meaningful work to the tiny AI capacity on 'AI PC' integrated CPUs.
 
So there is an insane demand for everything related to AI, but it's a bubble.
Sounds perfectly logical.

The one wrinkle that maybe makes it fit is that a lot of AI capacity is currently being offered well below cost. The current insane demand reflects that affordable price. It might be a very different story if all users were paying the full costs + profit margin that these companies hope to and eventually must achieve to stay afloat.
 
Back