So Apollo believes in the thesis strongly enough to put billions of its own capital behind it. That sounds more like putting their money where their mouth is than a rebuttal.
Or they invested billions and it behooves them to generate justification for that investment at a time when everyone else is also building railroads, I mean data centers.
That's possible, but it's still an argument about Apollo's incentives rather than the thesis itself. By all means scrutinize anyone talking their book, but then show where the analysis is wrong. Otherwise we've moved from “Apollo is conflicted” to “Apollo must be wrong because it invested,” which is not much of an argument.
Yes but how much of that compute shortage is from demand that is subsidized? We’ve seen companies like Uber drastically cut how much they are willing to spend on AI because they are paying actual usage costs, while at the same time OpenAI and Anthropic increase the limits on their fixed cost plans for individuals meaning people not paying usage costs are using it more and more… doesn’t this show that the compute shortage is because OpenAI and Anthropic are paying for it, not their customers? And the moment OpenAI and Anthropic stop paying for it, demand will collapse.
Anthropic is currently profitable, generating around $1B/quarter and ~$50B in ARR. About 75% to 85% of Anthropic's revenue comes from its usage-based API business, which has a gross margin that exceeds 80%. https://www.tradingkey.com/analysis/stocks/us-stocks/2620181...
Meanwhile, OpenAI is at ~$25B ARR, but is likely not yet profitable.
>The divergence in business models is directly reflected in financial data. SemiAnalysis estimates that Anthropic's overall gross margin has rebounded from negative 94% in 2024 to the mid-60% range, with the gross margin of its API business exceeding 80%.
Their link seems to claim semi analysis thinks it is 80%. It looks like it might be referencing this newer article from them, as the same picture is in both articles, but I didn't feel like paying to find out: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
Looks like it’s behind a paywall. I’ll take their word for it that semi analysis now estimate it to be 80%. That makes my point even stronger, if that number is true, where is the money? The report says that Anthropic generate over $50bn in revenue so at 80% margins that gives $40bn in profit. Where is that money? If they’re generating $40bn in profit, even after accounting for very high employee compensation and training costs… they should have tens of billions in profit, yet they’re out raising tens of billions instead. Where is the money going? And if only 20% is their actual inference costs, where are all these compute providers going to make their money? The world is at compute capacity on, what, $10bn in revenue?
If you figured out how to build a machine that turns electricity into gold with an 80% margin, of course you’d go out raising capital to build more machines.
Your contention is they're spending it on what, exactly? Leaks put OpenAI's training spend at single-digit billions so that can't be the machines they're building, and they (OpenAI + Anthropic) are famously renting/leasing/borrowing compute through varying-degrees-of-circular deals... so what's the machines they're building?
but your contention is they are profitable on inference. Why would they need to raise money to get their hands on compute if they're making money on compute? There's no upfront costs for Anthropic, Anthropic don't own or build compute infrastructure, they just rent access to compute owned by someone else, such as the $15bn/year SpaceX deal they signed recently (which they used to create more demand without increasing revenue).
no, even if we assume their margins are 90% (they are not) they are still losing money because the $200 plans allow for tens of thousands of dollars worth of inference and a huge number of users are milking every cent across multiple accounts. Every “reset” OpenAI and Anthropic do is setting money on fire.
If it were true that they’re making money hand over fist they wouldn’t need to raise tens of billions of dollars every few months.
> even if we assume their margins are 90% (they are not)
How do you know they are not?
It will be curious to see the cost of inference for these newly released open weight models and will help give an idea of the actual cost of inference. But for now, I think saying the $200 plans allows for "tens of thousands of dollars worth of inference" provides very little insight when you are measuring the inference cost in API pricing with an unknown margin.
The simple question to ask is, if it is so profitable, where is all the money going? If Anthropic have 90% margins on API usage and API usage is $50bn+ in revenue per year, where is the $45bn going? Why do they need to raise so much cash, constantly?
In Amodei’s Dwarkesh podcast he says that they are constantly estimating the increase in demand for the next leg up, then going out to raise money to build for it. So it’s not necessarily that they are raising money because they are unprofitable.
Anthropic aren't building out the data centres themselves, they're renting/leasing/borrowing from companies that are doing the actual spend on building out infrastructure. And the data centre companies aren't spending their own money, they're borrowing too (hence Apollo investing in data centres). Anthropic are paying SpaceX ~$1.25bn/month right now for access to more compute, that's $15bn a year, more than what these supposed margins would require in total spend (based on current revenue estimates).
The SpaceX deal is a great example of Anthropic creating demand, i.e:
> We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.
They committed to spending $15bn per year with SpaceX and then increased limits for customers on fixed cost plans, creating more demand without any increase in revenue.
So, sure, it's not necessarily that they are raising money because they are unprofitable, but no alternate explanation makes any sense. The argument that could maybe made in favor is based on announcements like this one:
> Today, we are announcing a $50 billion investment in American computing infrastructure, building data centers with Fluidstack in Texas and New York, with more sites to come. These facilities are custom built for Anthropic with a focus on maximizing efficiency for our workloads, enabling continued research and development at the frontier.
You might conclude from that, Anthropic are financing Fluidstack's build out, but they're not.
Just after that announcement, Fluidstack raised $830 million to build out data centres, none of the money coming from Anthropic. Fluidstack are currently rumored to be raising another $1bn. Anthropic's "$50 billion investment in American computing infrastructure" is just committed spend on renting compute from Fluidstack, a commitment that Fluidstack then use to raise money to actually deliver it. If Anthropic making money hand over fist, they wouldn't need to raise for committed spend.
And thus we return to the original question, how does future demand translate to spend? Actual handing over of dollars?
The compute demand is not fake. Non-coding industries have barely begun to deploy this technology. In the legal sector, I’ve been a tech pessimist my entire career, because it was uniformly quite bad. I’ve spent the last few months demoing legal tools backed by frontier models, and we’re definitely going to buy one of them. They’re real and they work and they address a bunch of needs.
You’re talking across the issue. The demand is real because it is cheap. The demand is being generated by OpenAI and Anthropic selling inference below cost on fixed price plans. If everyone was paying the actual costs then demand would fall through the floor. The legal tools you’re looking at use barely any compute. They’re not driving the compute demand. You can validate this by asking how much they are spending on API usage. A company spending $100,000 per month on a frontier model via an API is the equivalent of… 10 or so OpenAI and Anthropic fixed price plan customers. Are these legal tools spending hundreds of millions per year on the frontier models?
My point is that you’re projecting forward into the future where providers need to raise prices, but overlooking the demand that will arise when the rest of the economy starts using AI.
The entire economy could be entirely powered by AI and use less compute than is being used today. You're projecting the amount of compute used by coding while subsidized onto other industries, but that doesn't translate.
Speak to some of these legal technology companies and ask them 2 questions:
1. How much is your AI spend on coding?
2. How much is your AI spend on AI within your product?
The answer to #1 will dwarf #2 by orders of magnitude. And that's now, when these companies are still finding their feet, using the most expensive frontier models for their product that are likely overkill (as the product matures, they'll find the right mix of cost vs. capability, whether that's lower cost models from frontier labs, or open weight models).
Put simply, compute does not scale with economic value. A task you bill $500/hour for could be done in 2 seconds by an LLM vs. a task a developer bills $500/day for could take an LLM 10 minutes. Same dollar value, huge disparity in compute.
The only use-cases for AI that are comparable to coding on compute are image and video generation. Unless we end up with an economy primarily made up of companies producing code, image and video, there's literally no way compute needs can keep growing without subsidies.
I think it is hard to overstate just how much "work" is being done by Claude Code and Codex because it is "free" at the point of use. There are millions of newly minted developers prompting Claude and Codex to generate trillions of lines of code every single day because it has no marginal cost, not because it is driving any economic value. And as soon as they're exposed to the real cost, when economic value becomes a factor, they're going to stop doing it (as we're already seeing with companies like Uber).
Through your YC connection and legal work, you have access to a lot of very successful people working in every industry: ask those 2 questions, how much are they spending on coding with AI vs. how much are they spending on AI in their product? You're going to find that even the most aggressively AI-integrated products are spending pennies on their product's AI usage compared to the dollars they spend on their AI coding.
We're at peak compute demand, and it is almost entirely driven by coding, which is subsidized. The real economy isn't like the creative and chaotic make things and see what sticks world of coding. The real economy is boring, routine, regimented, task oriented, you hire people, train them, they do the tasks, you make some money. Most of the economy could be replaced with a few semi-intelligent macros.
Yes, AI is coming to every industry, you're right, but it isn't going to explode compute demand, it is going to make these industries more efficient, it is going to reduce costs, not shift them to compute, because $1 of compute can do more than $1,000 of a human in most industries.
We can come back to this comment in a couple of years. I bet we'll be using less compute then than we are today.
True, and even selling token price may not be illustrative of what it actually costs them to provide the service. Tokens may be sold at a loss if majority of spenders are running agents 24/7
The industry and users moves on from single chat-based to more and more "agentic" workflows that may generate longer workloads with multiple simultaneous agents (separate agents - separate contexts - separate KV caches).
My estimation is based on, say, running a Kimi3 on a 24 B200 GPUs - it is very easy to lose money when selling tokens at "market" prices.
Recent Kimi K3 release is supposedly frontier-grade, so would be interesting to see what it really costs to run inference for models of this class, when the weights get released and other independent providers pick it up (then we could reasonably expect competitive pricing with low-ish margins).
When gas prices went up in the 70s because of fuel shortage, smaller (more fuel efficient) cars became more popular to use less fuel to do the same thing. I wonder if the same will happen with compute, by making software more efficient, and do the same thing with less compute.
The only thing I really hope for if this continues for an extended period of time is everyone optimizing their programs to use a minimal amount of ram.
I bet everyone will soon use a thin client with 4GB of RAM, and all the compute happens in the cloud of some American corporation you pay subscription fees to.
And that will probably fail. I don't think on average companies are capable anymore to deliver software that can operate on only 4GB... Even if everything but presentation layer is cloud based...
This article is about silicon. But the other shortage that matters is meat-compute.
AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.
But.. "good design taste" has no compiler.
Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.
In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.
"A compiler" isn't the bottleneck. The bottleneck of RLVR is the rollout.
Before you even get the code for a compiler to do its thing, you need an AI to ingest and emit thousands of tokens autoregressively. That's what makes RLVR so expensive.
Likewise - I don't think your "design taste" argument holds? Areas where "taste" exists are far easier for AI to tackle than areas where no data exists. It's easier to teach an AI how to make an image that looks good than to teach an AI how to de-solder a BGA chip. "Tacit knowledge" yes, "taste" - not really?
No, it was "<linked article> from the beginning of 2026 claimed <x>." Also, if you open the article, you'll see the data goes back multiple years, but does not cover 2026 (as it was posted at the beginning of 2026).
> We fit a trend to the quarterly quantities of AI compute sold, as measured in H100-equivalents. We find that computing capacity has been growing by 3.3x per year, equivalent to a doubling time of 7 months.
This is just a graph of cash changing hands. I transferred a loonie from one hand to the other one trillion times this morning and have the largest data centre in the world.
1. There is a huge demand for compute, specifically GPU compute
2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills
3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying.
4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others).
All while advances in open weight models are making it appear that the major labs truly have no model moat.
Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.
Being first rarely matters in tech. Fast follow that’s “good enough” and cheaper eats “first” for lunch all day long, and that’s the pattern starting to play out.
Eh, I don’t think there’s any reason to think #3 is true, and the whole thing being a house of cards is predicated on that one.
If the massive demand is still present for compute at market rates (which I believe it is), then your second point is just investors spending cap ex to build out valuable and profitable assets, no problem there.
What evidence says otherwise? OpenAI is projected to have massive losses for years to come and all indications are that’s driven by the cost of compute being far higher than the revenue generated by said compute.
I feel this post is blind to many of the secondary side effects of this "shortage". The rapid increase in prices and delivery times is having deleterious effects on all things tech - everything from phones to smart appliances and all sorts of gadgets has moved into unreachable price levels.
Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range.
So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.
> Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this?
Phones are a bad example in the US at least-not sure how it works in the EU.
I think most people in the US simply finance new phones through their provider who is already providing them the plan, so it just looks like an expensive phone plan to them for a few years. Apple could offer a deal where a $3000 phone is simply a $200/mo. add-on to your bill over 3 or 4 years.
I bet this causes more things to be tied to a cellular provider and offered the same way, like game consoles and maybe even TVs. Imagine signing on to your cellular provider for a 10 year contract to get a few new phones, a game console or two, and a couple TVs.
The issue is that this will be another vision pro. The issue is that every year, we take device and services price increase (which of course nobody can do anything about because Apple-Google is a cartel based in the US where consumer protection is not even in the dictionary but that's a different topic probably) and nobody keeps track of just how much these things cost for a "not an influencer / person who works in tech / friend of the president" household.
What on earth is that H100 price trend graph. The equally spaced x-axis points are 2x6 months, 6x3 months, 9x1 month. The whole visualisation of the trend is ruined on the back of that.
There are liars, damned liars, and people who play silly buggers with scales.
I think this is a severely compromised visualization but not actually misleading, either by design or in effect. Seems like they wanted to include the full picture - namely, that the 2026 spike is still less than the 2023 price when new, but doing all the data monthly would have been hard to read. Note that doing it all monthly would have made the derivative of 2026 more visually stark, so the way the graph is presented actually weakens their argument (albeit inconsequentially).
I think the best way to fix the graph would be shading or line texture to indicate the scale. The biggest problem is that the inconsistent scale is a surprise you only see upon close reading, after you've visually digested the trend. So the scale needs to be apparent in this first visual digest. (Sort of like how many logarithmic charts include thin axis marker lines in the body of the graph itself, so as to immediately inform a quick glance.)
i dont see how compressing the past ruins the present day extrapolation? How does the conclusion change for you if the graph was 3x wider on the left? The title is "rates are rising" is that not true?
A graph is much more than one conclusion; in fact, almost the entire point of graphing is to allow the comparison of "shapes" and to easily hypothesise about associations across datasets.
This graph misrepresents the rate the price declines and the length of time it has been stable for, which throws off nearly all non-trivial conclusions.
That specific conclusion is unaffected, but that doesn't make it okay. If they used fake numbers, but the conclusion were still the same, you presumably wouldn't think that was okay either?
Meh. That’s like saying we have a parking shortage in cities. No we have more car journies than we need.
We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park)
Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam
It’s all choices
We just are making bad ones mostly
The choices now are would you like ads or more ads.
Shall I waste compute seeing if the web page you clicked on can be summarised ?
The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out.
If instead of us training the models to be human utility maximizers, the models had figured out how to train us to maximize their own utility, could we tell the difference?
I suspect that utility may not enter the picture at all, and what we have are LLMs that have been optimized for getting the sort of humans who decide what to spend money on to maximize spending on LLMs and related infrastructure. The LLMs in turn could be anything from Skynet to a very spicy cold-reading autocomplete.
While building Fabs takes time you can only imagine that the amount of capital flooding in means there will be an explosion in infrastructure supply over the next decade and eventually prices will crater
To preface, this article presents valid points that I agree with. I would have also liked to know their findings on computational memory (CXL) and not just HBM or DDR. Also, a theoretical background on the Von Neumann and Harvard architectures would be helpful.
Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.
The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over.
> At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027.
Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them.
Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows).
At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.
GPUs aren’t scarce. There is just a lot of demand, so prices have gone up.
Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.
Without realizing it, you are coming close to the economics definition of relative scarcity. If we went by uour definition, basically nothing is scarce, ever, because it's possible to provide it, but the prices are set in a way that is completely unrelated to the cost of production + risk related margin. Instead prices resemble what would be an auction. And it all only works out because demand is elastic, so there are prices that people refuse to pay, because for their use case, they'd lose money.
Housing, collectors items, access to the time of a top of the line doctor... all relatively scarce. And You'd be laughed out pf most rooms if you claimed those things aren't scarce.
"Shortage" means exactly that. It doesn't mean "there's no GPUs at all." In a gasoline shortage, prices go up and there are lines at the pump, but there are still cars on the road.
If prices go up there aren't lines. You get lines when demand goes up and supply doesn't or supply goes down and demand doesn't follow. Either way the "fix" is either to raise prices or have lines.
> GPUs aren’t scarce. There is just a lot of demand, so prices have gone up.
What does "scarce" mean to you?
I'd wager that to most reasonable people it means "available in a supply that, measured against demand, is low". Market forces typically react to such a state by pushing prices up. Saying "they're not scarce, they're just expensive" is just silly.
The exception to this would be during e.g. supply chain hiccups (lots of compute sitting there, unable to get to its operational destination) or when other forces artificially push up prices. But that's not what we're seeing here. Compute is much scarcer than it's been for years.
There is the type of scarce when you can't get something for love or money. Look at the gasoline lines in Russia. Those people sit in lines miles long for gas stations that aren't even open. No matter how much money they are willing to spend, they still can't get the gas.
I'm just grateful that I don't need to upgrade my computer for a while, and cross my fingers that I don't have a hardware failure in the next couple of years. I had an unpleasant laptop failure in 2021 that I don't wish to repeat.
The definition of scarce is rare, or insufficient to meet demand. Not just that supply is low in relation to demand, it means that you physically have trouble getting something.
As such, saying “GPUs are scarce” is an empty statement from a strictly economic perspective - so are all other physical goods. The correct economic phrasing would be “there is a shortage of GPUs” or “GPUs are in high demand”.
Of course, in colloquial English we interpret “GPUs are scarce” as meaning the same thing.
It will be fascinating to watch this play out. Particularly if AI-compute satellite constellations become a reality - We're trending towards a matrioshka brain and I'll be happy if I live to see the beginnings of that future
https://www.apollo.com/insights-news/pressreleases/2026/01/a...
https://www.apollo.com/insights-news/pressreleases/2025/11/a...
etc
Meanwhile, OpenAI is at ~$25B ARR, but is likely not yet profitable.
https://newsletter.semianalysis.com/p/anthropic-growth-and-b...
Their link seems to claim semi analysis thinks it is 80%. It looks like it might be referencing this newer article from them, as the same picture is in both articles, but I didn't feel like paying to find out: https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
https://www.anthropic.com/news/higher-limits-spacex
If it were true that they’re making money hand over fist they wouldn’t need to raise tens of billions of dollars every few months.
https://xcancel.com/i/article/2076078865060151465
How do you know they are not?
It will be curious to see the cost of inference for these newly released open weight models and will help give an idea of the actual cost of inference. But for now, I think saying the $200 plans allows for "tens of thousands of dollars worth of inference" provides very little insight when you are measuring the inference cost in API pricing with an unknown margin.
The simple question to ask is, if it is so profitable, where is all the money going? If Anthropic have 90% margins on API usage and API usage is $50bn+ in revenue per year, where is the $45bn going? Why do they need to raise so much cash, constantly?
But I do wonder how a 60% margin would be realistic when Sonnet costs 3-6x more than GLM 5.2 hosted by third party providers.
Anthropic aren't building out the data centres themselves, they're renting/leasing/borrowing from companies that are doing the actual spend on building out infrastructure. And the data centre companies aren't spending their own money, they're borrowing too (hence Apollo investing in data centres). Anthropic are paying SpaceX ~$1.25bn/month right now for access to more compute, that's $15bn a year, more than what these supposed margins would require in total spend (based on current revenue estimates).
https://www.anthropic.com/news/higher-limits-spacex
The SpaceX deal is a great example of Anthropic creating demand, i.e:
> We’ve agreed to a partnership with SpaceX that will substantially increase our compute capacity. This, along with our other recent compute deals, means that we’ve been able to increase our usage limits for Claude Code and the Claude API.
They committed to spending $15bn per year with SpaceX and then increased limits for customers on fixed cost plans, creating more demand without any increase in revenue.
So, sure, it's not necessarily that they are raising money because they are unprofitable, but no alternate explanation makes any sense. The argument that could maybe made in favor is based on announcements like this one:
https://www.anthropic.com/news/anthropic-invests-50-billion-...
> Today, we are announcing a $50 billion investment in American computing infrastructure, building data centers with Fluidstack in Texas and New York, with more sites to come. These facilities are custom built for Anthropic with a focus on maximizing efficiency for our workloads, enabling continued research and development at the frontier.
You might conclude from that, Anthropic are financing Fluidstack's build out, but they're not.
https://x.com/fluidstack/status/2079250004510728559
Just after that announcement, Fluidstack raised $830 million to build out data centres, none of the money coming from Anthropic. Fluidstack are currently rumored to be raising another $1bn. Anthropic's "$50 billion investment in American computing infrastructure" is just committed spend on renting compute from Fluidstack, a commitment that Fluidstack then use to raise money to actually deliver it. If Anthropic making money hand over fist, they wouldn't need to raise for committed spend.
And thus we return to the original question, how does future demand translate to spend? Actual handing over of dollars?
Speak to some of these legal technology companies and ask them 2 questions:
1. How much is your AI spend on coding? 2. How much is your AI spend on AI within your product?
The answer to #1 will dwarf #2 by orders of magnitude. And that's now, when these companies are still finding their feet, using the most expensive frontier models for their product that are likely overkill (as the product matures, they'll find the right mix of cost vs. capability, whether that's lower cost models from frontier labs, or open weight models).
Put simply, compute does not scale with economic value. A task you bill $500/hour for could be done in 2 seconds by an LLM vs. a task a developer bills $500/day for could take an LLM 10 minutes. Same dollar value, huge disparity in compute.
The only use-cases for AI that are comparable to coding on compute are image and video generation. Unless we end up with an economy primarily made up of companies producing code, image and video, there's literally no way compute needs can keep growing without subsidies.
I think it is hard to overstate just how much "work" is being done by Claude Code and Codex because it is "free" at the point of use. There are millions of newly minted developers prompting Claude and Codex to generate trillions of lines of code every single day because it has no marginal cost, not because it is driving any economic value. And as soon as they're exposed to the real cost, when economic value becomes a factor, they're going to stop doing it (as we're already seeing with companies like Uber).
Through your YC connection and legal work, you have access to a lot of very successful people working in every industry: ask those 2 questions, how much are they spending on coding with AI vs. how much are they spending on AI in their product? You're going to find that even the most aggressively AI-integrated products are spending pennies on their product's AI usage compared to the dollars they spend on their AI coding.
We're at peak compute demand, and it is almost entirely driven by coding, which is subsidized. The real economy isn't like the creative and chaotic make things and see what sticks world of coding. The real economy is boring, routine, regimented, task oriented, you hire people, train them, they do the tasks, you make some money. Most of the economy could be replaced with a few semi-intelligent macros.
Yes, AI is coming to every industry, you're right, but it isn't going to explode compute demand, it is going to make these industries more efficient, it is going to reduce costs, not shift them to compute, because $1 of compute can do more than $1,000 of a human in most industries.
We can come back to this comment in a couple of years. I bet we'll be using less compute then than we are today.
The industry and users moves on from single chat-based to more and more "agentic" workflows that may generate longer workloads with multiple simultaneous agents (separate agents - separate contexts - separate KV caches).
My estimation is based on, say, running a Kimi3 on a 24 B200 GPUs - it is very easy to lose money when selling tokens at "market" prices.
I think it more resembles the "content" of our era.
https://refactoringenglish.com/blog/why-i-stopped-creating-c...
AI is trained on human intelligence. The hyper-scalers are squeezing every last drop of automatically verified reward, and that may get us very far. A compiler passes or fails in milliseconds for free, forever.
But.. "good design taste" has no compiler.
Of course, taste isn't unverifiable. But it's is expensively verifiable. Noisy, slow, and orders of magnitude lower throughput. People with deep domain knowledge often can't articulate well _why_ one design works and the other doesn't. So, judgment arrives as a verdict, and not a crisp rationale. I guess we'll see if sample efficiency outpaces the cost of human judgement.
In the mean time, leverage will sit with whoever holds this tacit knowledge (incumbents). I.e., hospital systems, law firms, chip designers, studios, SaaS that are dominating their niche.. and not with the labs training on it. To me, this is why valuations of companies like Palantir could potentially make sense.
Before you even get the code for a compiler to do its thing, you need an AI to ingest and emit thousands of tokens autoregressively. That's what makes RLVR so expensive.
Likewise - I don't think your "design taste" argument holds? Areas where "taste" exists are far easier for AI to tackle than areas where no data exists. It's easier to teach an AI how to make an image that looks good than to teach an AI how to de-solder a BGA chip. "Tacit knowledge" yes, "taste" - not really?
That's definitely one way to say "it doubled this year"
> We fit a trend to the quarterly quantities of AI compute sold, as measured in H100-equivalents. We find that computing capacity has been growing by 3.3x per year, equivalent to a doubling time of 7 months.
This is just a graph of cash changing hands. I transferred a loonie from one hand to the other one trillion times this morning and have the largest data centre in the world.
https://www.bigtrends.com/education/lessons-from-the-past-10...
1. There is a huge demand for compute, specifically GPU compute
2. Infrastructure providers are building like crazy, including taking on massive debt to fund this because their own cash flow can’t cover the bills
3. The demand for that compute is broadly being paid for with investor dollars pumping up the valuation of AI companies, not cash flow from said companies. If those subsidies go away these companies can’t pay for the compute they’re buying.
4. Those that own a lot of compute are starting to offload it, looking for interested buyers (e.g., Meta looking to build a cloud biz or SpaceX selling its excess compute to others).
All while advances in open weight models are making it appear that the major labs truly have no model moat.
Put together those 4 things paint a very ugly business and financial picture that seems unlikely to just correct itself naturally. History tells us, very clearly, that “the way out” of such a scenario is a series of events that is likely to leave some of the current players severely damaged if not simply out of business.
Does it not make sense to rent out your compute if competitors have a better model and demand at higher prices?
About the moat, Mythos was first made available to customers at the beginning of April. Kimi K3 is still behind this.
Both OpenAI and Anthropic are expected to deploy significant upgrades in August.
If the massive demand is still present for compute at market rates (which I believe it is), then your second point is just investors spending cap ex to build out valuable and profitable assets, no problem there.
Time will tell.
Just imagine, if the rumours are true and Apple's 'foldable' phone costs 2500€ or more - who is going to buys this? So much of Apple's ecosystem depends on enough users consuming services, buying apps and using their phones to facilitate digital interactions. I've been waiting for 2 months to get a new mac mini for our office lab, and the Studios have moved into "we can't really justify this expense" price range.
So the framing that AI is inevitable or projections that put the world into a year or more of this "scarcity", where people can no longer afford non-entry level gadgets, are just naive IMO.
Phones are a bad example in the US at least-not sure how it works in the EU.
I think most people in the US simply finance new phones through their provider who is already providing them the plan, so it just looks like an expensive phone plan to them for a few years. Apple could offer a deal where a $3000 phone is simply a $200/mo. add-on to your bill over 3 or 4 years.
I bet this causes more things to be tied to a cellular provider and offered the same way, like game consoles and maybe even TVs. Imagine signing on to your cellular provider for a 10 year contract to get a few new phones, a game console or two, and a couple TVs.
As long as they sell an iPhone 18 at a more accessible price point, I do not see the issue.
While it will be at a premium price point, people use their phones for hours every day, while the Vision Pro is a niche gadget.
With videos, photos, web browsing, and reading being much better on a foldable, I can see this have mass appeal.
There are liars, damned liars, and people who play silly buggers with scales.
I think the best way to fix the graph would be shading or line texture to indicate the scale. The biggest problem is that the inconsistent scale is a surprise you only see upon close reading, after you've visually digested the trend. So the scale needs to be apparent in this first visual digest. (Sort of like how many logarithmic charts include thin axis marker lines in the body of the graph itself, so as to immediately inform a quick glance.)
This graph misrepresents the rate the price declines and the length of time it has been stable for, which throws off nearly all non-trivial conclusions.
B) be careful about telling other people what they think
I don’t believe in 5-10yrs we’ll be in a shortage anymore
We have societies designed by the default choice. If that choice is walkable streets, low cost electric buses, dense neighbourhoods (mostly I mean you can walk for miles along streets and parks and shops without crossing a car park)
Then you get far less car use. People aren’t stupid, but living in downtown Houston means you have far less choice about driving everywhere than living in suburban Amsterdam
It’s all choices
We just are making bad ones mostly
The choices now are would you like ads or more ads. Shall I waste compute seeing if the web page you clicked on can be summarised ?
The simple answer to AI is to charge it at cost - not subsidised. Then the market will start to shake out.
It might take the US stock market with it …
Many people designing these systems understand that there are vast shortages in almost all sectors in computing hardware. I'm more interested in AI strategy of a future where there's an oversupply. I speculate that by 2030 there will be both an overproduction in the factories that make these chips and a burgeoning second hand market of AI GPU's. The differences in the compute requirements for training versus inference of AI will explain the hindsight in oversupply.
No, it is not a moat.
The story here changes dramatically when you start to look at the facets of the industry that the poster is skipping over.
> At the semiconductor level, TSMC’s advanced-node capacity—particularly N3, which underpins much of the AI accelerator ecosystem—is approaching full utilization through at least 2027.
MS bought more GPU's than they had rack space for: https://www.datacenterdynamics.com/en/news/microsoft-has-ai-...
Open AI bought out the memory: https://x.com/kwharrison13/status/2029248559388746168 but they have no means to consume anywhere close to their order.
Meanwhile both google and amazon are consuming a bunch of TSMC capacity to build their own ai chips, bypassing NVIDIA ... And as for them, they seem to be addicted to burning power to keep scaling, and that is a massive problem - if the next gen chips burn more watts for the same amount of work that is only going to exacerbate the power issues were having not help them.
Tokens are just Gacha for business. https://en.wikipedia.org/wiki/Gacha_game - It is software you dont control and you are going to pay for a cache miss. That isnt a model that is sustainable (even more so in the authors multi agent flows).
At the point that prices come down, (and they will) you're going to see a lot of corporations move from the cloud to on premise or back into colocation.
Anyone can get GPUs right now and build out what they need; it’s just a matter of paying more for them than the next company. I’ve been looking at building a multi-TB HBM system lately and they’re readily available - they just cost $400k.
Housing, collectors items, access to the time of a top of the line doctor... all relatively scarce. And You'd be laughed out pf most rooms if you claimed those things aren't scarce.
What does "scarce" mean to you?
I'd wager that to most reasonable people it means "available in a supply that, measured against demand, is low". Market forces typically react to such a state by pushing prices up. Saying "they're not scarce, they're just expensive" is just silly.
The exception to this would be during e.g. supply chain hiccups (lots of compute sitting there, unable to get to its operational destination) or when other forces artificially push up prices. But that's not what we're seeing here. Compute is much scarcer than it's been for years.
I'm just grateful that I don't need to upgrade my computer for a while, and cross my fingers that I don't have a hardware failure in the next couple of years. I had an unpleasant laptop failure in 2021 that I don't wish to repeat.
As such, saying “GPUs are scarce” is an empty statement from a strictly economic perspective - so are all other physical goods. The correct economic phrasing would be “there is a shortage of GPUs” or “GPUs are in high demand”.
Of course, in colloquial English we interpret “GPUs are scarce” as meaning the same thing.