Outside a relatively small world of circular investment and FOMO feeding FOMO the general consensus seems to be “let it burn.”
It appears very unlikely we will ever see an IPO of OpenAI. Anthropic appears less doomed, but still iffy at best. Tons of other large, but little discussed, AI startups are just dead-companies-walking at this point.
The likes of AWS are showing good headline numbers but are taking out massive debt to build infrastructure that looks increasingly unneeded. Those with capacity are looking to offload it, quickly. Yes AWS has “committed contracts” for this capacity but if those commitments are with shaky AI startups then it’s mostly just fluff PR and these hyperscalers will get left holding the bag on all this debt.
Anthropic is almost purely a model company. They own close to no data centers.
If models are becoming commodities, and the main bottleneck is actually serving them at scale, Anthropic does not appear particularly well positioned.
If AWS had a ~60-80% margin for decades, I see no reason why inference can't have a ~60-80% margin for quite some time.
The problem is, if costs continue to drop ~90% for the same level of quality every 18 months, demand is unlikely to grow 10x to keep the revenue stable.
Who knows. Jevon's paradox. But the cost/quality is dropping too fast that it's hard for me to imagine demand keeps up long term to keep revenues (and profits) GROWING.
> If models are becoming commodities, and the main bottleneck is actually serving them at scale, Anthropic does not appear particularly well positioned.
That has been my thesis ever since Dario went on about how difficult it is to forecast capacity on the Dwarkesh podcast right about the time Claude started having 9's comparable to GitHub's.
Even as an outsider, seeing all the cloud providers bemoan their lack of capacity and their growing backlogs, I could tell compute capacity was already the biggest constraint on industry growth and that it would be very hard to come by in the future... which also probably made it the only moat.
And now we see Anthropic making very expensive deals with competitors to secure the necessary capacity. Claude likely still has enough momentum to make it worthwhile. But I can't believe Dario, probably the most AI-pilled of them all, would under-estimate demand so severely.
> But I can't believe Dario, probably the most AI-pilled of them all, would under-estimate demand so severely.
To be fair, this mess up actually increases my respect for him, as he focused on making sure that they could continue to run the business under reasonable assumptions rather than YOLOing compute into existence like OpenAI.
My surprise is rooted in the point that he seems to be the biggest believer in AI there is -- to the extent the label "zealot" gets bandied about at times -- and he still underestimated the demand they would get. I would not have classified Altman as a bigger believer than Dario.
My guess is OpenAI, having the first mover advantage and the bulk of AI traffic for a very long time, could gauge the trends much better than anyone else.
> I see no reason why inference can't have a ~60-80% margin for quite some time.
Same as anything with low margin: Viable competition and low switching costs.
Many sysadmins lazily do 100% AWS because they don't know any alternatives. CFO's might enforce using more economical inference providers if the cost is even 10% less.
Advantages: Lowest inference cost due to scale, experience, advanced datacenters and custom chips. Huge customer base. AI automation will make it much easier to switch from AWS to GCP. Serious organization that doesn't let it's product unknowingly hack companies.
Disadvantages: Top model is slightly behind the frontier for now. Bias against using their products because it's not cool/trendy.
Anthropic is perceived as the leading AI company in the world. Even if they started selling inference alone, people would prefer to pay them a premium for it rather than figuring out how to operate opencode or other replacements. Add to that government contracts, enterprise and other markets, and I think Anthropic would be totally fine even if models plateau.
Transistor count has increased exponentially for decades and so did demand for compute. I think we will see a similar phenomenon with AI.
> Are you saying people haven’t been throwing unneeded GPU capacity on the market? That is happening.
That, by itself, doesn't have to mean anything.
AWS is far more compute capacity than Amazon needs, but that's not because Amazon misjudged how much capacity they need. They built it out to sell it to others, leveraging know-how and economies of scale to do so very profitably.
SpaceX has put a ton of excess capacity on the market.
Meta tanked chip stocks by saying it was considering the same.
Worries about compute overcapacity, via folks saying they want to offload capacity, is literally the thing that nuked the “situational awareness” fund last week.
It’s long term demand that matters, not just “price.” High prices without true demand is the literal definition of a bubble.
SpaceX is selling compute to a willing buyer. You can't point to one side of the trade as evidence for under/over capacity. The way we usually judge such things is the price. If the price for compute is going up, which it is, then that's the thing to look at.
It's holding up the entire US stock market and therefore the entire US economy and the dollar value. You probably don't want this particular Atlas to shrug.
The “Let it burn” attitude doesn’t mean people think they are immune.
Yes the overall market will take a hit but, like a forest fire we need a healthy burn to just wipe out the weaker players so the older more mature trees can get on with it. Yes the big trees will get burned a bit but they’ll be fine in the long run.
We need a good brush fire to just wipe out all the iffy startups and investors that over-indexed here. Thats what people want with “let it burn.”
I would rather be market-corrected out of a job for a few weeks or months than have this inflate into an economy-wrecking bubble that derails my career and life for years. I'm worried that we're at that state already...
The thing about the past is that it's the past. Things change, big changes do not bring back the past. They change things.
If you are market-corrected out of a job that situation most software engineers will be out of a job for years and wages will not recover for decades, or just never.
The trajectory of feudalism has always bent towards low wages and high rent extraction. There is no possible future where we stay with feudalism but increase wages.
Let's see ... feudalism started with slave labour (which is what muslims are still fighting for, fighting to bring it back), there were many bills of rights. Feudalism more or less started with Christianity forcibly implementing the right to quit your job, any job, including slavery, and ended with equality for all.
So: nope. Quite the opposite. Feudalism was a very slow increase in rights for normal people over time.
I do. The faster it happens, the faster a recovery begins and we can devote resources back to doing more practical things. Get it over with as soon as possible, rip the bandaid off, insert your idiom of choice.
But what is the point of this discussion? That we deny a possible burst because we stand to lose personally? Or is it, we believe it could burst because reasons?
This guy makes a decent argument that the funny money party doesn't really get started until Q3 2027. To which I would respond probably, but also probably not if another coding agent sized market opens up between now and then. But predicting the future is hard and all that.
> It's holding up the entire US stock market and therefore the entire US economy and the dollar value. You probably don't want this particular Atlas to shrug.
So you're saying we want to keep this bubble going forever?
If's it's a bubble, it's gonna pop, and better sooner than later.
TSLA's 400 PE ratio (along with Elon Musk's trillionaire status) was never sustainable. Where I disagree is with companies running market average PE ratios supposedly being doomed as well.
But also, don't count out OpenAI and Anthropic, no matter how shaky the numbers. The market has also demonstrated how to price SPCX so if they IPO, they will see their true FMV and I'm sure it's >0 and much much less than what they believe. For even in the worst case scenarios, they have valuable personnel, experience, and IP deploying AI at scale for what is to come, even if it's based on Chinese open weight models. RIP lightcone of all future value though.
As for the hyperscalers embedded in FAANNG, they've been using profitable divisions to cover the expense of growing but unprofitable divisions for a while now. In fact, that's business as usual. They're all going to land on their feet with PE ratios of ~20 or higher. Now take your favorite tech stock you prefer to scapegoat and figure out its approximate bottom relative to today. You'll be glad you did.
> But also, don't count out OpenAI and Anthropic, no matter how shaky the numbers. The market has also demonstrated how to price SPCX so if they IPO, they will see their free market FMV and I'm sure it's >0 and much much less than what they believe. For even in the worst case scenarios, they have valuable personnel, experience, and IP deploying AI at scale for what is to come, even if it's based on Chinese open weight models. RIP lightcone of all future value though.
You know what I'd like to see? OpenAI going bankrupt, wiping out its investors, then being bought up by IBM for a nickel. Ditto with Anthropic, but ending up with Oracle (if it's still around after running up so much debt).
The karmic Schadenfreude of those outcomes would be delicious. And while Sam Altman has long been out as a lucky fool, until recently, I thought more of Dario Amodei. Not anymore.
In what way/shape/form is the infrastructure showing signs that are contrary to growth in need/demand? I'd really like to know if there's something I'm missing that I should keep my eye on, thanks
I mean, talk about circular investment and FOMO all you want, but if you use Fable every day you know what's coming. This is no bubble, we're just getting started.
More and more companies are signing up with OpenAI and Anthropic at near exponential growth to automate everyday tasks. If anything, these two should manage to IPO just fine. The rest of the downstream startups probably won't make it.
Even if we agree with this take (and I do think it's a likely take that you're right on the long term), it doesn't change that it seems likely we're in a bubble, and it probably will pop.
We see a similar paradigm with lots of revolutionary technology. The initial promise is high, people get very excited, lots of money pours in, and.... 15-30 years go by before we start seeing real impact across the economy at large.
It's just a real slog to actually implement and roll out new tech.
So take robots: I can promise you that you won't see robotics in every household in the next decade (especially so if we exclude the current market of robot vacuums). Even if a company makes an incredibly capable robot "today" (and to be clear - they are not) it won't have time to scale out production, reduce costs, generate a used market that's accessible to less wealthy consumers, deal with regulatory hurdles and quality problems that only pop up in real-world usage, etc...
It's just slower than you're implying.
The change very well will happen (I'm inclined to agree that things are going to shift). That doesn't mean that the current investment is sane and will pay off.
So many historical examples of this, just two here real quick:
- Ford built his first automobile in 1896, founded a company in 1901, went out of business, got sued by ALAM, didn't build more than 10k Model T's until 1910, then only finally hit real scale (of low hundred of thousands of units) in 1913: More than a decade to "basic scale". Household ownership didn't hit 60% until 1929... 30+ years later.
- The initial web enthusiasm, followed by the dot-com crash in early 2000s...
> So take robots: I can promise you that you won't see robotics in every household in the next decade (especially so if we exclude the current market of robot vacuums).
People in 2016 were adamant that the current frontier LLM capabilities would not be achievable in the next decade.
Modern home robots prototypes look awkward and mostly useless the same way GPT-2 was looking like a curious but mostly useless thing in 2016.
I can totally see a capable and affordable home robots that are worth buying for the majority of the population being a reality by 2036. Maybe not _every_ household, but a good number of them.
Again, hardware just doesn't roll out that fast, and you're talking about trying to distribute expensive hardware to a population with a 47k median yearly wage.
I've built robots. Even if the software expense is ~$0... you're not going to escape the fact that a robot that can do anything remotely similar to household work is just going to be expensive. Like, several thousand dollars in actuators and sensors alone, not counting any compute.
LLMs are a SaaS - Robots are delivered hardware. The graveyard is littered with software folks who thought they could deliver hardware like it was software.
To be useful they’re gonna have to be capable and powerful. If they’re good enough to be useful, they will be dangerous.
Advanced household robots are closer to fully autonomous cars and airplanes than roombas. The code is safety critical. Imagine something capable of flooding or burning down a building with hundreds of people if left to the vibe coders?
Robotics in the household has been the realm of science fiction for decades and it still hasn't happened. We get dedicated, compact machines like dishwashers and washing machine/dryers, that's it. Most recent innovation has been the automatic vacuum.
If you want someone else to do household chores, hire someone. You can pay someone to do your house chores for years for the cost that these things will have initially.
>Most recent innovation has been the automatic vacuum
Which are pretty useless for a lot of home layouts and degree of putting cords etc. away. I took a look a few years back and got a stick vac instead. (And have a monthly housekeeper who does a lot more than a robo-vac would.)
> Robotics in the household has been the realm of science fiction for decades and it still hasn't happened.
Wasnt AI science fiction (research) for decades but just took couple of years after Chatgpt to become mainstream. Why do you that wont happen to robotics?
We don't have sci-fi AI, and in any case "this thing is now possible, therefore [much more difficult thing] will be possible soon" is not a rational thought.
Current frontier models are absolutely sci-fi AI by any measure that existed up until the late 2010s. If you showed a current frontier agentic AI, with bidirectional speech, tool use etc. to a person from 2010, they'd say it can't be real, there must be a person inside that mechanical turk.
Sci-fi AI has always been characterised by learning on the fly, which these models do not do. I guess we've got pretty close to what sci fi has often called "VI", autonomous machines that can carry out pre determined tasks to a level that looks concious, but are incapable of learning further.
They learn again when the labs train the next version. The previous version often has an indispensable role in orchestrating that process, as a critic and from usage transcripts of the previous model that they have access to. And between two training runs the model can assemble a lot of md files, memory files, and load them in on-demand. I think this is enough to make it scifi-level, but I agree there should be some way to customize it at a deeper level and adapt to each person's context. But things are still moving too fast for that to make sense. If they fine tuned a specific LoRA for Claude just for your projects and workflows, it would get outdated quite fast and you'd have to retrain it at quite a cost for the new version.
> Current frontier models are absolutely sci-fi AI by any measure that existed up until the late 2010s.
Sci-fi is Lt. Cmdr. Data, HAL, Skynet, Culture Minds...
> Current frontier models are absolutely sci-fi AI by any measure that existed up until the late 2010s. If you showed a current frontier agentic AI, with bidirectional speech, tool use etc. to a person from 2010, they'd say it can't be real, there must be a person inside that mechanical turk.
It wasn't a natural language chat interface that makes stupid mistakes a noticeable fraction of the time and can execute command line programs.
>Sci-fi is Lt. Cmdr. Data, HAL, Skynet, Culture Minds...
The current state of things is closer (not close to by a long shot) to Fred Pohl's "AI" programs from his novel Gateway[0] than it is to the above examples.
But even those are decades (at least) beyond what we have now. What we have now is the foreseeable (from at least 30 or 40 years ago) evolution of what was once called "expert systems."[1]
Expert systems didn't use neural nets or in fact any machine learning.
They did interviews with experts and tried to do domain modeling, mapping the expertise to a formal symbolic logical description, ontologies, relations etc.
Both the knowledge base and the inference engine changed. They both changed to a unified machine learning paradigm. It's like you throw away your axe's handle, then the head, and then you grab a chainsaw. It's "the same thing" to that same degree.
Depends on what sci-fi you choose. We could definitely hook up an LLM to the controls for an elevator, swap out the panel of buttons with a voice interface and give it a depressed, snarky prompt and voilà we've got one of the AIs from Hitchhikers Guide to the Galaxy
The cost of a timeshare robot will come down first, so you'll be hiring a robot before you own one. But that might not take so long to occur as you imagine.
All robotic home appliances will feature cameras with which people will be watching you and microphones with which people will be listening to you, all the time, forever.
Yep, that is what the bet is. The build out will continue, even if it looks different.
The Internet was a 'bubble' at one point, and after it crashed in 2000, it didn't go away, it continued to build out. We're still using the Internet after the Internet bubble popped.
On a $100B AI datacenter, some 60-70% of the center needs to be replaced every 2-6 years. These companies claim the hardware lasts 6 years while simultaneously claiming they are relying on new hardware to lower inference costs (implying much more frequent updates).
Over 40 years, that's 7-20 hardware changes, that turns into between $420B and $1.4T of ongoing investment (not accounting for inflation). The $30B that is called "infrastructure" only accounts for 2-7% of the overall bill.
This is NOTHING like fiber buildouts because the fiber lasts the whole 40 years with ZERO replacements and very close to 100% of the cost is infrastructure rather than a tiny percentage.
I've been through several network upgrades. You are saying the fiber in the ground lasts a long time, but all the equipment on either end gets replaced every ~5 years. I agree this build out seems pretty crazy, but people were saying the same thing with the internet. Arguing if the equipment has a replacement rate of 2-6 years or 30-40 years, I think isn't the part that makes this a bubble.
For fiber, the main costs were not the fiber itself, but acquiring the rights to lay the fiber and the labor to do so. The equipment at either end has a lifespan of somewhat less that 50 years, but that's a minor cost.
Conversely, for the data centers we're talking about, the cost of the things that need replacing every five years is one of the primary costs.
I was responding to you mentioning fiber. But of course, the "internet" does run on servers, which also have a replacement rate. There were data centers before AI, and they had a lifespan well below 30 years.
I'm not arguing that the current build out isn't crazy, just that AI isn't going away. Just like everyone thought the internet was a fad, but then it continued to grow (remember web 2.0). AI will morph into something useful and will be sticking around, even if we can't envision what it will look like.
Of course we know it. It’s been obvious since at least 2023. Everyone in AI oversells, except a few companies that built an actual business with revenue, like Midjourney.
There is no AGI coming anytime soon no matter how much hype is being thrown around. We are not in the singularity. However, peak bullshit is NEAR.
The trick the frontier labs have done is define AGI as "better than humans at the vast majority of valuable knowledge work" which is definitely not what most people think it means.
There is something I have been pondering recently. If we compare the cost of AI subscriptions (let's say Claude's 100/month) to a median developer salary (let's say 100k/year to 200k/year), the difference is orders of magnitude. This fills like a gap that needs to close. I suspect llms are too cheap right now but will raise their prices to a point where only big companies will be able to afford subscriptions to use them. I think soon we will see models that are only sold at very high prices.
This is something I have been pondering recently. If I compare the cost of a plunger ($23.99 on Amazon) to a median plumber salary ($62,970 per year per BLS), the difference is orders of magnitude. This feels like a gap that needs to close. I suspect plungers are too cheap right now but will raise their prices to a point where only big companies will be able to afford subscriptions to use them. I think soon we will seen plungers that are only sold at very high prices.
And I'm sorry to you and the parent poster for being needlessly snarky, I wasn't in a great mood when I wrote my original comment. I'll avoid bothered-posting in the future ;).
Why is this a gap that needs to close? Can you elaborate more than feelings? Right behind these big models are local inference with AMD/Apple having a great hardware start and a lot of local sized models making great progress.
Open weight models tell a different story. Inference is not that much more expensive than what a 100$ plan would allow and will only get cheaper (for current capability models of course, frontier not so much).
Also, something else to add. At first I thought no one wants to build data centers in hotter areas in the middle of deserts (many places in the American continent). So nobody would spend money building a data center in the Chihuahuan Desert for instance. However, a game of latencies will either require cover llm access from these areas, or make people move closer to the other data centers. In the former, llm prices will go up; in the latter, there will be a migration towards data centers that increase the price of the areas around.
> I suspect llms are too cheap right now but will raise their prices to a point where only big companies will be able to afford subscriptions to use them.
Is this range just Silicon Valley or what is this? Even including just Europe, you're looking at a lower bracket of 10k. If you expand to the rest of the world... Or do you think rich cities in the USA, where developers make 100k+ per year, can alone sustain this industry?
salary cost to the company is a lot higher than gross salary (which itself is a lot higher than net salary). you can guestimate total salary cost to be about 1.5x gross salary, so for the lower bound: 100k / 1.5 = $66,666.67 = €57,813.34.
€57-114k p.a. is well within the order of magnitude of yearly gross developer salaries in Western Europe (e.g. Germany).
"The line separating investment and speculation, which is never bright and clear, becomes blurred still further when most market participants have recently enjoyed triumphs. Nothing sedates rationality like large doses of effortless money. After a heady experience of that kind, normally sensible people drift into behavior akin to that of Cinderella at the ball. They know that overstaying the festivities ¾ that is, continuing to speculate in companies that have gigantic valuations relative to the cash they are likely to generate in the future ¾ will eventually bring on pumpkins and mice. But they nevertheless hate to miss a single minute of what is one helluva party. Therefore, the giddy participants all plan to leave just seconds before midnight. There’s a problem, though: They are dancing in a room in which the clocks have no hands."
Warren Buffett 2000
https://www.berkshirehathaway.com/2000ar/2000letter.html
> The same thing happened with their massive investment in Anthropic. They accounted for $53.4 billion due to deals with Anthropic last quarter. He said if you follow one Anthropic dollar through the earnings release, it's counted in AI business revenue, chips business, and AWS segment revenue.
That is insane if that is true, is that even legal?
There's a bit of Chinese whispers happening here. If you read the original piece they're talking about -- https://www.theregister.com/paas-and-iaas/2026/07/31/amazon-... (definitely worth a read, it's hilarious) -- you'll realize it means that that Anthropic dollar is "counted" multiple times in the marketing of three different Amazon businesses.
As far as accounting of revenue is concerned, it would have been counted as appropriate. Else, like you, I would guess it's not very legal.
It is. However remember that when the bubble pops it all works in the reverse direction too. Suddenly you have to mark down investment losses, missed revenue, and written off commitments.
Companies can go from looking really good to a complete financial mess almost overnight when all that leverage and self-reinforcing stuff unwinds. See last weeks headlines for one such scenario.
> And the fundamentals here are OpenAI and Anthropic, which are massively valued companies. They have humongous commitments and are generating real revenue on the order of twenty billion a year.
I think the size of their commitments is predicated on demand. Anthropic's annualized revenue run rate is now close to $50 billion, a fivefold increase from a year before [1]. They are making big investments, like $200 billion on Google's TPUs over the next five years [2], but those numbers seem justified by their expected revenue this year alone. If Anthropic cannot capture that revenue, someone else will.
Stock market valuations are a different beast, I personally think we have been due for a correction for ages now. But criticism of AI investment and particularly betting that it will all come crashing soon appears misguided to me. I can see a future where AI expenditures shifts around, not a future where everyone simply stops spending in AI all of a sudden.
I mean we are due our 6-8 year financial crash that we won't learn from. Once again it'll be coming from the USA's feral financial investments all to be bailed out by the tax player whilst the rest of the world picks up the pieces. Maybe its time we moved away from the petro-dollar if the USA can't be trusted to keep its finances in order.
There's a bubble around data centers, mainly. That's fueled by projected demand of AI and assumptions companies make about how the pie for that revenue is going to be divided up.
What's very real is the rapidly growing amount of revenue for both OpenAI and Anthropic. That's already tens of billions per year and growing quite rapidly. Investments against that kind of revenue aren't completely horrible. To a point. But at the multi trillion dollar valuation level, of course there are going to be issues with living up to those expectations.
In my view some of the base assumptions are looking not so solid currently. It's not a given that OpenAI and Anthropic will end up with most of the revenue. The Chinese trust Silicon Valley just about as much as vice versa. Which is why they are doing their own models, chips, and data centers. This is driving a rapid commoditization for things like frontier models, open model weights, and chips. This in turn gives countries outside the US a lot of options to stay independent. Which burst the bubble that all that global revenue was going to flow towards Silicon Valley. Some of that still might. But that will have to happen based on cost and merit.
There are also geopolitical circumstances that cause most data center plans to be bottle necked on permitting, chip shortages, grid connectivity, availability of gas turbines, gas, solar panels, inverters, batteries, water, and other resources. As it turns out, you can't just willy nilly plan for hundreds of GW of data centers and expect those to pop into existence overnight along with all the needed infrastructure. Most of the announced/planned capacity for this will likely not be realized. Certainly not this decade. 5-10% by 2035 would be a lot given all the constraints and scarcity. No amount of reality distortion can change the physical constraints on this topic.
The good news is that most of the money needed for this hasn't been spent yet. And what has been spent won't be going to waste. Up and running data centers are a hot commodity right now. They won't be running idle if a bubble bursts. But probably investors dreaming of multi trillion dollar IPOs might be a bit more cautious now that SpaceX stock is trading well below its IPO value.
Just before coming to read this article and thread, I read how Hyatt got rid of something like 30% of their call centre staff, and how the industry is gearing for replacing human work with AI.
There is sadly ample fiscal headroom in mundane drone-like work that was being outsourced (still cheaply, I might add) that AI can replace and even do a marginally better job of. I suspect AI prices can even increase and it will still be profitable for enterprises.
Companies like this will certainly keep expanding their AI use, and once they commit to that, there's little stopping them from moving to open models or local inference if need be.
My concern is less over the bubble and more over the social cost of AI. Call centres and the like provide a tremendous number of jobs. As AI moves into enterprises more and more, where are all these people supposed to work? Become baristas? They certainly won't be "learning to code"... What sorts of social and other unrests will this cause?
These forces I think will muddy the waters and make predictions difficult. Say what you want about the AI bubble, but if it pops, it will be different than previous ones. The bubble doesn't even need to burst because of the insane economic model, if enough people are economically devastated by it, it will cause ripple effects of its own.
The thing is that companies that are investing in AI crap to replace mundane drone-like workers is that the AI will get better but more importantly we will get more used to the way behaves towards us. Right now we some janky-as-fuck implementations but we're currently being trained on how to react to AI just as much it is trained on us responding to it.
> If you look at most big tech earnings this quarter, with the exception of Amazon, almost all others lost significant value after reporting. Microsoft, Alphabet, Meta, and Apple did.
Microsoft spiked on earnings, and is now actually about ~24% above it's pre-earning level.
Alphabet did lose about 7% the day of the earnings, but it recovered and now, after MSFT earnings, is actually 9% above the pre-earnings level.
Meta dropped 10% on earnings but is now back to its pre-earnings level.
Apple dropped ~10% and has not recovered (yet) but it's also famously "sitting out the AI bubble", so not sure why it's included here other than a "tech stock that went down."
Oracle has been dropping forever but after Microsoft's earnings it's climbing again.
If you zoom out, the stories change, and as you keep zooming out, they keep changing all over again.
My point is, 1) reading stocks in isolation is like reading tea leaves, and 2) if you want to point to any stocks, you should make sure they support your narrative.
The title's clickbait (from The Register, of all people - colour me shocked!).
There's nothing really groundbreaking at all in there, just "chips are expensive, and open weights models hosted locally in enterprise could displace Claude/GPT"
> I caution anyone looking at API prices: they dropped the price, but is it actually less expensive?
> Yeah, they dropped the price, but count the number of tokens you're tossing into it and see if it's actually cheaper. It's good for marketing, but the rub is how much you're actually using.
The level of discourse is so horrible now, I don't have words. Are these the ones making predictions on AI bubble?
This is poorly stated in the podcast, but the underlying point is correct: while cost-per-token is going down, overall token use is way up. This is causing the cost of using LLMs in corporations to skyrocket.
No, cost per task is probably not going down on average. The underlying issue is that LLMs are still bad at most tasks they could be used for, so as their contexts get larger and they're able to run longer without becoming incoherent, more tokens are spent on tasks to improve the quality of output. So cost per task is probably going up on average, but so is the quality of the output.
I obviously mean cost per fixed task. Do you agree that if you fix a task, the cost is going down? Meaning you get more work done from the same cost over the years.
Yes, obviously. The same task at the same quality and the same speed got cheaper. That's not what anyone is disputing. The problem is that overall expenses are going up.
I think you're arguing with the wrong person. Please note what I actually said:
> This is poorly stated in the podcast, but the underlying point is correct: while cost-per-token is going down, overall token use is way up. This is causing the cost of using LLMs in corporations to skyrocket.
> If the model consumes four times as many tokens to deliver a result, it’s not cheaper
This is literally what the guy says in podcast. So the underlying point is not correct. He’s specifically talking about per task cost. There’s no “but actually they meant something else”.
I don’t discount your point about overall cost increasing but that’s not relevant here.
Let’s agree that the podcast is fundamentally wrong in their main claim.
I want to show that the podcasters should be discredited because they don’t understand a fundamental aspect of the economics so you can’t trust the main thesis.
Here's the actual quote from the podcast you're presumably referring to. I guess everybody can make up their own mind about what they're saying and whether they are "fundamentally wrong."
You gotta watch something, because harnesses have changed how you can look at the cost structure of these things. If you look at cost per million tokens, it might look lower, but if the model consumes four times as many tokens in order to deliver a meaningful result, it's not cheaper. And so, Tom Claburn, one of our senior software reporters, had an excellent piece looking at how Anthropic's latest models use a tremendous number of tokens in order to deliver the result. So, sure, OpenAI's latest models might look less expensive from an API standpoint, which is great for marketing, but if it's using twice as many tokens, that's not the same thing. And that's somewhat dependent on the harness, but it's also dependent on how much reasoning effort is put into it, how they're routing the models.
I did. And he implies that cost per task increases. Otherwise there’s literally no reason to bring up tokens - that is an internal implementation detail.
> Otherwise there’s literally no reason to bring up tokens - that is an internal implementation detail.
I'm not sure if you're not aware of what he's referencing there, but the specific example he brings up is that Anthropic ostensibly kept pricing identical with new models, but changed the tokenizer, which made the model use more tokens for the same input and output text. So that's a case where, prima facie, the cost per task increased. Of course, this also depends on how verbose the model is and what harness you use, and so on.
Correct me if I'm wrong; I think you believe that he makes an argument like "OpenAI decreased API pricing, but this decrease was actually secretly an increase in cost." But he does not. He's saying that even though cost-per-token is going down overall, the actual cost of using LLMs is in many cases going up, because there isn't a direct causal relationship between cost-per-token and the total cost of using LLMs.
Of course, the price of tokens going down makes it cheaper. You can now afford to throw millions of tokens at solving a problem that would have been impossible for AI to solve at all a year ago for any price.
People completely lack imagination about this stuff. The main problem right now with AI isn't even AI successfully producing code at a reasonable cost, it's human coordination and review that is the bottleneck.
I think it's both. It's easy to run up a massive bill with AI without much to show for it, which is why token-maxxing has now been replaced by cost-awareness and AI budgeting such as Uber's "max 10% of salary per employee".
Cost isn't just token price though - it's (number-of-tokens-used x price-per-token), and there are large differences in token efficiency between different models and harnesses. Increasingly we're seeing benchmark sites focusing on "cost per completed task" as a cost metric, and it's not always the cheapest tokens that win.
I agree that ultimately AI/coding cost is just part of the picture - at the end of the day it's about software development cost, which for time being involves humans.
It's entirely possible China has a little bit of a bubble too when it comes to development investment vs demand. But obviously the bubble in the US is basically a whale compared to whatever small fish China is.
Nvidia already seems to be pivoting to "Local inferencing" with stuff like RTX Spark laptops.
The problem is that local inference machines won't allow anywhere close to their current margins or gross sales figures. At the same time, the huge overbuild of GPUs is going to crash server sales and prices for around a half-decade as companies try to avoid hardware upgrades or buy cheap, used equipment to save costs.
I think Nvidia will survive, but they'll be back to something like a 1T (or lower) valuation.
Nvidia has hordes of raucous gamers on the prowl waiting to tear their chips out of their hands the moment hyperscaler demand falters; plus, China has made sure there will always be plenty of open models to run for hobbyists and neoclouds, even if training compute were to drop off a cliff. They'll be just fine.
I want to see the number of gamers that tear the chips from the dead hands of hyperscalers to make up for the implosion of that market.
If I had to guess, half of humanity does not have interest in dedicated gaming chips.
Memory, that's another story.
The gpus will be too expensive and a lot of them are not even close to useable for gaming. Any consumer gpus maybe but I suspect that nvidia will be able to pivot more quickly producing consumer gpus again.
The market for video game hardware is absolutely puny in comparison to the datacenter GPU boom.
I saw it put quite well in a comment on reddit:
> In 2020, the gaming segment was 47% of revenue at around 8 billion. Today it's doubled to 16 billion, or 7% of revenue. That's right, data center went from 6 billion 2020 to around 198 billion today.
Even if gaming revenue doubles again when the AI bubble pops, their total revenue will still drop by something like 80%. I'm neither smart nor dumb enough to be confident about whether that's something Nvidia can survive.
I think Nvidia surviving is pretty certain. They are not particularly leveraged to my understanding and their commitments are for their own products. Debt is at 12,34 billion so not that much even at their gaming revenue comparison.
Now their market cap is most likely destroyed for forever... Still I little doubts about them continuing to exist.
Totally agree. For a lot of tradefolks maps seem to be the actually more important location. Still Google obviously but it's less slopified in comparison to search. For now at least.
well there are just 2 sources for website profits:
1. money from advertisements (from views, visits): you need to be on google search results at the bear minimum.
2. money from non-advertisements, mostly-offline (eg. you sell preminum stuff/service): you CAN hand out pamphlets to your service website IF your ROI per visit is so good to be true...
but even in the 2nd case, it really helps to be on the google's search result
Note that the “AI Bubble” term used here is only defined in the context of market speculation. If you are not an investor of AI companies then there is nothing to worry about for you. If you are an investor, then you should know that people can’t really predict when bubbles burst. Every prediction in the markets is a speculation and some investors can simply bets against the popular expectations to make good money.
Movements of AI stocks shouldn’t be confused with “AI as a technology” and “AI as a business”. Market valuation is a different game.
It appears very unlikely we will ever see an IPO of OpenAI. Anthropic appears less doomed, but still iffy at best. Tons of other large, but little discussed, AI startups are just dead-companies-walking at this point.
The likes of AWS are showing good headline numbers but are taking out massive debt to build infrastructure that looks increasingly unneeded. Those with capacity are looking to offload it, quickly. Yes AWS has “committed contracts” for this capacity but if those commitments are with shaky AI startups then it’s mostly just fluff PR and these hyperscalers will get left holding the bag on all this debt.