This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
Existing CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
ASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support.
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
We already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.
Reducing everything to national aggregates provides no insight into the strong negative externalities, imposed on the immediate surrounding communities, of unregulated industrial facilities. That's literally the reason we have zoning laws in the first place,
There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
I mean, it will obviously crash at some point. With so much pressure on token price to go down that means way less opportunity for margin for AI providers. OpenAI is in a pretty bad situation
The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)
I would very much like it if anything that moves with appreciable mass is governed locally just in case the link drops and/or latency suddenly goes up. Motion is very unforgiving and accidents will happen if that's not taken into account.
this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.
GPUs and memory have gone up in price. It's more expensive to buy a 1-2 year old video card than it was at launch, sometimes by a shockingly large factor. Laptop vendors have recently shipped flagship models with less memory than the previous model, because they can't match price expectations for a laptop.
i figured out why this comment is so confusing: this is actually a message from the past, around 2020. either that or simianwords is a time traveler that arrived today and hasn't read the news yet.
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
I would posit there’s no way in hell they’re getting sufficiently cheaper on a short enough time frame vs how much demand is sky rocketing. AI companies are seeing quarterly doubling of revenue if not more.
If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.
All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).
Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.
I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.
With the corollary that old hardware valuations will plummet with them.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
Nvidia buys the memory it uses on its GPUs, same as all other ASICs.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.
Nah, DRAM is easy and very regular, it's a transistor and capacitor plus a massive decoder/encoder for addressing. The hard part is that it's a crushingly low margin business and without the added AI demand there were constant boom and bust cycles wiping out the manufacturers.
It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
I can totally see how ternary would work from a physical implementation perspective but I have a really hard time visualizing anything using base-e, can you explain how such a thing would work in practice?
I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
It isn't about bunk takes or not. It is about the motivation behind doing something.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
> "I'm not so worried about the motivation behind it besides how it biases their takes."
Troi oi. Read what you wrote again.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.
The guy that was part of FTX, fired from openai for alleged theft, got billions in a hedge fund somehow then lost billions. Why are all these people so scummy? It is like voting Trump three times in a row.
They're very intelligent people who do very clever things at a young age, which draws the attention of very rich people who can exploit them to get richer, and no one tells the young person they're being exploited. They're heaped with praise and 'wealth' (millions, but crumbs compared to what they're making for other people), and told they're geniuses who can do no wrong, mostly by the media that happens to be owned by the rich.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
> I love how now you have to consider the possible s*** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods
I mean, previously you could have said something much the same except substitute "frat boys".
The semianalysis people have scripts which incorrectly count their numerators and denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.
for example, their people think that GB300s are twice as fast as B300s, when really their benchmarks just incorrectly divide GB300 instances on azure by 8 instead of 4, since they don't read or verify any of the code that executes their benchmarks.
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
The "industry news and research" part of the AI industry feels very... suspect to me. My intuition is telling me that it's a bunch of people with influencer-y type social media skills and no actual credentials just grifting because there's so much money floating around.
When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from.
If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.
This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.
ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.
General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.
New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.
Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.
If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.
how much pollution do you believe modern gas-turbine engines to produce?
>Not to mention the water use controversy.
what percentage of US water usage do you believe is by AI data centers?
So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.
what makes you think this?
you really don't get it?
- gpus
- retail computers
- laptops
- ~gpu~ appliances like washing machines
- cloud computing
i think you don't get how economy usually works in tech
In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.
It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.
Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.
(I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)
This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.
Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets
Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.
The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.
They should GTFO with this crap.
Create barriers to computing for ordinary people while milking businesses for tokens.
There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.
To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.
If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.
Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.
If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.
One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)
They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.
The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.
I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.
Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.
Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.
Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.
I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.
SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?
Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.
Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.
It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.
I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.
> Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.
I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.
I enjoy the AI images done in their style though, it's pretty funny.
> AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.
NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.
Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.
And the cycle continues.
https://fortune.com/2026/03/16/peter-thiel-giving-pledge-bil...
Edit: OK, hn is removing one *
I mean, previously you could have said something much the same except substitute "frat boys".
Yeah, those guys aren't biased at all.
the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!
lol. lmao even.
Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.
Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.
If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.