NewsLab
Aug 28 16:22 UTC

OpenAI Jalapeño: Better than Nvidia Blackwell (newsletter.semianalysis.com)

583 points|by bmulholland||379 comments|Read full story on newsletter.semianalysis.com
https://www.bloomberg.com/news/articles/2026-08-25/openai-cl..., https://archive.ph/yCTrr

Comments (379)

120 shown|More comments
  1. 1. ChoosesBarbecue||context
    This is most impressive. The interesting question to me, is outside of the LLM accelerator space: will generalized chips have massive leaps in performance once LLM technology is used to create the next generation? In general, will we see rapid advances while we extract the value of these models in creating architectures? I'm so far removed from the space that this is a very naive interpretation of all this, but I'm curious.
  2. 2. wmf||context
    Existing CPUs have been extremely optimized by ~6 competing, well-funded teams. I expect AI to accelerate things somewhat but it's not clear that there is any low-hanging fruit available for AI to find.
  3. 3. manquer||context
    ASICs always do better than general purpose chips. General purpose chips is turtles and turtles of virtualization and have to consider 4+ decades of backward compatible instructions set support.

    ASICs are deployed when the application area is economically large enough to so there is return on the investment in developing one. Bitcoin mining few years ago or today inference or more mundane things like video decoding/encoding.

    General purpose chips on the other hand have to be general purpose first to be useful, i.e. support as many application domains and instruction sets as possible . It can be long tail of support which both slow your chip down and also slow development down. Apple's took a long time to develop M series to be general purpose enough and still need even now software tooling like Rosetta to make say virtualization work for a good reason.

    New tooling would always help and there is already lot of software emulation for developing chips today but you still need physical iterations to tap-out and have high enough yield, no LLM can help with that.

  4. 4. anthonypasq||context
    Continued hardware improvements really make it hard for me to believe token prices will not continue to plummet.
  5. 5. datakan||context
    Token prices coming down means nothing if the models keep wasting them
  6. 6. gwerbin||context
    Hopefully this also means billionaires can stop trying to drop data centers into residential neighborhoods with zero noise control and polluting on-site generators, signing local politicians on with NDAs, calling for eminent domain to seize homes to build power lines to data centers, etc. etc. etc. Not to mention the water use controversy.

    Token prices plummeting is probably a good thing, but not without the regulatory backstops that prevent these effectively industrial facilities from being operated with no regard for the externalities they impose on people who live near them.

  7. 7. tmp10423288442||context
    Nah, Jevon’s Paradox says that cheaper tokens will mean increased overall energy consumption.

    If we can’t even build data centers, the least disruptive industrial use possible, there’s no hope to reindustrialize the US or anywhere outside of China.

  8. 8. gwerbin||context
    We already had plenty of data centers in the US before the AI boom that weren't severely harmful to their neighbors. Cutting red tape is not the same as eliminating meaningful regulation. There are plenty of old industrial sites that could be repurposed as data centers. It turns out it's cheaper to bribe some small town government to give you a tax cut and discounted electricity and water rate.
  9. 9. vlyan||context
    >polluting on-site generators

    how much pollution do you believe modern gas-turbine engines to produce?

    >Not to mention the water use controversy.

    what percentage of US water usage do you believe is by AI data centers?

  10. 10. gwerbin||context
    Reducing everything to national aggregates provides no insight into the strong negative externalities, imposed on the immediate surrounding communities, of unregulated industrial facilities. That's literally the reason we have zoning laws in the first place,
  11. 11. dgellow||context
    There is just so much downward pressure on token price, from every direction. We would need a completely new understanding of economics to explain why the price shouldn’t go down. Or market collusion/regulatory manipulation.
  12. 12. simianwords||context
    The price has been going down for ages, its not clear what you are pointing at
  13. 13. phoghed||context
    Pointing at the nay sayers who say tokens are heavily subsidized and it’s all going to come crashing down soon, surely any moment now
  14. 14. dgellow||context
    I mean, it will obviously crash at some point. With so much pressure on token price to go down that means way less opportunity for margin for AI providers. OpenAI is in a pretty bad situation
  15. 15. simianwords||context
    What does this have to do with margins? It can remain the same once prices go down
  16. 16. dgellow||context
    At the price going down? And that it will continue to go down, even if the hardware improvements stop. Not sure what isn’t clear
  17. 17. dumberquestions||context
    The demand for them is growing _per person_, not just across the wider economy, if tokens cost half as much but you want to use 3 times as much you're going to have to pay more.
  18. 18. jazzyjackson||context
    Maybe 1000s of tokens per second unlocks realtime robotic decision making, and now every robot needs to continuously stream tokens to and from the cloud to operate. That could 1000x demand overnight, just to speculate :)
  19. 19. hypfer||context
    Think about the agents buying computers for their agents. /s
  20. 20. jacquesm||context
    I would very much like it if anything that moves with appreciable mass is governed locally just in case the link drops and/or latency suddenly goes up. Motion is very unforgiving and accidents will happen if that's not taken into account.
  21. 21. HDThoreaun||context
    Seems unsafe to make locomotive decisions remotely
  22. 22. dgellow||context
    I think you just found what we will see in the S-1 prospectus of OpenAI
  23. 23. mathisfun123||context
    this is a story about a proprietary accelerator being built/designed by a token provider. and you think they're going to return the efficiency gains to the customer instead of capture the value for themselves? interesting take.
  24. 24. simianwords||context
    Yes, I can bet on this happening. If anything, this is a net gain for consumers as it is a competitive market.
  25. 25. mathisfun123||context
    go ahead and bet: alibaba is a publicly traded company
  26. 26. spacephysics||context
    We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss

    So as much as i agree “more profits to stakeholders screw the customer”, i think its more of an emergency to get to profitability before the music stops.

  27. 27. anthonypasq||context
    > We should be mindful of the context that many of these providers VERY likely have been selling their subscriptions at a substantial loss.

    what makes you think this?

  28. 28. RealityVoid||context
    Because everyone keeps saying this so it must be true. Real "it is known" kind of vibe with these statements.
  29. 29. polski-g||context
    He's a subscription truther. There's loads of them. OpenAI's profit increases with each subscription that is cancelled. Pretty soon they'll have more profit than God.
  30. 30. anthonypasq||context
    OpenAI just dropped the price of Luna by 80% and Sol by 20-30%
  31. 31. mathisfun123||context
    and amazon shipping used to be free without prime, and uber used to be cheaper than taxis, and airbnb used to be cheaper than hotels.

    you really don't get it?

  32. 32. simianwords||context
    almost every pure tech commodity has gone down in price

    - gpus

    - retail computers

    - laptops

    - ~gpu~ appliances like washing machines

    - cloud computing

    i think you don't get how economy usually works in tech

  33. 33. zirkonit||context
    I'm especially enjoying how RAM and SSDs are going down in price.
  34. 34. fl4regun||context
    GPUs and laptops and memory and storage are all crazy expensive
  35. 35. thefreeman||context
    listing gpu's here is crazy considering the current prices
  36. 36. fer||context
    I was checking laptops today for an upgrade from the model I bought back in 2019 and it's not gonna happen from how cheap they are.
  37. 37. asveikau||context
    GPUs and memory have gone up in price. It's more expensive to buy a 1-2 year old video card than it was at launch, sometimes by a shockingly large factor. Laptop vendors have recently shipped flagship models with less memory than the previous model, because they can't match price expectations for a laptop.
  38. 38. mathisfun123||context
    i figured out why this comment is so confusing: this is actually a message from the past, around 2020. either that or simianwords is a time traveler that arrived today and hasn't read the news yet.
  39. 39. ilaksh||context
    Yeah but is it really even as good as Rubin? Seems just competitive.
  40. 40. jrflo||context
    This may just be a classic case of Jevons paradox: https://en.wikipedia.org/wiki/Jevons_paradox

    In short, better hardware will drive down token cost in the near-term, but will drive up the demand for tokens as it gets cheap enough for other sectors to start to use it heavily.

    It comes from steam engines where economists originally thought that coal demand would plummet with more efficient engines, but it actually just meant that we found more uses for steam engines.

  41. 41. kilroy123||context
    This is exactly what I see happening now.

    Codex keeps doing these usage resets. What do I do? Burn even more tokens than ever before. I know I'm not the only one.

  42. 42. CapsAdmin||context
    Is this a normal thing now? I remember seeing this talked about as a surprising thing, but now I'm seeing posts about this as if it's normal.

    (I switched to using local models as usage limits, api instability and the concept of paying per token stresses me out)

  43. 43. anthonypasq||context
    the total cost spent on tokens may go up, but i just cant imagine per token costs going up
  44. 44. jrflo||context
    Depends on compute capacity. If we become supply constrained on tokens, then prices will necessarily go up.
  45. 45. anthonypasq||context
    no they dont because inference stacks are getting more efficient and models are getting more intelligent per parameter.
  46. 46. vlovich123||context
    I would posit there’s no way in hell they’re getting sufficiently cheaper on a short enough time frame vs how much demand is sky rocketing. AI companies are seeing quarterly doubling of revenue if not more.
  47. 47. sobellian||context
    If we are applying Jevons paradox to this then the unit being consumed is not tokens but the inputs for token production - power, capex, something else. To draw an analogy to the steam engine, coal:electricity::mechanical-work:tokens. Jevons paradox does not talk about mechanical work becoming cheaper in the short term setting up a sort of rubber band of demand creating spiking prices for mechanical work. Compared to the renaissance, mechanical work was much cheaper throughout the industrial revolution and remains cheaper to this day. We can still definitely say that the easier it is to produce tokens, the cheaper they will be.
  48. 48. vlovich123||context
    All Jevon’s paradox says is that as a resource becomes cheaper total consumption of that resource increases. It applies equally well to the inputs of token production as it does to the tokens themselves. The former would describe the effect the sellers into AI companies see (energy, GPU chips, RAM etc - if they lower their prices they’ll have more overall consumption) while the latter describes what the AI companies see with their customers (if they lower token prices consumers will use more tokens overall).
  49. 49. sobellian||context
    Jevons paradox states it might increase. It is not an ironclad law and there are many many cases where increasing efficiency wrt. a certain resource really will decrease the total consumption of that resource. Yes tokens are an input themselves, but this thread is discussing hardware that is more efficient at generating tokens. To increase the efficiency by which tokens are converted into some other product would require innovation in some other area - harnesses, the models themselves, better skill from the prompters, etc.
  50. 50. altmanaltman||context
    I think you're reducing a very complex thing (the global economy) into a very simplistic model (Jevons' paradox) and thinking both are the same thing. This has no predictive power or rigor. You're just wishing things would happen as they did before, without considering that conditions and situations change significantly, and instead of Jevon's paradox, we look back at today 50 years from now and talk about Jensen's paradox.

    This doesn't mean the concept is BS, but one single concept cannot explain away everything in such a system.

  51. 51. holoduke||context
    That's when demand is higher than capacity. Now imagine places like Gigalab and Chinese labs are online and able to produce significant percentage of chips. That could cause real surge in prices.
  52. 52. cactusplant7374||context
    It is incredibly cheap now. What sectors are you thinking of?
  53. 53. m101||context
    With the corollary that old hardware valuations will plummet with them.

    Although given we have marginal pricing we need to push through to those lower prices in the face of increasing demand, so timing of this is uncertain and the key to the AI financial markets

  54. 54. stymaar||context
    > continue to plummet.

    Continue what? The cost per output token has kept going up for the past three years across the board, as thinking models keep leaning more on test-time scaling.

    The quality of the said output tokens obviously increased, and arguably increased more than their price, but the price still went up. Or, on the flip side, the price of combined tokens went down (a bit, it did not "plummet" at all though) but so did the average token quality if you count thinking tokens.

  55. 55. fg137||context
    Counter point: Many AWS services barely decreased their prices (if at all) in the past decade despite advancement in hardware
  56. 56. varispeed||context
    Why they don't research how to make their own RAM and they have to buy it from the common market?

    They should GTFO with this crap.

    Create barriers to computing for ordinary people while milking businesses for tokens.

  57. 57. petcat||context
    Building a custom-designed ASIC is much easier than producing state of the art memory chips.

    There's a reason why Micron and Nvidia are the crown jewels of American technology right now and for the foreseeable future.

  58. 58. brcmthrowaway||context
    NVIDIA produces memory?
  59. 59. fc417fc802||context
    Fabless AFAIK. And that's the actual problem - drawing up CAD diagrams doesn't help if the factories are fully booked out.
  60. 60. Cyph0n||context
    A state of the art GPU is much harder to design & produce at scale and than an internal ASIC.
  61. 61. JV00||context
    Nvidia does not make RAM
  62. 62. varispeed||context
    That doesn't excuse them from wrecking the market for ordinary person.
  63. 63. chris_money202||context
    Nvidia buys the memory it uses on its GPUs, same as all other ASICs.

    To give some context, Intel started making DRAM, I think they were actually the company that came up with modern memory techniques. They exited the market and pursued a more lucrative moat with CPUs.

  64. 64. datakan||context
    People keep saying stuff like this without understanding what it takes to make RAM. It's one of, if not the most, heavily patented things in the world. The second you dip your toes into those waters the lawsuits begin.

    If somehow you get around the patent issues, you're now faced with huge research and development costs, fabs to build, processes to sort out and all of that has very high failure rates.

    Last time I checked Micron was the largest patent holder in the world and even for them this is a hard area where they are number 3 in the market.

  65. 65. chris_money202||context
    RAM chips are not hard to produce compared to many other types of semiconductors; Intel started in the memory game and left because the margins weren't great and they were going to fold. The failure rates on these chips are actually very tolerable; you can have a very bad yield and still have a viable chip due to things like ECC.
  66. 66. datakan||context
    Intel entered the memory space because they partnered with Micron. They left the memory space when Micron pulled out of the partnership.
  67. 67. chris_money202||context
    Intel started making DRAM in 1970, Micron was founded in 1978.
  68. 68. varispeed||context
    Yes, it is difficult, but shafting working class is easy, therefor it is okay.

    If the rich decided to buy all drinking water, you would probably be saying that's okay, making water is difficult, shortly before dying.

  69. 69. imtringued||context
    Nah, DRAM is easy and very regular, it's a transistor and capacitor plus a massive decoder/encoder for addressing. The hard part is that it's a crushingly low margin business and without the added AI demand there were constant boom and bust cycles wiping out the manufacturers.
  70. 70. epistasis||context
    It's so funny to see FP4.... I remember 20 years ago being asked what sort of HPC we needed in genomics, and the answer was basically, "lower precision, faster" for the stuff I was working on. But FP4 is, well, almost comical.

    One thing not on that comparison table: die size. If I'm understanding that correctly, it's about the same as the Rubin, but at 1/3 the number of NVFP4 PFLOPs. (The text disagrees with the table, I'm taking the table as truth, perhaps that's wrong...)

  71. 71. nxtfari||context
    Agree, I remember when even half precision made its way into C# sometime around 2020 (I didn’t know much about ML then) and I thought, well I guess that’s a worthwhile tradeoff but I can’t imagine going lower. Lo and behold (1-bit Bonsai) how much lower you could go.
  72. 72. jacquesm||context
    Ternary?
  73. 73. jeffbee||context
    Knuth's base-e proposal enters the chat.

    They were right about everything 50+ years ago, but they didn't have the budget for the right hardware, had to write conference papers and books instead.

  74. 74. jacquesm||context
    I can totally see how ternary would work from a physical implementation perspective but I have a really hard time visualizing anything using base-e, can you explain how such a thing would work in practice?
  75. 75. jeffbee||context
    No it's impossible. But it would be optimal!
  76. 76. jacquesm||context
    Ah, the spherical cow of number bases :) Thanks for the response, that saved me a sleepless night.
  77. 77. Razengan||context
    Maybe we have to do what quaternions did for complex numbers and jump straight from 2 to 4
  78. 78. kroaton||context
    But Bonsai is garbage.
  79. 79. jimmySixDOF||context
    I love how now you have to consider the possible s** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods -- it's one of the best stories in AI that SemiAnalysis is not cut from the same cloth as Gartner McKinsey et al
  80. 80. verall||context
    semianalysis is pretty good
  81. 81. latchkey||context
  82. 82. 7thpower||context
    This would have been far more effective with 1/10 as many words.
  83. 83. latchkey||context
    Not really. It is a long story.
  84. 84. verall||context
    I really really disagree. Most of the prose is devoid of information. I find it hard to believe the author even read the whole thing one time.
  85. 85. verall||context
    I've been reading them since before all of the AI hype, and I've always thought they're pretty good. You a few spicy takes with the overview/opinions/benchmarks. Better than semiaccurate.

    The article you link says not a lot of criticisms with very many words, and the AI prose gets much worse towards the end, seemingly when the author also gave up on reading it. I am disappointing in the plagiarism though, especially of Ryan Smith.

    I am much more interested in what you think of the site though vs your own experiences running a GPU cloud. I've seen your comments on it for a long time, it's super interesting. So if you think their takes are mostly bunk I'd consider it way more than this hot aisle guy.

  86. 86. latchkey||context
    It isn't about bunk takes or not. It is about the motivation behind doing something.

    Their takes are fabricated in such a way as to drive clicks to their business, where they are printing money selling MNDA to the highest bidder.

    Dylan uses his influence as a service and it is borderline criminal. He just sued a whistleblower employee. It is so blatant, he even lives and works directly with people in power who feed him information.

    Kind of like how SBF used his altruism to cover up the fraud he was doing. Everyone thought he was a good guy, until they realized he wasn't.

  87. 87. verall||context
    I dunno, it's a blog, so I'm not so worried about the motivation behind it besides how it biases their takes. I think being close to people that feed you information might be prerequisite to the kind of information he sends out.

    I've seen the paid subscriber sections and it's nothing groundbreaking. I wouldn't/don't pay for it.

    SBF used his altruism to cover up fraud. If the SA benchmarks were fraudulent, that would be a big deal. If he's just "in bed with the AI companies", like, that's a big part of the reason it's such a popular blog?

    Also, I don't really have to think that Dylan is a good guy, and I certainly wasn't the only one that did not think SBF was a good guy. I get that it's really easy to call his implosion unsurprising after the fact, but it was truly unsurprising.

  88. 88. latchkey||context

      > "I'm not so worried about the motivation behind it besides how it biases their takes."
    
    Troi oi. Read what you wrote again.

    Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.

    SA benchmarks for inferencemax? Yea... AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast, it is about appearing better on the charts to sell more chips.

    It is a popular blog because it is an influence service. That's the whole point. Write things that get clicks.

  89. 89. verall||context
    > Troi oi. Read what you wrote again.

    I can get info from and even enjoy reading a biased take. It's not so hard to see where they are coming from. In the "semicon trash talk and rumors blogosphere", you take everything with a grain of salt, etc.

    > Dylan is a grifting fraud. I wrote a too long document explaining a ton of examples and you're handwaving it away.

    I've taken this to mean that it's your post so I went through it again. Some of the points are interesting but I don't think it's a very good case that he's a grifting fraud.

    I enjoy the AI images done in their style though, it's pretty funny.

    > AMD / NVidia put their best engineers on tuning, just for the benchmarks. It isn't about serving up inference fast.

    NV has enough "best engineers" to have a couple people tuning for one of the most popular public tok/$ benchmarks without sweating. IDK about amd.

  90. 90. tmp10423288442||context
    SemiAnalysis’ founder was roommates with Anthropic people, not OpenAI, so he may be slightly (very slightly) more objective here.
  91. 91. LogicFailsMe||context
    Along with Leopold Aschenbrenner so maybe not so much.
  92. 92. rustystump||context
    The guy that was part of FTX, fired from openai for alleged theft, got billions in a hedge fund somehow then lost billions. Why are all these people so scummy? It is like voting Trump three times in a row.
  93. 93. onion2k||context
    They're very intelligent people who do very clever things at a young age, which draws the attention of very rich people who can exploit them to get richer, and no one tells the young person they're being exploited. They're heaped with praise and 'wealth' (millions, but crumbs compared to what they're making for other people), and told they're geniuses who can do no wrong, mostly by the media that happens to be owned by the rich.

    Then the rich people pull the rug leaving them holding the bag, and they move on to the next young clever group.

    And the cycle continues.

  94. 94. thelastgallon||context
    The only effective altruism is making the mega-billionaires richer.

    https://fortune.com/2026/03/16/peter-thiel-giving-pledge-bil...

  95. 95. xyzsparetimexyz||context
    s** posting? sex posting?
  96. 96. msh||context
    shit posting
  97. 97. minimaltom||context
    Thats what I thought too but then it would be s**?
  98. 98. madspindel||context
    s**?

    Edit: OK, hn is removing one *

  99. 99. yjftsjthsd-h||context
    If it's trying to convert it to italics, you may have to use a backslash to escape them
  100. 100. masfuerte||context
    Or double them up: s****** gives s***.
  101. 101. 2001zhaozhao||context
    I like that to type s****** you had to type s************.
  102. 102. yjftsjthsd-h||context
    Or escape them;)
  103. 103. jareklupinski||context
    i see 'hunter2'
  104. 104. antonvs||context
    > I love how now you have to consider the possible s*** posting motivation behind analysis of a trillion dollar industry being conducted at a world-class level by a bunch of ex Reddit and 4Chan adjacent mods

    I mean, previously you could have said something much the same except substitute "frat boys".

  105. 105. FrustratedMonky||context
    "not cut from the same cloth as Gartner McKinsey et al"

    Yeah, those guys aren't biased at all.

  106. 106. Alifatisk||context
    Why censor yourself?
  107. 107. TiredOfLife||context
    Bots do that because other platforms remove or hide posts with bad words
  108. 108. onion2k||context
    Humans do it because they've been raised not to swear.
  109. 109. adabovehuman||context
    > because they've been raised to grant advertisers and their vile spawn more rights than humans
  110. 110. TiredOfLife||context
    People that have been raised to not swear either do not swear or replace words.
  111. 111. onion2k||context
    I f**ing do actually.
  112. 112. doctorpangloss||context
    The semianalysis people have scripts which incorrectly count their numerators and denominators all the time. All their benchmarks are flawed. It is such a slipshod operation and they charge exorbitant amounts of money for it.
  113. 113. ShrigmaMale||context
    Say more about this please
  114. 114. doctorpangloss||context
    for example, their people think that GB300s are twice as fast as B300s, when really their benchmarks just incorrectly divide GB300 instances on azure by 8 instead of 4, since they don't read or verify any of the code that executes their benchmarks.

    the problem is they're so cryptopilled, surprises are what they want. they don't look at surprises and think, "that's wrong." they look at surprises and double down!

  115. 115. subtlejellyfish||context
    The "industry news and research" part of the AI industry feels very... suspect to me. My intuition is telling me that it's a bunch of people with influencer-y type social media skills and no actual credentials just grifting because there's so much money floating around.
  116. 116. senordevnyc||context
    What credentials do you need to write a substack about an industry so it’s not grifting?
  117. 117. A_D_E_P_T||context
    > McKinsey

    lol. lmao even.

    Have you seen the quality of their output? I'd take Claude or ChatGPT Free Tier over advice from McKinsey these days.

  118. 118. empath75||context
    When people talk about the commodification of inferencing, they imagine a future where everyone has access to frontier models and can run them at the same cost, and what will actually happen is closer to the commodification of _oil_, where only a few companies have the scale to produce it at a competitive price, and advances like this are _why_.

    Once models are more or less interchangeable, the price of LLMs will drop to essentially the price of energy required to run them, and the big labs will be able to run them cheaper than anyone else.

  119. 119. simianwords||context
    I don't believe models will be commodified because each model is unique with strengths and weaknesses. Its not like Steel which is more or less the same no matter where you purchase it from.

    If what you said were true, you would hardly see people complaining about the quality of Opus 5 or good writing from Sol. But people do.

  120. 120. airspresso||context
    This depends heavily on what the use-case is. Yes, if it's a coder making software and having to read LLM output then writing style matters. If the LLM is used in an automated data processing pipeline with a capped level of complexity, entirely different aspects matter and LLMs become more interchangeable.