NewsLab
Aug 28 13:40 UTC

GLM-5.3-Flash (z.ai)

1,117 points|by Philpax||564 comments|Read full story on z.ai
https://news.ycombinator.com/item?id=49450353

Comments (564)

120 shown|More comments
  1. 1. rahimnathwani||context
    Related: https://news.ycombinator.com/item?id=49446422

    (281 points, 118 comments)

  2. 2. iamsyr||context
    Standard API Pricing for GLM-5.3-Flash (per 1M tokens)

    - Input: $0.15 - Output: $0.50 - Cached input: $0.03

  3. 3. Xunjin||context
    Is that cheaper than DS4 flash?
  4. 4. javier123454321||context
    All I can say is that even if it is, I was almost glad to go back to using DS4 Flash. Because 0XAlpha was just so friggin slow to complete a task because of the level of circular reasoning that it would go over and over into, sometimes even returning no output. If I just wanted something done I would switch from a free model to a paid one which is crazy.
  5. 5. denysvitali||context
    Tbh it was also slow because it was being hammered by everyone making use of the free tokens
  6. 6. javier123454321||context
    Possibly influenced by that, but I believe that is a different issue. I meant the way it processed a request. It went into so many more loops of thinking.
  7. 7. swiftcoder||context
    It's even cheaper than DS4's off-peak pricing. Seems like DeepSeek have some stiff competition now
  8. 8. arizen||context
    Few weeks ago, I wouldn't expect this statement to be true. Accelerate!
  9. 9. nateb2022||context
    Slightly more expensive than the (post-price hike) DS4 flash pricing, but in the ballpark.

    https://openrouter.ai/compare/deepseek/deepseek-v4-flash-073...

  10. 10. walrus01||context
    Comparison should be to 0731
  11. 11. nateb2022||context
    updated thanks
  12. 12. drob518||context
    Hm. GLM is more expensive in all dimensions than DS but it has a lower weighted average input? How is that?? Something seems off.

    EDIT: Looks like they are swizzling around the pricing dynamically on that page, on both the GLM and the DS sides, so who knows.

  13. 13. desterothx||context
    I mean, it implies that it has an even better cache hit rate that DS flash, which is impressive, as the chr on DS flash was already really good in my experience
  14. 14. nateb2022||context
    I think the weighted average takes into account all providers (some DS4 flash providers are 'premium' providers and offering higher speeds for higher pricing) and these are tilting the scale
  15. 15. epolanski||context
    I'm starting to think that this whole sanctioning China may motivate and prompt them to do more and better in every field.

    It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

  16. 16. esperent||context
    This has been clearly stated as what would happen going back several decades at least.
  17. 17. ricardobeat||context
    Starting? This was obvious way back in 2019, when the US decided to give China a little push developing their own silicon industry.
  18. 18. himata4113||context
    Well the big problem with china is that they do not respect international law when it comes to technology theft. But that argument is very weak when it appears that a lot of what they do is out in the open for anyone to replicate.
  19. 19. nananana9||context
    That's how you catch up when you're behind.

    Now the US is behind in EVs can you guess what they're doing? [1]

    [1] https://evwire.com/p/video-ford-ceo-jim-farley-says-they-fly...

  20. 20. himata4113||context
    "argument is very weak" regardless as I said.
  21. 21. cyanydeez||context
    yeah, America is totally out there respecting international law.

    "problem" indeed.

  22. 22. epolanski||context
    No major power respects nor cares about international law.

    Intellectual property is part of WTO agreements but enforcement is domestic.

    US companies do it too, regularly, they simply hire and poach staff from competitors.

    Proving it to be IP theft is difficult unless you can prove documents being passed. But often all you need is the know-how of the hired talent.

  23. 23. fwip||context
    There isn't one global "international law" for copyright. There are treaties that countries negotiate with each other.

    If the USA wanted a copyright treaty with China bad enough, we would negotiate one. China is not breaking any laws here, international or otherwise.

  24. 24. Natfan||context
    if the americans didn't want IP theft, they shouldn't have taught tens of millions of people how to make their IP
  25. 25. pshirshov||context
    > I'm starting to think

    That's good. Keep going.

  26. 26. abroszka33||context
    > It's too big, bright and resourceful of a country to choose confrontation instead of collaboration.

    It's not like we didn't try it. China first have to learn to make deals where both party benefits.

  27. 27. AshleyGrant||context
    I would say the Trump Administration needs to learn this as well.
  28. 28. abroszka33||context
    America knew how to do it. But they are learning quick from China.
  29. 29. sunbum||context
    > with all of this traffic served on Chinese AI chips

    RIP Nivida shareholders

  30. 30. ChoosesBarbecue||context
    God I wish I could’ve shorted NVIDIA right now
  31. 31. browningstreet||context
    It's earnings day for them...
  32. 32. re-thc||context
    Which 9/10 times hasn't been great anyway (stock reaction).
  33. 33. Bluestein||context
    Of course the release was not coincidental - with the earnings days - I am sure.-
  34. 34. vdfs||context
    And two models released same day + openai chip
  35. 35. kingstnap||context
    Whats stopping you? You could buy puts right now.

    Get a 210 strike put contract and if your thesis is that nvidias current 10 day slide continues you could make some money.

  36. 36. outworlder||context
    Unless NVidia craters you are likely to lose money given the IV crush that will happen today.
  37. 37. Bluestein||context
    This is the takeaway here: That's how they have been serving it at scale as Ox-Alpha. This is a definitional moment.-

    Further quote:

    "Compared with our initial baseline on the same hardware, we achieved a 3× improvement in end-to-end serving performance, reaching hardware efficiency and per-token cost comparable to mainstream NVIDIA GPUs. This demonstrates that Chinese chips can support frontier-model inference efficiently and economically at scale."

    https://z.ai/blog/glm-5.3-flash

  38. 38. xtracto||context
    Anyone knows what are those Chinese chips? Can they be bought? (Assuming im not i the US, And actually im in a 3rd world country).
  39. 39. gunalx||context
    They are high end really expensive Huawei ascend GPUs. It is kinda bruteforcing the performance on a older semiconductor processing tech, so total production is pretty low.
  40. 40. ayewo||context
  41. 41. bean469||context
    > comparable to mainstream NVIDIA GPUs

    By this they probably mean RTX series GPUs? If so, then they are not comparing the hardware efficiency with the A100 / H100, etc. that are commonly used for training models

  42. 42. rvz||context
    This is no surprise [0] [1].

    >> "They are already there on open weight models and Jensen knows that it is only a matter of time until China catches up with GPUs or other AI accelerators."

    It is also why Nvidia becoming a bank for other AI companies who are unable to find VCs to fund them isn't really a good thing and that is bearish.

    [0] https://news.ycombinator.com/item?id=49397204

    [1] https://news.ycombinator.com/item?id=49431231

  43. 43. ThouYS||context
    yay, I called it! :) (in the other thread)
  44. 44. dannyw||context
    Another self-inflicted own courtesy of US government policy.

    While I think China would always get to hardware self-sufficiency eventually, all export controls have done is (1) accelerate China's development, and (2) divert revenue that would've otherwise gone to NVIDIA/AMD/etc instead.

  45. 45. ignoramous||context
    The export controls were revoked before it triggered Chinese protectionism: https://www.silicon.co.uk/e-innovation/artificial-intelligen... / https://archive.vn/B2pah
  46. 46. mlinsey||context
    Revoked or not, just ever having those controls signals to the Chinese ecosystem that you're not necessarily a reliable supplier (Would you trust US export policy to remain stable for the next ~decade given the state of US politic?) and to the Chinese government just how strategically important you see these components.

    This isn't the kind of thing you can hash out in public and go back and forth on. Once you put it out there, the other party will take steps to make sure they don't have to rely on us in the long run.

  47. 47. re-thc||context
    > The export controls were revoked before

    Zai is on another "export control" list outside the broader 1. Doesn't help.

  48. 48. bigbadfeline||context
    The export controls were not revoked, only reduced, and not before, but after China refused to buy low performing chips. Top gear was and is still sanctioned, as is any EUVL equipment.
  49. 49. bigbadfeline||context
    And to add to the above: by building their own supply chain for chips, China is helping the unprivileged, those who can't front-run the market with long-term contracts. If China wasn't producing their own chips, the prices for us would be even higher.

    Similar to the war-pricing of oil, China's reduction of imports is actually helping to keep our inflation from going even higher.

  50. 50. anramon||context
    And it doesn't matter, it still pushed China to speed-up their AI related hardware development.
  51. 51. mrngld||context
    Long term it's irrelevant. The only relevant thing is that there's lots of money in chips that can do high performance inference. You see all kinds of competitor products in development or already on the market even here in the US where there are no such restrictions. Cerebras comes to mind. It's natural and expected that eventually Nvidia will either have to keep way ahead or competition will catch up with specialized products.

    That doesn't mean by any stretch of the imagination Nvidia will disappear. But the entire stock market valuation, not just tech, has had me scratching my head for a while.

  52. 52. bigyabai||context
    Cerebras "competes" with Nvidia in the same way a Vespa scooter competes with a Ford F-150. Groq and Tenstorrent are in a similar boat, ASICs don't really threaten CUDA.

    Curiously, there is not a single real CUDA competitor anywhere in the world. We almost had one with OpenCL, but all of the American stakeholders abandoned it right before the crypto/AI takeoff. All of which means that Nvidia sets their own margins, exploiting American investors and taxpayers while letting China avoid their dominance. So the American economy subsumes the bulk of Nvidia's arbitrarily-priced debt, and the Chinese economy can direct SOEs to pour billions in liquid cash into real GPGPU research.

    I'm an American and I'm pretty fond of Nvidia, but Jensen was right about this policy; it gives China everything they need to actually replace CUDA. It's reminiscent of America's attempts to deprive China of ARM and Texas Instruments IP, only to end up swimming in unlicensed clones after refusing to sign an IP deal.

  53. 53. redox99||context
    Not really a brag: it ran like shit. Very slow (~20tps, VERY high latency) and it would timeout all the time.

    I'm sure the chips are fine, but they clearly didn't have enough capacity for the demand they had (that 100T/day claim was asbolute bs)

  54. 54. nchmy||context
    seems unlikely that they'll get nearly as much demand now that it isnt free
  55. 55. redox99||context
    Sure, although I still expect it to become the most used model on openrouter.
  56. 56. Aurornis||context
    Ox Alpha is a smaller model and it was running very slowly. Chinese AI accelerators are coming along, but nVidia’s lead is huge.
  57. 57. knowaveragejoe||context
    Has there been any confirmation about what that model even is?

    Edit: Ah:

    > This stealth model was developed and operated by ZAI, revealed to be ZAI GLM-5.3-Flash.

  58. 58. cortesoft||context
    It's also in this very announcement, in the first paragraph:

    > Before release, we tested GLM-5.3-Flash anonymously as ox-alpha on OpenCode and OpenRouter to gather user feedback. It quickly became the most popular model of the week — with all of this traffic served on Chinese AI chips.

  59. 59. VulgarExigency||context
    It was being served for free. They were almost certainly being overloaded.
  60. 60. Aurornis||context
    Presumably the efficiency numbers they're quoting are for the high concurrency state they were serving.

    RAM was probably the bottleneck for the amount of context they were offering.

    I assume it would run a little faster with lower concurrency but "RIP nVidia" is a little premature. The cutting edge inference hardware is amazingly powerful

  61. 61. Implicated||context
    > and it was running very slowly

    ... I'm at a loss for words here. It was being served for free. To the entire world.

  62. 62. Aurornis||context
    GPT-5.6 Luna is also served for free to the entire world with a tokens per second rate nearly 10X higher.

    > ... I'm at a loss for words here

    No need to be so dramatic. I think it's great that they're developing chips, but the whole "RIP nVidia" claim was overly dramatic.

  63. 63. computerex||context
    Do you know how much traffic luna was getting vs Ox Alpha?
  64. 64. vitorgrs||context
    Are you really comparing chatbot to agentic/code work?

    Why is Luna not free on OpenRouter? :)

  65. 65. nkjvhb||context
    Lead doesn't really matter anymore. I just ported a very old cuda library to rocm, so it can be run on MI300s. 2 years ago this would have been a nightmare. Today it was an afternoon.
  66. 66. throwdbaaway||context
    Exactly. Coding for inference is solved. CUDA is no longer a moat.
  67. 67. HDBaseT||context
    Ox Alpha was also serving 10T+ tokens a day for free.

    When it first launched on OpenRouter I was getting nearly 70 Tokens/second.

  68. 68. WarmWash||context
    I don't see a situation where subscription payers move outside American LLMs (chatgpt, claude, gemini)

    And I don't see a situation where serious API payers are OK with handing the Chinese state all their data. Like manufactures of decades past did and learned a hard, even existential, lesson for it. The state mantra has been "Collect and Copy" for a long time now, tech just hasn't had that moment to experience it yet.

    So that leaves local hosting/leasing, but one of those has totally non-practical economics and the other doesn't have enough compute to meet any kind of real demand.

    I also have yet to meet a single person who isn't neck-deep in the tech space mention a Chinese LLM. It's 100% the big American three.

    If anything it's custom chips from the labs that threatens Nvidia.

  69. 69. Jcampuzano2||context
    Genuine question but who do you put as the "three" in big three.

    Because I genuinely can't tell if you mean Google or SpaceX/X.ai lol.

  70. 70. WarmWash||context
    Google probably serves more tokens then OAI and Anthropic combined, even if many of those tokens aren't from explicit gemini requests, but from AI overviews and other service integrations.

    xAI is already selling spare compute, and basically exists just to gas spacex's perceived valuation.

  71. 71. pianopatrick||context
    I can easily see a situation where most non American AI usage is on Chinese models on Chinese chips though.
  72. 72. uhfraid||context
    What about the current situation, where serious API payers are increasingly OK with using open-weight models running on US providers?

    https://www.ft.com/content/32a70a3c-7d28-40b4-808e-36edb58c7...

  73. 73. rapind||context
    These open models serve as price / performance pressure. Not all tasks require frontier models and cheap open models can be quite good for in-app assistants, if you're building that sort of thing. We also aren't sure the subscriptions will continue to be sustainable. They're currently subsidized to the tune of 50-70x. As someone who is hitting limits weekly that would easily cost me over $10k month per sub.
  74. 74. computerex||context
    Casual consumers are using American models because their usage is low. As usage scales, the economics heavily favor open weight models. The API pricing from American companies is absurd. This is particularly true in an enterprise setting.
  75. 75. WarmWash||context
    Open weight model hosts don't have the compute to meet enterprise demand. A large part of why these models are so cheap is because overall demand for them is incredibly low. Back in May, Gemini alone was doing about a month's worth of Openrouter tokens every day.
  76. 76. computerex||context
    I disagree totally. DeepSeek raised prices because they couldn’t serve the demand. But there are tons of American vendors ready to fulfill it. Many enterprises, including the one I work for, are swapping to open weights.

    Why wouldn’t you?

  77. 77. nkjvhb||context
    I am not ok with handing all my data to American companies that are best friends with the American surveillance state. I still remember the Snowden revelations. Chinese companies are a much better option in that regard.
  78. 78. efficax||context
    you don't have to hand them your data, the models are available so you can run them on bedrock yourself (or use another US housed inference service). and for what it's worth in my job i have access to data that gives a picture of the way companies are doing inference, and they're using a lot of chinese models (deepseek-v4 is a huge percentage of inference requests for example)
  79. 79. bityard||context
    Most US companies that have anything to do with government, finance, medical, etc. already have contractual or regulatory obligations which prevent them from using Chinese hardware or services, even before the AI boom. That's a huge market.

    Nvidia will do just fine. (Disclaimer: not a shareholder. At least, not directly.)

  80. 80. cheema33||context
    > Most US companies that have anything to do with government, finance, medical, etc... That's a huge market.

    Compared to the rest of the world?

  81. 81. bityard||context
    I don't have any pie charts in front of me, but yes, I would estimate it's a decently big slice of the world market.
  82. 82. itemize123||context
    bigger than RotW
  83. 83. saberience||context
    Not really. Chinese AI companies were never using NVidia AI chips.

    This announcement doesn't really mean anything at all. It means the very few people who are already using Z.ai's API will continue to do so, but the vast majority of money going to Nvidia is through the massive amount of business going to Anthropic, OpenAI, and other western cloud providers and inference providers, who are mostly using NVidia chips for inference.

    Also, NVidia chips are still sold out and supply constrained.

  84. 84. mmastrac||context
    Weights on HF here: https://huggingface.co/zai-org/GLM-5.3-Flash

    I decided to take the plunge and get myself four sparks at a decent price (and bought the QSFP cables from AliExpress because they are literally 1/2 the price of Amazon), even knowing Apple was going to release new hardware and there's probably a spark 2 on the horizon. It looks like this is going to be a decent fit for what I need. I've been experimenting with a two-node DS4 and it's _good_ at some tasks, but it really just spins its wheels when it hits the limit of what it can reason through.

    I can offload mundane/basic tasks to DS4 on two sparks, but I've been pushing it harder on some novel work and it just can't run on its own at all beyond a certain complexity level.

    I would love to see an Opus-4.8-level local model but TBH I just haven't got there yet. The models I've tried so far _are_ good but they aren't able to solve tough technical challenges, regardless of harness/prompting/etc.

  85. 85. kilroy123||context
    > get myself four sparks at a decent price

    Wow, if you don't mind me asking. How and where?

  86. 86. mmastrac||context
    I bought 4x Asus GX10 with the 1TB option. I don't understand why, but it's the only model in the whole lineup that isn't priced insanely.

    They were briefly on sale with a $200-off coupon, but they show up on warehouse deals from time-to-time as well.

  87. 87. swiftcoder||context
    > it's the only model in the whole lineup that isn't priced insanely

    $4,000 isn't priced insanely? ye gads

  88. 88. a3w||context
    I thought 4000 in sum. No wait, 4000 per, plus tax. Or EUR pricing to similar accord. Ouch.
  89. 89. swiftcoder||context
    Yeah, that little cluster costs about the same as a brand-new Dacia Sandero.
  90. 90. Bluestein||context
    Yeah, yeah. BUT, will the Sandero be ... load-bearing? :)
  91. 91. KptMarchewa||context
    It can bear the load of a few people, at least.
  92. 92. Bluestein||context
    Hey, I can at least say you will get more value (or, at least more predictable value) from a Dacia than from Anthropic's tokens: "Upgrade to [SUPER DUPER] for 5x the [TOKEN-SERF PACKAGE] token use!"

    Actual net work doable with/intelligence supplied by the [TOKEN-SERF PACKAGE]: Unknown. Fluctuating.-

  93. 93. swiftcoder||context
    > you will get more value (or, at least more predictable value) from a Dacia

    I picked up an older Dacia Sandero for cheap a few years back - it's the best money I've ever spent on a car, hands down. That car does not quit.

  94. 94. PcChip||context
    to save others from having to look up what that is like I did, it's a car
  95. 95. esafak||context
    Yes, but it was $200 off!
  96. 96. swatcoder||context
    Compare to the cost of professional-grade tools in other trades and craft hobbies.

    Sure, $4000 can be a lot of if you're a casual hobbyist or are struggle to meet everyday lifestyle costs, but it's definitely not "insane" if this is the trade you make your living from or if you've established a lifestyle that affords disposable income for your hobbies.

    And for some people, $4000 for a device you have complete control over and can repurpose and tinker with to your own needs and curiosities is a much much more justifiable expense than a $200/mo rental for some narrow-access tool that somebody else controls.

  97. 97. swiftcoder||context
    It's "insane" compared to the $1,500 it should have cost before the RAM crisis
  98. 98. chews||context
    I am pretty confident that given a $200 subscription on any of the big labs, you're getting $4000-$8000 per month in subsidized tokens... do what you wan't with your dough... and I too have a spark that I got really early (October 2025), but no, economically it does not compare to what's runnable locally in terms of quality from the frontier models. Economically, it looks like for as long as there are subscriber plans, you're better off renting.

    Before getting the spark, I was just using a google colab account, their $49 dollar plan allows you access to h100's and I can run qwen there in a Jupyter notebook... and if I really need that web front end I can just use cloudeflair/tailscale/the local ssh client to reverse tunnel it.

  99. 99. overgard||context
    The cloud stuff is definitely a much better economic value, but I would argue:

    1. You learn a lot more running this stuff yourself (especially since you can poke at its internals if you're interested or watch the reasoning chain.) Just being a consumer of this stuff doesn't really teach you much about it other than model & harness specific tricks that become obsolete pretty quickly. (IE, your Claude.md from 6 months ago probably needs a rewrite). Which is fine, I don't think you're going to be "left behind" if you're not a hardcore AI enthusiast or anything (I'm not), but as a guy that's always been interested in computer science I want to see how it ticks.

    2. You can't really depend on this subsidization lasting forever IMO. I know the financials thing has been beaten to death but I guess I'm in the camp that it's good to be in control of your tools so that you can go elsewhere if the economics change.

    I like to check in with ccusage pretty frequently, and honestly like if I were paying API prices for Claude I'd probably be paying thousands a month.

  100. 100. lukan||context
    3.privacy

    Any organisation or individuals not wanting to have their sensitive data flowing away (either because of trade secret or data protection laws)

  101. 101. robotresearcher||context
    Or good old fashioned privacy.

    There’s no law or business advantage preventing me giving my financial transaction and medical info to Google/Anthropic/OpenAI but I just don’t want to.

  102. 102. jrockway||context
    I am also not sure I would choose to use the cheap and easy to run at home model, given a choice. The marketing copy says this is a frontier model, but it's not. Sol and Mythos are the frontier right now. GLM 5.3 Flash simply isn't. I'd rather use the frontier model as they waste less of my time than even Opus.
  103. 103. booty||context
    Anecdotally, ~$500-1500/month token spend at API OpenAI/Anthropic pricing seems pretty realistic for full-time engineers at companies with "liberal but not unlimited" LLM spend policies.

    This is of course anecdata. I know plenty of outliers, too. I know a principal engineer who uses many multiples of the number I quoted above. I am sure we also know many people making do with much much smaller budgets as well, via all kinds of well-discussed methods.

    But, "$500-$1500 per month per full-time developer" is just kind of the personal mental baseline I use when making my decisions with regards to thinking about whether any of this makes any economic sense.

  104. 104. chasd00||context
    This should be obvious but with a model running on local hardware you can do your own RLHF and mod its behavior however you see fit. With cloud hosted models you can't. A few years ago when the models were smaller there were people undoing the guardrails, censorship, and general lobotomization with some form of a RLHF training. You can't do that on larger models unless you have the hardware like this person does.

    Notice all the comments saying like "omg why so expensive so just use the API??". It's a trick for lockin even with, so called, "open" models. Keep trying to run them locally, keep undoing the lobotomies, mod model behavior so that they work for you and do what you want vs only what someone else says they're allowed to do.

  105. 105. rudedogg||context
    > You can't do that on larger models unless you have the hardware like this person does.

    Or just rent something substantial for like $4/hr on runpod or w/e to do that.

    My gripe is this persons compute is wasteful and makes it harder for me to buy something with like 64gb ram to do normal work and run containers while I keep using cloud models.

    Someone else calculated the break even being 10 years, it’s just dumb. And I think it’s clear there won’t be a big rug pull anymore, there are too many open models and providers now.

  106. 106. girvo||context
    I love my Spark-like, but even for training you're better off using Vast or Runpod or whatever to rent cloud compute. Much faster and cheap as hell, to be honest.

    I do set up my initial runs and likes like quantisation-aware-distillation on my Spark-like to test it out and get it working, so it has value! But its not "worth" it other than its fun hardware to tinker with, IMO.

  107. 107. bredren||context
    With multiple 200 a month subs you are getting a multiple of those subsidized tokens. At least if you tabulate at retail api prices.

    This rent in the era of expensive hardware thing is not exclusive to inference.

    I’ve needed x86 architecture for windows builds recently and have just hemmed and hawed over buying a decent windows 11 box.

    I can’t make the math work against Azure instances.

    I can spin up a nice one for build deallocate,spin up something cheaper for QA and then turn that off.

    I can build all the devops around that, with a number of passes, with a skills based interface so working with the cloud is not too bad.

    The only thing that still has me thinking about it is the prospect of price is going up even more, which is acid as far as I know.

    And I’m hopefully going to need this x86 stuff enough that I don’t wanna wish I had gotten one for that high prices now.

  108. 108. vehemenz||context
    That's only half the reason it's expensive.

    The other reason is that it would likely take years to spend $4000 (plus the real cost of electricity) worth of tokens on a 3rd-party provider that's running a similar limited, DS Flash type model. By that time, the hardware will be obsolete, assuming it's still operational.

  109. 109. swiftcoder||context
    > it would likely take years to spend $4000 (plus the real cost of electricity)

    Since that cluster only yields 20-30 tok/s on that size of model, at least a decade before the hardware breaks-even with current token costs, and that's not counting electricity. Assuming continued downward pressure on token prices, and the cost of electricity, it never pays for itself.

  110. 110. dannyw||context
    As a counterpoint, my homelab/home-LLM hardware has appreciated in value by about 60% since I bought it.

    Of course, it's not real unless I sell, and the value will eventually go down, but so far I have significant paper profits.

    Also, DeepSeek token prices are continuing to _increase_, not decrease.

  111. 111. swiftcoder||context
    > DeepSeek token prices are continuing to _increase_

    One increase does not a trend make. And the current crop of models are now undercutting deepseek flash...

  112. 112. Implicated||context
    You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right? Have you seen what's happening with Codex/Claude subscriptions? Deepseek raising API prices.. We've been getting subsidized tokens for some time now and as the hardware costs skyrocket these labs/people with inference compute are going to continue to clamp down.
  113. 113. swiftcoder||context
    > You can't possibly think that it's going to get cheaper and cheaper to pay for tokens though. Right?

    Absolutely I do. Each generation of open-weight models has come with significant efficiency improvements, and there are significant hardware gains on the horizon: both increasing competition from Chinese chip manufacturers, and new custom silicon from the established players. And unlike Anthropic and OpenAI, most of the pure inference providers aren't massively leveraged - the more hardware they can bring online, the cheaper they can serve tokens.

  114. 114. maxglute||context
    $40,000 GPU is like few pennies in sand. Only mildly hyperbolic. But a GPU fresh out of fab is $2000 after ASML, TSMC and inputs get their 50-75% margin, then somehow $40k laundered through US financialization / Nvidia margins. Commoditized GPUs shouldn't cost more than 1-2% current price once there's competition.
  115. 115. selectodude||context
    We’re getting subsidized training. The inference is not a loss leader. And since providers can run hardware at 100 percent 24/7 their per token cost is going to be far below mine, regardless of how long I’m willing to wait for a token to come out.
  116. 116. nkozyra||context
    I don't understand how people don't consider this.

    Plus you're spec'd out of near-SOTA level in months.

    The only reasons to actually do this are a) you have a lot of dispensable income and are a hobbyist/tinkerer, b) you have real, legitimate privacy concerns or, relatedly, c) you're doing something you don't want to get flagged

  117. 117. billiam||context
    "you don't want to get flagged"

    Ding!

  118. 118. swiftcoder||context
    What is actually getting you flagged by the openweights inference providers? Thus far I haven't hit any of the reverse engineering or infosec guardrails that Anthropic is so keen on
  119. 119. nkozyra||context
    While I'm sure some of the open weight providers do this as well, I think the comparison is frontier labs v local inference.
  120. 120. swiftcoder||context
    I'm not sure that is the comparison - the OP is planning to run an open weights model locally, the obvious comparison would be paying a hosted provider to run the exact same model