One thing about these numbers that's absolutely shocking to me is how low the energy use is:
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
At some point the agents will be good enough that you can tell them "here's $100, go make me money" and they will, maybe not a lot and not all the time, but the EV given the cost of inference will be positive.
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
This seems somewhat obviously true, to some degree, but it all depends on the prompt, I think?
In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.
For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.
Those rewards seem like they would be captured by the providers and AI companies though who would use them first to make themselves money. You would be left with whatever they didn’t pursue with their first movers advantage.
I guess the question boils down to, what remote work could a human do to earn money over the internet?
So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)
"the talk of AI data centers' impact on the world" has been wildly exaggerated and you can see here that this is the least impact of any major new technology in the history of industrialization
The AI datacenter buildout is in the gigawatts. A single 1.21 GW data center multiplies that single-user 5.5 W average load (4 kWh per month) by as much as ~200 million. Obviously a widely distributed load is much less impactful on the surrounding environment.
It's running a normal central home heat pump/air conditioner for one hour, or eight hours of playing on a gaming PC. Yeah, the propaganda around datacenter resource usage has massively outrun the truth, which makes me think there are some very interested parties pushing behind the scenes.
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the "hyper" scaling of energy hungry datacenters in US part of a bubble?
Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
> Does it take into account the entire lifecycle, building the data centre
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you're at 2.5% of all-in cost of running these models is from electricity.
The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
Yeah the whole data center environmental panic seemed astroturfed to me. I took a look at the numbers and it’s not that bad. If you telework one day instead of commuting you make up for over a week of heavy AI use, and the water use is on par with an average golf course.
There are noise issues in some places. But the panic is excessive. Like nuts.
Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.
I'm running Qwen3.8-Flash-Next; other than speed on a 395+, it does most things that are properly planned out.
I can't believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.
And by defense, obviously, in America, we mean war. killing, etc.
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
I've been very excited with the most recent speed improvements for GLM5.3-Flash on DGX Spark clusters. It really feels close to what Opus ~4.5 was like to talk to. It's not quite there yet on consistency, but it's a really nice experience. Less guardrails and high quality abliterated versions further enhance its usefulness.
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
Yeah, I clicked expecting a review of GLM 5.3 Flash, but the article was more a retrospective about what he learned during his September challenge: "Only use GLM 5.3 Flash for one month"
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
Some vague commentary about performance with what appears to be assumptions about GPU availability, but no clarity about which inference provider is being used. If ZAI is assumed, I believe they aren't subject to the assumptions in the post based on what they've said publicly, but if they were using some other provider, perhaps.
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
> For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
I've been using GLM 5.3 Flash quite a lot for coding this month since they've had their promotion running. It's been tackling some difficult stuff - coding up a Julia version of the luminal GPU kernel optimization package, FPGA work with Verilog (including getting a BitNet 2B model running on FPGA - that one's been split between Claude and GLM), coding up a Julia version of DiffLUT. It's been handling these tasks pretty well. I do move between Claude Sonnet 5.5 & GLM 5.3-flash on the BitNet one based on what's available.
It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
Not OP, but I've used both quite a bit and I think GLM 5.3 (non-flash) is a vastly better model. The flash variant is good for general workhorse agents, but it doesn't seem to reason holistically about code and over-engineers solutions to each specific problem it solves. But if you use GLM 5.3 to write a detailed plan with little to no ambiguity, GLM 5.3 Flash executes it just fine for a fraction of the price.
I have a hard time justifying GLM 5.3 these days. It’s slightly better than Flash but rarely enough to justify the much steeper price. We chose to use usage-based billing only so are very sensitive to model price.
When text wuality or for pure but adwansed coding is concerned i always pick glm 5.3. The flash is awesome for everything that dosent really matter though.
Ill tell you from my personal use glm 5.3 flash was better then deepseek 4.1 flash, also deepseek liked to yap in his reasoning traces soo fucking much, the yapping was fast but the task was so slow to complete...
Interesting, I have roughly the opposite experience regarding speed: DeepSeek V4.1 Flash gets stuff done way more quickly for me than GLM 5.3 Flash. (I was getting 300+ tokens/sec with DS vs ~100 with GLM in my testing.) I agree that GLM 5.3 Flash is a slightly better model, but for me at least, it's not a huge difference, and I'd rather have DeepSeek's speed.
What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
The open model excitement isn't directly regarding cloud services.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
I agree with you 100%. That is the attraction, and the truth, and the marketing.
However, in groups, we are a deeply stupid species. Based on multiple previous conversations on this webstite, for example: z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of outside consultancies, the CCP is going to eat it all due to our outsourced laziness.
This is how we lost our manufacturing base, why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in this regard.
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
That number sounds about right, if a little low. According to OpenRouter, the weighted average input cost (which includes cache discounts) of GLM 5.3 is $0.2337 and the output cost is $3.291 per million tokens. If we assume 80% of the tokens are inputs, the cost of 450M tokens should be right around $300, which is the correct order of magnitude. And it depends highly on the provider(s) that the author is using, the ratio of inputs to outputs, etc.
It really makes you see how heavily subsidized the subscriptions are.
Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.
The high usage was due to omp in vibe mode overnight, probably working way too hard through things. The high cost, yes we’d rather pay extra to work with providers that provide other benefits than just lowest cost possible (open models, no training, EU DC, etc)
Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)
(tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)
There are some issue point not properly mentioned ...
* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
Edit: That are some interesting downvotes ...
To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.
Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).
Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.
For what it’s worth the energy numbers I get from Neuralwatt are within the ballpark of what is available elsewhere like https://cleerdash.sustainableaigroup.com/. What’s your source that there is profit / capex in there? Their published methodology seems pretty transparent
I'm skeptical that the DeepSeek official API actually nets out cheaper right now, but I avoid it anyway because they store and train off of your prompts.
I think NW's profits are mostly between what they pay for electricity and what they charge you for electricity. I don't think there's any need to look to conspiracies to explain billing errors.
One thing about these numbers that's absolutely shocking to me is how low the energy use is:
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
[delayed]
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
At some point the agents will be good enough that you can tell them "here's $100, go make me money" and they will, maybe not a lot and not all the time, but the EV given the cost of inference will be positive.
If I’m the AI company why would I let you do that, when I could just do it myself and get all of the money?
Regulation. If it got proven to the point that it was scaled out en mass, it would get hit hard and fast.
???
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
If that does become possible then the market will naturally establish efficiency again and the opportunity will disappear.
This seems somewhat obviously true, to some degree, but it all depends on the prompt, I think?
In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.
For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.
If the market would be fully efficient how would it be possible to make money in the market? I font think the market is that fast at rebalancing.
it'll basically be scam vs scam if that's every what happens.
Those rewards seem like they would be captured by the providers and AI companies though who would use them first to make themselves money. You would be left with whatever they didn’t pursue with their first movers advantage.
Good enough at what, fraud? Robotic Ponzi schemes would be an easy way to generate "income".
More likely they bot will do research on how to make make money and do experiments given the resources available to it.
I guess the question boils down to, what remote work could a human do to earn money over the internet?
So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)
"the talk of AI data centers' impact on the world" has been wildly exaggerated and you can see here that this is the least impact of any major new technology in the history of industrialization
The AI datacenter buildout is in the gigawatts. A single 1.21 GW data center multiplies that single-user 5.5 W average load (4 kWh per month) by as much as ~200 million. Obviously a widely distributed load is much less impactful on the surrounding environment.
It's running a normal central home heat pump/air conditioner for one hour, or eight hours of playing on a gaming PC. Yeah, the propaganda around datacenter resource usage has massively outrun the truth, which makes me think there are some very interested parties pushing behind the scenes.
Related to this: see the difference between China and US here: https://aidatacenterindex.com/compare/united-states-vs-china...
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the "hyper" scaling of energy hungry datacenters in US part of a bubble?
Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
> Does it take into account the entire lifecycle, building the data centre
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you're at 2.5% of all-in cost of running these models is from electricity.
The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
No those numbers are GPU only. See https://cleerdash.sustainableaigroup.com/ for a fuller model
> AI datacenter buildout is an overbuild
a.) we’re supply constrained
b.) only 3% of households pay for AI
Inference amounts will continue to grow heavily.
It would be a pretty big deal if the average person's car became 50% less energy efficient, no?
Yeah the whole data center environmental panic seemed astroturfed to me. I took a look at the numbers and it’s not that bad. If you telework one day instead of commuting you make up for over a week of heavy AI use, and the water use is on par with an average golf course.
There are noise issues in some places. But the panic is excessive. Like nuts.
Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.
I'm running Qwen3.8-Flash-Next; other than speed on a 395+, it does most things that are properly planned out.
I can't believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.
And by defense, obviously, in America, we mean war. killing, etc.
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
I've been very excited with the most recent speed improvements for GLM5.3-Flash on DGX Spark clusters. It really feels close to what Opus ~4.5 was like to talk to. It's not quite there yet on consistency, but it's a really nice experience. Less guardrails and high quality abliterated versions further enhance its usefulness.
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
That’s just a horrible article, very vague, wrong reasons, waste of time really.
Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.
Yeah, I clicked expecting a review of GLM 5.3 Flash, but the article was more a retrospective about what he learned during his September challenge: "Only use GLM 5.3 Flash for one month"
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
Some vague commentary about performance with what appears to be assumptions about GPU availability, but no clarity about which inference provider is being used. If ZAI is assumed, I believe they aren't subject to the assumptions in the post based on what they've said publicly, but if they were using some other provider, perhaps.
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
Might do more of that next time! Inference was with TensorX and Neuralwatt.
GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.
GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan.
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
Do you switch model mid session, or do you use subagent to do the switch after planning?
Neither, I use different sessions for each step.
> For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
I've been using GLM 5.3 Flash quite a lot for coding this month since they've had their promotion running. It's been tackling some difficult stuff - coding up a Julia version of the luminal GPU kernel optimization package, FPGA work with Verilog (including getting a BitNet 2B model running on FPGA - that one's been split between Claude and GLM), coding up a Julia version of DiffLUT. It's been handling these tasks pretty well. I do move between Claude Sonnet 5.5 & GLM 5.3-flash on the BitNet one based on what's available.
> BitNet 2B model running on FPGA Very cool! What kind of FPGA & What kind of TPS are you hitting?
It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
Considering they were your top two models, how did the flash and non-flash versions compare? Did you use them for different tasks?
Not OP, but I've used both quite a bit and I think GLM 5.3 (non-flash) is a vastly better model. The flash variant is good for general workhorse agents, but it doesn't seem to reason holistically about code and over-engineers solutions to each specific problem it solves. But if you use GLM 5.3 to write a detailed plan with little to no ambiguity, GLM 5.3 Flash executes it just fine for a fraction of the price.
I have a hard time justifying GLM 5.3 these days. It’s slightly better than Flash but rarely enough to justify the much steeper price. We chose to use usage-based billing only so are very sensitive to model price.
When text wuality or for pure but adwansed coding is concerned i always pick glm 5.3. The flash is awesome for everything that dosent really matter though.
Worth a comparison with DeepSeek v4.1 flash, if you've got another month to spare!
Yep, I think next month will be on that. It feels slightly better from a few days of use, and in our WIP benchmarking it scores way higher
Ill tell you from my personal use glm 5.3 flash was better then deepseek 4.1 flash, also deepseek liked to yap in his reasoning traces soo fucking much, the yapping was fast but the task was so slow to complete...
Interesting, I have roughly the opposite experience regarding speed: DeepSeek V4.1 Flash gets stuff done way more quickly for me than GLM 5.3 Flash. (I was getting 300+ tokens/sec with DS vs ~100 with GLM in my testing.) I agree that GLM 5.3 Flash is a slightly better model, but for me at least, it's not a huge difference, and I'd rather have DeepSeek's speed.
DS v4.1 Flash is roughly equivalent. It's going to get some things right/better that GLM flash doesnt and vice versa.
They've got an image of their homegrown benchmark in the post that lists DS4.1.
What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
The open model excitement isn't directly regarding cloud services.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
I agree with you 100%. That is the attraction, and the truth, and the marketing.
However, in groups, we are a deeply stupid species. Based on multiple previous conversations on this webstite, for example: z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of outside consultancies, the CCP is going to eat it all due to our outsourced laziness.
This is how we lost our manufacturing base, why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in this regard.
> 450M tokens / $150 / 5kWh
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
That number sounds about right, if a little low. According to OpenRouter, the weighted average input cost (which includes cache discounts) of GLM 5.3 is $0.2337 and the output cost is $3.291 per million tokens. If we assume 80% of the tokens are inputs, the cost of 450M tokens should be right around $300, which is the correct order of magnitude. And it depends highly on the provider(s) that the author is using, the ratio of inputs to outputs, etc.
It really makes you see how heavily subsidized the subscriptions are.
Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.
The high usage was due to omp in vibe mode overnight, probably working way too hard through things. The high cost, yes we’d rather pay extra to work with providers that provide other benefits than just lowest cost possible (open models, no training, EU DC, etc)
How did you measure energy usage?
Edit: I found a linked article that mentions the inference provider who does the measurements.
Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.
Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.
What would you use if you wanted to stay (at least) open weight?
qwen3.8, kimi3, kimi2.7, GLM-5.3 are all good families I use in my coding team
I'm mainly using flash varients, at least as the default, bump.up to stronger model as needed (less often these days)
qwen3.8 27B or 2.4T? they are completely different models with completely different pricing
I use all the qwen!
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)
(tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)
Deepseek 4.1 Flash and Mimo 2.6 Flash.
Flash is plenty good for planning and reviewing, for my needs. In fact, I use it for that because it's too slow for execution, despite the name.
edit: I subscribe to z.ai, I don't host.
Will agree on this. From z.ai i have found the non flash to have way more consistent performance.
> I use it for that because it's too slow for execution
What kind of hardware and what particular quant?
There are some issue point not properly mentioned ...
* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
Edit: That are some interesting downvotes ...
To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.
Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).
Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.
For what it’s worth the energy numbers I get from Neuralwatt are within the ballpark of what is available elsewhere like https://cleerdash.sustainableaigroup.com/. What’s your source that there is profit / capex in there? Their published methodology seems pretty transparent
I'm skeptical that the DeepSeek official API actually nets out cheaper right now, but I avoid it anyway because they store and train off of your prompts.
I think NW's profits are mostly between what they pay for electricity and what they charge you for electricity. I don't think there's any need to look to conspiracies to explain billing errors.