Rendered at 15:51:49 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
epistasis 20 hours ago [-]
One thing about these numbers that's absolutely shocking to me is how low the energy use is:
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
scottcha 15 hours ago [-]
On our journey at Neuralwatt (we are likely the provider he's using as we are the one that does all the energy observability and reporting in our cloud and I'm the CTO there) we quickly discovered that there is a large disconnect between what is in the press (like the rest of the AI narrative very subject to worst case but possibly unlikley future extrapolation of current trends) and what we see on the ground. We actually spend most of our time focused on various way energy constraints (including maximizing tokens/joule while maintaining perf) manifest in current datacenters and provide better ways to get more tokens out of the energy that are already in these data centers or already available but hard to utilize on the grid. So its really a technical constraint problem <it>today</it> rather than an impact problem and I think the press narrative could be better about this.
Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).
ThibWeb 6 hours ago [-]
(OP) yes all energy / carbon footprint numbers are from Neuralwatt. ty for your work! Switched from Standard to Pro last week, it’s great
polytely 6 hours ago [-]
Can you say anything about how much variance is there in energy use between models and effort levels. I'm very curious about the difference between the fronteer and things like GLM 5.3 flash.
12 hours ago [-]
stkdump 9 hours ago [-]
I have a computer with a 5090 on a smart plug at home running Qwen3.8 27B for agentic coding. On a busy day it can use around 5kWh, though on most days it is around 2kWh. I am sure cloud is more efficient because there is probably efficiency in running many parallel streams, some of the models have fever than 27B active parameters and the power draw of the non-GPU components is also spread over more GPUs. But still I also believe that the energy use numbers you get from the inference providers are a bit "beautified". After all, they still fight a political battle and have to show that it isn't all so bad.
Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can't believe that I am that far away.
jmiskovic 3 hours ago [-]
Inference vs training. It's training that takes enormous amount of power, both electrical and the compute. I suspect many many models get simply thrown away because they end up being too low on benchmarks by the time they are done. And some are not released to the public. So we learn only about tiny percentage of trained models and their environmental impact.
mapontosevenths 32 minutes ago [-]
Hey guys, has anyone seen my goalposts?
More seriously, that needs to be done once and then it can be used by millions of people. Divide the cost by all the users and it's trivial. It's certainly not enough to lose sleep over.
rhdunn 8 hours ago [-]
I suspect that some/a large portion of the build out is due to OpenAI/Anthropic/etc. just throwing hardware at the problem -- why bother trying to make training and inference efficient when you can just throw hardware/money at the problem. I remember NVIDIA talking about their supercomputer clusters with a large number of interconnected GPUs.
I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:
1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.
2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.
[1] Though the original idea for mixture of experts comes from a 1991 research paper (https://huggingface.co/blog/moe), so maybe a third area is research from Universities, etc.
yorwba 3 hours ago [-]
The incentives are rather in the opposite direction. Enthusiasts are trying to get models to run at all on the hardware they have, even if the resulting efficiency is poor, whereas a big company spending billions on hardware can afford to hire hundreds of performance engineers to tune their systems end-to-end for maximum efficiency, since the expense pays for itself even if they only manage to eke out a 1% improvement. I.e. why not make training and inference efficient when you can just throw money at the problem?
darkwater 45 minutes ago [-]
Because at the moment they have basically infinite money and it's almost always faster to throw hardware at the problem than optimizing the solution, given infinite money.
holbrad 2 hours ago [-]
This is such a hilarious comment when we've literally just had Sol 6 and Opus 5.5 drop that are massive wins for efficiency. I mean, how out of touch can you be? Especially if you think power saving and efficiency come from local AI enthusiasts. What a fucking joke.
rpdillon 7 hours ago [-]
Similarly with my Strix Halo running DSV4. Got a meter for it a couple of weeks back, and, during inference, it's pulling about 160W, and I run it about 15 hours a day on a busy day, or about 2.4kW. I'd say an average day is more like 2 hours for me. It's a small impact on my monthly bill. Owning a hot tub is orders of magnitude worse, in my experience.
gchamonlive 16 hours ago [-]
> With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
This was going so well until this. Everything at scale has environmental impact because you centralize the downside and distribute the upside. This is an important alienation, but it hides the amount of heat, noise and impact on distribution that datacenters have on local infrastructure and environment.
hedora 15 hours ago [-]
Yes, but, I can run qwen 3.8 flash next (> 100b parameters; ~ opus 4.6) on a 200W tdp desktop.
So, assume 8 of those, and it’s a space heater per house. I have never been disturbed by my neighbor’s space heater.
The problem is centralization, not the absolute energy usage. (And also that LLMs are trending to 100x more efficient than the data center sizing assumed).
epistasis 14 hours ago [-]
Centralization is vastly more energy efficient than decentralization, by orders of magnitude. LLMS love batching, to an insane degree. Running a single conversation through a GPU is about the same cost as doing ~100 in parallel.
Centralization is a huge huge environmental win, far far more than even cloud computing was compared to tons of inefficient, under utilized racks spread though our office buildings.
hypfer 8 hours ago [-]
> Centralization is a huge huge environmental win
In a vacuum, yes. Assuming everyone is a good actor.
rpdillon 7 hours ago [-]
Totally true. But then it introduces a much greater degree of control, and it creates the aforementioned local concentration of energy usage, noise, etc. So it's a tradeoff, more than a strict win.
z0r 13 hours ago [-]
Huge environmental win if considering a fixed supply & demand, but Jevons Paradox kicks in with tokens in the cloud
rpdillon 7 hours ago [-]
This is very true. I'm more careful with what I ask my local model, since it's about 1/3 as fast as the cloud models I'm used to.
gchamonlive 14 hours ago [-]
Decentralization can also cause havoc because of back-EMF for instance, in any case the public infrastructure needs time to adapt to demand.
epistasis 14 hours ago [-]
What is your point?
Calculate the scale. Add it up. Do it.
Look at the numbers and come back to me.
gchamonlive 12 hours ago [-]
My point is the impact is real, measurable, but nobody will do it because nobody can stop the hype train lest they be labeled opponents to the progress
epistasis 2 hours ago [-]
These days even the AI companies are urging a slowdown if pace.
If somebody is afraid of speaking their mind, because the are afraid of being labelled an opponent to progress, that's some personal issue to work through.
It is popular and encouraged to be skeptical of AI, there no social opprobrium about it, unless you're in an extremist political cult, in which case you probably call it SI.
gchamonlive 2 hours ago [-]
> there no social opprobrium about it
You can't really say this categorically, because that's not a universal experience.
Maybe you are one of the lucky ones not having immense pressure at work to adopt AI at all costs, but the reality of it is that if I open my mike in a "debate" where I work with all devs and managers, to speak in favor of taking things easy and to give time for people to learn these tools properly, I'd be frowned upon.
So while it's true that's encouraged to be skeptical about the tech, this really depends on the context and saying it doesn't clashes hard with my experience of the world.
ls612 13 hours ago [-]
He is a Marxist (the alienation rhetoric gives it away) he is primarily interested in causing social revolution not in environmentalism or anything else.
12 hours ago [-]
gchamonlive 12 hours ago [-]
Labels are nice to fit what you don't understand in terms you feel like you understand. The local impact is real. Alienation is important. You are stuck in late 1800s rethoric.
Fidelix 4 hours ago [-]
You didn't deny anything he said
gchamonlive 4 hours ago [-]
If you think long and hard about it you'll see I did.
But if you must you can call me a moderately progressive Marxist.
ericd 17 hours ago [-]
It's running a normal central home heat pump/air conditioner for one hour, or eight hours of playing on a gaming PC. Yeah, the propaganda around datacenter resource usage has massively outrun the truth, which makes me think there are some very interested parties pushing behind the scenes.
cyberrock 13 hours ago [-]
I don't discount the possibility, but I think part of it is that no one was interested in knowing how many resources normal computing costs. Any one of us who's worked on a video platform or cloud storage providers already knew how the sausage was made, but as long as it provided unlimited videos and photo storage, the public didn't care.
Though I do suspect that golf course owners in Arizona and alfalfa farmers in California must be ecstatic.
ThibWeb 20 hours ago [-]
I agree but it does worry me how fast my usage is increasing. Two months ago I was using 10x less tokens and probably not much more than 5kWh on inference. This month about 30kWh on inference. If it becomes more affordable, is there going to be another jump? Not quite sure
driverdan 13 hours ago [-]
To put that in perspective my house uses an average of 1300kWh per month. Adding 30kWh per month would be an increase of 2.3% at a cost of around $4.20, not enough to notice.
4 hours ago [-]
ThibWeb 6 hours ago [-]
for me it’s more an environmental consideration than financial? Passivhaus Standard for houses is about 50kWh/m²/year, so 30kWh is getting close to 10% of total
oblio 4 hours ago [-]
Do you live in the US? You should know in that case that the US is already one of the most wasteful energy users on the planet, far ahead of most other developed countries, even.
So potentially increasing energy usage even more, especially with likely token usage acceleration, is hardly reason for celebration.
driverdan 2 hours ago [-]
No one is celebrating and using LLMs isn't necessarily a waste of energy. Where I live isn't relevant, LLMs use the same amount of power regardless of where their users are. You also know nothing about how I use electricity in my home, how I monitor my overall energy use and carbon output, and how I offset it.
nicoburns 1 hours ago [-]
> You also know nothing about how I use electricity in my home
I don't
> Where I live isn't relevant,
But where you live does matter. Average monthly household energy usage in the EU is 200-400 kWh. In the US it's more like 800-900kWh. So what you consider "normal energy usage" may vary considerably by where you live.
At 200-400kWh total usage, 30kwh starts to look a lot more significant. So you should at least consider the possibility that it's not that energy usage from AI is insignificant, but that your existing energy usage is wasteful.
oblio 1 hours ago [-]
You're American and the same 2.3% of energy you use would probably represents 5-10% of what the average Japanese person uses, and likely 15-20% of what the average Indian or Ethiopian uses.
My point isn't to attack you personally (I'd rather attack your country's disastrous energy policies) but to point out that LLMs will constitute a huge increase in our energy consumption at a time where we haven't transitioned to renewable energy and we're probably 20-30 away from it.
spiffytech 13 hours ago [-]
Neuralwatt prices proportional to the energy used (to cover all the non-energy costs). 30 kWh is $200 of service.
driverdan 12 hours ago [-]
Sure, they have to cover their other costs and make money. I was trying to put the energy use in perspective.
rapatel0 16 hours ago [-]
Yeah but performance / watt will decrease
- Quadratically with improved chip scaling
- orders of magnitude with improved model efficiency
skew-aberration 10 hours ago [-]
Quadratically? As in, reducing power used for (external) communication/memory busses?
rapatel0 4 hours ago [-]
It's a very very very rough approximation, but scaling laws (admittedly weaking over time) yield ~x^2 from a size, ~x^2 from a dynamic power, and more of the data movement is moving from PCB interconnect to on package which looks more similar to chip scaling.
It's a finger in the wind approximation with lots of confounders, but historically it seems to work out that way for most circuit things.
CuriouslyC 19 hours ago [-]
At some point the agents will be good enough that you can tell them "here's $100, go make me money" and they will, maybe not a lot and not all the time, but the EV given the cost of inference will be positive.
MisterMunchkin 18 hours ago [-]
If I’m the AI company why would I let you do that, when I could just do it myself and get all of the money?
CuriouslyC 18 hours ago [-]
Regulation. If it got proven to the point that it was scaled out en mass, it would get hit hard and fast.
monkpit 18 hours ago [-]
???
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
rhyperior 17 hours ago [-]
If that does become possible then the market will naturally establish efficiency again and the opportunity will disappear.
gunalx 17 hours ago [-]
If the market would be fully efficient how would it be possible to make money in the market? I font think the market is that fast at rebalancing.
epistasis 16 hours ago [-]
This seems somewhat obviously true, to some degree, but it all depends on the prompt, I think?
In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.
For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.
swiftcoder 4 hours ago [-]
> which does seem possible
But strictly time-limited, because as soon as someone else figures out what you are doing, there's no moat to prevent them from telling their AI to do it too.
Markets reach equilibrium very fast when everyone has access to the exact same tools and information.
cyanydeez 17 hours ago [-]
it'll basically be scam vs scam if that's every what happens.
mannanj 18 hours ago [-]
Those rewards seem like they would be captured by the providers and AI companies though who would use them first to make themselves money. You would be left with whatever they didn’t pursue with their first movers advantage.
pphysch 18 hours ago [-]
Good enough at what, fraud? Robotic Ponzi schemes would be an easy way to generate "income".
CuriouslyC 18 hours ago [-]
More likely they bot will do research on how to make make money and do experiments given the resources available to it.
lukan 17 hours ago [-]
I guess the question boils down to, what remote work could a human do to earn money over the internet?
So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)
a3dds 16 hours ago [-]
Bro.. your brain is fried.
Take a vacation.
timmmmmmay 18 hours ago [-]
"the talk of AI data centers' impact on the world" has been wildly exaggerated and you can see here that this is the least impact of any major new technology in the history of industrialization
DragonStrength 6 hours ago [-]
Exaggerated or poorly represented? One candidate for governor in my state likens it to the railroad bubble. Satya Nadella makes the same comparison. What happens when the bubble pops? We're seeing data centers outbid industrial plants for sites. Plants which would provide 100x as many jobs as a data center and not be abandoned in 2028 when the chips still aren't available and contracts come due.
Oh right, just some rubes in flyover states that don't know what's good for them.
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the "hyper" scaling of energy hungry datacenters in US part of a bubble?
zozbot234 17 hours ago [-]
The AI datacenter buildout is in the gigawatts. A single 1.21 GW data center multiplies that single-user 5.5 W average load (4 kWh per month) by as much as ~200 million. Obviously a widely distributed load is much less impactful on the surrounding environment.
BTW, according to the article the full GLM 5.3 model is in the "10x to 100x" ballpark. So this is not outside the realm of plausibility.
epistasis 13 hours ago [-]
Average US electrical grid load is 500GW.
As we electrify and decarbonize, we are looking at going to 1000-1500GW average load (clean energy sources are 2x-6x more energy efficient than fossil fuels at delivering energy services , so looking at the primary energy flows here you'll not only see elimination of the "rejected energy" category, but heating by fossil fuels gets counted as ~100% effect when heat pumps are 200%-600% effect on those terms https://flowcharts.llnl.gov/ )
Going to 50GW of AI would be an amazing jump for which we don't have fab capacity anytime.
etdznots 6 hours ago [-]
Inference needs very little compute, it’s mostly IO. hence the low power draw, training is compute-heavy and uses lots of power
apitman 1 hours ago [-]
Honest question: if that's the case why do my GPUs draw their full TDP during inference?
etdznots 15 minutes ago [-]
I am not an expert on why, but my vague answer is you are still running near max clock speeds, and you can significantly downclock your gpu and/or set power limits on your GPU without seeing much loss in performance, e.g. you can cap a 3090 to 60% of it’s normal power draw and it will lose like 5% of tg performance and 10% of pp performance (https://www.reddit.com/r/LocalLLaMA/comments/1hg6qrd/relativ...)
I think theoretically this is still very wasteful with lots of compute that’s getting powered and sitting unused but you are limited either by the firmware or by the GPU architecture from pushing power usage even lower without cratering performance. (no one anticipated the demand for relatively low-compute devices with lots of super fast memory)
killingtime74 20 hours ago [-]
Your only talking about variable direct energy. Does it take into account the entire lifecycle, building the data centre, running the cooling, building the chips, the % the chips are not utilised.
epistasis 18 hours ago [-]
> Does it take into account the entire lifecycle, building the data centre
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you're at 2.5% of all-in cost of running these models is from electricity.
The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
It would be a pretty big deal if the average person's car became 50% less energy efficient, no?
api 18 hours ago [-]
Yeah the whole data center environmental panic seemed astroturfed to me. I took a look at the numbers and it’s not that bad. If you telework one day instead of commuting you make up for over a week of heavy AI use, and the water use is on par with an average golf course.
There are noise issues in some places. But the panic is excessive. Like nuts.
Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.
epistasis 13 hours ago [-]
Don't need any astroturfing to explain the data center panic, it's well within organic NIMBYism. I've been trying to get housing built in my town for a decade, the crap stuff say about housing and the fervor with which they block housing is even greater than the data center panic.
Add in 1) all job destruction that Sam Altman and other leaders tout, 2) Musk operating an environmental disaster in the most offensive way possible for xAI, and 3) general fear bait a new technology that has such uncertain consequences, and I'm surprised the backlash isn't stronger.
Go talk to community members about any change in the buildings in their area and the data center backlash fits right in line with normal responses.
duskdozer 7 hours ago [-]
How is it excessive and nuts? If anywhere else is like it is by me, they've been ramming a bunch of giant warehouses at the least, and the increased noise both from the operation itself and the loss of natural barriers blocking other noise, light, and air pollution is really noticeably degrading things. And from what I understand, data centers are even worse due to their increased use of cooling and things like gas generators. It shouldn't come as a surprise to anyone that people are pissed off and increasingly so.
gmerc 3 hours ago [-]
But bro we need to shoot the DC to space because UNLIMITED POWER
christkv 8 hours ago [-]
Not only an overbuild but one that’s not real as there is a huge difference between talking about building and actually doing the building. Announcements are not the same as execution.
jjcm 18 hours ago [-]
> AI datacenter buildout is an overbuild
a.) we’re supply constrained
b.) only 3% of households pay for AI
Inference amounts will continue to grow heavily.
teaearlgraycold 14 hours ago [-]
> With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
Well there's also training to consider. But yeah, most of what I see online about datacenters is nonsense. People worried about water use as a top concern are misinformed.
Look, I absolutely get it if you're worried your boss is just waiting for the day to replace you with an LLM. Fucking get organized with your fellow laborers instead of getting distracted by datacenter water use. Call your politicians to rein in billionaires. If anything the ownership class is probably directing online discourse towards electricity and water user to keep people away from class-focused debates.
etdznots 6 hours ago [-]
No, because it’s a viable strategy that has teeth because of existing regs and agencies, not the battle to fight until the end of the earth, but it’s a pretty big hammer
redanddead 15 hours ago [-]
> My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
I never understand the mindset that leads to these massive leaps in reasoning. I was recruiting a guy, he was a VC GP, he says the same thing. I think he’s wrong. The more we use the more we want to use.
Do we seriously think our species will jinx the Kardashev energy requirements
epistasis 13 hours ago [-]
The unjustified leap here seems to be talking about Kardashev every requirements with relation to data centers.
Every massive technological leap that have big private buildouts results in massive overbuild. It's the nature of FOMO when a big civilization changing tech gets introduced.
Will AI change everything? Are too many data centers planned? The answer is almost certainly yes to both.
rpdillon 7 hours ago [-]
Didn't the overbuild of fiber directly lead to Google buying up dark fiber, and then lighting it up later on to power YouTube? I ran across this when I was trying to figure out why YouTube has no real competitors, and the explanations I ran across were basically “Google has an immense economic advantage because of their control over all this extra bandwidth capacity". I have not further validated that in the couple of years since I read it, but it does make me wonder what excess AI capacity could lead to.
swiftcoder 4 hours ago [-]
> Didn't the overbuild of fiber directly lead to Google buying up dark fiber, and then lighting it up later
Yes, but the lag there is decade-long, and a bunch of heavily fibre-invested companies went bust before anyone saw the upside.
In the same vein, it's very possible that early movers-and-shakers in the LLM space will overextend and end up bankrupt before LLMs find a long-term profitable niche
rpdillon 3 hours ago [-]
Yep, the decade-long timeline about matches what I recall. Agree on the culling that's coming in the AI space as well. Temporal's comment from a couple of days ago about inside vs. outside AI has been thought-provoking as I think about who is going to go bust...
redanddead 13 hours ago [-]
Regardless, we'll need the processing power and the energy
Is that best spent on LLMs, world models, physical AI, who knows. But we'll need the infrastructure, what's being built is conservative
cyanydeez 17 hours ago [-]
I'm running Qwen3.8-Flash-Next; other than speed on a 395+, it does most things that are properly planned out.
I can't believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.
And by defense, obviously, in America, we mean war. killing, etc.
dash-44 1 hours ago [-]
"Setting a challenge to spend the whole of September on only one efficient open model felt like a great idea at the time. Turns out not so much in practice."
This is just such a bad opening statement, almost pure clickbait. It makes it sound like the rest of the article is going to be on how it was a mistake to go with cheap models because their performance isn't good enough etc; the same usual stuff we see when someone tries to replace a frontier model with a cheap/flash model.
But the reality is that GLM 5.3 flash if a fantastic model; very cheap/performant and really a new era for this type of model, and the blog is just about some completely unrelated mistakes in development processes. Give us our time back.
aktenlage 20 hours ago [-]
> Unfortunately there are still consequences to it. I chose the 'wrong' model for the prototype, and we spent 450M tokens / $150 / 5kWh of energy use almost overnight. The MCP server itself works well and we now have a great demo of the capabilities, so it’s not for nothing:
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
ThibWeb 20 hours ago [-]
Hmm I might need to rephrase. Initial challenge was to use GLM 5.3 Flash and I was on the non-Flash version for the whole vibe coded build. Just wasn’t paying attention and I didn’t realize that one session was a quarter of the month’s spend and 150% of the budget (big price difference between models)
andai 7 hours ago [-]
That pareto graph isn't very helpful because it shows cost per token, but some models are way more token hungry than others. There are also massive differences in speed to complete a task.
My two favorite graphs are AA's A Index vs Time Per Task and AA Index vs Cost Per Task.
ah my bad, I’ve switched the visual to A index by cost per task. Note the blog post’s plot is filtered further than AA’s, based on what models are available with inference providers we can actually work with (Europe only). Vibe-coded filtering is here (very rough): https://pareto-eco-sum.netlify.app/
disiplus 20 hours ago [-]
Was this post generated with LLM, did he properly mention anywhere why exactly did it fail with example or i have trouble reading.
desmondl 19 hours ago [-]
Yeah, I clicked expecting a review of GLM 5.3 Flash, but the article was more a retrospective about what he learned during his September challenge: "Only use GLM 5.3 Flash for one month"
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
jensb1 6 hours ago [-]
Yeah agree, misunderstood the article intent as well.
raggi 19 hours ago [-]
Some vague commentary about performance with what appears to be assumptions about GPU availability, but no clarity about which inference provider is being used. If ZAI is assumed, I believe they aren't subject to the assumptions in the post based on what they've said publicly, but if they were using some other provider, perhaps.
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
ThibWeb 18 hours ago [-]
Might do more of that next time! Inference was with TensorX and Neuralwatt.
ydj 11 hours ago [-]
Agree, this post feels like it said very little. “Failed to use glm for a month because sometimes it’s slow and want to try other models too.” Okay, so I guess the only complaint about the model (or the service rather) is that it’s a bit slow.
cloudengineer94 3 hours ago [-]
I really love GLM 5.3 Flash, I have it working on my own Coding Agent which gets called by my Central Orchestrator in Hermes.
I created it's his own bot profile even...
shieldagent 3 hours ago [-]
A month of agentic coding for $68 is the headline for me. Model choice is becoming a budgeting decision as much as a quality one.
Aboutplants 3 hours ago [-]
In two years (or less) the frontier of today will be this cheap, which kind of blows my mind
dash-44 2 hours ago [-]
GLM 5.3 flash matches Opus 4.8 in general intelligence, and somewhere in the Opus 4.7–4.8 band for coding.
There was 4.3 months between Opus 4.7 and GLM 5.3 flash. Shorten your timelines ;).
hgoel 16 hours ago [-]
I've been very excited with the most recent speed improvements for GLM5.3-Flash on DGX Spark clusters. It really feels close to what Opus ~4.5 was like to talk to. It's not quite there yet on consistency, but it's a really nice experience. Less guardrails and high quality abliterated versions further enhance its usefulness.
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
vardalab 13 hours ago [-]
Yeah, I finally got a decent version of GLM 5.3 flash running on my dual sparks. It's not that fast (faster that sol 6.1 though, lol), but it's manageable. Qwen 3.8 on the other hand, that thing really performs on dual sparks. You can get consistent multiple (3-4) concurrent sessions at 40+ tokens per second. It's pretty nice. Good progress, but definitely not that cost effective. I mean, OPUS 5.5 has been pretty magnificent.
But nothing compares to when 4.5 came out. I still remember, it was just about Thanksgiving and that just was awesome.
monksy 20 hours ago [-]
GLM5.3-flash has been fantastic for me to make minor fixes in ambigious ways. "Fix x feature, whats going wrong. " It does the job.
girvo 14 hours ago [-]
Between GLM 5.3 Flash on my legacy Z.ai coding plan, and Qwen 3.8 Flash locally on my DGX Spark-like, I'm barely using my Anthropic/OpenAI subscriptions, likely to cancel them soon.
vardalab 13 hours ago [-]
It's been slow like molasses on the coding plan. I ended up using it more on fireworks. But DeepSeek is so much cheaper because the cache cost is much better.
drob518 2 hours ago [-]
Cache cost is the primary metric, imo. Most people look at output cost, but you only pay that once. Cache cost you pay every turn and it grows each turn as the context length grows.
girvo 13 hours ago [-]
I’m kind of lucky that most of the time I don’t have to use it through peak times so it’s not that bad speed wise
The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha
surgical_fire 20 hours ago [-]
GLM-5.3-flash is my implementation model after GLM-5.3 writes the plan.
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
drob518 2 hours ago [-]
Those are my two workhorses as well.
monksy 15 hours ago [-]
Same. It's barely even touching the credit I have on openrouter, it's great.
finnjohnsen2 18 hours ago [-]
Do you switch model mid session, or do you use subagent to do the switch after planning?
drob518 2 hours ago [-]
I sometimes switch mid-session. Depends what I’m doing. Sometimes a model is slow or it’s not giving me what I want and I don’t want to dump the context. But if there’s a natural break, I’ll start a new session with a new model.
surgical_fire 16 hours ago [-]
Neither, I use different sessions for each step.
finnjohnsen2 2 hours ago [-]
So your planner leaves markdown files I assume...?
surgical_fire 50 minutes ago [-]
Correct. I customized my harness so agents communicate with each other through markdown files.
I find that it makes coding and review at the same time less prone to errors and cheaper
Tepix 17 hours ago [-]
That’s just a horrible article, very vague, wrong reasons, waste of time really.
ThibWeb 1 days ago [-]
It was a bit of a silly challenge, wasn’t sure how workable, learned a lot in the process about what actually drives usage / costs, and how to keep both under control
cvburgess 1 days ago [-]
Considering they were your top two models, how did the flash and non-flash versions compare? Did you use them for different tasks?
RussianCow 19 hours ago [-]
Not OP, but I've used both quite a bit and I think GLM 5.3 (non-flash) is a vastly better model. The flash variant is good for general workhorse agents, but it doesn't seem to reason holistically about code and over-engineers solutions to each specific problem it solves. But if you use GLM 5.3 to write a detailed plan with little to no ambiguity, GLM 5.3 Flash executes it just fine for a fraction of the price.
ThibWeb 20 hours ago [-]
I have a hard time justifying GLM 5.3 these days. It’s slightly better than Flash but rarely enough to justify the much steeper price. We chose to use usage-based billing only so are very sensitive to model price.
gunalx 21 hours ago [-]
When text wuality or for pure but adwansed coding is concerned i always pick glm 5.3. The flash is awesome for everything that dosent really matter though.
sampullman 21 hours ago [-]
Worth a comparison with DeepSeek v4.1 flash, if you've got another month to spare!
ThibWeb 20 hours ago [-]
Yep, I think next month will be on that. It feels slightly better from a few days of use, and in our WIP benchmarking it scores way higher
nowittyusername 20 hours ago [-]
Ill tell you from my personal use glm 5.3 flash was better then deepseek 4.1 flash, also deepseek liked to yap in his reasoning traces soo fucking much, the yapping was fast but the task was so slow to complete...
sampullman 12 hours ago [-]
I've only compared it in a few tool assisted design tasks, and deepseek performed significantly better (and faster). I haven't had them do enough coding to get a feel for that, though.
RussianCow 19 hours ago [-]
Interesting, I have roughly the opposite experience regarding speed: DeepSeek V4.1 Flash gets stuff done way more quickly for me than GLM 5.3 Flash. (I was getting 300+ tokens/sec with DS vs ~100 with GLM in my testing.) I agree that GLM 5.3 Flash is a slightly better model, but for me at least, it's not a huge difference, and I'd rather have DeepSeek's speed.
18 hours ago [-]
samtheprogram 20 hours ago [-]
DS v4.1 Flash is roughly equivalent. It's going to get some things right/better that GLM flash doesnt and vice versa.
lnenad 21 hours ago [-]
They've got an image of their homegrown benchmark in the post that lists DS4.1.
nxobject 15 hours ago [-]
This wasn’t the author’s big point, but it did hit me when he said “this isn’t what we usually aspire to, but it works well for prototypes”. It captures how I feel about vibe coding - you can deal with incredibly complex tasks! But, boy, am I not about to use (say) an AI produced GPU driver on a daily basis! Perhaps we should normalize (again) “you’ll throw away your first attempt”.
fallinditch 18 hours ago [-]
> For day-to-day developer work, it’s totally viable to focus on one or two flash-tier cheap models
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
UncleOxidant 18 hours ago [-]
I've been using GLM 5.3 Flash quite a lot for coding this month since they've had their promotion running. It's been tackling some difficult stuff - coding up a Julia version of the luminal GPU kernel optimization package, FPGA work with Verilog (including getting a BitNet 2B model running on FPGA - that one's been split between Claude and GLM), coding up a Julia version of DiffLUT. It's been handling these tasks pretty well. I do move between Claude Sonnet 5.5 & GLM 5.3-flash on the BitNet one based on what's available.
buildbot 17 hours ago [-]
> BitNet 2B model running on FPGA
Very cool! What kind of FPGA & What kind of TPS are you hitting?
5 hours ago [-]
sheepscreek 19 hours ago [-]
> 450M tokens / $150 / 5kWh
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
RussianCow 18 hours ago [-]
That number sounds about right, if a little low. According to OpenRouter, the weighted average input cost (which includes cache discounts) of GLM 5.3 is $0.2337 and the output cost is $3.291 per million tokens. If we assume 80% of the tokens are inputs, the cost of 450M tokens should be right around $300, which is the correct order of magnitude. And it depends highly on the provider(s) that the author is using, the ratio of inputs to outputs, etc.
It really makes you see how heavily subsidized the subscriptions are.
Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.
ThibWeb 18 hours ago [-]
The high usage was due to omp in vibe mode overnight, probably working way too hard through things. The high cost, yes we’d rather pay extra to work with providers that provide other benefits than just lowest cost possible (open models, no training, EU DC, etc)
Edit: I found a linked article that mentions the inference provider who does the measurements.
ThibWeb 20 hours ago [-]
Yes, all from Neuralwatt, GPU energy use only. Makes models’ “efficiency” much more visible than tokens.
consumer451 18 hours ago [-]
What's really fun about this is that when people hear "open models," they rarely question where the inference happens, and who owns it, and how that can feed RL.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
hgoel 16 hours ago [-]
The open model excitement isn't directly regarding cloud services.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
consumer451 16 hours ago [-]
I believe that I might agree with you. That is the attraction, and the truth, and the marketing. Near frontier open models are truly amazing.
Avoiding 2 labs having control of knowledge work at the true frontier is not just good for nerd reasons. If only 2 labs control the frontier of knowledge work, well.. that's pretty much the end of the USA's market/service economy. I mean if one or two labs run it, does that make a country's economy?
Meanwhile, we are a deeply stupid species. Based on multiple previous conversations on this website, it appears that for example, z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of external consultancies, the CCP is going to eat it all due to our outsourced laziness. Shareholder value, amirite!!?
This is how we lost our manufacturing base. Why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in the last few years. I guess changing the name to SI and tariffs on Canada and the EU will solve these problems.
I think most posts are best read when there is a solution at the end, but is there one? What would that look like?
consumer451 8 hours ago [-]
I want to add to this: we didn't lose our manufacturing base just for shareholder value. The consumers also got super cheap and yet exponentially more amazing technology.
I am an old-ish person. My first PC cost as much as a used car. Thanks to global supply chains, we all get those in our pockets, all the time! Capitalism at its best!
If we could agree to stop killing each other, there is literally no end to this hockey stick.
This is the true contrarian POV in 2026, fight me.
What I mean is, Thiel/Yarvin followers. Let's have a talk about this.
zozbot234 16 hours ago [-]
> You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
consumer451 12 hours ago [-]
Yes, but there is only one reason that one might buy truly private inference. That is true privacy.
The price one pays for this is that lower than frontier models are dumber. At the current pace of advancement, is that a good idea? Isn't it better to just go with with ZDR contracts, and not lose in the software product game, which now moves at relativistic speed?
everlastingemai 15 hours ago [-]
I am excited and worried at the same time when it comes to these open weight models coming out of china..
throw930rmdkdk 21 hours ago [-]
Flash is pretty decent coder, but it should be paired with good planner and reviewer. I would pick astra low for planning and sol 6.1 medium for reviews.
eikenberry 21 hours ago [-]
What would you use if you wanted to stay (at least) open weight?
verdverm 21 hours ago [-]
qwen3.8, kimi3, kimi2.7, GLM-5.3 are all good families I use in my coding team
I'm mainly using flash varients, at least as the default, bump.up to stronger model as needed (less often these days)
Geof25 20 hours ago [-]
qwen3.8 27B or 2.4T? they are completely different models with completely different pricing
verdverm 20 hours ago [-]
I use all the qwen!
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)
(tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)
esafak 21 hours ago [-]
Deepseek 4.1 Flash and Mimo 2.6 Flash.
esafak 21 hours ago [-]
Flash is plenty good for planning and reviewing, for my needs. In fact, I use it for that because it's too slow for execution, despite the name.
edit: I subscribe to z.ai, I don't host.
gunalx 20 hours ago [-]
Will agree on this. From z.ai i have found the non flash to have way more consistent performance.
jminnl 20 hours ago [-]
> I use it for that because it's too slow for execution
What kind of hardware and what particular quant?
theunclej 8 hours ago [-]
[flagged]
lin7c 14 hours ago [-]
[flagged]
benjiro29 19 hours ago [-]
There are some issue point not properly mentioned ...
* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
Edit: That are some interesting downvotes ...
To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.
Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).
Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.
ThibWeb 19 hours ago [-]
For what it’s worth the energy numbers I get from Neuralwatt are within the ballpark of what is available elsewhere like https://cleerdash.sustainableaigroup.com/. What’s your source that there is profit / capex in there? Their published methodology seems pretty transparent
eli 19 hours ago [-]
I'm skeptical that the DeepSeek official API actually nets out cheaper right now, but I avoid it anyway because they store and train off of your prompts.
I think NW's profits are mostly between what they pay for electricity and what they charge you for electricity. I don't think there's any need to look to conspiracies to explain billing errors.
> That model’s usage was well within our budget ($68, about 4kWh of energy use / 365 grams of carbon emissions).
The energy cost is literally 1% of the total cost. For context, 4kWh of energy would drive you about 15 miles in an EV, about half of the average person's daily driving miles. It's boiling 10 gallons of water.
With the talk of AI data centers' impact on the world, you'd think this would be 10x to 100x the amount of energy in order to get the effects they're using here.
My takeaway: the AI data center buildout is an overbuild probably at least as large as the fiber buildout that left us with so much dark fiber. If not even bigger. The only thing that will save the economy is the inability of NVIDIA and chip fabs to produce enough chips to match the buildout planned.
Regarding total costs relative to the pure energy costs it is multiple orders of magnitude different but also realize in the datacenter the energy is the pure commodity while almost every other component has huge margins driven by lack of supply. I do think over time this might get closer together (more competition on HW might lower margins) while energy might become more of a bottle neck (raising the energy prices).
Having said that, seeing the incredible progress of models throughout this year, I also strongly believe that the planned buildout is overeager. Even I with my gaming hardware often run out of instructions to give. And the smarter the models get that I can run, the less I will be able to saturate my hardware. Is it because of my lack of creativity of which kinds of tasks I can give to AI? Maybe a bit, but currently I can't believe that I am that far away.
More seriously, that needs to be done once and then it can be used by millions of people. Divide the cost by all the users and it's trivial. It's certainly not enough to lose sleep over.
I think the power savings and efficiencies (that the large companies have also benefited from) have come from 2 areas:
1. open source and local AI enthusiasts -- think things like llama.cpp, quantization, etc.
2. Chinese labs and other smaller/research companies like Mistral that are using constrained hardware -- see the various advancements in the various models to reduce compute complexity such as mixture of experts [1], sharing key/value data between a group of layers, etc.
[1] Though the original idea for mixture of experts comes from a 1991 research paper (https://huggingface.co/blog/moe), so maybe a third area is research from Universities, etc.
This was going so well until this. Everything at scale has environmental impact because you centralize the downside and distribute the upside. This is an important alienation, but it hides the amount of heat, noise and impact on distribution that datacenters have on local infrastructure and environment.
So, assume 8 of those, and it’s a space heater per house. I have never been disturbed by my neighbor’s space heater.
The problem is centralization, not the absolute energy usage. (And also that LLMs are trending to 100x more efficient than the data center sizing assumed).
Centralization is a huge huge environmental win, far far more than even cloud computing was compared to tons of inefficient, under utilized racks spread though our office buildings.
In a vacuum, yes. Assuming everyone is a good actor.
Calculate the scale. Add it up. Do it.
Look at the numbers and come back to me.
If somebody is afraid of speaking their mind, because the are afraid of being labelled an opponent to progress, that's some personal issue to work through.
It is popular and encouraged to be skeptical of AI, there no social opprobrium about it, unless you're in an extremist political cult, in which case you probably call it SI.
You can't really say this categorically, because that's not a universal experience.
Maybe you are one of the lucky ones not having immense pressure at work to adopt AI at all costs, but the reality of it is that if I open my mike in a "debate" where I work with all devs and managers, to speak in favor of taking things easy and to give time for people to learn these tools properly, I'd be frowned upon.
So while it's true that's encouraged to be skeptical about the tech, this really depends on the context and saying it doesn't clashes hard with my experience of the world.
But if you must you can call me a moderately progressive Marxist.
Though I do suspect that golf course owners in Arizona and alfalfa farmers in California must be ecstatic.
So potentially increasing energy usage even more, especially with likely token usage acceleration, is hardly reason for celebration.
I don't
> Where I live isn't relevant,
But where you live does matter. Average monthly household energy usage in the EU is 200-400 kWh. In the US it's more like 800-900kWh. So what you consider "normal energy usage" may vary considerably by where you live.
At 200-400kWh total usage, 30kwh starts to look a lot more significant. So you should at least consider the possibility that it's not that energy usage from AI is insignificant, but that your existing energy usage is wasteful.
My point isn't to attack you personally (I'd rather attack your country's disastrous energy policies) but to point out that LLMs will constitute a huge increase in our energy consumption at a time where we haven't transitioned to renewable energy and we're probably 20-30 away from it.
It's a finger in the wind approximation with lots of confounders, but historically it seems to work out that way for most circuit things.
Then why would any AI company exist when they could just use their money to buy tokens from another AI company and make money for zero effort? There would be no incentive to be a provider.
Not to mention inflation would grow to match or outpace the rate you could earn on these guaranteed AI gains.
In the equilibrium, profits fall to zero in capitalism, but there's never equilibrium, everything is changing. Profit comes from insights that others haven't seen, from access to opportunities that others don't have, or through monopolistic control generating economic rents.
For that AI to make money, it has to have a harness that gives it one of the above things, which does seem possible.
But strictly time-limited, because as soon as someone else figures out what you are doing, there's no moat to prevent them from telling their AI to do it too.
Markets reach equilibrium very fast when everyone has access to the exact same tools and information.
So how many companies/individuals pay a random spam bot contacting them 100€ to upgrade their website? I guess start with targeting gullible and trying to scam them will be the more successful experiment for them. (If run unrestricted)
Take a vacation.
Oh right, just some rubes in flyover states that don't know what's good for them.
I have the feeling China is somehow ahead when it comes to energy (and cost) efficiency for AI usage. After all, the two are in a direct competition, and this difference is significant. Or is the "hyper" scaling of energy hungry datacenters in US part of a bubble?
BTW, according to the article the full GLM 5.3 model is in the "10x to 100x" ballpark. So this is not outside the realm of plausibility.
As we electrify and decarbonize, we are looking at going to 1000-1500GW average load (clean energy sources are 2x-6x more energy efficient than fossil fuels at delivering energy services , so looking at the primary energy flows here you'll not only see elimination of the "rejected energy" category, but heating by fossil fuels gets counted as ~100% effect when heat pumps are 200%-600% effect on those terms https://flowcharts.llnl.gov/ )
Going to 50GW of AI would be an amazing jump for which we don't have fab capacity anytime.
I think theoretically this is still very wasteful with lots of compute that’s getting powered and sitting unused but you are limited either by the firmware or by the GPU architecture from pushing power usage even lower without cratering performance. (no one anticipated the demand for relatively low-compute devices with lots of super fast memory)
My entire point is that 99% of the dollar cost of running these models goes to things other than the GPU power. The capex cost to building cost to GPU cost to storage/networking/chasses/wiring plus the other operation costs dwarf the electricity. Even the other electricity costs, lets say double it for all the supporting compute, plus another 25% for a 1.25 PUE, and you're at 2.5% of all-in cost of running these models is from electricity.
The non-electricity costs are massive and the constraints on fabs, etc. will drive the amount of the AI build far more than energy availability.
There are noise issues in some places. But the panic is excessive. Like nuts.
Maybe environmental panics are to the left what moral panics about stuff like trans people are to the right.
Add in 1) all job destruction that Sam Altman and other leaders tout, 2) Musk operating an environmental disaster in the most offensive way possible for xAI, and 3) general fear bait a new technology that has such uncertain consequences, and I'm surprised the backlash isn't stronger.
Go talk to community members about any change in the buildings in their area and the data center backlash fits right in line with normal responses.
a.) we’re supply constrained
b.) only 3% of households pay for AI
Inference amounts will continue to grow heavily.
Well there's also training to consider. But yeah, most of what I see online about datacenters is nonsense. People worried about water use as a top concern are misinformed.
Look, I absolutely get it if you're worried your boss is just waiting for the day to replace you with an LLM. Fucking get organized with your fellow laborers instead of getting distracted by datacenter water use. Call your politicians to rein in billionaires. If anything the ownership class is probably directing online discourse towards electricity and water user to keep people away from class-focused debates.
I never understand the mindset that leads to these massive leaps in reasoning. I was recruiting a guy, he was a VC GP, he says the same thing. I think he’s wrong. The more we use the more we want to use.
Do we seriously think our species will jinx the Kardashev energy requirements
Every massive technological leap that have big private buildouts results in massive overbuild. It's the nature of FOMO when a big civilization changing tech gets introduced.
Will AI change everything? Are too many data centers planned? The answer is almost certainly yes to both.
Yes, but the lag there is decade-long, and a bunch of heavily fibre-invested companies went bust before anyone saw the upside.
In the same vein, it's very possible that early movers-and-shakers in the LLM space will overextend and end up bankrupt before LLMs find a long-term profitable niche
Is that best spent on LLMs, world models, physical AI, who knows. But we'll need the infrastructure, what's being built is conservative
I can't believe the TAM requires more than it at a 2x speed up. OpenAI, Anthropic built models whose only purpose now is things like research and defense.
And by defense, obviously, in America, we mean war. killing, etc.
This is just such a bad opening statement, almost pure clickbait. It makes it sound like the rest of the article is going to be on how it was a mistake to go with cheap models because their performance isn't good enough etc; the same usual stuff we see when someone tries to replace a frontier model with a cheap/flash model.
But the reality is that GLM 5.3 flash if a fantastic model; very cheap/performant and really a new era for this type of model, and the blog is just about some completely unrelated mistakes in development processes. Give us our time back.
> Nonetheless, it’s a good reminder to be careful with model selection and with agentic patterns. We could have achieved similar results for most likely 5x less cost with not that much more effort. Lessons learned! We need to budget for this, and be more careful. Could have seen it coming, but now we know.
I don't get it. Why was it wrong? Which one would have been better? What was the lesson and how could you have foreseen it?
My two favorite graphs are AA's A Index vs Time Per Task and AA Index vs Cost Per Task.
https://artificialanalysis.ai/#intelligence-comparison-tabs
https://artificialanalysis.ai/?intelligence-comparison=intel...
He said his experiment was a failure because:
1. He accidentally spent 450M tokens vibe coding with the wrong model, instead of GLM 5.3 Flash.
2. When he used GLM 5.3 Flash, it was sometimes slow. So he switched to other models (Deepseek / Qwen) instead. His guess to why it was slow: GLM 5.3 Flash was so good that the providers were congested.
3. He still needed to use other models besides GLM 5.3 Flash, for R&D and benchmarking.
His takeaways from doing the experiment were:
1. Measure local usage more.
2. Experiment with agent orchestration, with bounded goals.
3. Don't count other models that are used for R&D.
4. Play with Jev.
5. Include experiments with flagship models to compare with cheap open models.
His conclusion about GLM 5.3 Flash: Probably viable for day to day work, but he'll have more thoughts next month.
The second reason appeared to be simply "because we chose not to". The post seems to be pretty much content-less in any practical sense. I clicked on it because I do quite like this models average performance and I was hoping to see some kind of review content.
I created it's his own bot profile even...
There was 4.3 months between Opus 4.7 and GLM 5.3 flash. Shorten your timelines ;).
Though, Qwen3.8-Flash-Next is very close to that level while requiring fewer resources to run, so I'm really looking forward to Qwen4.
The legacy plan I have is so good as to be basically unlimited usage for my workloads, so I’m kind of stuck with it til they stop renewing it haha
It's an excellent workhorse. When I am running out of my GLM quota I switch GLM-5.3-flash to DS-4.1-flash.
I find that it makes coding and review at the same time less prone to errors and cheaper
GLM 5.3 flash has been good as my default profile Hermes bot, after I readjusted its memory to point to a couple of key skills.
I got some great coding results with GLM 5.2, and 5.3 Flash is supposedly almost as good, so I will be trying it out soon for day to day tasks as the post advises.
Makes me appreciate my ChatGPT subscription. I’ve had multiple days between 1B-2B tokens (now less so, models have indeed become token efficient) and regularly in the > 100M range. Even then, $150 sounds excessive. I wonder if their cache is getting nuked for some reason, or maybe they decide to use Cerebras that doesn’t subsidize cached tokens.
It really makes you see how heavily subsidized the subscriptions are.
Edit: Fixed my math. Edit 2: I was looking at the wrong model on OR. Either way, the math is within the correct ballpark.
https://imgur.com/a/mSXGzzJ
Edit: I found a linked article that mentions the inference provider who does the measurements.
I am a huge Anthropic fan, USA fan, but are we cooked with AI? Sorry SI, that's the important thing.
The direct excitement regarding open models is that many of these Flash Mixture-of-Expert models run reasonably well on hardware a tech employee in the West, and businesses in less affluent countries, can realistically afford.
The indirect excitement is that the models are so efficient, so cloud prices also end up being very low.
You don't see the same scale of excitement surrounding the open weight trillion+ parameter models because, while it's neat they're open weight, it doesn't mean a lot if you need $50k worth of computers to just barely run them.
Avoiding 2 labs having control of knowledge work at the true frontier is not just good for nerd reasons. If only 2 labs control the frontier of knowledge work, well.. that's pretty much the end of the USA's market/service economy. I mean if one or two labs run it, does that make a country's economy?
Meanwhile, we are a deeply stupid species. Based on multiple previous conversations on this website, it appears that for example, z.ai hosting is not a big deal.
Yes, anyone doing real due diligence on client data will face reality, maybe. However, given the long tail of external consultancies, the CCP is going to eat it all due to our outsourced laziness. Shareholder value, amirite!!?
This is how we lost our manufacturing base. Why would the token manufacturing base be any different? As far as I can tell, we have gotten even stupider in the last few years. I guess changing the name to SI and tariffs on Canada and the EU will solve these problems.
I think most posts are best read when there is a solution at the end, but is there one? What would that look like?
I am an old-ish person. My first PC cost as much as a used car. Thanks to global supply chains, we all get those in our pockets, all the time! Capitalism at its best!
If we could agree to stop killing each other, there is literally no end to this hockey stick.
This is the true contrarian POV in 2026, fight me.
What I mean is, Thiel/Yarvin followers. Let's have a talk about this.
This looks like a rather outdated POV. With increasingly pervasive use of SSD offload, there nothing particularly stopping you from running even the largest open models on ordinary local hardware. Sure, it will be really slow, but if you need the smarts for e.g. a one-off planning role it's a no-brainer.
The price one pays for this is that lower than frontier models are dumber. At the current pace of advancement, is that a good idea? Isn't it better to just go with with ZDR contracts, and not lose in the software product game, which now moves at relativistic speed?
I'm mainly using flash varients, at least as the default, bump.up to stronger model as needed (less often these days)
It's my favorite model family to interact with, it's prose is the best imo, it makes me laugh from time-to-time (like when it said it would "crib" some code from another project, lul)
I currently have qwen-flash working on an NES emulator harness so qwen-little can play my first RPG (ff1)
(tho I have used all the others I mentioned, happenstance I'm using qwen this iteration/task)
edit: I subscribe to z.ai, I don't host.
What kind of hardware and what particular quant?
* Models like DeepSeek V4.1 Flash are much cheaper on DeepSeek their API directly because of the cache handeling is better. Neuralwatt can hit up to 98% but DeepSeek can do 99.x... That may not sound like a big difference but it quickly widens the gap on long tasks to grow 2x a 3x in price. DeepSeek their cache handeling is S-tier (with a ton of features, for instance 24h caching).
* The same issue is also present if you compare GLM 5.3 Flash with z.ai vs Neuralwatt. Its just way more cheaper from the source, then from Neuralwatt.
* The energy numbers from Neuralwatt are ... to be taken with a ton of salt. Past energy numbers had the same models (for instance) GLM 5.2 up to 6x cheaper in energy usage, then after they "fixed" issues with the energy numbers. In reality, those energy numbers are just a different form of billing, but not a actual representation of the energy usage of AI models. Things like profits are inside those energy numbers. So seeing 4kWH used for a model, does not mean that it uses 4Kwh.
Edit: That are some interesting downvotes ...
To answer the questions. It was stated by the CEO himself in one of the video blogs that the energy prices inc their profit margins. Regarding their published numbers ... I like to point out that this is the same company that had up to 6x cheaper energy numbers at the start of the year until they got updated. Again, its in one of those video blogs the CEO did. Its around the same time when they increased the price from $5/1kwh to $10/1kwh.
Yes, DeepSeek API is cheaper then Neuralwatt. I have done way too many comparisons between NW and other providers, regarding their prices. Over long sessions, that gap grows because of the differences in caching. You need to use the NW Flex option to reduce the impact but then your constantly waiting on responses (good for overnight work, not great in prime time).
Edit 2: I am getting a little bit fed up with the people who downvote and do not give their reasons for the downvotes.
I think NW's profits are mostly between what they pay for electricity and what they charge you for electricity. I don't think there's any need to look to conspiracies to explain billing errors.