No datacenter needed now, just run it on your computer. Anything with 12gb+ of vram can run it comfortably at low quants. If they can come out with a good MOE model it will crush the competition. At this current rate of progress, China will out-compete foreign AI development within a year or two.
The speed at which open weight models are progressing is crazy. They must be having sleepless nights in Silicon Valley lately.

Anything with 12gb+ of vram can run it comfortably at low quants.
27B dense models need like somewhere between 14 and 16 GB of VRAM at 4-bit quantization, and that’s not even accounting for KV cache. Even on a 16 GB card that would be uncomfortably tight, and wouldn’t give you any room for long context lengths. You definitely need a 24 GB card to run this properly. That’s still damn impressive. I didn’t expect frontier model performance on a single consumer GPU so soon.
I think when they release the 35B-A3B MoE model, it will be the final knockout punch. That one you can technically even run on an 8 or 6 GB card at decent speeds with some smart offloading if you have enough system RAM. Being able to run a frontier level model on a low power home server definitely changes the game.
Can you explain what the 35B-A3B MoE model is? I’ve seen so much jargon like this before that I mentioned in a previous thread on why people ignore self-hosted AI (it’s like listening to aliens speak). I am a little familiar with some of the terminology, but don’t really understand why a 35B model would be easier to run than a 27B one and so on. Or what the quantization is.
I’ve got a 16GB vram card myself and find most models like the recent Qwen to respond far too slow or cause out of memory issues and crash.
35b-a3b means it only actively calls 3b relevant parameters while generating, it offers similar performance as a base model like 27b but for less resources (you can hold some of the model in ram, but at the cost of storing more of the model on your drive). if youre able to run it entirely in vram somehow it can run blazingly fast in comparison to other models as well. I personally hope they’re more ambitious than a 35b a3b, something that works more like a6b, a9b, a13b, etc would be better for most lower end computers.
These data centers got to be part grift right? Like I’m a dolt when it comes to tech and even I know eventually powerful ai will be able to be launched off consumer hardware. So who are these datacenters for?
As has been mentioned, private equity grifting, but also

How do I do the emojis
bug your admins about it
:panopticon:

:c
:smile:
lol, yeah, I think you just have to know the url for the emojis you want to use.
:sicko:
Damn it…
Come on just give me one.
They’re not really -for- anyone, and they may or may not even get built. It’s a front for the circular financing that’s propping up the bubble. NVidia invests $3bn in OpenAI, who goes and buys $3bn worth of future datacenter capacity from a cloud provider, who pre-purchase $3bn of GPUs from NVidia.
Everybody just made $3bn and the line keeps going up.
Obviously if you can run frontier quality models on consumer hardware then the quality of models you can run in a data center will let you hit AGI. Unless there’s any sort of diminishing returns in quality when increasing model size.
But in seriousness Qwen3.8 was trained in a data center. It takes way more VRAM and compute to train a model than to run one.
Unless there’s any sort of diminishing returns in quality when increasing model size.
There is actually.
ZWQbpkzl is being sarcastic, the line after says “but in seriousness”
(edit pronouns sorry)
Thank you for your service.

Nah, they redefined getting 20% better results for 10x the cost as “not diminishing returns actually” years ago.
Obviously if you can run frontier quality models on consumer hardware then the quality of models you can run in a data center will let you hit AGI

Obviously
why is this obvious?
ZWQbpkzl is being sarcastic, the line after says “but in seriousness”
(edit pronouns sorry)
thank you
I try
The big tech companies have more money than God and have had historically high revenue growth for decades. They achieve that high growth rate by re-investing that giant stream of money into more growth. Their stocks are entirely valued on that high growth rate - for whatever reason, growth is valued more than dividends in the current stock market climate. But they’ve basically saturated their markets - Google can’t spend more money to make more people use search, Facebook can’t spend more money to make more people use social media, approximately everybody with an internet connection is already using those. They can’t just sit on the money, because if they do, the expected thing is to pay a portion of it out to the shareholders as a dividend. But dividend stocks are valued less than growth stocks, so their stock price will tank, and the board (comprised of rich people who own a lot of that company’s stock) don’t want that to happen. So they need something else to spend the giant stream of cash on.
Data centers full of GPUs are the most expensive possible thing that a tech company could conceivably make a profit off of, so it solves the problem of what to do with the giant stream of cash. And the boards and investors are required to believe in it, because otherwise they’re wasting the giant stream of cash.
its a real estate bubble
Basically the top tier datacenter models right now are extreme generalists and store quite a lot of humanity’s total knowledge on them, even the ones that are capable of being run locally are 2-4tb in size and would require thousands of gigabytes of ram to run well, a cluster that could run that would cost 50k-100k usd. The local 27b models try to rip out the “intelligence” of these titanic large models. They then augment the model’s intelligence with toolcalls (looking at say, github or a search engine) instead of having all the knowledge baked into the model. This approach results in less hallucinations and far less resource usage, but it also isn’t as generally useful because you may not have access to some niche information, information you couldn’t find with a google search.
The big benefit for local models is you can use them to crush through a lot of google searches in one go and find you a good source, its great for people doing scientific research because you can search all the journals you have access to very quickly for relevant info. Its also great for doing up small bits of code quickly and iterating on that code for your various experiments.
The big benefit for local models is you can use them to crush through a lot of google searches in one go and find you a good source, its great for people doing scientific research because you can search all the journals you have access to very quickly for relevant info.
This is the only thing I’ve tried to really get an AI model working for, I want it to do some large scale searching and format the info with links to the sources for me to look through. None of the text it writes will end up in my final product but it should be able to do the searching a lot more efficiently than I can myself.
Unfortunately I’ve completely failed to get this to work
You might like Proton Mail’s Lumo project, I find it useful for quicker searches without needing an account, they use open-weights models only. Unsloth with EXA or setting up searxng-mcp is your other best bet, Unsloth has a desktop app now.
Honestly it’s all up in the air at this point, I don’t think anybody can predict what kind of AI will be ran on data centres and on consumer hardware in the next 5-10 years. Maybe it all deflates because we hit a wall with it, maybe your phone will run a model more powerful than today’s top ones.
My guess is the really good models will keep running on data centres but will be prohibitively expensive and/or will have their use regulated by their government so only big institutions and companies get access to it and the users will be running dumber models on their own hardware.
Surveillance and squeezing stones.
Surveillance
I also cannot deny the military applications something like a fast running qwen 3.8 27b will have for guerrilla groups. I think we’re going to see a revolution in warfare soon, and it might become so deadly that no country will be interested in fucking with anyone because even individual civilians will be able to destroy major targets. Drones will be able to be trained on specific individuals faces, they can be used as smart missiles, they can be mobile defensive mines to protect against aerial bombardment, you can constantly surveil an area with automated drone swarms, the possibilities are endless and its frankly pretty dangerous tech.
What does it mean when any opposition to anything will paint a lethal target on your head? Are we going to have to go through our lives acting out personas that are severable as soon as we offend someone enough? Will there be an arms race over knowing someone’s true identity? Will we all have to further atomize ourselves in our homes with slaughterbot-proof bars on the windows? Maybe even more draconian surveillance measures to try in vain to prevent non-state actors from acquiring slaughterbots that weigh 200g and can be 3D printed?
For most of human history, there were people that were locally important to you, and there were people that were inaccessible to you, with no in-between. Killing someone casually at your leisure was something that was mostly within the domain of a small and well-known group of (assassinateable) powerful people.
Everyone is thinking
“wow imagine what it could do for meeeee” and no one is thinking “what would the ecosystem of this look like with universal access”.You’d better hope that you have the definitive model for society up your sleeve to prevent these from being used as they get introduced, otherwise humanity is going to eat itself alive.
I say the biggest danger is that given this tech is built off of our knowledge, there is a lot of things that we think we know are for certain that just isn’t. I just do not see a way around false targets for things like this. That of course, doesn’t mean it will stop it from being made and tried, but…it’s going to be a fucking rough.
I also cannot deny the military applications something like a fast running qwen 3.8 27b will have for guerrilla groups.

Not naming any names ofc
Training a frontier model is like Karsus’ casting a 12th level spell. it takes a massive amount of everything to get it to go. Even the enigma machine was huge before it could be one shot by claude. So if it’s going to become something that can use your phone’s worth of compute then it’s gotta get big and dense and then you take the 10% of it that works and do it over again. But in the mean time they’ll keep pushing the frontier, it’ll get distilled and open sourced, and that will be whittled down. Like a giant, electric-steel, writhing eldritch monster
One of the nice things about Deepseek being open-source is that it significantly mitigates the training costs involved as well as use costs (since it’s so efficient). If you only one need to train one model that the whole world can then use, that’s a lot better environmentally than multiple companies all making their own proprietary models.
Also Deepseek is apparently free on open opencode right now lol
Huh explain
opencode has a hosting org attached to it that provides free preconfigured llm runtime to anyone with opencode installed. it can be a way to access deepseek v4 flash without having good computer specs. but its got a pretty long queue during the day when everyone wants to use it.
















