Open-weight AI models are having a moment.
This week alone, we saw three notable releases:
- Meta's Muse-Glimmer-30B
- NVIDIA Nemotron 3.5 Lightning
- Qwen3.8-2.4T-A95B
Two weeks earlier, DeepSeek released DeepSeek-V4-Flash-0731, a 284B-parameter mixture-of-experts model with 13B active parameters and a one-million-token context window.
I ran the Muse Glimmer, DeepSeek and Nemotron on an ASUS GX10 with 128GB of unified memory and got useful output from all of them.
The trend is moving past text. MiniMax-H3's open base model generates 768p video with native stereo audio on local hardware. I tested it in ComfyUI with this prompt:
"A cinematic all white spitz dog walking towards the camera through the desert at dawn, soft golden rays, detailed fur, gentle camera movement, natural ambience and light breeze dusting up the sand with soft sand blowing, no text or watermark"
It produced the full clip locally on the GX10 within 6 minutes.
The direction is clear.
Frontier-level capability will not stay locked inside large cloud platforms. Cloud models keep their place, especially where speed and top quality matter.
What mattered to me is local AI works and it gives something just as valuable: control of your data and predictable cost.
The future will not be cloud or local. It will be both.
If you want to learn how to run cloud and local models using Hermes Agent, find out more here: https://lnkd.in/g6qirhv7