14 August 2026 · Local AI · Benjamin Cheng

Open-weight AI models are having a moment

Open-weight AI models are having a moment.

This week alone, we saw three notable releases:

Two weeks earlier, DeepSeek released DeepSeek-V4-Flash-0731, a 284B-parameter mixture-of-experts model with 13B active parameters and a one-million-token context window.

I ran the Muse Glimmer, DeepSeek and Nemotron on an ASUS GX10 with 128GB of unified memory and got useful output from all of them.

The trend is moving past text. MiniMax-H3's open base model generates 768p video with native stereo audio on local hardware. I tested it in ComfyUI with this prompt:

"A cinematic all white spitz dog walking towards the camera through the desert at dawn, soft golden rays, detailed fur, gentle camera movement, natural ambience and light breeze dusting up the sand with soft sand blowing, no text or watermark"

It produced the full clip locally on the GX10 within 6 minutes.

MiniMax-H3 generated this 768p clip with native stereo audio locally on the ASUS GX10, using the prompt in the article.

The direction is clear.

Frontier-level capability will not stay locked inside large cloud platforms. Cloud models keep their place, especially where speed and top quality matter.

What mattered to me is local AI works and it gives something just as valuable: control of your data and predictable cost.

The future will not be cloud or local. It will be both.

If you want to learn how to run cloud and local models using Hermes Agent, find out more here: https://lnkd.in/g6qirhv7

← Back to the blog