Nvidia has kicked off a month-long "Local AI" campaign highlighting new open-weight models and tools designed to run agentic AI directly on local hardware rather than in the cloud, the company said in a rolling blog series published from 11 August 2026.
Among the releases, Nvidia introduced Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for always-on agents, offering up to four times faster token generation than comparable open models. It also launched NeMo Switchyard, an open-source routing library that directs each step of an agent's workflow to the most cost-effective suitable model.
Meta released Muse Glimmer, a 30-billion-parameter model optimised for local coding and agentic tasks, capable of over 200 tokens per second on Nvidia's RTX 5090 GPU. Other partners including Poolside AI, DeepSeek, Alibaba and Thinking Machines Lab also shipped new open-weight models optimised for Nvidia hardware this week.
Nvidia separately updated its Sync application with a Cluster Assistant tool to simplify linking multiple DGX Spark systems for running larger models, and said DGX Spark will gain a native Chrome browser build and a new system resource monitor later in August.
