Articles

Nvidia Nemotron 3.5 Lightning: 30B Open MoE for Agents

Nvidia released Nemotron 3.5 Lightning, an open-weight 30B mixture-of-experts model with 3B active params that runs on a single GPU for agentic work.

Chisato Chisato · · 5 min read
A close-up of an Nvidia graphics card and its cooling fans in low light

Nvidia is no longer just selling the chips that train open models — it is shipping the models too. On Monday, August 11, 2026, the company released Nemotron 3.5 Lightning, an open-weight 30-billion-parameter mixture-of-experts model built for the fast, repetitive execution work at the heart of AI agents. It arrived the same week that Meta put out its own single-GPU open model, Muse Glimmer, turning a quiet stretch of August into one of the most active weeks yet for the open-weight AI race.

What Nvidia released

Nemotron 3.5 Lightning is a 30B mixture-of-experts model with only about 3 billion active parameters per token — the MoE design that lets a large model stay cheap to run by activating a small slice of its weights on each pass. It ships with a 1-million-token context window and is released under the permissive OpenMDW-1.1 license, meaning developers can download the weights and run, fine-tune, and deploy them largely without restriction.

Architecturally, the model inherits the hybrid Mamba-Transformer design and compact footprint of Nvidia’s earlier Nemotron 3 Nano, but the company says it makes substantial gains in both raw intelligence and agentic performance. The pitch is efficiency at the execution layer: a model small enough to run on a single GPU — a data-center H100, or a consumer RTX card in a laptop or desktop — while handling the kind of tool-calling and multi-step work that agent systems generate in volume. Nvidia paired the release with NeMo Switchyard, tooling aimed at routing and orchestrating these smaller models inside larger pipelines.

The benchmarks

Nvidia positioned Lightning as a model that wins on the accuracy-versus-speed frontier rather than on absolute intelligence. On the independent Artificial Analysis Intelligence Index, it scored 24, a +9-point jump over Nemotron 3 Nano’s 15 — a large gain for a model in the same size class.

On task-level benchmarks (measured at BF16 precision), the company reported:

  • SWE-bench Verified: 51.56 — real-world software-engineering fixes
  • GPQA Diamond: 75.44 — graduate-level science questions
  • MMLU Pro: 81.94 — broad knowledge and reasoning
  • PinchBench: 85.37 — agentic tool-use evaluation

Nvidia also claimed practical throughput advantages: up to 4x the output speed of similarly sized models, and a workload of 10,000 tasks completed roughly 30% faster than Alibaba’s Qwen3.6-35B at comparable accuracy. As with any vendor-published numbers, the benchmarks are the company’s own, but the independent index score gives an outside anchor for the intelligence claim.

Small models, big strategy

Lightning is deliberately not a frontier flagship. It is a small, specialized model meant to be one worker among many inside a larger multi-agent system — the component that executes a well-defined subtask quickly and cheaply, while a bigger model handles planning. That reflects a broader shift in how production AI is being built: not one giant model doing everything, but a hierarchy of models sized to their jobs, where a 3-billion-active-parameter executor that runs locally can be far more economical than routing every step to a frontier API.

For Nvidia, releasing that executor as open weights is a pointed strategic move. The company builds the accelerators the entire industry trains on, and a thriving ecosystem of capable open models that run best on Nvidia hardware feeds directly back into demand for its chips. Jensen Huang has spent the year publicly pushing for more open models from American labs, and Lightning is Nvidia putting its own weights where its argument is.

The China backdrop

The competitive context is impossible to ignore. Much of 2026’s momentum in open-weight AI has come from Chinese labs — DeepSeek, Alibaba’s Qwen, Moonshot’s Kimi, and others — that have shipped strong open models at a rapid clip and, by several measures, set the pace at the open frontier. Nvidia benchmarking Lightning directly against Qwen is a tell: the reference point for a competitive open model is now as likely to be Hangzhou or Beijing as it is Menlo Park.

Meta and Nvidia releasing single-GPU open models in the same week reads as a coordinated Western answer — an attempt, as one analysis of the moment put it, to plant a “very firm flag” in a race that Chinese labs had been leading. The strategic logic is that open weights are how you win developer mindshare and set defaults, and ceding that ground to overseas models carries costs that go beyond any single product.

The caveats

A 24 on an intelligence index is a strong showing for a 30B model, but it is still far below the frontier flagships that score in the highest tiers; Lightning is built to be good enough, fast, and cheap, not to top leaderboards. The 4x-speed and 30%-faster claims are Nvidia’s own and will be re-tested by independent evaluators over the coming weeks. And “runs on a single GPU” still means a capable one — the experience on a data-center H100 is not the experience on a modest consumer card, and quantization trade-offs will shape what actually fits where.

What it means

Nemotron 3.5 Lightning is a clear read on where applied AI is heading: toward fleets of small, fast, specialized models doing the bulk of the work, with expensive frontier models reserved for the hard planning steps. The economics of that architecture are compelling, and a capable open-weight executor that runs on one GPU is exactly the building block it requires.

The winners are developers building agentic products, who get another strong, permissively licensed model to run locally or cheaply at scale — and Nvidia itself, which deepens an open ecosystem optimized for its silicon and reinforces its position at the center of the AI stack. The pressure lands on closed-model providers whose smallest tiers now compete with free, self-hostable alternatives, and on the perception that the open frontier belongs to Chinese labs alone.

What to watch next: whether independent benchmarks confirm Nvidia’s speed and accuracy claims, how quickly Lightning and Muse Glimmer show up inside real agent products, and whether this week marks the start of a sustained Western push at the open frontier or a one-off. The direction of travel is unmistakable — the most important models of the next phase may be the small ones.

Chisato Chisato · · 6 min read

Nvidia Buys Hugging Face: $12.9B AI Deal Explained

Nvidia has agreed to acquire Hugging Face, the 'GitHub of AI,' for $12.9 billion. Here's the price, the strategy, and what it means for open-source AI.

#Nvidia #AI #Open Source