Why Nvidia Nemotron 3.5 Lightning is built for speed

Nvidia has released Nemotron 3.5 Lightning, an open-weights model designed around throughput rather than maximum benchmark intelligence. It matches OpenAI's gpt-oss-120b on the Intelligence Index with fewer active parameters and reaches nearly 670 tokens per second in pre-release tests.

WTF Index NEUTRAL
◄ Terminator 1 Idiocracy 0 ►

This is mostly a routine model launch focused on faster, more efficient inference rather than new autonomy or social degradation.

Why Nvidia Nemotron 3.5 Lightning is built for speed

Nvidia's Nemotron 3.5 Lightning is a clear bet on speed. The new open-weights model does not lead every intelligence benchmark, but it aims at a different goal: delivering strong reasoning performance with very fast inference and a much smaller active parameter footprint.

The model is the first release in Nvidia's Nemotron 3.5 lineup. It follows Nemotron 3 Nano 30B A3B and keeps the same hybrid Mamba-Transformer architecture, with 31.6 billion total parameters and only 3.6 billion active at any given time.

A smaller active model with competitive benchmark results

According to the independent benchmarking platform Artificial Analysis, Nemotron 3.5 Lightning scores 24 on the Intelligence Index. That is a nine-point increase from its predecessor, Nemotron 3 Nano, which scored 15.

The score puts Lightning level with OpenAI's gpt-oss-120b, which also scores 24. It sits just behind Nvidia's own Nemotron 3 Super, which scores 26 and is about four times larger.

That result matters because Lightning is not trying to win by brute scale. Its design keeps the total parameter count at 31.6 billion, while activating only 3.6 billion parameters at a time. In practical terms, the model is positioned for use cases where response speed and operating efficiency matter as much as raw benchmark rank.

The source data also shows the tradeoff clearly. Small models in the same size class can still score higher on intelligence. Qwen3.6 35B A3B scores 32, while Meta's new Muse Glimmer scores 35. Lightning therefore looks less like an all-purpose benchmark leader and more like a fast, capable model for workloads where throughput is central.

Inference speed is the main story

Nvidia is placing Lightning on a specific part of the efficiency frontier. In pre-release tests using the final NVFP4 weights, the model reaches nearly 670 tokens per second. Artificial Analysis describes that as the highest measured throughput among all compared models.

That speed is almost twice as fast as Google's Gemini 3.5 Flash-Lite, which is listed at 386 tokens/s. On the Intelligence Index task timing, Lightning completes a task in about 0.5 minutes. By comparison, Qwen3.6 35B A3B needs around 3.5 minutes, and Gemma 4 31B takes roughly 5.8 minutes.

For developers and teams building agent-based pipelines, this distinction is important. A model that responds quickly can support repeated tool calls, multi-step workflows, and high-volume applications more comfortably than a slower system with a higher score in some categories. The source does not claim Lightning is the smartest model available. It shows that Nvidia is optimizing for fast useful output.

Proprietary systems still set a higher bar on the overall efficiency frontier. Gemini 3.5 Flash-Lite scores 37 on the Intelligence Index with a similar time per task, while GPT-5.6 Luna (max) reaches 52 points in under two minutes. Lightning's role is different: it brings high throughput to an open-weights model that can compete closely with gpt-oss-120b on the cited intelligence measure.

Agentic benchmarks show the strongest jump

The most notable gains appear in agentic benchmarks, according to Artificial Analysis. On GDPval-AA v2, Nemotron 3.5 Lightning reaches an Elo rating of 824. That is a 334-point gain over Nemotron 3 Nano.

On the same benchmark, Lightning beats both gpt-oss-120b, which scores 800, and the larger Nemotron 3 Super, which scores 698. That is a meaningful result for Nvidia's positioning, because the model is being framed as a high-throughput workhorse for agent-based pipelines rather than as a maximum-intelligence system.

Terminal-Bench v2.1 shows another large improvement. The score moves from 7 to 24.3 percent, nearly matching gpt-oss-120b at 26.2 percent.

These results suggest that the model's speed is not the only relevant point. Within the limits of the source data, Lightning also appears much stronger than its predecessor on tasks associated with agentic behavior. That combination is what makes the release significant: fast inference would be less compelling if the model could not also handle useful reasoning work.

Availability and model options

Nvidia provides Nemotron 3.5 Lightning in both BF16 and NVFP4 weights. The NVFP4 version also scores 24 on the Intelligence Index, with minimal quality loss compared to the higher-precision version, according to Artificial Analysis.

The model is text-only and supports a context window of one million tokens. Nvidia ships it under the permissive OpenMDW-1.1 license.

Weights are available now. Serverless inference is offered by DeepInfra, Fireworks, FriendliAI, CoreWeave, GMI Cloud, Nebius, and Crusoe, among others.

Artificial Analysis also reports that Nvidia worked with partners including CodeRabbit and Harvey on post-training to improve performance in specific domains. The source does not give further details on those domains beyond naming the partners, but it does connect that post-training work to the model's stronger benchmark behavior.

What Lightning says about Nvidia's model strategy

Nvidia's push toward smaller, faster agent models did not begin with this release. In a widely discussed paper last year, Nvidia researchers argued that models under 10 billion parameters can handle most agent workloads as well as 70- to 175-billion-parameter models at one-tenth to one-thirtieth the cost.

Nemotron 3.5 Lightning is larger than that threshold in total size, with 31.6 billion parameters. But because it activates only 3.6 billion per step, it fits the same lightweight direction described in the source.

The result is a model built around a practical claim: many agent workloads may benefit more from speed and efficiency than from chasing the highest intelligence score. With nearly 670 tokens per second, an Intelligence Index score matching gpt-oss-120b, and strong gains on agentic benchmarks, Lightning gives Nvidia a concrete product example of that approach.