ComputeLabs Research
NVIDIA Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model, delivers up to 4× faster output and 30% faster agentic-task completion.
· ComputeLabs Research · from the August 11, 2026 edition
NVIDIA expanded its Nemotron 3 family with Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model. A mixture-of-experts architecture selectively activates model components for a given request rather than necessarily using every parameter for every token.
NVIDIA said the model can produce output as much as four times faster than comparable models and complete agentic tasks 30% faster. These are NVIDIA-reported performance claims, and the supplied sources do not provide the complete benchmark configurations behind every comparison.
The model targets long-running agents that perform high-volume tool calls, validate results and delegate tasks to subagents. NVIDIA described it as the highest-efficiency model in its class for long-running agentic workloads.
Nemotron 3.5 Lightning is available as an NVIDIA NIM inference microservice through Hugging Face, ModelScope, OpenRouter and build.nvidia.com. NVIDIA also introduced NeMo Switchyard for routing agent workloads among models with different capabilities and cost profiles.
Sources
- NVIDIA Nemotron 3.5 Lightning Delivers Fast, Accurate Specialized Task Execution for Long-Running Agents(developer.nvidia.com)
- NVIDIA Nemotron 3.5 Lightning and NeMo Switchyard Deliver Faster, Smarter, More Efficient Agentic AI(blogs.nvidia.com)
- Route AI Agent Workloads Across Models with NVIDIA NeMo Switchyard(developer.nvidia.com)
Additional reporting
- Financial_Express(t.me)
- NVIDIA Nemotron 3.5 Lightning

