NVIDIA agentic AI models hit 4x speed with Nemotron 3.5 Lightning

NVIDIA is pushing further into agentic AI models with the launch of Nemotron 3.5 Lightning, a new open model built specifically for the kind of always-on, high-volume tasks that autonomous AI agents now handle around the clock. Alongside it, the company released NeMo Switchyard, an open source routing tool designed to send each AI request to whichever model can handle it fastest and cheapest. Together, the two releases mark NVIDIA’s latest attempt to make agentic AI systems not just smarter, but genuinely efficient to run at scale.
Key takeaways
- NVIDIA released Nemotron 3.5 Lightning, a 30-billion-parameter open mixture-of-experts model built for high-volume agentic AI workloads.
- The model delivers up to 4x faster output speed and 30% faster task completion than comparable models in its class.
- NVIDIA also launched NeMo Switchyard, an open source routing library that can cut task completion costs to nearly one-third of using a single frontier model alone.
- Nemotron 3.5 Lightning runs locally on NVIDIA RTX PCs, DGX Spark, DGX Station and Jetson, or scales to data centers and the cloud.
- Both releases are open, customizable and already backed by ecosystem partners including CrowdStrike, Harvey, CodeRabbit and Lila Sciences.
NVIDIA Launches Nemotron 3.5 Lightning for Agentic AI Workloads
Nemotron 3.5 Lightning is the newest addition to NVIDIA’s Nemotron 3 family, and the company describes it as the most efficient model in its class for long-running agentic tasks. It arrives after Nemotron 3 Nano and continues a pattern NVIDIA has followed with every Nemotron release: prioritize open weights, customization and raw speed over sheer size.
… Continue reading the full article at the original source below.

