18 August 2026
NVIDIA releases model optimized for faster, cheaper inference
- Nemotron 3.5 Lightning uses sparse mixture of experts, a technique where only parts of the model activate per query, reducing computational cost.
- The model combines multiple efficiency methods built into its core design, rather than applying speed improvements as an afterthought to an existing model.