18 August 2026

NVIDIA releases model optimized for faster, cheaper inference

  • Nemotron 3.5 Lightning uses sparse mixture of experts, a technique where only parts of the model activate per query, reducing computational cost.
  • The model combines multiple efficiency methods built into its core design, rather than applying speed improvements as an afterthought to an existing model.

How it was covered