18 August 2026

Nvidia releases efficient model with fewer active parameters

  • Nvidia released Nemotron 3.5 Lightning, a model designed to run efficiently by activating only 3 billion of its 30 billion total parameters at any given time.
  • The model can predict multiple tokens simultaneously, reducing the number of computational steps needed to generate text.

How it was covered