18 August 2026

AI labs shift focus to model design for faster inference

  • Nvidia released Nemotron 3.5 Lightning, a model with 30 billion total parameters but only 3 billion active at once, reducing computational demands.
  • Efficiency improvements now come from fundamental architecture choices and training methods, not just compression techniques applied after models are built.

How it was covered