NVIDIA releases open model 'Nemotron 3.5 Lightning,' a 30B model that runs on just 3B active parameters
The MoE architecture activates only 3B of its total 30B parameters per token, and NVIDIA claims up to 4x faster output than comparable models
00
One email each morning — yesterday's AI, sortedGet it in your inbox
Tag
The MoE architecture activates only 3B of its total 30B parameters per token, and NVIDIA claims up to 4x faster output than comparable models
MoE architecture with 37B active parameters, FP8 quantized version uploaded to Hugging Face
That's the last story.