NVIDIA releases open model 'Nemotron 3.5 Lightning,' a 30B model that runs on just 3B active parameters
The MoE architecture activates only 3B of its total 30B parameters per token, and NVIDIA claims up to 4x faster output than comparable models
10
Tag
The MoE architecture activates only 3B of its total 30B parameters per token, and NVIDIA claims up to 4x faster output than comparable models
MoE architecture with 37B active parameters, FP8 quantized version uploaded to Hugging Face
That's the last story.