NVIDIA Releases Nemotron-Labs-3-Puzzle-75B-A9B: A Compressed Hybrid MoE LLM Delivering 2.03x Server Throughput at Matched User Throughput
Large hybrid MoE fashions like Nemotron-3-Super are correct however costly to serve. Their energetic parameters, KV cache, and Mamba state cap what number of customers a node can maintain at a given per-user token fee. NVIDIA AI group has launched Nemotron-Labs-3-Puzzle-75B-A9B, a compressed variant of Nemotron-3-Super. The guardian mannequin has 120.7B complete and 12.8B energetic…
