AMD Releases Instella-MoE-16B-A3B: A Fully Open Mixture-of-Experts LLM With 2.8B Active Parameters Trained On Instinct GPUs
AMD launched Instella-MoE-16B-A3B, a completely open Mixture-of-Experts language mannequin skilled from scratch on Instinct MI300X and MI325X GPUs. The mannequin holds 16B complete parameters however prompts solely 2.8B per token. AMD is publishing weights from each coaching stage, together with information mixtures, coaching configs, and inference code. Two systems-level decisions carry the discharge: Gated Multi-head…
