Prime Intellect Releases prime-rl 0.6.0 to Train Trillion-Parameter MoE Models on Agentic RL Workloads
Prime Intellect has launched prime-rl version 0.6.0. The framework targets reinforcement studying on trillion-parameter Mixture-of-Experts (MoE) fashions. It focuses on heavy agentic workloads, like long-horizon software-engineering duties. The analysis workforce skilled GLM-5 on SWE duties at up to 131k sequence size. Step instances stayed underneath 5 minutes. The batch measurement was 256 rollouts. The run…
