UCSD and Together AI Research Introduces Parcae: A Stable Architecture for Looped Language Models That Achieves the Quality of a Transformer Twice the Size
The dominant recipe for constructing higher language fashions has not modified a lot since the Chinchilla period: spend extra FLOPs, add extra parameters, prepare on extra tokens. But as inference deployments eat an ever-growing share of compute and mannequin deployments push towards the edge, researchers are more and more asking a tougher query — are…
