Stochastic Gradient Descent (SGD’s) Frequency Bias and How Adam Fixes It
Modern language fashions are educated on information with extraordinarily uneven token distributions. A small variety of phrases seem in nearly each sentence, whereas many uncommon however significant tokens happen solely sometimes. This creates a hidden optimization problem: parameters related to frequent tokens obtain fixed gradient updates, whereas parameters tied to uncommon tokens could go tons…
