OpenAI Researchers Train Weight Sparse Transformers to Expose Interpretable Circuits
If neural networks at the moment are making selections in all places from code editors to security techniques, how can we really see the particular circuits inside that drive every habits? OpenAI has launched a brand new mechanistic interpretability research study that trains language fashions to use sparse inner wiring, in order that mannequin habits…
