Catastrophic Forgetting is Low-Rank: A Function-Space Theory for Continual Adaptation

June 16, 2026 ยท Grace Period ยท ๐Ÿ› the ICML 2026 Workshop on Continual Adaptation at Scale: Towards Sustainable AI

โณ Grace Period
This paper is less than 90 days old. We give authors time to release their code before passing judgment.
Authors Ido Nitzan Hidekel, Dan Raviv arXiv ID 2606.18024 Category cs.LG: Machine Learning Cross-listed cs.AI Citations 0 Venue the ICML 2026 Workshop on Continual Adaptation at Scale: Towards Sustainable AI
Abstract
Catastrophic forgetting in continual adaptation is usually studied through parameter drift, replay, or distillation, but these views do not identify which output-space directions are vulnerable. We give a function-space account in the NTK regime: new-task training induces old-task prediction drift through the cross-task kernel, yielding a closed-form predictor for the forgetting vector before any new-task gradient step. In frozen-backbone linear-head PEFT-CL, where the model is linear in the trainable parameters, the predictor is exact up to numerical precision; for nonlinear adapters/full fine-tuning, it is a local NTK approximation. The same expression reveals that forgetting concentrates in a small number of old-task NTK eigenmodes and under frozen linear heads gives a Kronecker scaling rule for the vulnerable rank. These results clarify the relation to prior NTK-overlap theory, explain why parameter-space regularizers can miss output-space interference, and motivate a targeted spectral regularizer.
Community shame:
Not yet rated
Community Contributions

Found the code? Know the venue? Think something is wrong? Let us know!

๐Ÿ“œ Similar Papers

In the same crypt โ€” Machine Learning