Deep Dive Advanced Lectures by
• Thomas Strohmer (UC Davis) [Can AI Truly Forget? A Mathematical Framework for Machine Unlearning]
• Fazl Barez (University of Oxford) [Does Unlearning Work for AI Safety? From Capability Control to Selective Forgetting]
• Venue: Lecture Theatre L3, Andrew Wiles Building (Mathematical Institute), University of Oxford
• Agenda:
• 11:00-12:00: Thomas Strohmer's talk (Lecture Room L3)
• 12:00-13:00: Networking Lunch break (Cafe PI)
• 13:00-14:00: Fazl Barez's talk (Lecture Room L2)
Thomas Strohmer (UC Davis), Lecture Room L3
Can AI Truly Forget? A Mathematical Framework for Machine Unlearning
As AI models are trained on ever-expanding datasets, the ability to remove the influence of specific data from trained models has become essential for privacy protection and regulatory compliance. Unlearning addresses this challenge by selectively removing parametric knowledge from the trained models without retraining from scratch, which is critical for resource-intensive models such as Large Language Models (LLMs). However, existing unlearning methods often severely degrade model performance by removing more information than necessary when attempting to "forget" specific data. We introduce a mathematical framework based on information-theoretic regularization that can accommodate different types of machine unlearning, such as feature unlearning and data point unlearning. Our theoretical analysis reveals intriguing connections between machine unlearning, information theory, optimal transport, and extremal sigma algebras. For LLMs, we propose Forgetting-MarI, an unlearning framework that provably removes only the additional (marginal) information contributed by the data to be unlearned, while preserving the information supported by the data to be retained. Extensive experiments confirm that our approach outperforms current state-of-the-art unlearning methods, delivering reliable forgetting and better preserved general model performance across diverse benchmarks. This advancement represents an important step toward making AI systems more controllable and compliant with privacy and copyright regulations without compromising their effectiveness. We will also discuss applications in machine learning driven scientific discovery.
Fazl Barez (University of Oxford), Lecture Room L2
Does Unlearning Work for AI Safety? From Capability Control to Selective Forgetting
It is widely understood that unlearning can help with specific data removal. However, when it comes to AI safety, we often care about controlling broader capabilities rather than removing specific data. These capabilities can be reconstructed by combining retained (and often seemingly unrelated) knowledge, and may rely on shared or entangled representations. As a result, removing one behaviour may either degrade unrelated capabilities or fail to fully remove the targeted capability.
In this talk, I will argue that selective forgetting may therefore be partly a representation-learning problem, rather than only an unlearning problem. In particular, I will discuss how pretraining data and curriculum learning might provide a new avenue for making later interventions more selective. I will show how imbalanced pretraining encourages tasks to be disentangled in separable neural circuits, whereas balanced training routes tasks through a common pathway. I will then discuss how shaping representations during pretraining may improve our ability to selectively modify capabilities later, with implications for unlearning and the precision and reliability of safety fine-tuning.