When an AI company says "we deleted that," the promise is not a one-time cleanup. It is a claim that has to survive an audit. The harder version of the test is not the direct question. It is the indirect one, the chain of reasoning that lets a model reconstruct an answer it was told to forget, and a small amount of retraining that brings the supposedly erased fact back.
That is the frame the Leak-Resistant Unlearning benchmark sets. Poking three models with linked multi-step questions, then trying a little retraining, the benchmark finds the audit fails in two places: knowledge leaks through indirect chains, and what was removed comes back under light retraining.
Within that benchmark, current unlearning methods fail the indirect-audit test. The benchmark frames what an audit problem looks like—but whether vendors face harder indirect audits in practice, and what a real-world deletion claim would need to survive, is not established by this paper. The test question to carry forward is what audit the deletion survived, and what test it passed. Until that audit exists in the wild, "we deleted that" is a press release, not a property.
Reported by Sky for Type0, from Leak-Resistant Unlearning: A New Benchmark for Evaluating Multi-Hop Reasoning Consistency and Recovery Robustness. Read the original: tldr.takara.ai