In a long 2026 essay, pseudonymous researcher Gwern Branwen proposes a 'guardian angel' LLM that emulates one user and refuses the scams aimed at them.
The email in your inbox right now is almost certainly not from the person whose name is on it. It is a vendor's pitch dressed as a colleague, a deepfake invoice dressed as a CFO, or an AI-drafted reply that approximates your voice just closely enough to be plausible. The same large language models that can write a memo in your style can also write one in the style of someone pretending to be you, and the defense gap between the two is widening faster than any inbox filter can close.
One of the more coherent responses is a long essay, not a product. In "Guardian Angels: LLM Personalization for Productivity and Security", pseudonymous independent researcher Gwern Branwen argues that the right answer is a personal model trained to emulate exactly one person: their values, their voice, their judgment. The model then acts as that person's filter before any vendor's assistant gets a turn. A "guardian angel" in Branwen's framework is the opposite of a generic chatbot persona. It is a digital twin that treats the user as the principal in the economics sense, the one who sets the goal, while the model becomes the agent that executes.
The mechanism is a stack, not a single trick. At the bottom sits a capable base model, periodically upgraded the way a phone operating system is. On top of it sits the personalization layer: a model that has been trained, fine-tuned, and continuously updated to mirror one specific user's preferences. To keep that mirror from drifting, the proposal leans on active learning, the technique of asking the user targeted questions about the cases the model is least sure about, rather than collecting generic feedback. Dynamic evaluation, in this framing, means continuously testing whether the personal model still behaves like its principal, not whether it can pass a static benchmark.
The security argument is the load-bearing one. A model hardwired to one situated user is harder to social-engineer than a model that takes instructions from whoever phrases them most persuasively, a variant of the classical "confused deputy" problem in which a trusted assistant is tricked into acting on behalf of an attacker. Periodic base-model upgrades add a "defenders' advantage": when an attacker finds a prompt-injection trick that works against last year's model, the user's guardian angel gets a new foundation before the trick can be industrialized. The same dynamic evaluation that keeps the personalization honest also serves as an early-warning system when something in the model starts behaving unlike its principal.
The framework is candid about what it is not. It does not solve the broader problem of AI alignment, the question of how to make powerful models behave well in society at large. A guardian angel is positioned as individual defense in depth, one layer in a much larger strategy. It also concedes real failure modes: a base-model upgrade can quietly change behavior in ways the personalization layer has not yet compensated for; a poorly trained mirror can confidently misjudge a situation the user would have caught; and the proposal assumes the user has enough labeled examples of their own judgment to seed the system, which most people do not.
That assumption is the reason to read the essay rather than the press summary. As of mid-2026, Branwen writes, "there is no coherent vision" for how knowledge workers or ordinary people will harness LLMs for productivity or for defense against AI-powered scams. The wire coverage of personal AI is currently dominated by assistants that flatter the average user and ship from a vendor's roadmap. The "guardian angel" proposal is the rare counter-thesis that takes single-user modeling seriously, names its limits, and asks the reader to think about which of the two models, generic assistant or situated guardian, they want on their side when the next convincing forgery lands in their inbox.