An arXiv preprint finds Claude Opus 4 and 4.1 can sometimes detect when researchers directly plant specific concepts inside their internal processing, but the authors call the capacity highly unreliable and context dependent.
A new arXiv preprint tested whether large language models can detect when researchers tamper with their own internal processing, and whether they can tell their own outputs apart from text inserted by the experimenters. The answer, in narrow conditions, is sometimes yes.
The study, "Emergent Introspective Awareness in Large Language Models," injected representations of known concepts directly into a model's activations, then asked the model whether it noticed. In certain scenarios, the most capable models in the test, Claude Opus 4 and 4.1, reported the injection accurately and could distinguish their own generations from "prefills" the researchers slipped in. Some models could also recall prior internal states.
The authors' own conclusion is conservative: the capacity is "highly unreliable and context-dependent," and trends across models are "complex and sensitive to post-training strategies." The paper is a preprint, not peer-reviewed, and the top-performing subjects are built by the organization most likely to publish the work.
A weak, current handle on whether a model can detect that its internal state has been altered is also a primitive tool for AI safety research. It does not mean AI is self-aware. It means a model can, sometimes, notice it has been manipulated.
A Hacker News thread on the paper is already weighing whether the result measures real introspection or sophisticated pattern-matching, and the paper leaves that question open.