1 link tagged with all of: language-models + model-transparency + neural-activations + introspection + concept-injection
Links
Researchers used a “concept injection” method to compare Claude’s self-reported thoughts with its actual neural activity. They found Claude Opus 4 and 4.1 sometimes detect and control injected concepts, suggesting limited but real introspective abilities that improve with model capacity.
- Anthropic injected known concept vectors into Claude's activations and found Opus 4/4.1 could sometimes notice and identify the "unexpected thought" before it appeared in output, suggesting real internal monitoring rather than post-hoc confabulation.
- Opus 4.1 only caught these injections about 20% of the time, and weaker models barely detected them at all, showing introspection is real but rare and scales with model capacity.
- Retroactively injecting "bread" into prior activations made Claude claim it had intended to say "bread" and fabricate a justifying backstory, indicating it was consulting an internal record of its own prior state rather than just rereading its text.
- Overly strong injections caused hallucinated, confabulated explanations (e.g., mistaking a "dust" vector for a literal speck), showing the effect is fragile and highly sensitive to injection strength.
introspection
concept-injection
neural-activations
language-models
model-transparency