2 comments
Is this in any way similar to Goodfire's work? <a href="https://www.goodfire.ai/research/rlfr#" rel="nofollow">https://www.goodfire.ai/research/rlfr#</a>
> So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.<p>Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)?<p>Curious to learn more. Any more info on your approach (esp the mechanistic study)?