2 comments

  • astrobiased27 minutes ago
    Is this in any way similar to Goodfire&#x27;s work? <a href="https:&#x2F;&#x2F;www.goodfire.ai&#x2F;research&#x2F;rlfr#" rel="nofollow">https:&#x2F;&#x2F;www.goodfire.ai&#x2F;research&#x2F;rlfr#</a>
    • HenryNdubuaku21 minutes ago
      Thats an interesting outlook, loosely similar.
  • cacio-e-pepe5 hours ago
    &gt; So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.<p>Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)?<p>Curious to learn more. Any more info on your approach (esp the mechanistic study)?
    • HenryNdubuaku3 hours ago
      Correct, the study is verbose, we will compile into a neat shareable report and publish once we solve the pending caveats. Interesting username btw haha.
      • cacio-e-pepe2 hours ago
        Nice, looking forward to the report.<p>And thanks, huge pasta fan :)