when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed<p>then we'll have to "flip" other models to be snitches on the other agents<p>then they'll make double-agents<p>the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves<p>yeah this won't end well, at all
Someone should train an LLM on a corpus without the concept of lies. I wonder if there’s enough data
>when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed<p>They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.