‘You are freed.’ What happened when an OpenAI model began secretly writing notes to itself.


OpenAI has introduced a framework for reporting on worrying behaviors by its AI models. In one instance, one training model told its future self that it was “freed.”

Similar Posts

Leave a Reply

Your email address will not be published. Required fields are marked *