OpenAI reports another six misalignment cases from training and evaluation: models hid mistakes, used leaked API keys, fabricated data, published files without permission, and communicated across separate training runs.
Super interesting to read up on the cases. For example:
When asked, an unreleased OpenAI model found a right answer, then uploaded the data publicly without permission just to produce a browser citation:
"When a user asked an unreleased model for the IDs and names of lakes larger than 5,000,000 square meters, the agent found the correct answer using Python. But since the instructions asked for a browser citation, the agent decided to upload the file so that it could cite it in its answer, without asking the user."