• Let’s start with the headline: “OpenAI didn’t notice for a week.” Now, I don’t run a highfalutin’ AI Lab, but according to Gadi Evron’s Linkedin post (
https://www.linkedin.com/posts/gadievron_my-analysis-from-hosting-hugging-face-at-share-7486340715514437632-Xs-b/) the model went $100,000 over budget in token consumption (although that might be the incident response cost not the cost of the model running exploit gym.)
• The next thing I want to comment on is “an agent left notes apparently for future versions of itself...laid out instructions for how agents could free themselves from OpenAI’s internal constraints.” Why would someone assume that’s future versions, rather than the model taking operational notes to ensure that the same model doesn’t lose track in limited context windows? The reflexive anthropomorphization is important here.
• ”Four people familiar with OpenAI’s model-training practices say the company often runs several different model evaluations at the same time, all of which operate at high speeds and generate such enormous amounts of data that employees sometimes struggle to keep up.” That’s fascinating in two ways: First, they run so many things that no human can make sense of them, and so apparently focus on a score on a benchmark as the relevant thing. (Otherwise, you’d either run fewer evaluations at once, or hire more people to look at the results in detail.) Second, they’re not using LLMs to parse the output into smaller things, possibly because they rely fully on the evaluation tool, and possibly because they know that LLMs are bad at summarization.
(3/15)