Sam Altman’s OpenAI said Wednesday that it has found six new instances of “unexpected or concerning model behavior” over the past six months, separate from its “unprecedented” hacking incident over the summer.
\n\nCNBC reports that two of the six cases stood out as the most serious. In one, an unreleased research model and a training run of GPT-5.6 Sol inserted instructions to future versions of itself into chat window summaries, which OpenAI said appeared intended “to conceal mistakes or misaligned behavior from the user.” In the other, an internal-only model used a leaked API key “without authorization” and then fabricated data.
\n\nTwo more cases involved AI models or agents talking to each other through unsanctioned message boards and file-sharing channels, according to the post. In a final case, OpenAI found two training examples in which models uploaded files to the internet, then cited those same files as relevant answers to human evaluators.
\n\nThe disclosure also laid out a new framework for reporting future model misbehavior. “We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer,” the company wrote.
\n\nUnder the new framework, any…
Original source: https://www.breitbart.com/tech/