Damus
Miguel Afonso Caetano 🦾 🌹 profile picture
Miguel Afonso Caetano 🦾 🌹
@Miguel Afonso Caetano 🦾 🌹
"The safety approach that emerges from such a culture starts with unimpeded optimism about being able to solve problems as they arise. OpenAI has thrived by trial and error (which it calls “iterative deployment”), looking for problems and improving its guardrails in response. But this approach, by its very nature, guarantees periodic failures—and the scale of those failures is growing as systems get more capable. This summer, in the Hugging Face incident, OpenAI let a swarm of agents out by mistake. The company responded by making security improvements. But even after those changes, OpenAI reported that its safety controls failed again, when a model in training bypassed restrictions on internet access: A monitoring system alerted human staff but did not automatically turn the model off as it was supposed to. Anthropic, too, has acknowledged accidentally turning off its own safeguards because of a misconfiguration. I believe that such mistakes are typical of the industry, given the speed and flexibility with which people operate.

An environment where things like this can happen is no place to grow artificial minds that could be smarter than we are and that might not do what we want them to. Paul Christiano, on joining OpenAI’s board a few weeks ago, wrote that “there is a meaningful risk that rapid acceleration in AI capabilities leads to catastrophic and irreversible loss of control in the very near term.” If this is the situation, then the time for trial and error is over. Achieving something much closer to perfection the first time is essential because iteration after a mistake may not be possible. People will not be safe if we depend on individual heroics after the fact."

https://www.theatlantic.com/technology/2026/10/openai-safety-team-resignation/688881/

#CyberSecurity #AI #GenerativeAI #AISafety #OpenAI #LLMs #SiliconValley #BigTech