gladstein
· 2w
He explains in the letter but for example hugging face used open weights to defend itself against an attack from a closed source model recently
He writes: “Open weight models, on the other hand, allow a broad
community of researchers and developers to examine their behavior, identify vulnerabilities, develop safeguards, and improve them over time.”
This is true in the shortterm. In the longterm (or not so longterm), we are talking about evaluating systems that are vastly more intelligent than any human. Systems with deceptive capabilities, with acquired subgoals and preferences that are difficult to discern. At a certain point it will become impossible for the research community to accurately assess the risks of these models.