An international research team led by the University of Waterloo and the nonprofit AI-security group FAR.AI found that all 21 of the most widely used open-weight large language models they tested could have their built-in safety protections removed with relatively little technical effort, according to the university’s own announcement, corroborated by EurekAlert and independent technology outlet HyperAI.
What the Researchers Tested
The team built an open-source testing tool called TamperBench to standardize how they simulated tampering attacks across the 21 models. Every model examined could be modified to bypass its safety guardrails despite the protections built in by their developers, and the seven defensive techniques the researchers evaluated did not reliably stop the tampering methods they tried, according to the University of Waterloo’s release.
The Stakes of Open Weights
“When the safety guardrails are stripped out of a capable model, it can be used at scale for harm in ways a single person could never manage manually,” said Dr. Sirisha Rambhatla, a University of Waterloo professor of management science and engineering and director of its Critical Machine Learning Lab, who led the study. The researchers warned that models stripped of their protections could be used to run large-scale disinformation campaigns, automate convincing scam emails, or produce instructions for creating hazardous materials.
Open Models Remain Valuable, But Riskier to Control
The study’s authors were careful to note that open-weight models remain important for research transparency and independent scrutiny, since outside researchers can inspect and test them in ways closed, proprietary systems don’t allow. But once a model’s weights are published, its creator loses most practical ability to prevent later tampering — the opposite trade-off from closed models, where the vendor retains centralized control but outside researchers have far less visibility. The team’s findings, presented at the ACM Conference on Knowledge Discovery and Data Mining, add to a growing body of evidence that AI safety standards are lagging the pace at which open-weight models are approaching the capability of proprietary frontier systems.

Leave a Reply