Back to News Hub
🟢TechCrunch AI
August 4, 2026
Regulation & Policy

Open-weight AI models are catching up to the frontier. The safety gap remains.

Overview

2 approaches frontier AI capabilities while lacking key safety mitigations, renewing concerns that powerful open models could outpace governance and safeguards. As policymakers debate how to govern increasingly powerful AI systems like OpenAI's GPT-5. 6 Sol and Anthropic's Mythos, a Chinese open-weight model has narrowed the gap with the industry's leaders.

Key Takeaways

  • 2, the open-weight AI model from China's Z.
  • 2 refused none of the offensive cyber or biology tasks it was given.

    7 "refused so consistently that SaferAI could not complete CyberGym on it at all.

  • ai could apply safety measures to its hosted API, those protections become unenforceable once someone runs the weights on their own hardware, where they can remove or modify any safeguards, fine-tune the models, or change system prompts.

    Frontier developers like OpenAI and Anthropic tend to rely on safeguards like classifiers, refusal training, and API-level controls to limit dangerous cyber and biological assistance.

  • But the safeguards in place for closed models don't work at all on open-weight models, which are designed to run on any infrastructure with any set of safeguards - or lack thereof.

    "The objective should clearly be that the good capabilities - the safe ones - are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion," Papadatos said.

  • Because of that, frontier developers have increasingly relied on other mitigations instead.

2, the open-weight AI model from China's Z. ai, is only a few months behind OpenAI's GPT-5. 5 and Anthropic's Claude Opus 4.

7 on cyber and bio capabilities, according to a new report from AI safety nonprofit SaferAI. But the divide between frontier capabilities and safety practices is growing. According to SaferAI's evaluation, which the nonprofit ran via Z.

2 refused none of the offensive cyber or biology tasks it was given. 7 "refused so consistently that SaferAI could not complete CyberGym on it at all. " (CyberGym is a benchmark that evaluates cybersecurity capabilities.

For more details please read the original article at TechCrunch AI.

Continue Learning

Originally published by TechCrunch AI
Read the original

Comments

Sign in to join the conversation