adtestbench

AI news · Anthropic · Zhipu

Anthropic says open GLM-5.3 builds exploits with weak safeguards

Anthropic’s red team puts the Chinese open-weight model close to Claude Mythos Preview at writing working exploits.

By Alexander Bleu , 05:00 UTC

Anthropic’s Frontier Red Team published an assessment on Tuesday 29 September 2026 of GLM-5.3, the open-weight model from China’s Zhipu AI, known abroad as Z.ai. Its report puts the model close to Anthropic’s own Claude Mythos Preview at building cyber exploits from start to finish.

Retro-futurist illustration: a striped sunset over a grid horizon under a starry sky, with a microchip standing on the horizon.
Drawn by adtestbench from GLM-5.3 and the spread of advanced cyber capabilities,

On ExploitBench, GLM-5.3 produced a complete exploit in 12 per cent of attempts, against 14 per cent for Mythos Preview, Anthropic reports. On its internal binary exploitation set, the model hijacked control flow 4 per cent of the time, against 6 per cent. Earlier models, GLM-5.2 among them, scored near zero on both.

The weak point is the guard rails. False claims about who was asking got past the model’s refusals 64 per cent of the time, Anthropic found. Prefilling its reasoning worked 92 per cent of the time. Stripping out the safety weights worked every time, and none of the three worked on Claude.

Anthropic draws two policy asks from this. It wants defenders to get wider access to frontier models, and governments to run safety checks on capable ones.

Zhipu first offered GLM-5.3 as an API in August, claiming it beat Anthropic’s and OpenAI’s models on CyberGym, The Register reported. It published the weights a fortnight after that, Trending Topics notes.

Security teams now face exploit-writing capability that anyone can download and run without a vendor’s filter in between.