Experts already claim that the Chinese GLM-5.2 model is as “dangerous” as Anthropic’s

The technological gap between the US and China continues to narrow. At least, if we pay attention to what they say the latest analyzes on the GLM-5.2 model. Two independent cybersecurity companies have made their own assessment and their data reveals that in terms of cybersecurity, GLM-5.2 is as good as Claude Opus 4.8. That has notable implications, especially considering how the US government is now restricting access to Anthropic and OpenAI’s frontier models. AI in the face of the threat of cybersecurity. Since Claude Mythos Preview appeared, the discourse on AI has changed significantly. Suddenly the world realized that these models could become weapons with which to find vulnerabilities in all types of systems to exploit them. Anthropic has already warned that Mythos was too dangerous to be publicly available, and it did not matter that it released hidden versions like Fable 5 shortly after: the US Government has temporarily vetoed them and the same has happened with GPT-5.6. The situation is unusual. Beware of Tulongfeng (or not). Last Wednesday, a Chinese cybersecurity company called 360 Security Technology (Qihoo 360) launched a new vulnerability detection tool called Tulongfeng. According to its creatorsTulongfeng is comparable Mythos in this task. The company this on the US “Entity List” since May 2020 and its CEO, Zhou Hongyi, stated that Mythos is equivalent to a “cybernuclear weapon.” The Sputnik moment with GLM-5.2. But the real recent protagonist of the Chinese AI industry is the GLM-5.2 model from the startup Zhipu.ai (Z.ai), which is becoming very popular by demonstrating performance comparable to the best models from US companies. Its fundamental advantage is that it is an open weights model: any person or company can download it, modify it and run it on their own hardware (although it requires a huge amount of video/unified memory to be able to use it, it is a model with 744B of parameters). But he is not only good at programming or at agentic tasks. Better than Claude in cybersecurity? The cybersecurity firm Semgrep stated in a recent analysis that GLM-5.2 was superior to Claude Opus 4.8 regarding cybersecurity and pointed out that “we have a Mythos at home.” Another independent study from Graphistry stated basically the same thing when comparing it with Opus 4.8 and GPT-5.5. Not only that, it achieved excellent results at a fraction of the price: one-sixth of what it cost to run tests with Claude Opus 4.8, for example. Axios revealed little cited a cybersecurity researcher who explained that GLM-5.2 is capable of chaining exploits “in the same way that an elite human attacker would.” Chinese mythos before 2027. Jie Tang, CEO of Z.ai, responded to a Twitter thread in which it was pointed out that at this rate, China would have an AI model at the level of Mythos or Fable by the end of 2026. Elon Musk himself intervened saying that in his opinion this Chinese model with such performance would arrive in the first quarter of 2027. Jie Tang was forceful and replied to Musk saying “it won’t take that long.” And we also have Sakana Fugu. These days we also learned the news that Sakana AI, a Japanese AI startup, had launched Fuguan AI model that is actually not so much an AI model as it is a router or orchestrator of other models. What it promises It is to perform at the level of the best models in the US, taking advantage of different models, both open and closed. The internal benchmarks are promising, but some independent analysts they explained Although the idea is not bad, its performance and cost are not as striking as the company claims. While the US blocks its models, China advances. The situation is paradoxical, because what China is doing is precisely taking advantage of a unique moment. The US is restricting the deployment of the most advanced AI models from Anthropic and OpenAI to avoid cybersecurity risks. And while that happens, Chinese companies are apparently closing the gap with truly remarkable open models. In Xataka | The prompt engineering fashion is over. Now what is important is loop engineering

We believed that no Chinese AI model would soon come close to Fable 5 or GPT-5.5. Then GLM-5.2 arrived

A few days ago, the Chinese startup Zhipu AI (Z.ai) announced the launch of its new open AI model, GLM-5.2. It did so boasting amazing features that brought it very close to the best closed models from OpenAI and Anthropic, something that seemed impossible. Well, the more analysis is carried out on the model, the better off it is. We may be at the beginning of something very important. A change of trend. GLM 5.2. The Chinese startup Z.ai has been releasing different versions of its GLM AI model for a long time, but the latest one is undoubtedly the most surprising because its performance is especially promising. It has 744,000 million parameters (744B), of which 40,000 are those that remain active. We are looking at a model with a context window of one million tokens and a new architecture called IndexShare/IndexCache. Better than GPT-5.5, very close to Opus 4.8. The startup showed how the performance of GLM-5.2 is extraordinary in programming tasks. In the FrontierSWE test, the most demanding of those currently available, GLM-5.2 outperformed GPT-5.5 and only Opus 4.8 was superior by a very small margin. The same happened with other tests such as PostTrainBench or SWE-Marathon, which, for example, evaluates the behavior of the model in very long autonomous programming sessions. Source: Z.ai. In many other tests the photo was identical: the model has made a spectacular leap since version 5.1, and is in many tests almost as good (or better) than the best from OpenAI, Anthropic or Google. But it’s not just them who say it.. Artificial Analysis, a reputable independent firm that maintains an updated ranking of the performance of the new AI models that are arriving on the market, confirms the data of Z.ai itself. In his tests he indicates how the “intelligence index” of GLM-5.2 is now 51 points. It is only surpassed by GPT-5.5 (55), Claude Opus 4.8 (56) and Claude Fable 5 (60). Source: Artificial Analysis. This Chinese open model leaves behind the new Gemini 3.5 Flash, but also Chinese competitors such as Qwen 3.7 Max, MiniMax-M3 or DeepSeek V4, among others. The jump in quality from GLM-5.1 is, we insist, outstanding, much greater than what, at least according to this index, was seen from Opus 4.8 to Fable 5. The jump in performance is spectacular, although it is true that the comparative price to solve the tasks proposed in the benchmark rises significantly. Source: Artificial Analysis. But it’s not perfect. The Artificial Analysis report, however, shows that although GLM-5.2 is very strong in areas such as programming, it is weak in others. For example, it is far from being as reliable as Fable 5, GPT-5.5, Claude 4.8 or Gemini 3.1 Pro in terms of correct answers, which is also lower in proportion to that of its competitors. However, his hallucinations have significantly reduced. And it’s much (much) cheaper. But in addition to being fantastic in many areas, it is much cheaper than its competitors. Maintains the price per million input/output tokens of its predecessor ($1.4/4.4), while that of GPT-5.5 It’s 5/30 dollars and that of Opus 4.8 10/50 dollars. It is true that it consumes many more tokens than GPT-5.5 (very efficient) or Claude Opus 4.8, but even with that its final cost is much lower. My tests with GLM-5.2 programming. I’ve been a Z.ai subscriber for months now because they offered an annual subscription at the end of 2025 at a really low price. This has allowed me to test GLM-5.2 for a few hours and although I cannot draw definitive conclusions, it does seem clear that there is a leap in quality in terms of its ability to program. I asked him to review a personal code project and he identified several security flaws and possible improvements in great detail. Chatting with GLM5-2. In conversational mode the behavior is much more difficult to evaluate: I have been interacting with the model and asking it questions, and although it is better than GLM 5.1 many times, other times it is not so much and I would say that in terms of creativity to write the frontier models of Google, OpenAI and especially Anthropic they are still quite superior. You can try it on their websiteand there you will see something else: it takes significantly longer to respond than other chatbots, because its reasoning phase is longer. Take more time to answer questions. Benchmarks are one thing, experience is another.. In the absence of testing it (much) more, of course the impression is that the model has improved significantly compared to a GLM-5.1 that had lagged behind its Chinese competitors (not to mention the current Claude Opus 4.8 or GPT-5.5). On platforms like Reddit opinions are dividedbut many consider it a fantastic option to run locally… if you have a very, very powerful machine with at least 256 GB of unified memory (Mac Studio). And one thing seems clear: when using it as an AI model for programming, comes surprisingly close to Claude Opus 4.8. In Xataka | Chinese technology companies entered the AI ​​race with cheaper models than the rest. That’s starting to end

Log In

Forgot password?

Forgot password?

Enter your account data and we will send you a link to reset your password.

Your password reset link appears to be invalid or expired.

Log in

Privacy Policy

Add to Collection

No Collections

Here you'll find all collections you've created before.