China’s Z.AI Releases GLM-5.2: A Model That Rivals Claude Opus—Using Zero Nvidia Chips

by shayaan

In short

  • GLM-5.2 trails Claude Opus 4.8 by just 1% on FrontierSWE – a benchmark that measures multi-hour autonomous engineering projects – and beats GPT-5.5 in the same test. It ships under an MIT license with no regional restrictions.
  • The model is built entirely on Huawei Ascend chips, with no NVIDIA hardware involved.
  • Unsloth AI has already released 2-bit GGUF quantizations that shrink the model from 1.51 TB to 238 GB. You still need 256 GB of RAM or VRAM, but at that point you can use it.

Z.ai dropped GLM-5.2 on June 16, promises top-level performance and surpasses the already advanced GLM 5.1.

The Beijing-based laboratory, which has been on the US Entity List since January 2025, appears to be benefiting from growing concerns about the US approach to AI. Over the past week, the ban of Anthropic Fable and the release of this new model have caused zAI’s shares to rise 90%, pushing it to a new all-time high.

GLM 5.2 has the numbers to back up the hype.

On FrontierSWE – a benchmark that evaluates an AI agent’s ability to complete open-ended technical projects, measured in hours, including system optimization, large-scale code construction, and applied ML research, scored by dominance percentage – GLM-5.2 scored 74.4 to Claude Opus 4.8’s 75.1. It beat GPT-5.5 at 72.6. On SWE-bench Pro, which tests GitHub’s autonomous solution to real-world problems, GLM-5.2 scored 62.1 versus GPT-5.5’s 58.6 – clearing its predecessor GLM-5.1’s 58.4 by a wide margin.

The leap in quality makes it the best open source model yet in the Artificial Analysis Intelligence Index, which aggregates the results of 9 different scores to assess the overall quality of an AI model. OpenRouter’s benchmarks put it in the same category as the now banned Claude Fable 5.

The hardware used to achieve this feat is another interesting part of the story. GLM-5.2 is trained on Huawei Ascend chips – there are no Nvidia in the pipeline. Emad Mostaque, founder of Stability AI, estimated The total training cost is approximately $25 million, 80% of that after training, which would make it extremely cheap compared to its peers.

See also  Bitcoin Owner Claims Claude AI Cracked Lost Wallet Password, Netting $400K in BTC

If Decrypt reported earlier this yearZ.ai was already training image models on Huawei’s Ascend Atlas servers without a single US chip. GLM-5.2 takes it one step further: a mix-of-experts model with 744 billion parameters, a real context window of 1 million tokens, five times the 200K limit on GLM-5.1, and an MIT license that means no government directive can flip the access switch.

Tokens are the pieces of tet that a model can read and generate, while parameters are the set of internal settings and values ​​that determine how a model processes information and generates responses

Who is it for and what does it cost?

For developers, the context window is the operational shift. Navigation across entire repositories, multi-file refactoring, and long agentic pipelines that previously required splitting into chunks become single-call workflows. API prices are $1.40 per million input tokens and $4.40 per million output – versus Claude Opus 4.8’s $5 input and $25 output. The Coding plan starts at around $18 per month and works directly in Claude Code, Cline, Kilo Code, and most popular agent environments.

Local deployment is also technically possible. Onloth AI pushed 2-bit GGUF quantizations that compress the model from 1.51 TB to 238 GB while maintaining ~82% accuracy.

Don’t get too excited, though. That still means 256 GB of unified memory or a matching RAM/VRAM combination is needed: a maxed out M4 Ultra Mac Studio or a workstation with a mid-range GPU and 256 GB of system RAM with a combination of experts. It’s still a lot of money, but at least something you can buy and run on your house if you really want to.

See also  Pavel Durov Wants to Give a Billion Telegram Users a Crypto Wallet

We ran a quick test and asked GLM-5.2 to build our standard type-mixing game mechanics with a shooter. The user interface wasn’t the prettiest: other models generated more polished-looking interfaces, but the experience was the most varied: different scenarios across waves, enemy types that changed, bosses that appeared later in the race.

It generated more diverse game states than anything else we tested for the same task in a zero shot setup.

If you want to play it, it’s live in our Itch.io profile.

This variance points in the direction of where GLM-5.2 makes the most economic sense. For multi-shot generation workflows and agentic pipelines where output diversity is more important than finish, the math on open source pricing levels is hard to argue with. For the toughest tasks – the SWE-Marathon, where it scores 13.0 versus Opus 4.8’s 26.0 – the gap with the closed border is still real, and 13 points wide.

Open source weights are live Hugging Face under the MIT license. The quantized weights are also available at Hugging Face. GLM Coding Plan subscribers can now transition with the GLM-5.2 model series, and it is also available for free testing on z.AI with some usage restrictions.

Daily debriefing Newsletter

Start every day with today’s top news stories, plus original articles, a podcast, videos and more.

Source link

You may also like

Latest News

Copyright © Sovereign Wealth Signals