DevOps / SRE / Platform · 26.08.2026, 16:49 UTC
Z.ai’s GLM-5.3-Flash is cheap, good, and served on Chinese chips
| Schweregrad | info |
|---|---|
| Kategorie | DevOps / SRE / Platform |
| Quelle | The New Stack ↗ |
| Veröffentlicht | 26.08.2026 UTC |
Sicherheitsmeldung mit Schweregrad noch nicht bewertet. Technische Details im Tab „Originaltext“; empfohlene Schritte in der Checkliste.
Ox-alpha, the stealth model that quickly became the most popular model on OpenRouter in the last few days, is actually Z.ai’s GLM-5.3-Flash, a 320 billion-parameter hybrid model (with 18 billion active parameters) that the team specifically trained for ultra-low-cost inference.
On Wednesday, Z.ai unmasked the stealth model and made its weights available on Hugging Face under the MIT license.
It’s also already available on a number of inference platforms, including OpenRouter, where it’s currently available at $0.075 per million input tokens and $0.25 per million output tokens (though those prices reflect a 50% discount).
Benchmarks: It’s good, but not Fable
While some of the early hype put the model at the level of Anthropic’s Claude Fable 5, the benchmarks don’t bear this out. But the model can, for the most part, keep up with a Claude Opus 4.8 and OpenAI’s GPT-5.6 Terra when set to its max-effort reasoning mode.
On the Artificial Analysis Intelligence Index, GLM-5.3-Flash sits at 57 points, in line with GPT-5.6 Terra, Google’s Gemini 3.7 Flash, Meta’s Muse Spark 1.2, and Qwen 3.8 2.4T A95B.
When it comes to its performance in driving AI agents, which may be a better indication of how it will perform in real-world use cases, it’s doing even better than its competitors.
The model is able to understand multimodal inputs, including images, videos, and files. Z.ai also stresses that it trained the model to do better at visual tasks like building presentations and websites, as well as at standard knowledge work tasks like working with documents, spreadsheets, and …
Maßnahmen
⬇ Als MarkdownVerwandte Beiträge
- info Claude Desktop can now easily run Qwen, DeepSeek and Kimi models — after Ollama’s first effort stalled
- info Google’s new legal AI exposes a bigger battle over the enterprise stack
- info Perplexity just separated reasoning from authority. Here’s why it matters for enterprises.
- info Z.ai’s GLM-5.3 Flash is cheap, good, and served on Chinese chips