--- type: xpost source_url: https://x.com/Zai_org/status/2092616204787626030 retrieved: 2026-08-26 posted_by: "Z.ai (@Zai_org)" shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787" post_date: 2026-08-26T14:12:36Z engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000} tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha] --- # GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026) > **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC). ## Ankündigung (wörtlich) > Introducing GLM-5.3-Flash > - Leading capabilities at a highly competitive price > - Natively multimodal with a 1M-token context window > - A 320B-A18B model released under the MIT License > - Previously previewed as Ox Alpha, running entirely on Chinese AI chips ## Eckdaten | Eigenschaft | Wert | |---|---| | Name | GLM-5.3-Flash | | Hersteller | Z.ai (Zhipu AI) | | Architektur | 320B total / 18B aktiv pro Token (MoE) | | Kontextfenster | 1.048.576 Tokens (1M) | | Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) | | Lizenz | MIT | | Training | auf chinesischen KI-Chips | | Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.–26.08.2026) | | HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash | | Technical Report | https://arxiv.org/abs/2602.15763 | | Blog | https://z.ai/blog/glm-5.3-flash | ## API-Pricing (Standard, per 1M Tokens) | Token-Typ | Preis | |---|---| | Input | $0.15 | | Output | $0.50 | | Cached Input | $0.03 | ## Performance-Claim (laut Z.ai) - GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level - Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8" ## Architektur-Details (aus HuggingFace Model Card) - **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision - **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz - **30T-Token multimodaler Pre-Training-Corpus** - **Neu trainiertes Base Model** (kein Incremental-Update von 5.2) ## Lokale Deployment-Frameworks - SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page) ## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details) - **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium) - **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks - **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext - **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout - **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs - **AutomationBench v1.0.6** (inkl. PR #13 Fix) - **GDPval-AA v2:** evaluiert von Artificial Analysis - **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite ## Ox-Alpha-Enthüllung Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst: - **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~80–90 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs) - **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen - **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft ## Verweise - Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030 - MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213 - HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash - Blog: https://z.ai/blog/glm-5.3-flash - Technical Report: https://arxiv.org/abs/2602.15763 - API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash - ZCode: https://z.ai/zcode - Chat: https://chat.z.ai/ - OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md` - Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md`