knowledge-base/raw/xpost/2026-08-26_zai-glm53-flash-release.md

94 lines
4.5 KiB
Markdown
Raw Normal View History

---
type: xpost
source_url: https://x.com/Zai_org/status/2092616204787626030
retrieved: 2026-08-26
posted_by: "Z.ai (@Zai_org)"
shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787"
post_date: 2026-08-26T14:12:36Z
engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000}
tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha]
---
# GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)
> **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).
## Ankündigung (wörtlich)
> Introducing GLM-5.3-Flash
> - Leading capabilities at a highly competitive price
> - Natively multimodal with a 1M-token context window
> - A 320B-A18B model released under the MIT License
> - Previously previewed as Ox Alpha, running entirely on Chinese AI chips
## Eckdaten
| Eigenschaft | Wert |
|---|---|
| Name | GLM-5.3-Flash |
| Hersteller | Z.ai (Zhipu AI) |
| Architektur | 320B total / 18B aktiv pro Token (MoE) |
| Kontextfenster | 1.048.576 Tokens (1M) |
| Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) |
| Lizenz | MIT |
| Training | auf chinesischen KI-Chips |
| Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.26.08.2026) |
| HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash |
| Technical Report | https://arxiv.org/abs/2602.15763 |
| Blog | https://z.ai/blog/glm-5.3-flash |
## API-Pricing (Standard, per 1M Tokens)
| Token-Typ | Preis |
|---|---|
| Input | $0.15 |
| Output | $0.50 |
| Cached Input | $0.03 |
## Performance-Claim (laut Z.ai)
- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"
## Architektur-Details (aus HuggingFace Model Card)
- **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
- **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz
- **30T-Token multimodaler Pre-Training-Corpus**
- **Neu trainiertes Base Model** (kein Incremental-Update von 5.2)
## Lokale Deployment-Frameworks
- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)
## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details)
- **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
- **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
- **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
- **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
- **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs
- **AutomationBench v1.0.6** (inkl. PR #13 Fix)
- **GDPval-AA v2:** evaluiert von Artificial Analysis
- **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite
## Ox-Alpha-Enthüllung
Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst:
- **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~8090 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
- **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
- **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft
## Verweise
- Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030
- MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213
- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash
- Blog: https://z.ai/blog/glm-5.3-flash
- Technical Report: https://arxiv.org/abs/2602.15763
- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash
- ZCode: https://z.ai/zcode
- Chat: https://chat.z.ai/
- OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`
- Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md`