knowledge-base/raw/xpost/2026-08-26_zai-glm53-flash-release.md

94 lines
No EOL
4.5 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: xpost
source_url: https://x.com/Zai_org/status/2092616204787626030
retrieved: 2026-08-26
posted_by: "Z.ai (@Zai_org)"
shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787"
post_date: 2026-08-26T14:12:36Z
engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000}
tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha]
---
# GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)
> **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).
## Ankündigung (wörtlich)
> Introducing GLM-5.3-Flash
> - Leading capabilities at a highly competitive price
> - Natively multimodal with a 1M-token context window
> - A 320B-A18B model released under the MIT License
> - Previously previewed as Ox Alpha, running entirely on Chinese AI chips
## Eckdaten
| Eigenschaft | Wert |
|---|---|
| Name | GLM-5.3-Flash |
| Hersteller | Z.ai (Zhipu AI) |
| Architektur | 320B total / 18B aktiv pro Token (MoE) |
| Kontextfenster | 1.048.576 Tokens (1M) |
| Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) |
| Lizenz | MIT |
| Training | auf chinesischen KI-Chips |
| Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.26.08.2026) |
| HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash |
| Technical Report | https://arxiv.org/abs/2602.15763 |
| Blog | https://z.ai/blog/glm-5.3-flash |
## API-Pricing (Standard, per 1M Tokens)
| Token-Typ | Preis |
|---|---|
| Input | $0.15 |
| Output | $0.50 |
| Cached Input | $0.03 |
## Performance-Claim (laut Z.ai)
- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"
## Architektur-Details (aus HuggingFace Model Card)
- **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
- **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz
- **30T-Token multimodaler Pre-Training-Corpus**
- **Neu trainiertes Base Model** (kein Incremental-Update von 5.2)
## Lokale Deployment-Frameworks
- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)
## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details)
- **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
- **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
- **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
- **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
- **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs
- **AutomationBench v1.0.6** (inkl. PR #13 Fix)
- **GDPval-AA v2:** evaluiert von Artificial Analysis
- **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite
## Ox-Alpha-Enthüllung
Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst:
- **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~8090 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
- **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
- **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft
## Verweise
- Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030
- MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213
- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash
- Blog: https://z.ai/blog/glm-5.3-flash
- Technical Report: https://arxiv.org/abs/2602.15763
- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash
- ZCode: https://z.ai/zcode
- Chat: https://chat.z.ai/
- OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`
- Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md`