94 lines
4.5 KiB
Markdown
94 lines
4.5 KiB
Markdown
|
|
---
|
|||
|
|
type: xpost
|
|||
|
|
source_url: https://x.com/Zai_org/status/2092616204787626030
|
|||
|
|
retrieved: 2026-08-26
|
|||
|
|
posted_by: "Z.ai (@Zai_org)"
|
|||
|
|
shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787"
|
|||
|
|
post_date: 2026-08-26T14:12:36Z
|
|||
|
|
engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000}
|
|||
|
|
tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha]
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)
|
|||
|
|
|
|||
|
|
> **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).
|
|||
|
|
|
|||
|
|
## Ankündigung (wörtlich)
|
|||
|
|
|
|||
|
|
> Introducing GLM-5.3-Flash
|
|||
|
|
> - Leading capabilities at a highly competitive price
|
|||
|
|
> - Natively multimodal with a 1M-token context window
|
|||
|
|
> - A 320B-A18B model released under the MIT License
|
|||
|
|
> - Previously previewed as Ox Alpha, running entirely on Chinese AI chips
|
|||
|
|
|
|||
|
|
## Eckdaten
|
|||
|
|
|
|||
|
|
| Eigenschaft | Wert |
|
|||
|
|
|---|---|
|
|||
|
|
| Name | GLM-5.3-Flash |
|
|||
|
|
| Hersteller | Z.ai (Zhipu AI) |
|
|||
|
|
| Architektur | 320B total / 18B aktiv pro Token (MoE) |
|
|||
|
|
| Kontextfenster | 1.048.576 Tokens (1M) |
|
|||
|
|
| Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) |
|
|||
|
|
| Lizenz | MIT |
|
|||
|
|
| Training | auf chinesischen KI-Chips |
|
|||
|
|
| Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.–26.08.2026) |
|
|||
|
|
| HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash |
|
|||
|
|
| Technical Report | https://arxiv.org/abs/2602.15763 |
|
|||
|
|
| Blog | https://z.ai/blog/glm-5.3-flash |
|
|||
|
|
|
|||
|
|
## API-Pricing (Standard, per 1M Tokens)
|
|||
|
|
|
|||
|
|
| Token-Typ | Preis |
|
|||
|
|
|---|---|
|
|||
|
|
| Input | $0.15 |
|
|||
|
|
| Output | $0.50 |
|
|||
|
|
| Cached Input | $0.03 |
|
|||
|
|
|
|||
|
|
## Performance-Claim (laut Z.ai)
|
|||
|
|
|
|||
|
|
- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
|
|||
|
|
- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"
|
|||
|
|
|
|||
|
|
## Architektur-Details (aus HuggingFace Model Card)
|
|||
|
|
|
|||
|
|
- **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
|
|||
|
|
- **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz
|
|||
|
|
- **30T-Token multimodaler Pre-Training-Corpus**
|
|||
|
|
- **Neu trainiertes Base Model** (kein Incremental-Update von 5.2)
|
|||
|
|
|
|||
|
|
## Lokale Deployment-Frameworks
|
|||
|
|
|
|||
|
|
- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)
|
|||
|
|
|
|||
|
|
## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details)
|
|||
|
|
|
|||
|
|
- **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
|
|||
|
|
- **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
|
|||
|
|
- **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
|
|||
|
|
- **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
|
|||
|
|
- **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs
|
|||
|
|
- **AutomationBench v1.0.6** (inkl. PR #13 Fix)
|
|||
|
|
- **GDPval-AA v2:** evaluiert von Artificial Analysis
|
|||
|
|
- **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite
|
|||
|
|
|
|||
|
|
## Ox-Alpha-Enthüllung
|
|||
|
|
|
|||
|
|
Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst:
|
|||
|
|
|
|||
|
|
- **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~80–90 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
|
|||
|
|
- **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
|
|||
|
|
- **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft
|
|||
|
|
|
|||
|
|
## Verweise
|
|||
|
|
|
|||
|
|
- Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030
|
|||
|
|
- MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213
|
|||
|
|
- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash
|
|||
|
|
- Blog: https://z.ai/blog/glm-5.3-flash
|
|||
|
|
- Technical Report: https://arxiv.org/abs/2602.15763
|
|||
|
|
- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash
|
|||
|
|
- ZCode: https://z.ai/zcode
|
|||
|
|
- Chat: https://chat.z.ai/
|
|||
|
|
- OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`
|
|||
|
|
- Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md`
|