| type |
source_url |
retrieved |
posted_by |
shared_by |
post_date |
engagement |
tags |
| xpost |
https://x.com/Zai_org/status/2092616204787626030 |
2026-08-26 |
Z.ai (@Zai_org) |
Kai (@PWeber, 617724210) in OME Topic 13, #12787 |
2026-08-26T14:12:36Z |
| likes |
reposts |
quotes |
replies |
bookmarks |
views |
| 2870 |
464 |
383 |
212 |
380 |
138000 |
|
| glm-5.3-flash |
| z-ai |
| zhipu-ai |
| release |
| open-weights |
| mit-license |
| multimodal |
| 1m-context |
| moe |
| chinese-chips |
| ox-alpha |
|
GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)
Quelle: X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).
Ankündigung (wörtlich)
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Eckdaten
API-Pricing (Standard, per 1M Tokens)
| Token-Typ |
Preis |
| Input |
$0.15 |
| Output |
$0.50 |
| Cached Input |
$0.03 |
Performance-Claim (laut Z.ai)
- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"
Architektur-Details (aus HuggingFace Model Card)
- Hybrid-Attention: Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
- Manifold-Constrained Hyper-Connections (mHC): verbessert Scaling-Effizienz
- 30T-Token multimodaler Pre-Training-Corpus
- Neu trainiertes Base Model (kein Incremental-Update von 5.2)
Lokale Deployment-Frameworks
- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)
- HLE w/ tools (full set): temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
- NL2Repo: temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
- DeepSWE: mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
- Terminal-Bench 2.1: Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
- Toolathlon Verified: offizieller Evaluation-Service, pass@1 über 3 Runs
- AutomationBench v1.0.6 (inkl. PR #13 Fix)
- GDPval-AA v2: evaluiert von Artificial Analysis
- BabyVision: temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite
Ox-Alpha-Enthüllung
Der X-Post bestätigt explizit: „Previously previewed as Ox Alpha". Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md) offiziell aufgelöst:
- Owner: Z.ai (Zhipu AI) — wie von der Community bereits zu ~80–90 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
- Stealth-Playbook: Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
- 4.096 Max-Output-Cap aus dem EP106-Listing war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft
Verweise