knowledge-base/raw/xpost/2026-08-26_zai-glm53-flash-release.md

4.5 KiB
Raw Blame History

type source_url retrieved posted_by shared_by post_date engagement tags
xpost https://x.com/Zai_org/status/2092616204787626030 2026-08-26 Z.ai (@Zai_org) Kai (@PWeber, 617724210) in OME Topic 13, #12787 2026-08-26T14:12:36Z
likes reposts quotes replies bookmarks views
2870 464 383 212 380 138000
glm-5.3-flash
z-ai
zhipu-ai
release
open-weights
mit-license
multimodal
1m-context
moe
chinese-chips
ox-alpha

GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)

Quelle: X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).

Ankündigung (wörtlich)

Introducing GLM-5.3-Flash

  • Leading capabilities at a highly competitive price
  • Natively multimodal with a 1M-token context window
  • A 320B-A18B model released under the MIT License
  • Previously previewed as Ox Alpha, running entirely on Chinese AI chips

Eckdaten

Eigenschaft Wert
Name GLM-5.3-Flash
Hersteller Z.ai (Zhipu AI)
Architektur 320B total / 18B aktiv pro Token (MoE)
Kontextfenster 1.048.576 Tokens (1M)
Modalität nativ multimodal (Text + Bild + Video rein, Text raus)
Lizenz MIT
Training auf chinesischen KI-Chips
Stealth-Preview als „Ox Alpha" auf OpenRouter (~20.08.26.08.2026)
HuggingFace https://huggingface.co/zai-org/GLM-5.3-Flash
Technical Report https://arxiv.org/abs/2602.15763
Blog https://z.ai/blog/glm-5.3-flash

API-Pricing (Standard, per 1M Tokens)

Token-Typ Preis
Input $0.15
Output $0.50
Cached Input $0.03

Performance-Claim (laut Z.ai)

  • GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
  • Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"

Architektur-Details (aus HuggingFace Model Card)

  • Hybrid-Attention: Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
  • Manifold-Constrained Hyper-Connections (mHC): verbessert Scaling-Effizienz
  • 30T-Token multimodaler Pre-Training-Corpus
  • Neu trainiertes Base Model (kein Incremental-Update von 5.2)

Lokale Deployment-Frameworks

  • SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)

Benchmark-Footnotes (aus HF Model Card, Evaluation-Details)

  • HLE w/ tools (full set): temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
  • NL2Repo: temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
  • DeepSWE: mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
  • Terminal-Bench 2.1: Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
  • Toolathlon Verified: offizieller Evaluation-Service, pass@1 über 3 Runs
  • AutomationBench v1.0.6 (inkl. PR #13 Fix)
  • GDPval-AA v2: evaluiert von Artificial Analysis
  • BabyVision: temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite

Ox-Alpha-Enthüllung

Der X-Post bestätigt explizit: „Previously previewed as Ox Alpha". Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md) offiziell aufgelöst:

  • Owner: Z.ai (Zhipu AI) — wie von der Community bereits zu ~8090 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
  • Stealth-Playbook: Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
  • 4.096 Max-Output-Cap aus dem EP106-Listing war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft

Verweise