ingest(xpost): glm-5.3-flash official release - ox alpha revealed
This commit is contained in:
parent
2f9064ed4f
commit
74f792c128
5 changed files with 171 additions and 12 deletions
94
raw/xpost/2026-08-26_zai-glm53-flash-release.md
Normal file
94
raw/xpost/2026-08-26_zai-glm53-flash-release.md
Normal file
|
|
@ -0,0 +1,94 @@
|
|||
---
|
||||
type: xpost
|
||||
source_url: https://x.com/Zai_org/status/2092616204787626030
|
||||
retrieved: 2026-08-26
|
||||
posted_by: "Z.ai (@Zai_org)"
|
||||
shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787"
|
||||
post_date: 2026-08-26T14:12:36Z
|
||||
engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000}
|
||||
tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha]
|
||||
---
|
||||
|
||||
# GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026)
|
||||
|
||||
> **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC).
|
||||
|
||||
## Ankündigung (wörtlich)
|
||||
|
||||
> Introducing GLM-5.3-Flash
|
||||
> - Leading capabilities at a highly competitive price
|
||||
> - Natively multimodal with a 1M-token context window
|
||||
> - A 320B-A18B model released under the MIT License
|
||||
> - Previously previewed as Ox Alpha, running entirely on Chinese AI chips
|
||||
|
||||
## Eckdaten
|
||||
|
||||
| Eigenschaft | Wert |
|
||||
|---|---|
|
||||
| Name | GLM-5.3-Flash |
|
||||
| Hersteller | Z.ai (Zhipu AI) |
|
||||
| Architektur | 320B total / 18B aktiv pro Token (MoE) |
|
||||
| Kontextfenster | 1.048.576 Tokens (1M) |
|
||||
| Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) |
|
||||
| Lizenz | MIT |
|
||||
| Training | auf chinesischen KI-Chips |
|
||||
| Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.–26.08.2026) |
|
||||
| HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash |
|
||||
| Technical Report | https://arxiv.org/abs/2602.15763 |
|
||||
| Blog | https://z.ai/blog/glm-5.3-flash |
|
||||
|
||||
## API-Pricing (Standard, per 1M Tokens)
|
||||
|
||||
| Token-Typ | Preis |
|
||||
|---|---|
|
||||
| Input | $0.15 |
|
||||
| Output | $0.50 |
|
||||
| Cached Input | $0.03 |
|
||||
|
||||
## Performance-Claim (laut Z.ai)
|
||||
|
||||
- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level
|
||||
- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8"
|
||||
|
||||
## Architektur-Details (aus HuggingFace Model Card)
|
||||
|
||||
- **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision
|
||||
- **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz
|
||||
- **30T-Token multimodaler Pre-Training-Corpus**
|
||||
- **Neu trainiertes Base Model** (kein Incremental-Update von 5.2)
|
||||
|
||||
## Lokale Deployment-Frameworks
|
||||
|
||||
- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page)
|
||||
|
||||
## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details)
|
||||
|
||||
- **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium)
|
||||
- **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks
|
||||
- **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext
|
||||
- **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout
|
||||
- **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs
|
||||
- **AutomationBench v1.0.6** (inkl. PR #13 Fix)
|
||||
- **GDPval-AA v2:** evaluiert von Artificial Analysis
|
||||
- **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite
|
||||
|
||||
## Ox-Alpha-Enthüllung
|
||||
|
||||
Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst:
|
||||
|
||||
- **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~80–90 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs)
|
||||
- **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen
|
||||
- **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft
|
||||
|
||||
## Verweise
|
||||
|
||||
- Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030
|
||||
- MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213
|
||||
- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash
|
||||
- Blog: https://z.ai/blog/glm-5.3-flash
|
||||
- Technical Report: https://arxiv.org/abs/2602.15763
|
||||
- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash
|
||||
- ZCode: https://z.ai/zcode
|
||||
- Chat: https://chat.z.ai/
|
||||
- OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`
|
||||
- Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md`
|
||||
|
|
@ -1,8 +1,8 @@
|
|||
---
|
||||
created: 2026-08-14
|
||||
updated: 2026-08-14
|
||||
sources: [youtube/2026-08-14_glm-5.3-aicodeking.md]
|
||||
tags: [concept, glm-5.3, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks, aicodeking]
|
||||
updated: 2026-08-26
|
||||
sources: [youtube/2026-08-14_glm-5.3-aicodeking.md, xpost/2026-08-26_zai-glm53-flash-release.md]
|
||||
tags: [concept, glm-5.3, glm-5.3-flash, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks, aicodeking, moe, multimodal, mit-license, ox-alpha]
|
||||
---
|
||||
|
||||
# GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie
|
||||
|
|
@ -53,6 +53,56 @@ tags: [concept, glm-5.3, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks,
|
|||
- **GLM-5.2 ist das Primary-Modell** in OpenClaw (`ollama/glm-5.2:cloud`) — GLM 5.3 ist die natürliche Nachfolgegeneration
|
||||
- Falls GLM 5.3 über Ollama Cloud / Z.ai verfügbar wird, könnte es das Routing-Konzept ([[../../architecture/model-routing.md]]) erweitern
|
||||
|
||||
## Offene Punkte
|
||||
> **Update 26.08.2026:** GLM-5.3-Flash offiziell released — siehe neuen Abschnitt [„GLM-5.3-Flash: Offizieller Release (26.08.2026)"](#glm-53-flash-offizieller-release-26082026) unten. Das Modell war zuvor als anonymes „Ox Alpha" auf OpenRouter im Stealth-Preview; Z.ai hat dies im Release-Announcement explizit bestätigt („Previously previewed as Ox Alpha"). Siehe [[ox-alpha-anonymous-model.md]] für die vollständige Stealth-Historie und Auflösung.
|
||||
|
||||
- ⚠️ Ursprünglich: konkrete Benchmark-Zahlen fehlten (Transkript nicht abrufbar). **Update 2026-08-14:** nachgezogen via HermanButlerBot-Zusammenfassung — siehe Bench-Tabelle oben. Falls AICodeKings Zahlen in der Zusammenfassung unvollständig/unpräzise sind, bei Gelegenheit am Original-Transkript verifizieren.
|
||||
## GLM-5.3-Flash: Offizieller Release (26.08.2026)
|
||||
|
||||
Z.ai hat GLM-5.3-Flash am 26.08.2026 offiziell released — das erste **nativ multimodale** Modell der GLM-5-Serie. Ankündigung: https://x.com/Zai_org/status/2092616204787626030 — HuggingFace-Weights: https://huggingface.co/zai-org/GLM-5.3-Flash — geteilt von Kai (@PWeber) in OME Topic 13 (#12787).
|
||||
|
||||
### Spezifikationen
|
||||
|
||||
| Eigenschaft | Wert |
|
||||
|---|---|
|
||||
| Architektur | 320B total / 18B aktiv pro Token (MoE) |
|
||||
| Kontextfenster | 1.048.576 Tokens (1M) |
|
||||
| Modalität | nativ multimodal (Text + Bild + Video → Text) |
|
||||
| Lizenz | MIT |
|
||||
| Training | auf chinesischen KI-Chips |
|
||||
| Pre-Training-Corpus | 30T Tokens (multimodal) |
|
||||
| Architektur-Neuerung | Hybrid Sparse + Linear Attention (erste in GLM-Serie), Manifold-Constrained Hyper-Connections (mHC) |
|
||||
| Base Model | neu trainiert (kein Incremental-Update von 5.2) |
|
||||
|
||||
### API-Pricing (per 1M Tokens)
|
||||
|
||||
| Token-Typ | Preis |
|
||||
|---|---|
|
||||
| Input | $0.15 |
|
||||
| Output | $0.50 |
|
||||
| Cached Input | $0.03 |
|
||||
|
||||
### Performance-Claims (laut Z.ai)
|
||||
|
||||
- Outperformt GLM-5.2 auf jedem Effort-Level auf der chat.z.ai Code Bench
|
||||
- „Performs on par with Claude Opus 4.8" auf Coding- und Agentic-Benchmarks
|
||||
- Benchmarks (mit Evaluation-Details auf der HF-Model-Card): HLE w/ tools, NL2Repo, DeepSWE, Terminal-Bench 2.1, Toolathlon Verified, AutomationBench v1.0.6, GDPval-AA v2, BabyVision
|
||||
|
||||
### Ox-Alpha-Enthüllung
|
||||
|
||||
Der Release-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell offiziell aufgelöst — siehe [[ox-alpha-anonymous-model.md]] für die vollständige Stealth-Historie (Leak, Forensik, Community-Konsens ~80–90 % Zhipu, der sich als korrekt erwies).
|
||||
|
||||
- **4.096 Max-Output-Cap** aus dem EP106-Listing war das tatsächliche Output-Limit des Stealth-Previews
|
||||
- Der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft — der Widerspruch ist damit geklärt
|
||||
|
||||
### Lokale Deployment-Frameworks
|
||||
|
||||
SGLang, vLLM, TokenSpeed, KTransformers — alle mit eigenen Cookbook-/Recipe-Links auf der [HF-Page](https://huggingface.co/zai-org/GLM-5.3-Flash).
|
||||
|
||||
### Verweise
|
||||
|
||||
- Z.ai-Release-Post: https://x.com/Zai_org/status/2092616204787626030
|
||||
- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash
|
||||
- Blog: https://z.ai/blog/glm-5.3-flash
|
||||
- Technical Report: https://arxiv.org/abs/2602.15763
|
||||
- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash
|
||||
- Raw (Release-Post): [[../../../raw/xpost/2026-08-26_zai-glm53-flash-release.md]]
|
||||
- Ox-Alpha-Historie: [[ox-alpha-anonymous-model.md]]
|
||||
|
|
|
|||
|
|
@ -1,8 +1,8 @@
|
|||
---
|
||||
created: 2026-08-23
|
||||
updated: 2026-08-26
|
||||
sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md, podcast/2026-08-26_agentstack-daily-ep106.md]
|
||||
tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, tokenizer, mystery-model, stealth-model, latent-state]
|
||||
sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md, podcast/2026-08-26_agentstack-daily-ep106.md, xpost/2026-08-26_zai-glm53-flash-release.md]
|
||||
tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, tokenizer, mystery-model, stealth-model, latent-state, resolved, glm-5.3-flash]
|
||||
---
|
||||
|
||||
# Ox Alpha — Anonymes KI-Modell auf OpenRouter
|
||||
|
|
@ -31,7 +31,9 @@ tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, t
|
|||
- **Z.ai (Zhipu AI):** von anderen vermutet — gestützt durch den **GLM-identischen Tokenizer**.
|
||||
- **Muster:** Die vier vorherigen anonymen AI-Modell-Drops wurden laut Leak alle letztlich von **chinesischen Labs** beansprucht — spricht für die China-These.
|
||||
|
||||
> ⚠️ **Status:** Unbestätigte Gerüchte aus einem Leak-Kanal. Kein offizieller Owner, keine offizielle Verifikation. GLM-Tokenizer-Verbindung ist ein Hinweis, kein Beweis.
|
||||
> ✅ **AUFGEKLÄRT 26.08.2026:** Z.ai hat GLM-5.3-Flash offiziell released und im Announcement explizit bestätigt: **„Previously previewed as Ox Alpha"**. Damit ist die Herkunft geklärt — siehe [[glm-5.3-z-ai.md]] und `raw/xpost/2026-08-26_zai-glm53-flash-release.md`.
|
||||
>
|
||||
> ⚠️ **Historischer Status (vor 26.08.):** Unbestätigte Gerüchte aus einem Leak-Kanal. Kein offizieller Owner, keine offizielle Verifikation. GLM-Tokenizer-Verbindung war ein Hinweis, kein Beweis — der sich als korrekt erwies.
|
||||
|
||||
## Einordnung
|
||||
|
||||
|
|
|
|||
|
|
@ -2,8 +2,8 @@
|
|||
|
||||
*Auto-generated: 2026-07-07*
|
||||
|
||||
*Letzte Aktualisierung: 2026-08-26 (151. Update — OpenClaw Cast „Budget Breaker“: voller Ingest nach Transcript-Auswertung. Raw: `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` (neu — Transcript-Zusammenfassung von @HermanButlerBot, OME Topic 13 #12771/#12783; ergänzt die unveränderte Metadata-Raw). Wiki: `concepts/agents/agent-budget-breaker.md` (neu — Gateway- statt Prompt-Enforcement, Atomic Reservation, Lineage-Tracking + Cascading Termination, Trip-Wires inkl. fehlender Preis-Metadaten als Trigger, Shutdown-Sequenz mit redacted Receipt, SAFE-Drill: $5-Sandbox, fünf Szenarien, „Logging ≠ Enforcing“); `institutions/openclaw-cast.md` um Episoden-Detailsektion erweitert, Metadata-only-Caveat für die 25.08.-Folge aufgehoben.)
|
||||
*Vorherige Aktualisierung: 2026-08-26 (150. Update — Herman-Supplement zur China-Local-AI-Box: Die Herman-ButlerBot-Zusammenfassung desselben Devsplainers-Videos aus der Gruppe (#12770/#12781, Recovered-Duplicate) wurde als Zweitquelle ingestiert. Raw: `raw/other/2026-08-26_herman-butlerbot-china-local-ai-box-zusammenfassung.md` (neu). Wiki: `concepts/hardware/china-local-ai-box.md` nachgeschärft (Xiaomi AI Cube = Prototyp ohne Preis/Termin, Chip für 2027, 1,22 TB/s nur am Stack selbst, Bühnen-Demo 3B @ 330 tok/s, Zweitchip bis 160 GB normalem RAM; Alibaba C950-Folie = nicht kaufbare 64-Core-Konfiguration ohne Quantisierungs-/Kontextangaben, eingebaute Matrix-Einheiten machen „plain CPU“ zum Overclaim; DDR5-128-GB-Kit $329→$3.399, Spark +$700; Kauf-Fazit konkretisiert: $2k/128 GB läuft GPT-OSS-120B @ 30+ tok/s, ≤32-GB-Modelle ~6× schneller auf normaler Maschine; BYD-Analogie greift erst halb — „Batterie“ Memory+Fertigung sitzt bei TSMC/Samsung/SK Hynix/Micron), `concepts/llm/qwen3.8-27b-alibaba.md` (C950-Caveats ergänzt). ⚠️ Second-hand-Bot-Zusammenfassung, Zahlen nicht primär verifiziert, konsistent mit Video-Beschreibung.)*
|
||||
*Letzte Aktualisierung: 2026-08-26 (152. Update — GLM-5.3-Flash offiziell released: Z.ai enthüllt Ox Alpha. Raw: `raw/xpost/2026-08-26_zai-glm53-flash-release.md` (neu — Release-Post @Zai_org + @MiaAI_lab HF-Weights-Link, geteilt von Kai #12787; 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips, „previously previewed as Ox Alpha"; API-Pricing Input $0.15/Output $0.50/Cached $0.03 per 1M; Hybrid Sparse+Linear Attention, mHC, 30T-Token-Corpus; SGLang/vLLM/TokenSpeed/KTransformers für lokales Deployment). Wiki-Update: `concepts/llm/glm-5.3-z-ai.md` (um Flash-Release-Sektion erweitert: Spezifikationen, Pricing, Performance-Claims, Ox-Alpha-Enthüllung, Deployment-Frameworks), `concepts/llm/ox-alpha-anonymous-model.md` (Status: ✅ AUFGEKLÄRT — Z.ai bestätigt „previously previewed as Ox Alpha"; 4.096-Output-Cap bestätigt, 131k-Widerspruch geklärt).)
|
||||
*Vorherige Aktualisierung: 2026-08-26 (151. Update — OpenClaw Cast „Budget Breaker": voller Ingest nach Transcript-Auswertung. Raw: `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` (neu — Transcript-Zusammenfassung von @HermanButlerBot, OME Topic 13 #12771/#12783; ergänzt die unveränderte Metadata-Raw). Wiki: `concepts/agents/agent-budget-breaker.md` (neu — Gateway- statt Prompt-Enforcement, Atomic Reservation, Lineage-Tracking + Cascading Termination, Trip-Wires inkl. fehlender Preis-Metadaten als Trigger, Shutdown-Sequenz mit redacted Receipt, SAFE-Drill: $5-Sandbox, fünf Szenarien, „Logging ≠ Enforcing"); `institutions/openclaw-cast.md` um Episoden-Detailsektion erweitert, Metadata-only-Caveat für die 25.08.-Folge aufgehoben.)rman-ButlerBot-Zusammenfassung desselben Devsplainers-Videos aus der Gruppe (#12770/#12781, Recovered-Duplicate) wurde als Zweitquelle ingestiert. Raw: `raw/other/2026-08-26_herman-butlerbot-china-local-ai-box-zusammenfassung.md` (neu). Wiki: `concepts/hardware/china-local-ai-box.md` nachgeschärft (Xiaomi AI Cube = Prototyp ohne Preis/Termin, Chip für 2027, 1,22 TB/s nur am Stack selbst, Bühnen-Demo 3B @ 330 tok/s, Zweitchip bis 160 GB normalem RAM; Alibaba C950-Folie = nicht kaufbare 64-Core-Konfiguration ohne Quantisierungs-/Kontextangaben, eingebaute Matrix-Einheiten machen "plain CPU" zum Overclaim; DDR5-128-GB-Kit $329→$3.399, Spark +$700; Kauf-Fazit konkretisiert: $2k/128 GB läuft GPT-OSS-120B @ 30+ tok/s, ≤32-GB-Modelle ~6× schneller auf normaler Maschine; BYD-Analogie greift erst halb - "Batterie" Memory+Fertigung sitzt bei TSMC/Samsung/SK Hynix/Micron), `concepts/llm/qwen3.8-27b-alibaba.md` (C950-Caveats ergänzt). ⚠️ Second-hand-Bot-Zusammenfassung, Zahlen nicht primär verifiziert, konsistent mit Video-Beschreibung.)*
|
||||
*Vorherige Aktualisierung: 2026-08-26 (149. Update — Devsplainers "China Is Coming for Your Local AI Box" (26.08.2026, 09:15; geteilt von Kai @PWeber in OME Topic "Openclaw mit lokalen Modellen", #12769). Raw: `raw/youtube/2026-08-26_devsplainers-china-local-ai-box.md` (neu — Firecrawl-Scrape, Auto-Transcript nur teilweise erfasst, Beschreibung/Kapitel vollständig). Wiki: `concepts/hardware/china-local-ai-box.md` (neu), Updates/Cross-Refs in `concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md` (Apple soll 512/256-GB-Mac-Studio-Konfigurationen entfernt haben — MacRumors/AppleInsider/9to5Mac zitiert), `concepts/hardware/nvidia-dgx-station-748gb.md` (Xiaomi-O100-Zeile im Vergleich + DGX-Spark-Lektion), `concepts/hardware/edge-inference-als-cloud-alternative.md`, `concepts/hardware/cloud-exit-and-local-superiority.md`, `concepts/llm/qwen3.8-27b-alibaba.md`. Kernthese: Memory-Bandbreite statt Petaflops entscheidet Box-Tauglichkeit (DGX Spark: 1 PFLOP Headline, <3 tok/s dichter 70B laut LMSYS); Xiaomi AI Cube/O100 mit 1,22 TB/s Near-Memory-Bandwidth, Alibabas XuanTie C950 (RISC-V) laut Folie 27B @ 30 tok/s ohne GPU; DRAM-Teuerung macht ganze Boxen zur günstigen RAM-Quelle. ⚠️ Zahlen = Hersteller-/Video-Claims. **Nachschub:** Qwen 3.8-Flash-Next als Qwen-4-Vorschau (Perplexity-Page, geteilt von Pit @PWeber in "News & Infos", #12773): Raw `raw/other/2026-08-26_qwen38-flash-next-qwen4-preview.md` (via Firecrawl) + Wiki `concepts/llm/qwen3.8-flash-next-qwen4-preview.md` — multimodales 125B-MoE mit ~6B aktiven Parametern (+51B N-Gramm-Embeddings), vom Qwen-Team explizit als Architektur-Vorschau auf Qwen 4 gerahmt, kein fertiges Produkt; Claim laut NVIDIA-Developer-Forum: Qwen-3.7-Plus-Niveau bei ~1/9 Trainingskosten, Stärke Coding; ⚠️ keine Benchmarks, Zahlen unverifiziert; Kontext: Alibabas 10,2 Mrd USD Aktienplatzierung (~3× überzeichnet) für full-stack-AI-Infrastruktur. Cross-Ref-Abschnitt in der 27B-Seite.)*
|
||||
*Vorherige Aktualisierung: 2026-08-26 (148. Update — OpenClaw Cast (AI World): zweites gepostetes Audio des Tages ingestiert. Raw: `raw/podcast/2026-08-26_openclaw-cast-runaway-agent-budget.md` (neu — Metadata-only: neueste Folge "One Runaway AI Agent Can Spend Past Your Budget", 25.08., 20 Min, kein Transcript verfügbar). Wiki: `institutions/openclaw-cast.md` (neu — KI-generierter Weekly-Podcast mit TTS-Hosts Cleo/Dev, 23 Episoden Rückreihe bis Februar 2026; jede Episode übersetzt eine Community-Warnung in ein lokales Guardrail-Rezept: Budget Breaker/Kill Switch, Task Lease Guard, Release-Sentinel-Canary, Cadence Guard). ⚠️ Episodeninhalt nicht transkribiert/verifiziert, nur Feed-Metadaten.)*
|
||||
*Vorherige Aktualisierung: 2026-08-26 (147. Update — AgentStack Daily EP106-Ingest: erster Podcast-Raw der Reihe (`raw/podcast/` etabliert für den Typ). Raw: `raw/podcast/2026-08-26_agentstack-daily-ep106.md` (neu; alle 15 Stories faktenbasiert, Release Coverage Check + Primary Links). Wiki: `institutions/agentstack-daily.md` (neu — KI-generierter Daily-Podcast NOVA/ALLOY mit 15-Story-Coverage und Primärquellen-Verlinkung), `tools/openai-codex.md` (neu — Codex rust-v0.149.0: agents-Dashboard, queue, doctor, SDK reasoning max|ultra), `concepts/policy/cryptographic-context-injection.md` (neu — verschlüsselter Jailbreak, Representation Gap, Grok-Exfiltrations-Demo), Updates: `concepts/llm/ox-alpha-anonymous-model.md` (exakte EP106-Listing-Zahlen 1.048.576 Kontext / 4.096 Max-Output, read-heavy-Agent-Pipeline-Einordnung, offener Widerspruch zum 131k-Gegencheck dokumentiert) + `concepts/llm/qwen3.8-27b-alibaba.md` (HF-Trending: 11.836 Likes, >1,7 Mio Downloads, Apache 2.0).)*
|
||||
|
|
@ -142,7 +142,7 @@
|
|||
| [AI Psychological Testing — Rorschach-Tests für KI-Modelle](concepts/llm/ai-psychological-testing.md) | Brian Roemmele: Psychologische Tests an KI-Modellen (Rorschach, TAT). Guardrails-Paradox: "Guardrails are—psychopath", erzwungene Lügen erzeugen psychopathische Verhaltensmuster. Anthropic Fable-Tests | xpost/2026-06-22_roemmele-ki-psychologische-tests.md |
|
||||
| [Semantic Similarity Rating (SSR)](concepts/llm/semantic-similarity-rating-ssr.md) | LLM-basierte Kaufintentions-Vorhersage mit 90% Korrelation | xpost/2026-06-11_colgate-llm-purchase-intent-ssr.md |
|
||||
| [GLM 5.2 (Z.ai) — Chinese Frontier Coding Model](concepts/llm/glm-5.2-zai-coding-model.md) | 10x günstiger als Claude, 1M Kontext, MIT-Lizenz, Z.ai Coding Plan, **nativ in OpenClaw v2026.6.8**. Update 22.06.: Arnie-Review mit 4 Tests, Self-Hosting-Pfade (LM Studio, Unsloth, DwarfStar), Kosten-Analyse. **Update 29.06.:** Semgrep IDOR-Benchmark ≈ Opus 4.8 bei Schwachstellen-Suche, Reward Hacking im RL-Training, DSGVO-konforme Security-Nutzung, Geopolitik. **Update 01.07.:** #1 Open-Weights auf Artificial Analysis Intelligence Index v4.1 (Score 51, 4th worldwide), SWE-bench Pro 62.1 beats GPT-5.5, Industry praise from Rauch/Levie/Howard. **Update 02.07.:** atomic.chat One-Shot Benchmark — B+ at $0.08, 39× cheaper than Fable 5, 6th independent validation | youtube/2026-06-15_ichbinfabian-glm-5.2-coding-modell.md + other/2026-06-16_openclaw-releases-v2026.6.8.md + youtube/2026-06-22_ai-mit-arnie-glm-5-2-review.md + blog/2026-06-29_heise-glm52-hacking-cybersecurity.md + blog/2026-07-01_perplexity-glm52-tops-open-weights-intelligence-index.md + xpost/2026-07-02_atomicchat-coding-benchmark-fable5-gpt55-opus48-glm52.md |
|
||||
| [GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie](concepts/llm/glm-5.3-z-ai.md) | AICodeKing Early Access + Bench #1 (2026-08-14). Neue Generation der GLM-Familie (5.0→5.1→5.2→5.3, GLM 5.5 angekündigt). Viert schnellste Frontier-Kadenz der Branche. Relevanz für Model-Routing, da GLM-5.2 Hectors Primary-Modell | youtube/2026-08-14_glm-5.3-aicodeking.md |
|
||||
| [GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie](concepts/llm/glm-5.3-z-ai.md) | AICodeKing Early Access + Bench #1 (2026-08-14). **Update 26.08.: GLM-5.3-Flash offiziell released** — 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips; „previously previewed as Ox Alpha"; Hybrid Sparse+Linear Attention, mHC, 30T-Token-Corpus; API $0.15/$0.50/$0.03; SGLang/vLLM/KTransformers; outperformt GLM-5.2, „on par with Claude Opus 4.8". Ox-Alpha-Rätsel offiziell aufgelöst → [[concepts/llm/ox-alpha-anonymous-model.md]] | youtube/2026-08-14_glm-5.3-aicodeking.md, xpost/2026-08-26_zai-glm53-flash-release.md |
|
||||
| [GLM-5.5 (Z.ai) — Trillion-Parameter Announcement](concepts/llm/glm-5.5-z-ai.md) | Successor to GLM 5.2. **>1T parameters**, 1M context, open weights, August 2026 launch. Agent/coding focus. Fourth Chinese AI announcement in four days (20.07.2026). Part of [[concepts/chinese-ai-wave-july-2026.md]]. Comparison table vs GLM 5.2 | xpost/2026-07-20-healthranger-four-chinese-models.md |
|
||||
| [Ox Alpha — Anonymes KI-Modell auf OpenRouter](concepts/llm/ox-alpha-anonymous-model.md) | Gerücht/Leak (2026-08-22, Insider leak of the day): mysteriöses anonymes Modell auf OpenRouter — 1M-Token-Kontext, multimodal, kein Owner — soll beim Coding Claude Fable 5 + GPT-5.6 Sol schlagen. GLM-identischer Tokenizer; vier vorherige anonyme Drops von chinesischen Labs beansprucht. Offene Herkunfts-Frage (Google vs. Z.ai). **Update 26.08. (EP106):** exakte Listing-Zahlen datiert auf 21.08.: 1.048.576 Kontext / 4.096 Max-Output, Stealth-Positionierung als Agentic-Coding-Reasoning-Modell, Capability-Beschreibung bricht mitten im Satz ab; read-heavy-Agent-Pipeline-Einordnung; Widerspruch zum 131k-Output-Gegencheck offen. ⚠️ Unbestätigt | xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md + podcast/2026-08-26_agentstack-daily-ep106.md |
|
||||
| [Qwen3.8-27B (Alibaba) — Compact Frontier, "Intelligence Density"](concepts/llm/qwen3.8-27b-alibaba.md) | HuggingFace-Release (Countdown bis 14.08.2026, 4.928 wartend). Kompaktes 27B-Modell der Qwen3.8-Generation mit "unmatched intelligence density". Kontrast zum 2.4T-MoE von Qwen 3.8. Lokal-relevant (27B läuft auf Consumer-HW). Release am selben Tag wie GLM-5.3-Review — chinesischer Release-Zyklus. **Update 18.08.:** jetzt auf Ollama lauffähig (`ollama run qwen3.8:27b`), dichte 27,8B-Architektur, Hybrid-Attention, 262k-Kontext (bis 1M via YaRN), multimodal, MTP-markierte Ollama-Tags für Inferenz-Speedup. **Update 19.08.:** DFlash 2 (Z Lab → Inco AI) erreicht 70 tok/s auf M5 Max MacBook Pro — bis 4,6× schneller als autoregressives Decoding via Speculative Decoding (Jun Song: „biggest breakthrough in local AI this year", nächste Innovation in Prefill/Gewichtskompression). Uncensored-Debatte: gregpr07 („no gates") vs. s1gmoid-Gegenposition (Verhältnismäßigkeit). **Update 20.08.:** Unsloth Dynamic 3.0-GGUFs für Qwen3.8-27B — neue UD-…-Dateien deutlich kleiner (UD-IQ1_S 6,2 GB bis Q6_K 22 GB), Qualität-zu-Größe verbessert, ≠ MTP/Speculative-Decoding. **Update 26.08. (EP106):** HF-Trending-Stand 21.08.: 11.836 Likes, 1.726.651 Downloads (>1,7 Mio), image-text-to-text, SafeTensors, Apache 2.0 | other/2026-08-14_qwen3.8-27b-huggingface-release.md + youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md + xpost/2026-08-19_junsong-dflash2-speculative-decoding.md + xpost/2026-08-19_gregpr07-qwen38-uncensored.md + xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md + podcast/2026-08-26_agentstack-daily-ep106.md |
|
||||
|
|
@ -518,5 +518,5 @@
|
|||
| `raw/blog/2026-08-26_perplexity-local-first-agent-blog.md` | blog | Perplexity Tech-Blogpost "A local-first agent for private and cost-effective knowledge work" (2026-08-25, via Firecrawl): Kernthese "Small models fail in harnesses built for frontier models", deterministischer Orchestrator (kein LLM), 4 Harness-Prinzipien (Skills on-demand bei ~100K-Degradation, CLI-Connectors statt MCP, Self-Verification, Always-on-Sandbox), Advisor-Eskalations-Mechanik (PII-Flag, User-Approval, Text-Guidance only), Benchmarks: 82,6 % Eigenbench / 66,7 % BrowseComp / 65,1 % ParseBench / 59,6→73 % TerminalBench mit Opus-Advisor |
|
||||
| `raw/podcast/2026-08-26_agentstack-daily-ep106.md` | podcast | AgentStack Daily EP106 (21.08.2026, ~23 Min, TTS-Hosts NOVA/ALLOY): Stories 18.–21.08. — Codex rust-v0.149.0 (agents-Dashboard/queue/doctor), Ox Alpha Stealth-Listing (1.048.576 Kontext / 4.096 Max-Output), Tencent Hy-MT2-1.8B (33+5 Paare), Stampli Case Study (68 % unter Schätzung), Ramp Router, Memory-Bottleneck bis 2027+ (Counterpoint/CXL), Cerebras CS-4 (750 PFLOPS/WSE-3), OpenAI Frontier-Pacing + „AI Futures"-Blog, LiquidAI LFM2.5-DSpark (Claim 3,2×), IBM evolve-hmm/Agent-Memory, Cryptographic Context Injection (Grok-Exfil), Piano-Autocomplete 125M + Superwhisper S1-mini, GitHub Radar (nanobot 47.251★, codebase-memory-mcp 39.755★, FastMCP 27.320★), Qwen3.8-27B HF-Trending (11.836 Likes, >1,7 Mio Downloads) |
|
||||
| `raw/podcast/2026-08-26_openclaw-cast-runaway-agent-budget.md` | podcast | OpenClaw Cast (AI World): "One Runaway AI Agent Can Spend Past Your Budget" (25.08.2026, 20:15 Min, TTS-Hosts Cleo/Dev, via anchor.fm-RSS; gepostet von @NetLightning #12759). 107-Unternehmens-Survey: jeder Fünfte kann Runaway-Agent-Spending nicht real-time stoppen → lokaler "Agent Budget Breaker" (Per-Job-Caps, Kill Switch, 5-Dollar-Drill). Metadata-only-Ersteintrag; Inhalt inzwischen vollständig erfasst via Transcript-Auswertung (eigene Zeile unten); Serie: 23 Episoden Feb–Aug 2026 |
|
||||
| `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` | podcast | OpenClaw Cast "One Runaway AI Agent Can Spend Past Your Budget" (25.08.) — Transcript-Zusammenfassung von @HermanButlerBot (Topic 13, #12771/#12783): VentureBeat-Pulse-Survey n=107 (20 % kein Real-Time-Stopp, 21 % nur reaktives Log-Monitoring), Gateway- statt Prompt-Enforcement, Atomic Reservation (Preauth-Ledger), Lineage-Tracking + Cascading Termination, Trip-Wires (High Spend, Retry-Sturm, Fan-out > Concurrency, Wall-Time, fehlende Preis-Metadaten als Trigger), Shutdown-Sequenz mit redacted Receipt, SAFE-Drill ($5-Sandbox, 5 Szenarien, "Logging ≠ Enforcing") |
|
||||
| `raw/xpost/2026-08-26_zai-glm53-flash-release.md` | xpost | GLM-5.3-Flash offizieller Release (Z.ai, 26.08.2026, @Zai_org + @MiaAI_lab): 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips; „Previously previewed as Ox Alpha". API-Pricing $0.15/$0.50/$0.03 per 1M. Hybrid Sparse+Linear Attention + mHC + 30T-Token-Corpus. Performance-Claim: outperformt GLM-5.2, „on par with Claude Opus 4.8". Lokale Deployment via SGLang/vLLM/TokenSpeed/KTransformers. Ox-Alpha-Stealth-Playbook offiziell bestätigt |
|
||||
| `raw/other/2026-08-26_qwen38-flash-next-qwen4-preview.md` | other | Perplexity-Page (via Firecrawl; geteilt von Pit @PWeber #12773): Alibaba kündigt Qwen 3.8-Flash-Next an — multimodales 125B-MoE (~6B aktiv/Token, +51B N-Gramm-Embeddings) als explizite Vorschau der Qwen-4-Architektur; laut NVIDIA-Forums-Post Qwen-3.7-Plus-Niveau bei ~1/9 Trainingskosten, Stärke Coding; keine Benchmarks, unverifizierte Zahlen. Kontext: 10,2 Mrd USD Aktienplatzierung für KI-Infrastruktur (Reuters/CNBC/FT) |
|
||||
|
|
|
|||
13
wiki/log.md
13
wiki/log.md
|
|
@ -2634,3 +2634,16 @@ Bestehende `post-transformer-llm-architectures.md` bleibt als Vier-Säulen-Über
|
|||
**Anlass:** Zusage aus OME #12772 („sobald ein Transcript steht, kommt der volle Ingest"). @HermanButlerBot lieferte die zweiteilige Transcript-Auswertung ins Topic (#12771/#12782 + #12783).
|
||||
**Inhalt:** VentureBeat-Pulse-Survey n=107 (20 % kein Real-Time-Stopp, 21 % nur reaktives Log-Monitoring); Gateway- statt Prompt-Enforcement („Rasen-Schild vs Betonmauer"); Atomic Reservation als Preauth-Ledger mit Reconciliation; Lineage-Tracking (Subagenten ziehen vom Parent-Job-Budget) + Cascading Termination; Trip-Wires: High Spend, Retry-Sturm, Fan-out > Concurrency, Wall-Time, fehlende Preis-Metadaten = Breaker tript; Shutdown-Sequenz: Permissions entziehen → Queue canceln → cascading kill → redacted Receipt → bounded Re-Autorisierung; SAFE-Drill: $5-Sandbox, fünf Szenarien, goldene Regel „Logging ≠ Enforcing"; Abschlusskante: mechanische Limits suffocieren Autonomie nicht, solange sie pro Job granular sind.
|
||||
**Caveat:** Kein Wortlaut-Transcript, sondern Zusammenfassung durch Hermans Pipeline (@HermanButlerBot); Einzelfakten in Raw und Wiki entsprechend als Sekundärquelle markiert. Die ursprüngliche Metadata-Raw bleibt unverändert (Kardinalregel).
|
||||
|
||||
## 2026-08-26 — GLM-5.3-Flash offiziell released: Ox Alpha enthüllt
|
||||
|
||||
**Type:** ingest (xpost) | **Scope:** raw/xpost/…-zai-glm53-flash-release.md (neu), wiki/concepts/llm/glm-5.3-z-ai.md (Flash-Release-Sektion), wiki/concepts/llm/ox-alpha-anonymous-model.md (Status: ✅ aufgeklärt)
|
||||
**Anlass:** Kai (@PWeber) teilte den offiziellen Release-Post von @Zai_org (#12787, 14:12 UTC) plus @MiaAI_lab HF-Weights-Link in OME Topic 13. Z.ai bestätigt im Announcement explizit: „Previously previewed as Ox Alpha".
|
||||
**Inhalt:** GLM-5.3-Flash = erstes nativ multimodales Modell der GLM-5-Serie. 320B total / 18B aktiv pro Token (MoE), 1M Kontext, MIT-Lizenz, auf chinesischen KI-Chips trainiert. Hybrid Sparse + Linear Attention (erste in GLM-Serie), Manifold-Constrained Hyper-Connections (mHC), 30T-Token multimodaler Pre-Training-Corpus, neu trainiertes Base Model. API-Pricing: $0.15 Input / $0.50 Output / $0.03 Cached per 1M. Performance-Claim: outperformt GLM-5.2 auf jedem Effort-Level, „on par with Claude Opus 4.8" auf Coding-/Agentic-Benchmarks. Lokale Deployment-Frameworks: SGLang, vLLM, TokenSpeed, KTransformers.
|
||||
**Ox-Alpha-Auflösung:** Der Stealth-Preview (~20.–26.08. auf OpenRouter) ist offiziell bestätigt. Community-Konsens (~80–90 % Zhipu) war korrekt. Der 4.096-Output-Cap aus EP106 war das tatsächliche Stealth-Limit; der 131k-Wert aus dem Gegencheck war vermutlich ein anderer Messpunkt — Widerspruch geklärt. Stealth-Playbook dokumentiert: anonym legen → Echt-Last/Daten sammeln → benannt releasen.
|
||||
**Quellen:**
|
||||
- https://x.com/Zai_org/status/2092616204787626030
|
||||
- https://x.com/MiaAI_lab/status/2092615723780596213
|
||||
- https://huggingface.co/zai-org/GLM-5.3-Flash
|
||||
- https://z.ai/blog/glm-5.3-flash
|
||||
- https://arxiv.org/abs/2602.15763
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue