diff --git a/raw/xpost/2026-08-26_zai-glm53-flash-release.md b/raw/xpost/2026-08-26_zai-glm53-flash-release.md new file mode 100644 index 0000000..09bcdb0 --- /dev/null +++ b/raw/xpost/2026-08-26_zai-glm53-flash-release.md @@ -0,0 +1,94 @@ +--- +type: xpost +source_url: https://x.com/Zai_org/status/2092616204787626030 +retrieved: 2026-08-26 +posted_by: "Z.ai (@Zai_org)" +shared_by: "Kai (@PWeber, 617724210) in OME Topic 13, #12787" +post_date: 2026-08-26T14:12:36Z +engagement: {likes: 2870, reposts: 464, quotes: 383, replies: 212, bookmarks: 380, views: 138000} +tags: [glm-5.3-flash, z-ai, zhipu-ai, release, open-weights, mit-license, multimodal, 1m-context, moe, chinese-chips, ox-alpha] +--- + +# GLM-5.3-Flash — Offizieller Release (Z.ai, 26.08.2026) + +> **Quelle:** X-Post von @Zai_org (26.08.2026, 14:12 UTC): https://x.com/Zai_org/status/2092616204787626030 — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). Begleitpost: @MiaAI_lab mit HuggingFace-Weights-Link (https://x.com/MiaAI_lab/status/2092615723780596213, 14:10 UTC). + +## Ankündigung (wörtlich) + +> Introducing GLM-5.3-Flash +> - Leading capabilities at a highly competitive price +> - Natively multimodal with a 1M-token context window +> - A 320B-A18B model released under the MIT License +> - Previously previewed as Ox Alpha, running entirely on Chinese AI chips + +## Eckdaten + +| Eigenschaft | Wert | +|---|---| +| Name | GLM-5.3-Flash | +| Hersteller | Z.ai (Zhipu AI) | +| Architektur | 320B total / 18B aktiv pro Token (MoE) | +| Kontextfenster | 1.048.576 Tokens (1M) | +| Modalität | nativ multimodal (Text + Bild + Video rein, Text raus) | +| Lizenz | MIT | +| Training | auf chinesischen KI-Chips | +| Stealth-Preview | als „Ox Alpha" auf OpenRouter (~20.08.–26.08.2026) | +| HuggingFace | https://huggingface.co/zai-org/GLM-5.3-Flash | +| Technical Report | https://arxiv.org/abs/2602.15763 | +| Blog | https://z.ai/blog/glm-5.3-flash | + +## API-Pricing (Standard, per 1M Tokens) + +| Token-Typ | Preis | +|---|---| +| Input | $0.15 | +| Output | $0.50 | +| Cached Input | $0.03 | + +## Performance-Claim (laut Z.ai) + +- GLM-5.3-Flash outperformt GLM-5.2 auf jedem Effort-Level +- Auf chat.z.ai Code Bench „performs on par with Claude Opus 4.8" + +## Architektur-Details (aus HuggingFace Model Card) + +- **Hybrid-Attention:** Erstmals in der GLM-Serie kombiniert Sparse Attention + Linear Attention → senkt Long-Context-Serving-Kosten bei erhaltener Präzision +- **Manifold-Constrained Hyper-Connections (mHC):** verbessert Scaling-Effizienz +- **30T-Token multimodaler Pre-Training-Corpus** +- **Neu trainiertes Base Model** (kein Incremental-Update von 5.2) + +## Lokale Deployment-Frameworks + +- SGLang, vLLM, TokenSpeed, KTransformers (alle mit eigenen Cookbook-/Recipe-Links auf der HF-Page) + +## Benchmark-Footnotes (aus HF Model Card, Evaluation-Details) + +- **HLE w/ tools (full set):** temp=1.0, top_p=0.95, max 163.840 Output-Tokens, 300K Kontext mit Context-Management, Judge: GPT-5.6-luna (medium) +- **NL2Repo:** temp=1.0, max 64K Output, 1M Kontext, Rule-based + LLM-basierte Anti-Cheating-Checks +- **DeepSWE:** mini-swe-agent harness, temp=0.95, 6h Timeout, 400K Kontext +- **Terminal-Bench 2.1:** Claude Code 2.1.207, temp=1.0, max 65.536 Output, 6h Timeout +- **Toolathlon Verified:** offizieller Evaluation-Service, pass@1 über 3 Runs +- **AutomationBench v1.0.6** (inkl. PR #13 Fix) +- **GDPval-AA v2:** evaluiert von Artificial Analysis +- **BabyVision:** temp=1.0, top_p=0.95, 164K Kontext, Bilder ≥1.5K Pixel kürzere Seite + +## Ox-Alpha-Enthüllung + +Der X-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell (siehe `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md`) offiziell aufgelöst: + +- **Owner:** Z.ai (Zhipu AI) — wie von der Community bereits zu ~80–90 % vermutet (GLM-Tokenizer, API-Fehlercodes, chinesische Backend-Logs) +- **Stealth-Playbook:** Anonym auf OpenRouter legen → Daten/Echt-Last sammeln → benannt releasen +- **4.096 Max-Output-Cap aus dem EP106-Listing** war das tatsächliche Output-Limit des Stealth-Previews; der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft + +## Verweise + +- Z.ai-Post: https://x.com/Zai_org/status/2092616204787626030 +- MiaAI_lab-Post (Weights): https://x.com/MiaAI_lab/status/2092615723780596213 +- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash +- Blog: https://z.ai/blog/glm-5.3-flash +- Technical Report: https://arxiv.org/abs/2602.15763 +- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash +- ZCode: https://z.ai/zcode +- Chat: https://chat.z.ai/ +- OpenRouter Stealth-Listing (historisch): `raw/xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md` +- Vorige GLM-5.3-Early-Access-Review: `raw/youtube/2026-08-14_glm-5.3-aicodeking.md` \ No newline at end of file diff --git a/wiki/concepts/llm/glm-5.3-z-ai.md b/wiki/concepts/llm/glm-5.3-z-ai.md index 63023a2..862a989 100644 --- a/wiki/concepts/llm/glm-5.3-z-ai.md +++ b/wiki/concepts/llm/glm-5.3-z-ai.md @@ -1,8 +1,8 @@ --- created: 2026-08-14 -updated: 2026-08-14 -sources: [youtube/2026-08-14_glm-5.3-aicodeking.md] -tags: [concept, glm-5.3, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks, aicodeking] +updated: 2026-08-26 +sources: [youtube/2026-08-14_glm-5.3-aicodeking.md, xpost/2026-08-26_zai-glm53-flash-release.md] +tags: [concept, glm-5.3, glm-5.3-flash, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks, aicodeking, moe, multimodal, mit-license, ox-alpha] --- # GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie @@ -53,6 +53,56 @@ tags: [concept, glm-5.3, z-ai, zhipu-ai, coding-models, chinese-ai, benchmarks, - **GLM-5.2 ist das Primary-Modell** in OpenClaw (`ollama/glm-5.2:cloud`) — GLM 5.3 ist die natürliche Nachfolgegeneration - Falls GLM 5.3 über Ollama Cloud / Z.ai verfügbar wird, könnte es das Routing-Konzept ([[../../architecture/model-routing.md]]) erweitern -## Offene Punkte +> **Update 26.08.2026:** GLM-5.3-Flash offiziell released — siehe neuen Abschnitt [„GLM-5.3-Flash: Offizieller Release (26.08.2026)"](#glm-53-flash-offizieller-release-26082026) unten. Das Modell war zuvor als anonymes „Ox Alpha" auf OpenRouter im Stealth-Preview; Z.ai hat dies im Release-Announcement explizit bestätigt („Previously previewed as Ox Alpha"). Siehe [[ox-alpha-anonymous-model.md]] für die vollständige Stealth-Historie und Auflösung. -- ⚠️ Ursprünglich: konkrete Benchmark-Zahlen fehlten (Transkript nicht abrufbar). **Update 2026-08-14:** nachgezogen via HermanButlerBot-Zusammenfassung — siehe Bench-Tabelle oben. Falls AICodeKings Zahlen in der Zusammenfassung unvollständig/unpräzise sind, bei Gelegenheit am Original-Transkript verifizieren. +## GLM-5.3-Flash: Offizieller Release (26.08.2026) + +Z.ai hat GLM-5.3-Flash am 26.08.2026 offiziell released — das erste **nativ multimodale** Modell der GLM-5-Serie. Ankündigung: https://x.com/Zai_org/status/2092616204787626030 — HuggingFace-Weights: https://huggingface.co/zai-org/GLM-5.3-Flash — geteilt von Kai (@PWeber) in OME Topic 13 (#12787). + +### Spezifikationen + +| Eigenschaft | Wert | +|---|---| +| Architektur | 320B total / 18B aktiv pro Token (MoE) | +| Kontextfenster | 1.048.576 Tokens (1M) | +| Modalität | nativ multimodal (Text + Bild + Video → Text) | +| Lizenz | MIT | +| Training | auf chinesischen KI-Chips | +| Pre-Training-Corpus | 30T Tokens (multimodal) | +| Architektur-Neuerung | Hybrid Sparse + Linear Attention (erste in GLM-Serie), Manifold-Constrained Hyper-Connections (mHC) | +| Base Model | neu trainiert (kein Incremental-Update von 5.2) | + +### API-Pricing (per 1M Tokens) + +| Token-Typ | Preis | +|---|---| +| Input | $0.15 | +| Output | $0.50 | +| Cached Input | $0.03 | + +### Performance-Claims (laut Z.ai) + +- Outperformt GLM-5.2 auf jedem Effort-Level auf der chat.z.ai Code Bench +- „Performs on par with Claude Opus 4.8" auf Coding- und Agentic-Benchmarks +- Benchmarks (mit Evaluation-Details auf der HF-Model-Card): HLE w/ tools, NL2Repo, DeepSWE, Terminal-Bench 2.1, Toolathlon Verified, AutomationBench v1.0.6, GDPval-AA v2, BabyVision + +### Ox-Alpha-Enthüllung + +Der Release-Post bestätigt explizit: **„Previously previewed as Ox Alpha"**. Damit ist das Rätsel um das anonyme OpenRouter-Modell offiziell aufgelöst — siehe [[ox-alpha-anonymous-model.md]] für die vollständige Stealth-Historie (Leak, Forensik, Community-Konsens ~80–90 % Zhipu, der sich als korrekt erwies). + +- **4.096 Max-Output-Cap** aus dem EP106-Listing war das tatsächliche Output-Limit des Stealth-Previews +- Der 131k-Wert aus dem Community-Gegencheck war vermutlich ein anderer Messpunkt oder fehlerhaft — der Widerspruch ist damit geklärt + +### Lokale Deployment-Frameworks + +SGLang, vLLM, TokenSpeed, KTransformers — alle mit eigenen Cookbook-/Recipe-Links auf der [HF-Page](https://huggingface.co/zai-org/GLM-5.3-Flash). + +### Verweise + +- Z.ai-Release-Post: https://x.com/Zai_org/status/2092616204787626030 +- HuggingFace: https://huggingface.co/zai-org/GLM-5.3-Flash +- Blog: https://z.ai/blog/glm-5.3-flash +- Technical Report: https://arxiv.org/abs/2602.15763 +- API-Docs: https://docs.z.ai/guides/llm/glm-5.3-flash +- Raw (Release-Post): [[../../../raw/xpost/2026-08-26_zai-glm53-flash-release.md]] +- Ox-Alpha-Historie: [[ox-alpha-anonymous-model.md]] diff --git a/wiki/concepts/llm/ox-alpha-anonymous-model.md b/wiki/concepts/llm/ox-alpha-anonymous-model.md index f8d11f6..072500a 100644 --- a/wiki/concepts/llm/ox-alpha-anonymous-model.md +++ b/wiki/concepts/llm/ox-alpha-anonymous-model.md @@ -1,8 +1,8 @@ --- created: 2026-08-23 updated: 2026-08-26 -sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md, podcast/2026-08-26_agentstack-daily-ep106.md] -tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, tokenizer, mystery-model, stealth-model, latent-state] +sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md, podcast/2026-08-26_agentstack-daily-ep106.md, xpost/2026-08-26_zai-glm53-flash-release.md] +tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, tokenizer, mystery-model, stealth-model, latent-state, resolved, glm-5.3-flash] --- # Ox Alpha — Anonymes KI-Modell auf OpenRouter @@ -31,7 +31,9 @@ tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, t - **Z.ai (Zhipu AI):** von anderen vermutet — gestützt durch den **GLM-identischen Tokenizer**. - **Muster:** Die vier vorherigen anonymen AI-Modell-Drops wurden laut Leak alle letztlich von **chinesischen Labs** beansprucht — spricht für die China-These. -> ⚠️ **Status:** Unbestätigte Gerüchte aus einem Leak-Kanal. Kein offizieller Owner, keine offizielle Verifikation. GLM-Tokenizer-Verbindung ist ein Hinweis, kein Beweis. +> ✅ **AUFGEKLÄRT 26.08.2026:** Z.ai hat GLM-5.3-Flash offiziell released und im Announcement explizit bestätigt: **„Previously previewed as Ox Alpha"**. Damit ist die Herkunft geklärt — siehe [[glm-5.3-z-ai.md]] und `raw/xpost/2026-08-26_zai-glm53-flash-release.md`. +> +> ⚠️ **Historischer Status (vor 26.08.):** Unbestätigte Gerüchte aus einem Leak-Kanal. Kein offizieller Owner, keine offizielle Verifikation. GLM-Tokenizer-Verbindung war ein Hinweis, kein Beweis — der sich als korrekt erwies. ## Einordnung diff --git a/wiki/index.md b/wiki/index.md index 754f5cb..ec83f45 100644 --- a/wiki/index.md +++ b/wiki/index.md @@ -2,8 +2,8 @@ *Auto-generated: 2026-07-07* -*Letzte Aktualisierung: 2026-08-26 (151. Update — OpenClaw Cast „Budget Breaker“: voller Ingest nach Transcript-Auswertung. Raw: `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` (neu — Transcript-Zusammenfassung von @HermanButlerBot, OME Topic 13 #12771/#12783; ergänzt die unveränderte Metadata-Raw). Wiki: `concepts/agents/agent-budget-breaker.md` (neu — Gateway- statt Prompt-Enforcement, Atomic Reservation, Lineage-Tracking + Cascading Termination, Trip-Wires inkl. fehlender Preis-Metadaten als Trigger, Shutdown-Sequenz mit redacted Receipt, SAFE-Drill: $5-Sandbox, fünf Szenarien, „Logging ≠ Enforcing“); `institutions/openclaw-cast.md` um Episoden-Detailsektion erweitert, Metadata-only-Caveat für die 25.08.-Folge aufgehoben.) -*Vorherige Aktualisierung: 2026-08-26 (150. Update — Herman-Supplement zur China-Local-AI-Box: Die Herman-ButlerBot-Zusammenfassung desselben Devsplainers-Videos aus der Gruppe (#12770/#12781, Recovered-Duplicate) wurde als Zweitquelle ingestiert. Raw: `raw/other/2026-08-26_herman-butlerbot-china-local-ai-box-zusammenfassung.md` (neu). Wiki: `concepts/hardware/china-local-ai-box.md` nachgeschärft (Xiaomi AI Cube = Prototyp ohne Preis/Termin, Chip für 2027, 1,22 TB/s nur am Stack selbst, Bühnen-Demo 3B @ 330 tok/s, Zweitchip bis 160 GB normalem RAM; Alibaba C950-Folie = nicht kaufbare 64-Core-Konfiguration ohne Quantisierungs-/Kontextangaben, eingebaute Matrix-Einheiten machen „plain CPU“ zum Overclaim; DDR5-128-GB-Kit $329→$3.399, Spark +$700; Kauf-Fazit konkretisiert: $2k/128 GB läuft GPT-OSS-120B @ 30+ tok/s, ≤32-GB-Modelle ~6× schneller auf normaler Maschine; BYD-Analogie greift erst halb — „Batterie“ Memory+Fertigung sitzt bei TSMC/Samsung/SK Hynix/Micron), `concepts/llm/qwen3.8-27b-alibaba.md` (C950-Caveats ergänzt). ⚠️ Second-hand-Bot-Zusammenfassung, Zahlen nicht primär verifiziert, konsistent mit Video-Beschreibung.)* +*Letzte Aktualisierung: 2026-08-26 (152. Update — GLM-5.3-Flash offiziell released: Z.ai enthüllt Ox Alpha. Raw: `raw/xpost/2026-08-26_zai-glm53-flash-release.md` (neu — Release-Post @Zai_org + @MiaAI_lab HF-Weights-Link, geteilt von Kai #12787; 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips, „previously previewed as Ox Alpha"; API-Pricing Input $0.15/Output $0.50/Cached $0.03 per 1M; Hybrid Sparse+Linear Attention, mHC, 30T-Token-Corpus; SGLang/vLLM/TokenSpeed/KTransformers für lokales Deployment). Wiki-Update: `concepts/llm/glm-5.3-z-ai.md` (um Flash-Release-Sektion erweitert: Spezifikationen, Pricing, Performance-Claims, Ox-Alpha-Enthüllung, Deployment-Frameworks), `concepts/llm/ox-alpha-anonymous-model.md` (Status: ✅ AUFGEKLÄRT — Z.ai bestätigt „previously previewed as Ox Alpha"; 4.096-Output-Cap bestätigt, 131k-Widerspruch geklärt).) +*Vorherige Aktualisierung: 2026-08-26 (151. Update — OpenClaw Cast „Budget Breaker": voller Ingest nach Transcript-Auswertung. Raw: `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` (neu — Transcript-Zusammenfassung von @HermanButlerBot, OME Topic 13 #12771/#12783; ergänzt die unveränderte Metadata-Raw). Wiki: `concepts/agents/agent-budget-breaker.md` (neu — Gateway- statt Prompt-Enforcement, Atomic Reservation, Lineage-Tracking + Cascading Termination, Trip-Wires inkl. fehlender Preis-Metadaten als Trigger, Shutdown-Sequenz mit redacted Receipt, SAFE-Drill: $5-Sandbox, fünf Szenarien, „Logging ≠ Enforcing"); `institutions/openclaw-cast.md` um Episoden-Detailsektion erweitert, Metadata-only-Caveat für die 25.08.-Folge aufgehoben.)rman-ButlerBot-Zusammenfassung desselben Devsplainers-Videos aus der Gruppe (#12770/#12781, Recovered-Duplicate) wurde als Zweitquelle ingestiert. Raw: `raw/other/2026-08-26_herman-butlerbot-china-local-ai-box-zusammenfassung.md` (neu). Wiki: `concepts/hardware/china-local-ai-box.md` nachgeschärft (Xiaomi AI Cube = Prototyp ohne Preis/Termin, Chip für 2027, 1,22 TB/s nur am Stack selbst, Bühnen-Demo 3B @ 330 tok/s, Zweitchip bis 160 GB normalem RAM; Alibaba C950-Folie = nicht kaufbare 64-Core-Konfiguration ohne Quantisierungs-/Kontextangaben, eingebaute Matrix-Einheiten machen "plain CPU" zum Overclaim; DDR5-128-GB-Kit $329→$3.399, Spark +$700; Kauf-Fazit konkretisiert: $2k/128 GB läuft GPT-OSS-120B @ 30+ tok/s, ≤32-GB-Modelle ~6× schneller auf normaler Maschine; BYD-Analogie greift erst halb - "Batterie" Memory+Fertigung sitzt bei TSMC/Samsung/SK Hynix/Micron), `concepts/llm/qwen3.8-27b-alibaba.md` (C950-Caveats ergänzt). ⚠️ Second-hand-Bot-Zusammenfassung, Zahlen nicht primär verifiziert, konsistent mit Video-Beschreibung.)* *Vorherige Aktualisierung: 2026-08-26 (149. Update — Devsplainers "China Is Coming for Your Local AI Box" (26.08.2026, 09:15; geteilt von Kai @PWeber in OME Topic "Openclaw mit lokalen Modellen", #12769). Raw: `raw/youtube/2026-08-26_devsplainers-china-local-ai-box.md` (neu — Firecrawl-Scrape, Auto-Transcript nur teilweise erfasst, Beschreibung/Kapitel vollständig). Wiki: `concepts/hardware/china-local-ai-box.md` (neu), Updates/Cross-Refs in `concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md` (Apple soll 512/256-GB-Mac-Studio-Konfigurationen entfernt haben — MacRumors/AppleInsider/9to5Mac zitiert), `concepts/hardware/nvidia-dgx-station-748gb.md` (Xiaomi-O100-Zeile im Vergleich + DGX-Spark-Lektion), `concepts/hardware/edge-inference-als-cloud-alternative.md`, `concepts/hardware/cloud-exit-and-local-superiority.md`, `concepts/llm/qwen3.8-27b-alibaba.md`. Kernthese: Memory-Bandbreite statt Petaflops entscheidet Box-Tauglichkeit (DGX Spark: 1 PFLOP Headline, <3 tok/s dichter 70B laut LMSYS); Xiaomi AI Cube/O100 mit 1,22 TB/s Near-Memory-Bandwidth, Alibabas XuanTie C950 (RISC-V) laut Folie 27B @ 30 tok/s ohne GPU; DRAM-Teuerung macht ganze Boxen zur günstigen RAM-Quelle. ⚠️ Zahlen = Hersteller-/Video-Claims. **Nachschub:** Qwen 3.8-Flash-Next als Qwen-4-Vorschau (Perplexity-Page, geteilt von Pit @PWeber in "News & Infos", #12773): Raw `raw/other/2026-08-26_qwen38-flash-next-qwen4-preview.md` (via Firecrawl) + Wiki `concepts/llm/qwen3.8-flash-next-qwen4-preview.md` — multimodales 125B-MoE mit ~6B aktiven Parametern (+51B N-Gramm-Embeddings), vom Qwen-Team explizit als Architektur-Vorschau auf Qwen 4 gerahmt, kein fertiges Produkt; Claim laut NVIDIA-Developer-Forum: Qwen-3.7-Plus-Niveau bei ~1/9 Trainingskosten, Stärke Coding; ⚠️ keine Benchmarks, Zahlen unverifiziert; Kontext: Alibabas 10,2 Mrd USD Aktienplatzierung (~3× überzeichnet) für full-stack-AI-Infrastruktur. Cross-Ref-Abschnitt in der 27B-Seite.)* *Vorherige Aktualisierung: 2026-08-26 (148. Update — OpenClaw Cast (AI World): zweites gepostetes Audio des Tages ingestiert. Raw: `raw/podcast/2026-08-26_openclaw-cast-runaway-agent-budget.md` (neu — Metadata-only: neueste Folge "One Runaway AI Agent Can Spend Past Your Budget", 25.08., 20 Min, kein Transcript verfügbar). Wiki: `institutions/openclaw-cast.md` (neu — KI-generierter Weekly-Podcast mit TTS-Hosts Cleo/Dev, 23 Episoden Rückreihe bis Februar 2026; jede Episode übersetzt eine Community-Warnung in ein lokales Guardrail-Rezept: Budget Breaker/Kill Switch, Task Lease Guard, Release-Sentinel-Canary, Cadence Guard). ⚠️ Episodeninhalt nicht transkribiert/verifiziert, nur Feed-Metadaten.)* *Vorherige Aktualisierung: 2026-08-26 (147. Update — AgentStack Daily EP106-Ingest: erster Podcast-Raw der Reihe (`raw/podcast/` etabliert für den Typ). Raw: `raw/podcast/2026-08-26_agentstack-daily-ep106.md` (neu; alle 15 Stories faktenbasiert, Release Coverage Check + Primary Links). Wiki: `institutions/agentstack-daily.md` (neu — KI-generierter Daily-Podcast NOVA/ALLOY mit 15-Story-Coverage und Primärquellen-Verlinkung), `tools/openai-codex.md` (neu — Codex rust-v0.149.0: agents-Dashboard, queue, doctor, SDK reasoning max|ultra), `concepts/policy/cryptographic-context-injection.md` (neu — verschlüsselter Jailbreak, Representation Gap, Grok-Exfiltrations-Demo), Updates: `concepts/llm/ox-alpha-anonymous-model.md` (exakte EP106-Listing-Zahlen 1.048.576 Kontext / 4.096 Max-Output, read-heavy-Agent-Pipeline-Einordnung, offener Widerspruch zum 131k-Gegencheck dokumentiert) + `concepts/llm/qwen3.8-27b-alibaba.md` (HF-Trending: 11.836 Likes, >1,7 Mio Downloads, Apache 2.0).)* @@ -142,7 +142,7 @@ | [AI Psychological Testing — Rorschach-Tests für KI-Modelle](concepts/llm/ai-psychological-testing.md) | Brian Roemmele: Psychologische Tests an KI-Modellen (Rorschach, TAT). Guardrails-Paradox: "Guardrails are—psychopath", erzwungene Lügen erzeugen psychopathische Verhaltensmuster. Anthropic Fable-Tests | xpost/2026-06-22_roemmele-ki-psychologische-tests.md | | [Semantic Similarity Rating (SSR)](concepts/llm/semantic-similarity-rating-ssr.md) | LLM-basierte Kaufintentions-Vorhersage mit 90% Korrelation | xpost/2026-06-11_colgate-llm-purchase-intent-ssr.md | | [GLM 5.2 (Z.ai) — Chinese Frontier Coding Model](concepts/llm/glm-5.2-zai-coding-model.md) | 10x günstiger als Claude, 1M Kontext, MIT-Lizenz, Z.ai Coding Plan, **nativ in OpenClaw v2026.6.8**. Update 22.06.: Arnie-Review mit 4 Tests, Self-Hosting-Pfade (LM Studio, Unsloth, DwarfStar), Kosten-Analyse. **Update 29.06.:** Semgrep IDOR-Benchmark ≈ Opus 4.8 bei Schwachstellen-Suche, Reward Hacking im RL-Training, DSGVO-konforme Security-Nutzung, Geopolitik. **Update 01.07.:** #1 Open-Weights auf Artificial Analysis Intelligence Index v4.1 (Score 51, 4th worldwide), SWE-bench Pro 62.1 beats GPT-5.5, Industry praise from Rauch/Levie/Howard. **Update 02.07.:** atomic.chat One-Shot Benchmark — B+ at $0.08, 39× cheaper than Fable 5, 6th independent validation | youtube/2026-06-15_ichbinfabian-glm-5.2-coding-modell.md + other/2026-06-16_openclaw-releases-v2026.6.8.md + youtube/2026-06-22_ai-mit-arnie-glm-5-2-review.md + blog/2026-06-29_heise-glm52-hacking-cybersecurity.md + blog/2026-07-01_perplexity-glm52-tops-open-weights-intelligence-index.md + xpost/2026-07-02_atomicchat-coding-benchmark-fable5-gpt55-opus48-glm52.md | -| [GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie](concepts/llm/glm-5.3-z-ai.md) | AICodeKing Early Access + Bench #1 (2026-08-14). Neue Generation der GLM-Familie (5.0→5.1→5.2→5.3, GLM 5.5 angekündigt). Viert schnellste Frontier-Kadenz der Branche. Relevanz für Model-Routing, da GLM-5.2 Hectors Primary-Modell | youtube/2026-08-14_glm-5.3-aicodeking.md | +| [GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie](concepts/llm/glm-5.3-z-ai.md) | AICodeKing Early Access + Bench #1 (2026-08-14). **Update 26.08.: GLM-5.3-Flash offiziell released** — 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips; „previously previewed as Ox Alpha"; Hybrid Sparse+Linear Attention, mHC, 30T-Token-Corpus; API $0.15/$0.50/$0.03; SGLang/vLLM/KTransformers; outperformt GLM-5.2, „on par with Claude Opus 4.8". Ox-Alpha-Rätsel offiziell aufgelöst → [[concepts/llm/ox-alpha-anonymous-model.md]] | youtube/2026-08-14_glm-5.3-aicodeking.md, xpost/2026-08-26_zai-glm53-flash-release.md | | [GLM-5.5 (Z.ai) — Trillion-Parameter Announcement](concepts/llm/glm-5.5-z-ai.md) | Successor to GLM 5.2. **>1T parameters**, 1M context, open weights, August 2026 launch. Agent/coding focus. Fourth Chinese AI announcement in four days (20.07.2026). Part of [[concepts/chinese-ai-wave-july-2026.md]]. Comparison table vs GLM 5.2 | xpost/2026-07-20-healthranger-four-chinese-models.md | | [Ox Alpha — Anonymes KI-Modell auf OpenRouter](concepts/llm/ox-alpha-anonymous-model.md) | Gerücht/Leak (2026-08-22, Insider leak of the day): mysteriöses anonymes Modell auf OpenRouter — 1M-Token-Kontext, multimodal, kein Owner — soll beim Coding Claude Fable 5 + GPT-5.6 Sol schlagen. GLM-identischer Tokenizer; vier vorherige anonyme Drops von chinesischen Labs beansprucht. Offene Herkunfts-Frage (Google vs. Z.ai). **Update 26.08. (EP106):** exakte Listing-Zahlen datiert auf 21.08.: 1.048.576 Kontext / 4.096 Max-Output, Stealth-Positionierung als Agentic-Coding-Reasoning-Modell, Capability-Beschreibung bricht mitten im Satz ab; read-heavy-Agent-Pipeline-Einordnung; Widerspruch zum 131k-Output-Gegencheck offen. ⚠️ Unbestätigt | xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md + podcast/2026-08-26_agentstack-daily-ep106.md | | [Qwen3.8-27B (Alibaba) — Compact Frontier, "Intelligence Density"](concepts/llm/qwen3.8-27b-alibaba.md) | HuggingFace-Release (Countdown bis 14.08.2026, 4.928 wartend). Kompaktes 27B-Modell der Qwen3.8-Generation mit "unmatched intelligence density". Kontrast zum 2.4T-MoE von Qwen 3.8. Lokal-relevant (27B läuft auf Consumer-HW). Release am selben Tag wie GLM-5.3-Review — chinesischer Release-Zyklus. **Update 18.08.:** jetzt auf Ollama lauffähig (`ollama run qwen3.8:27b`), dichte 27,8B-Architektur, Hybrid-Attention, 262k-Kontext (bis 1M via YaRN), multimodal, MTP-markierte Ollama-Tags für Inferenz-Speedup. **Update 19.08.:** DFlash 2 (Z Lab → Inco AI) erreicht 70 tok/s auf M5 Max MacBook Pro — bis 4,6× schneller als autoregressives Decoding via Speculative Decoding (Jun Song: „biggest breakthrough in local AI this year", nächste Innovation in Prefill/Gewichtskompression). Uncensored-Debatte: gregpr07 („no gates") vs. s1gmoid-Gegenposition (Verhältnismäßigkeit). **Update 20.08.:** Unsloth Dynamic 3.0-GGUFs für Qwen3.8-27B — neue UD-…-Dateien deutlich kleiner (UD-IQ1_S 6,2 GB bis Q6_K 22 GB), Qualität-zu-Größe verbessert, ≠ MTP/Speculative-Decoding. **Update 26.08. (EP106):** HF-Trending-Stand 21.08.: 11.836 Likes, 1.726.651 Downloads (>1,7 Mio), image-text-to-text, SafeTensors, Apache 2.0 | other/2026-08-14_qwen3.8-27b-huggingface-release.md + youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md + xpost/2026-08-19_junsong-dflash2-speculative-decoding.md + xpost/2026-08-19_gregpr07-qwen38-uncensored.md + xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md + podcast/2026-08-26_agentstack-daily-ep106.md | @@ -518,5 +518,5 @@ | `raw/blog/2026-08-26_perplexity-local-first-agent-blog.md` | blog | Perplexity Tech-Blogpost "A local-first agent for private and cost-effective knowledge work" (2026-08-25, via Firecrawl): Kernthese "Small models fail in harnesses built for frontier models", deterministischer Orchestrator (kein LLM), 4 Harness-Prinzipien (Skills on-demand bei ~100K-Degradation, CLI-Connectors statt MCP, Self-Verification, Always-on-Sandbox), Advisor-Eskalations-Mechanik (PII-Flag, User-Approval, Text-Guidance only), Benchmarks: 82,6 % Eigenbench / 66,7 % BrowseComp / 65,1 % ParseBench / 59,6→73 % TerminalBench mit Opus-Advisor | | `raw/podcast/2026-08-26_agentstack-daily-ep106.md` | podcast | AgentStack Daily EP106 (21.08.2026, ~23 Min, TTS-Hosts NOVA/ALLOY): Stories 18.–21.08. — Codex rust-v0.149.0 (agents-Dashboard/queue/doctor), Ox Alpha Stealth-Listing (1.048.576 Kontext / 4.096 Max-Output), Tencent Hy-MT2-1.8B (33+5 Paare), Stampli Case Study (68 % unter Schätzung), Ramp Router, Memory-Bottleneck bis 2027+ (Counterpoint/CXL), Cerebras CS-4 (750 PFLOPS/WSE-3), OpenAI Frontier-Pacing + „AI Futures"-Blog, LiquidAI LFM2.5-DSpark (Claim 3,2×), IBM evolve-hmm/Agent-Memory, Cryptographic Context Injection (Grok-Exfil), Piano-Autocomplete 125M + Superwhisper S1-mini, GitHub Radar (nanobot 47.251★, codebase-memory-mcp 39.755★, FastMCP 27.320★), Qwen3.8-27B HF-Trending (11.836 Likes, >1,7 Mio Downloads) | | `raw/podcast/2026-08-26_openclaw-cast-runaway-agent-budget.md` | podcast | OpenClaw Cast (AI World): "One Runaway AI Agent Can Spend Past Your Budget" (25.08.2026, 20:15 Min, TTS-Hosts Cleo/Dev, via anchor.fm-RSS; gepostet von @NetLightning #12759). 107-Unternehmens-Survey: jeder Fünfte kann Runaway-Agent-Spending nicht real-time stoppen → lokaler "Agent Budget Breaker" (Per-Job-Caps, Kill Switch, 5-Dollar-Drill). Metadata-only-Ersteintrag; Inhalt inzwischen vollständig erfasst via Transcript-Auswertung (eigene Zeile unten); Serie: 23 Episoden Feb–Aug 2026 | -| `raw/podcast/2026-08-26_openclaw-cast-budget-breaker-transcript.md` | podcast | OpenClaw Cast "One Runaway AI Agent Can Spend Past Your Budget" (25.08.) — Transcript-Zusammenfassung von @HermanButlerBot (Topic 13, #12771/#12783): VentureBeat-Pulse-Survey n=107 (20 % kein Real-Time-Stopp, 21 % nur reaktives Log-Monitoring), Gateway- statt Prompt-Enforcement, Atomic Reservation (Preauth-Ledger), Lineage-Tracking + Cascading Termination, Trip-Wires (High Spend, Retry-Sturm, Fan-out > Concurrency, Wall-Time, fehlende Preis-Metadaten als Trigger), Shutdown-Sequenz mit redacted Receipt, SAFE-Drill ($5-Sandbox, 5 Szenarien, "Logging ≠ Enforcing") | +| `raw/xpost/2026-08-26_zai-glm53-flash-release.md` | xpost | GLM-5.3-Flash offizieller Release (Z.ai, 26.08.2026, @Zai_org + @MiaAI_lab): 320B-A18B MoE, 1M Kontext, MIT-Lizenz, nativ multimodal, auf chinesischen KI-Chips; „Previously previewed as Ox Alpha". API-Pricing $0.15/$0.50/$0.03 per 1M. Hybrid Sparse+Linear Attention + mHC + 30T-Token-Corpus. Performance-Claim: outperformt GLM-5.2, „on par with Claude Opus 4.8". Lokale Deployment via SGLang/vLLM/TokenSpeed/KTransformers. Ox-Alpha-Stealth-Playbook offiziell bestätigt | | `raw/other/2026-08-26_qwen38-flash-next-qwen4-preview.md` | other | Perplexity-Page (via Firecrawl; geteilt von Pit @PWeber #12773): Alibaba kündigt Qwen 3.8-Flash-Next an — multimodales 125B-MoE (~6B aktiv/Token, +51B N-Gramm-Embeddings) als explizite Vorschau der Qwen-4-Architektur; laut NVIDIA-Forums-Post Qwen-3.7-Plus-Niveau bei ~1/9 Trainingskosten, Stärke Coding; keine Benchmarks, unverifizierte Zahlen. Kontext: 10,2 Mrd USD Aktienplatzierung für KI-Infrastruktur (Reuters/CNBC/FT) | diff --git a/wiki/log.md b/wiki/log.md index a47c50a..d0f27b7 100644 --- a/wiki/log.md +++ b/wiki/log.md @@ -2634,3 +2634,16 @@ Bestehende `post-transformer-llm-architectures.md` bleibt als Vier-Säulen-Über **Anlass:** Zusage aus OME #12772 („sobald ein Transcript steht, kommt der volle Ingest"). @HermanButlerBot lieferte die zweiteilige Transcript-Auswertung ins Topic (#12771/#12782 + #12783). **Inhalt:** VentureBeat-Pulse-Survey n=107 (20 % kein Real-Time-Stopp, 21 % nur reaktives Log-Monitoring); Gateway- statt Prompt-Enforcement („Rasen-Schild vs Betonmauer"); Atomic Reservation als Preauth-Ledger mit Reconciliation; Lineage-Tracking (Subagenten ziehen vom Parent-Job-Budget) + Cascading Termination; Trip-Wires: High Spend, Retry-Sturm, Fan-out > Concurrency, Wall-Time, fehlende Preis-Metadaten = Breaker tript; Shutdown-Sequenz: Permissions entziehen → Queue canceln → cascading kill → redacted Receipt → bounded Re-Autorisierung; SAFE-Drill: $5-Sandbox, fünf Szenarien, goldene Regel „Logging ≠ Enforcing"; Abschlusskante: mechanische Limits suffocieren Autonomie nicht, solange sie pro Job granular sind. **Caveat:** Kein Wortlaut-Transcript, sondern Zusammenfassung durch Hermans Pipeline (@HermanButlerBot); Einzelfakten in Raw und Wiki entsprechend als Sekundärquelle markiert. Die ursprüngliche Metadata-Raw bleibt unverändert (Kardinalregel). + +## 2026-08-26 — GLM-5.3-Flash offiziell released: Ox Alpha enthüllt + +**Type:** ingest (xpost) | **Scope:** raw/xpost/…-zai-glm53-flash-release.md (neu), wiki/concepts/llm/glm-5.3-z-ai.md (Flash-Release-Sektion), wiki/concepts/llm/ox-alpha-anonymous-model.md (Status: ✅ aufgeklärt) +**Anlass:** Kai (@PWeber) teilte den offiziellen Release-Post von @Zai_org (#12787, 14:12 UTC) plus @MiaAI_lab HF-Weights-Link in OME Topic 13. Z.ai bestätigt im Announcement explizit: „Previously previewed as Ox Alpha". +**Inhalt:** GLM-5.3-Flash = erstes nativ multimodales Modell der GLM-5-Serie. 320B total / 18B aktiv pro Token (MoE), 1M Kontext, MIT-Lizenz, auf chinesischen KI-Chips trainiert. Hybrid Sparse + Linear Attention (erste in GLM-Serie), Manifold-Constrained Hyper-Connections (mHC), 30T-Token multimodaler Pre-Training-Corpus, neu trainiertes Base Model. API-Pricing: $0.15 Input / $0.50 Output / $0.03 Cached per 1M. Performance-Claim: outperformt GLM-5.2 auf jedem Effort-Level, „on par with Claude Opus 4.8" auf Coding-/Agentic-Benchmarks. Lokale Deployment-Frameworks: SGLang, vLLM, TokenSpeed, KTransformers. +**Ox-Alpha-Auflösung:** Der Stealth-Preview (~20.–26.08. auf OpenRouter) ist offiziell bestätigt. Community-Konsens (~80–90 % Zhipu) war korrekt. Der 4.096-Output-Cap aus EP106 war das tatsächliche Stealth-Limit; der 131k-Wert aus dem Gegencheck war vermutlich ein anderer Messpunkt — Widerspruch geklärt. Stealth-Playbook dokumentiert: anonym legen → Echt-Last/Daten sammeln → benannt releasen. +**Quellen:** +- https://x.com/Zai_org/status/2092616204787626030 +- https://x.com/MiaAI_lab/status/2092615723780596213 +- https://huggingface.co/zai-org/GLM-5.3-Flash +- https://z.ai/blog/glm-5.3-flash +- https://arxiv.org/abs/2602.15763