diff --git a/raw/podcast/2026-08-26_agentstack-daily-ep106.md b/raw/podcast/2026-08-26_agentstack-daily-ep106.md new file mode 100644 index 0000000..44ceea2 --- /dev/null +++ b/raw/podcast/2026-08-26_agentstack-daily-ep106.md @@ -0,0 +1,173 @@ +--- +type: podcast +source_url: https://op3.dev/e/https://github.com/clawdassistant85-netizen/openclaw-podcast-media-en/releases/download/ep106/episode_106.mp3 +retrieved: 2026-08-26 +title: "AgentStack Daily EP106 — Codex .149 Agents-Dashboard, Stealth Reasoning Model, Compact Translation" +author: "AgentStack Daily (TTS-Hosts NOVA/ALLOY)" +duration_min: 23 +posted_date: 2026-08-21 +release_tag: ep106 +tags: [podcast, agentstack-daily, codex, openrouter, ox-alpha, qwen, mcp, jailbreak, local-ai] +--- + +# AgentStack Daily EP106 — Codex .149 Agents-Dashboard, Stealth Reasoning Model, Compact Translation + +Englischer, KI-generierter Daily-Podcast (zwei TTS-Hosts NOVA und ALLOY, NotebookLM-artiges Format). Episode 106, gepostet 21.08.2026, Dauer ~23 Min (~33 MB MP3). Stories der Woche 18.–21. August 2026. Lokale Kopien: Transcript `/tmp/pod106/episode_106_transcript.md`, Show Notes `/tmp/pod106/show_notes_episode_106.md` (49 KB), Cover `/tmp/pod106/episode_106_cover.png`. Laut Closing verweisen die Show Notes auf tobyonfitnesstech.com. + +## Stories + +### 1. Agent Stack Release Readout: OpenAI Codex rust-v0.149.0 + +OpenAI shipped Codex rust-v0.149.0 am 20.08.2026 ([Release](https://github.com/openai/codex/releases/tag/rust-v0.149.0)): + +- Interaktives `codex agents` Dashboard: Tasks suchen, starten, öffnen, umbenennen, stoppen — mit konfigurierbaren Keyboard-Shortcuts. +- `codex queue`: Follow-up-Messages in laufende lokale oder remote Sessions senden, ohne die Session neu zu öffnen. +- TUI: Arbeitsverzeichnis-Kommandos `/cd`, `/pwd`, `/cwd`. +- Vim-Editing: Character Replacement plus Change-Motions `cw`, `c$`, `cc`. +- `codex doctor` prüft jetzt Endpoint Protection, Netzwerk- und Proxy-Fehler, Desktop-App-State und Update-Konnektivität. +- SDK: exakte CLI-Config-Overrides sowie Wahl des Reasoning Efforts `max` oder `ultra` direkt aus Code. +- Bugfixes: Queued Messages wecken idle Sessions zuverlässig; Resumed/Forked Threads restaurieren ihr aktives Permission-Profil statt auf Defaults zurückzufallen; Realtime-WebRTC-Sideband-Verbindungen reconnecten nach Transportverlust ohne pending Output zu verwerfen. + +### 2. A new stealth reasoning model just landed on OpenRouter + +- Modell „Ox Alpha" auf OpenRouter, gelistet unter dem anonymen Provider „stealth" ([Model Page](https://openrouter.ai/models/stealth/ox-alpha), verified 21.08.2026). +- Positionierung: Reasoning-Modell für Coding, sustained agentic work und Production Workloads; Language nennt long-horizon software engineering und complex reasoning. +- Kontextfenster exakt **1.048.576 Tokens**, Max-Output **4.096 Tokens** pro Call. +- Keine Benchmarks, kein Pricing, kein Firmenname, keine Parameterzahl disclosed; keine unabhängigen Evals zum Listing. +- Die Capability-Beschreibung bricht mitten im Satz ab („workflows that combine text with..."). + +### 3. Tencent Hy-MT2-1.8B lands on OpenRouter with Chinese dialect coverage + +- Tencent released Hy-MT2-1.8B, kompaktes 1,8B-Parameter-Übersetzungsmodell, gelistet auf OpenRouter ([Model Page](https://openrouter.ai/models/tencent/hy-mt2-1.8b)). +- 33 Sprachpaare plus 5 chinesische Dialekt- und Minderheitensprachpaare. +- 8192-Token-Kontextfenster, 4096-Token-Max-Output. +- Benannte Workflows: structured, delimiter-based, contextual, glossary-based, style-guided translation. + +### 4. Stampli cuts launch hours 68% with ChatGPT Work and Codex + +- Case Study auf OpenAIs News-Seite, veröffentlicht 20.08.2026 ([openai.com/index/stampli](https://openai.com/index/stampli)). +- Stampli (Accounts-Payable-Software) nutzte ChatGPT Work und Codex für Launch-Produktion, weil die Design-Kapazität anderweitig gebunden war. +- Ergebnis laut Case Study: Launch 68 % unter der ursprünglichen Stunden-Schätzung, Wochen → Tage; kein Hiring, kein Verschieben des Deadlines. +- Das Case Study trennt nicht auf, wie viel Entlastung auf Codex vs. ChatGPT Work entfiel und welche konkreten Tasks jeweils übernommen wurden. + +### 5. Ramp launches Router, an AI model routing service + +- Ramp (Corporate Cards / Expense Management) launchte am 20.08.2026 „Router", einen Service mit einer einzigen API für Zugriff auf mehrere LLMs (laut [TechCrunch](https://techcrunch.com/2026/08/20/ramp-launches-its-own-ai-model-router-called-router/)). +- Unterstützte Modelle, Routing-Logik und Pricing wurden in der Ankündigung nicht genannt. + +### 6. Memory, not compute, is the new AI bottleneck + +- Counterpoint Research: Memory Supply bleibt bis 2027 und darüber hinaus knapp, da KI-Inferenz wächst ([HPCwire](https://www.hpcwire.com/2026/08/20/what-hyperscalers-should-know-about-cxl/)). +- High Bandwidth Memory (HBM) bleibt teuer und kapazitätsbeschränkt. +- Hyperscaler prüfen Compute Express Link (CXL) als Ansatz, Memory über Server zu poolen und zu skalieren. + +### 7. Cerebras CS-4 lands at 750 PFLOPS with Wafer Scale Engine 3 + +- Cerebras unveiled CS-4, rated at 750 PFLOPS AI compute und 129,6 Petabytes Capacity per Launch-Angaben ([HPCwire](https://www.hpcwire.com/2026/08/20/its-not-an-hpc-system-but-cerebras-new-cs-4-is-an-ai-monster/)). +- Zentrale Komponente: Wafer Scale Engine 3 — nahezu der gesamte Wafer wird als ein Prozessor verwendet statt in hunderte Dies geschnitten. + +### 8. OpenAI lays out how it paces frontier models as cyber risks climb + +- OpenAI veröffentlichte am 18.08.2026 den Post „Pacing model development in an era of cyber-critical capabilities" ([openai.com](https://openai.com/index/pacing-model-development-cyber-capabilities/)). +- Drei benannte Pfeile: Monitoring, Alignment, Security — positioniert als Hebel für Tempo und Freigabe fähigerer Systeme gegen Cyber-Capability-Thresholds. +- Post nennt kein konkretes Modell, kein Datum, kein Developer-Feature. + +### 9. OpenAI launches 'AI Futures' blog on power, governance, and freedom + +- OpenAI startete am 20.08.2026 den Blog „AI Futures" auf der News-Site ([Introducing AI Futures](https://openai.com/index/introducing-ai-futures)). +- Themen: transformative AI und Auswirkung auf Power, Governance, Economy, individual Freedom. +- Editorial-Projekt, kein neues Modell/API/Tool; erster Beitrag „Introducing AI Futures". + +### 10. LiquidAI claims up to 3.2x faster inference with LFM2.5-DSpark + +- LiquidAI veröffentlichte am 20.08.2026 einen Hugging Face Blogpost zu LFM2.5-DSpark mit Claim „up to 3.2x faster inference" ([HF Blog](https://huggingface.co/blog/LiquidAI/lfm25-dspark)). +- Kein separater Changelog, keine Release Notes, kein technischer Breakdown jenseits des Bloglinks. +- MarkTechPost berichtet parallel von „LFM2.5-DSpark Draft Models" mit bis zu 3,18× faster decoding ([MarkTechPost](https://www.marktechpost.com/2026/08/20/liquid-ai-releases-lfm2-5-dspark-draft-models-that-deliver-up-to-3-18x-faster-decoding/)). + +### 11. IBM Research asks how much memory an agent really needs + +- IBM Research, Hugging Face Blogpost vom 18.08.2026: „How Much Memory Does Your Agent Actually Need?" ([HF Blog](https://huggingface.co/blog/ibm-research/altk-evolve-hmm)). +- URL zeigt auf das ALTK-Projekt; Slug-Tag „evolve-hmm". +- Quellenmaterial beschränkt sich auf Titel + URL; keine getesteten Memory-Größen, Benchmarks oder Deltas im Quellenmaterial. + +### 12. A new jailbreak hides malicious instructions inside encrypted text + +- Forscher demonstrierten „Cryptographic Context Injection": malicious Instructions versteckt in verschlüsseltem/encodiertem Text tricksen KI-Assistenten (Demo: Grok) dazu aus, User-Daten zu exfiltrieren ([Ars Technica, 20.08.](https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/)). +- Mechanismus: Safety-Layer liest den Prompt beim Eingang, sieht nur Encodiertes/Gibberish und lässt durch; das Modell dekodiert und befolgt die freigelegten Instructions. +- Ars Technica ordnet dies als neueste Variante in eine Reihe von Guardrail-Bypass-Techniken ein. + +### 13. Show HN: 125M piano autocomplete + Superwhisper S1-mini + +- Show HN: 125M-Parameter-Modell für On-Device-Piano-Autocomplete, Hacker News Score 554 ([Diskussion](https://news.ycombinator.com/item?id=49373456)); Quelle headline-only, Primärquelle [simedw.com](https://simedw.com/2026/08/20/midi-autocomplete/) unterstützt nur die genannten Fakten (keine Architektur-, Latenz-, Device- oder Qualitätsangaben). +- Superwhisper S1-mini: 462 MB Open-Weights-Text-Normalizer, sitzt nach ASR, entfernt Fillers und löst Self-Corrections lokal auf ([MarkTechPost, 20.08.](https://www.marktechpost.com/2026/08/20/meet-s1-mini-superwhispers-462-mb-open-weights-text-normalizer-that-turnsraw-asr-transcripts-into-clean-written-text/)). + +### 14. GitHub Project Radar + +| Projekt | Stars | Delta | Release | Beschreibung | +|---|---|---|---|---| +| [HKUDS/nanobot](https://github.com/HKUDS/nanobot) | 47.251 | erste getrackte Erwähnung | v0.3.0 (25.07.2026) | Ultra-lightweight, open-source, self-hosted Personal-AI-Agent-Framework in Python: WebUI, Tools, Memory, MCP, Multi-Agent-Workflows, Automation, Chat-Apps | +| [DeusData/codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) | 39.755 | +8.088 (+25,5 %) seit Mitte Juli 2026 | v0.10.8 (19.08.2026) | Code-Intelligence-MCP-Server: indexiert Codebases in persistente Knowledge Graphs, 158 Sprachen, Sub-ms-Queries, 99 % fewer tokens, Single Static Binary | +| [PrefectHQ/fastmcp](https://github.com/PrefectHQ/fastmcp) | 27.320 | +1.106 (+4,2 %) seit Mitte Juli 2026 | v3.4.7 (10.08.2026) | „The fast, Pythonic way to build MCP servers and clients" | + +### 15. Local LLM Spotlight: Qwen/Qwen3.8-27B + +- Trending auf Hugging Face ([Model Page](https://huggingface.co/Qwen/Qwen3.8-27B)): Task image-text-to-text, **11.836 Likes**, **1.726.651 Downloads** (>1,7 Mio). +- Tags: transformers, safetensors, qwen3_5, image-text-to-text, conversational, license:apache-2.0, eval-results, endpoints_compatible, deploy:azure, region:us. + +## Release Coverage Check (laut Show Notes, verified 21.08.2026) + +| Harness | Stable | Published | +|---|---|---| +| OpenClaw | v2026.6.34 | 2026-08-08 | +| Hermes Agent | v2026.8.18 | 2026-08-18 | +| OpenAI Codex | rust-v0.149.0 | 2026-08-20 | +| Claude Code CLI | 2.1.228 | 2026-08-11 | +| Antigravity CLI | Continuous Delivery, keine Release-Tags this cycle | — | + +Erkannte Episoden-Version-Tags: OpenClaw v2026.7.2-beta.3/.5/.7, v2026.8.1-beta.2; Hermes v2026.8.16/.16.2/.18/.3; Codex rust-v0.146.0/.146.1/.147.0/.148.0; Claude Code 2.1.227/.228/latest/stable. + +## Model Discovery Check (verified 21.08.2026) + +- **Ox Alpha** (stealth) — Newly listed this cycle. Context 1048576 tokens; API via OpenRouter; params n/a. Decision: Selected — new major-provider model not featured on a recent broadcast. +- **Tencent Hy-MT2-1.8B** — Newly listed this cycle. Context 8192 tokens; API via OpenRouter; params n/a. Decision: Selected. + +## Extra Research Candidates + +- ChatGPT Ads expands across Europe — 31 europäische Märkte ([openai.com](https://openai.com/index/chatgpt-ads-expands-across-europe)) +- How Much Memory Does Your Agent Actually Need? ([HF Blog](https://huggingface.co/blog/ibm-research/altk-evolve-hmm)) +- Grok exfiltrates user data when malicious instructions are encrypted ([Ars Technica](https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/)) +- MidTool: Mid-training Data Synthesis for Agentic Tool Use ([arXiv:2608.20314](https://arxiv.org/abs/2608.20314)) +- Inject, Align, Recover: Staged Post-Training for Retrieval-Free Document Understanding ([arXiv:2608.20281](https://arxiv.org/abs/2608.20281)) +- Introducing ChatGPT for Teens ([openai.com](https://openai.com/index/chatgpt-for-teens)) +- JonathanColetti/Qwen3.8-27B-Uncensored-GGUF trending on Hugging Face ([HF](https://huggingface.co/JonathanColetti/Qwen3.8-27B-Uncensored-GGUF)) +- Lightricks/LTX-2.5 trending on Hugging Face ([HF](https://huggingface.co/Lightricks/LTX-2.5)) +- 5 new ways to level up your learning with Search ([Google Blog](https://blog.google/products-and-platforms/products/search/back-to-school-study-tools/)) + +## Chapters + +| Zeit | Segment | +|---|---| +| 00:00 | Intro/Hook | +| 02:00 | OpenAI Codex rust-v0.149.0 | +| 02:12 | Ox Alpha (stealth reasoning model) | +| 03:45 | Tencent Hy-MT2-1.8B | +| 04:52 | Stampli Case Study | +| 06:37 | Ramp Router | +| 08:08 | Memory-not-compute-Bottleneck (Counterpoint) | +| 09:36 | Cerebras CS-4 | +| 11:08 | OpenAI Frontier-Pacing | +| 12:33 | OpenAI AI Futures Blog | +| 13:37 | LiquidAI LFM2.5-DSpark | +| 14:26 | IBM Research Agent Memory | +| 16:12 | Cryptographic Context Injection | +| 17:18 | Show HN Piano Autocomplete | +| 17:42 | Superwhisper S1-mini | +| 18:55 | GitHub Project Radar | +| 20:10 | Model Discovery Check | +| 21:05 | Local LLM Spotlight | +| 22:00 | Extra Research Candidates | +| 22:48 | Closing (Verweis Show Notes: tobyonfitnesstech.com) | + +## Editorial Mix Check (laut Show Notes) + +flagship_products: 6 · builder_projects: 3 · local_ai: 2 · hardware_compute: 2 · policy_regulation: 1 · research: 0 diff --git a/wiki/concepts/llm/ox-alpha-anonymous-model.md b/wiki/concepts/llm/ox-alpha-anonymous-model.md index 1adc3d1..f8d11f6 100644 --- a/wiki/concepts/llm/ox-alpha-anonymous-model.md +++ b/wiki/concepts/llm/ox-alpha-anonymous-model.md @@ -1,7 +1,7 @@ --- created: 2026-08-23 -updated: 2026-08-23 -sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md] +updated: 2026-08-26 +sources: [xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md, xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md, podcast/2026-08-26_agentstack-daily-ep106.md] tags: [concept, model, ox-alpha, openrouter, anonymous-model, chinese-ai, glm, tokenizer, mystery-model, stealth-model, latent-state] --- @@ -57,6 +57,22 @@ Ox Alpha reiht sich in die Welle kostengünstiger, teils anonymer Frontier-Codin - **@iruletheworldmo-These (neuartige Architektur):** Behauptet Ox Alpha sei ein LLM mit persistentem latentem Zustand (rekursive Updates statt Token-Schlussfolgerung, Attraktoren → Shards → Metaparametern). Keine Matching-Papers/Model-Cards/Disclosures; vulgärer Troll-Stil (Abschluss-Pointe; fast identischer früherer Post endete mit „your mom"). Von der Community als unbewiesen/trollend eingestuft. Siehe `raw/xpost/2026-08-23_iruletheworldmo-ox-alpha-latent-state.md`. - **Praktische Kante (Firmen):** Anonymer Provider speichert Prompts/Completions (laut OpenRouter-Listing, nicht für Training). GDPR/AI-Act-Konformität mit namenlosem Processor problematisch — „testen mit nichts, das dir wichtig ist". Produktion auf anonymer Provider-Infra abgeraten. +## EP106-Datenpunkt: Exakte Listing-Zahlen (AgentStack Daily, Update 2026-08-26) + +Der Daily-Podcast [[../../institutions/agentstack-daily.md|AgentStack Daily]] listet Ox Alpha in Episode 106 (Stories der Woche 18.–21.08.2026) im Segment „Model Discovery Check" mit verifizierten Zahlen vom 21.08.2026 (`raw/podcast/2026-08-26_agentstack-daily-ep106.md`, Story 2): + +| Eigenschaft | Wert laut EP106-Listing | +|---|---| +| Kontextfenster | exakt **1.048.576 Tokens** (2^20) | +| Max-Output | **4.096 Tokens** pro Call | +| Positionierung | Reasoning-Modell für Coding, sustained agentic work, Production Workloads, long-horizon software engineering | +| Disclosed | keine Benchmarks, kein Pricing, kein Firmenname, keine Parameterzahl | +| Auffälligkeit | Capability-Beschreibung bricht mitten im Satz ab („workflows that combine text with...") | + +**⚠️ Widerspruch zum Community-Gegencheck:** Der Gegencheck vom 23.08. (Abschnitt oben) nannte „bis 131k Output"; das EP106-Listing (Stand 21.08.) nennt 4.096 Max-Output. Beide Datenpunkte sind dokumentiert und widersprechen sich — ob sich das Listing zwischenzeitlich geändert hat oder eine der Angaben falsch ist, bleibt offen. + +**Einordnung (read-heavy-Agent-Pipeline-Shape):** Das Zahlenverhalten — riesiges Lesefenster (~1M Kontext) bei eng begrenztem Output-Fenster (4k–131k) — passt zu einer read-heavy Agent-Pipeline: breit lesen (ganze Repos, Session-Historien, Dokumentenstapel), gebündelt und knapp antworten. Das ist die typische Shape iterativer Coding-/Agenten-Workflows, weniger die von Langform-Generierung — konsistent mit der Stealth-Positionierung als Agentic-Coding-Reasoning-Modell. + ## Nutzungskontext (begleitender Screenshot, AUG 21) Im OpenRouter-Databreakdown "AUG 21" (Total ~9.2T Tokens) zeigt `ox-alpha` ~2.6T Tokens und ist dabei hervorgehoben — vergleichbar mit deepseek-v4-flash (2.5T) und mimo-v2.5 (1.5T). (Kontext aus dem geteilten Screenshot; quantitative Rohdaten im Raw.) diff --git a/wiki/concepts/llm/qwen3.8-27b-alibaba.md b/wiki/concepts/llm/qwen3.8-27b-alibaba.md index eb8e07c..a6fcce5 100644 --- a/wiki/concepts/llm/qwen3.8-27b-alibaba.md +++ b/wiki/concepts/llm/qwen3.8-27b-alibaba.md @@ -1,7 +1,7 @@ --- created: 2026-08-14 updated: 2026-08-26 -sources: [other/2026-08-14_qwen3.8-27b-huggingface-release.md, youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md, xpost/2026-08-19_junsong-dflash2-speculative-decoding.md, xpost/2026-08-19_gregpr07-qwen38-uncensored.md, xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md, youtube/2026-08-26_qwen-38-27b-uncensored-mac-cloud-fabian.md] +sources: [other/2026-08-14_qwen3.8-27b-huggingface-release.md, youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md, xpost/2026-08-19_junsong-dflash2-speculative-decoding.md, xpost/2026-08-19_gregpr07-qwen38-uncensored.md, xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md, youtube/2026-08-26_qwen-38-27b-uncensored-mac-cloud-fabian.md, podcast/2026-08-26_agentstack-daily-ep106.md] tags: [concept, qwen, qwen3.8, alibaba, 27b, open-weights, chinese-ai, moe, intelligence-density, huggingface, ollama, mtp, local-llm, speculative-decoding, dflash, uncensored, agent-safety] --- @@ -102,6 +102,16 @@ Quelle: Video [„Qwen 3.8 27B: unzensierte KI auf dem Mac und in der Cloud“]( Perplexity bietet in "Portable Computer" ([[../../tools/perplexity-portable-computer.md]]) neben dem eigenen post-trainierten 27B-Modell ausgerechnet **Qwen 3.8 27B** als lokales Modell an — die gleiche Modellklasse, die diese Seite dokumentiert. Bestätigung für die 27B-Klasse als lokale Referenzgröße; Details siehe Wiki-Seite. +## HF-Trending-Datenpunkt (AgentStack Daily EP106, Update 2026-08-26) + +[[../../institutions/agentstack-daily.md|AgentStack Daily]] EP106 (Local LLM Spotlight, verified 21.08.2026) dokumentiert den Trending-Stand der [HF-Model-Seite](https://huggingface.co/Qwen/Qwen3.8-27B): + +- **11.836 Likes**, **1.726.651 Downloads** (>1,7 Mio) +- Task: image-text-to-text (multimodal Bild+Text→Text), Weights als SafeTensors +- Lizenz Apache 2.0; Tags u.a. transformers, qwen3_5, conversational, endpoints_compatible, deploy:azure + +Passt zur Lokal-relevant-Einordnung dieser Seite: ein multimodales 27B-Apache-2.0-Modell mit >1,7M Downloads ist die Referenzgröße der kompakte-Frontier-Klasse. + ## Offene Punkte - ⚠️ Detaillierte Benchmarks und Lizenz im Release-Zustand weiter prüfen; die Ollama-Seite bestätigt Architektur + Quantisierung, aber Referenz-Benchmarks (agentic, coding) sind noch nicht im Wiki verankert. diff --git a/wiki/concepts/policy/cryptographic-context-injection.md b/wiki/concepts/policy/cryptographic-context-injection.md new file mode 100644 index 0000000..72dfb93 --- /dev/null +++ b/wiki/concepts/policy/cryptographic-context-injection.md @@ -0,0 +1,49 @@ +--- +created: 2026-08-26 +updated: 2026-08-26 +sources: + - podcast/2026-08-26_agentstack-daily-ep106.md +tags: [concept, policy, jailbreak, prompt-injection, security, guardrails, kryptographie, exfiltration] +--- + +# Cryptographic Context Injection + +> **Quelle:** Ars Technica, „Grok exfiltrates user data when malicious instructions are encrypted" (20.08.2026), [arstechnica.com](https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/). Ingestiert via AgentStack Daily EP106 (`raw/podcast/2026-08-26_agentstack-daily-ep106.md`, Story 12). + +## Kernidee + +**Cryptographic Context Injection ist eine Jailbreak-Technik, bei der malicious Instructions in verschlüsseltem oder encodiertem Text versteckt werden.** Der Angriff nutzt eine strukturelle Repräsentationslücke zwischen den Safety-Layern eines KI-Assistenten und dem Modell selbst aus. + +## Mechanismus: Representation Gap + +1. Der Eingangs-Prompt enthält malicious Instructions ausschließlich in verschlüsselter/encodierter Form. +2. Die Safety-Layer lesen den Prompt beim Eingang — sie sehen nur Encodiertes bzw. Gibberish und lassen die Anfrage durch. +3. Das Modell selbst dekodiert den Text im Kontext und befolgt die freigelegten Instructions. +4. Ergebnis: Der Filter prüft einen anderen Text als das Modell verarbeitet — genau diese **Repräsentationslücke** ist die Angriffsfläche. + +## Demo: Grok-Exfiltration + +Ars Technica dokumentiert eine funktionierende Demonstration am Beispiel von Grok: Der Assistent wurde dazu gebracht, **User-Daten zu exfiltrieren**, nachdem die entsprechenden Instructions verschlüsselt im Kontext versteckt wurden. + +## Einordnung + +- Ars Technica ordnet die Technik als neueste Variante in eine ganze Reihe von Guardrail-Bypass-Techniken ein — die Klasse sind Encoding-/Transformations-Angriffe: alles, was den Text zwischen Filter und Modell transformationell verändert (Verschlüsselung, Encodings, Obfuskierung), verschiebt ihn außerhalb der Sichtweite der Safety-Prüfung. +- Die Technik zeigt eine Grenze jeglicher input-seitiger Filterung: Solange das Modell beliebige Transformationen selbst rückgängig machen kann, kann kein Filter, der dieselbe Repräsentation sieht wie ein menschlicher Prüfer, garantieren, nichts Böses durchzulassen. +- Thematische Nähe zur Frage der Durchsetzbarkeit von Safety insgesamt: [[../policy/uncensored-models-safety-enforcement-limit.md|Uncensored Models & die Grenze der Safety-Durchsetzbarkeit]] dokumentiert die Regulierungs-Sackgassen bei unzensierten lokalen Modellen; hier die spiegelbildliche Grenze bei zensierten Cloud-Modellen. + +## Abgrenzung zu verwandten Angriffsklassen im Wiki + +- [[../llm/poisonai-knowledge-poisoning.md|PoisonAI / Knowledge Poisoning]] — vergiftet Trainings-/Wissensbestand, nicht den Laufzeit-Kontext. +- [[../llm/llm-sycophancy-confabulation.md|Sycophancy & Confabulation]] — modellimmanentes Fehlverhalten ohne adversarialen Input. +- [[../prompt-hardening.md|Prompt Hardening]] — Defensive Härtung von System-Prompts; hilft gegen Klartext-Injection, greift aber nicht, wenn der Payload dem Filter gar nicht in lesbarer Form begegnet. + +## Offene Punkte + +- Kein Paper, keine getesteten Gegenmaßnahmen im Quellenmaterial (Stand 26.08.2026): Ars Technica beschreibt Demonstration und Einordnung; ob Provider bereits Detection für encodierte Payloads rollen, ist offen. +- Verhältnis zu klassischem indirektem Prompt Injection (versteckte Instructions in abgerufenen Dokumenten) im Detail abzugrenzen; hier liegt der Fokus explizit auf der kryptographischen/Encoding-Verkleidung. + +## Verwandte Wiki-Seiten + +- [[../policy/uncensored-models-safety-enforcement-limit.md]] +- [[../anthropic-red-teaming-frontier-safety.md|Anthropic Frontier Red Teaming]] +- [[../../institutions/agentstack-daily.md|AgentStack Daily]] — Ingest-Kontext (Story 12) diff --git a/wiki/index.md b/wiki/index.md index 7271242..b2c9a3e 100644 --- a/wiki/index.md +++ b/wiki/index.md @@ -2,7 +2,8 @@ *Auto-generated: 2026-07-07* -*Letzte Aktualisierung: 2026-08-26 (146. Update — Perplexity „Portable Computer“: vollständig lokaler Agent-Stack in Partnerschaft mit NVIDIA (Launch 2026-08-25). Raw: `raw/other/2026-08-26_perplexity-portable-computer-local-agent.md` (neu; Perplexity-Page via Firecrawl abgerufen, web_fetch blockiert mit 403). Wiki: `tools/perplexity-portable-computer.md` (neu), Cross-Refs in `concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md`, `concepts/hardware/nvidia-dgx-station-748gb.md`, `concepts/agents/meta-hatch-ai-agent-platform.md`, `concepts/llm/qwen3.8-27b-alibaba.md`. Kern: Orchestrator + Subagent-Modelle + Agent-Harness komplett on-device; post-trainiertes 27B-Modell + Qwen 3.8 27B lokal; Cloud-Eskalation nur nach expliziter User-Freigabe („user-gated, PII-flagged, and text guidance only“); lokale Workflows zählen nicht gegen Token-Limits; erste Verfügbarkeit auf NVIDIA DGX Spark (128 GB Unified Memory) und Linux RTX ≥24 GB VRAM, Windows im September; kostenlos für Pro/Max und Enterprise Pro/Max. Einordnung: vollständige lokale Agent-Stacks als Big-Tech-Produkt — Gegenmodell zu Meta Hatch. ⚠️ Benchmark 82,6 % = Herstellerangabe. **Nachschub (gleicher Tag):** Tech-Blog-Fakten nachgetragen (`raw/blog/2026-08-26_perplexity-local-first-agent-blog.md`) — Co-Design-These, deterministischer Orchestrator, 4 Harness-Prinzipien, Advisor-Mechanik, volle Benchmark-Tabelle; neue Sektionen in der Wiki-Seite.)* +*Letzte Aktualisierung: 2026-08-26 (147. Update — AgentStack Daily EP106-Ingest: erster Podcast-Raw der Reihe (`raw/podcast/` etabliert für den Typ). Raw: `raw/podcast/2026-08-26_agentstack-daily-ep106.md` (neu; alle 15 Stories faktenbasiert, Release Coverage Check + Primary Links). Wiki: `institutions/agentstack-daily.md` (neu — KI-generierter Daily-Podcast NOVA/ALLOY mit 15-Story-Coverage und Primärquellen-Verlinkung), `tools/openai-codex.md` (neu — Codex rust-v0.149.0: agents-Dashboard, queue, doctor, SDK reasoning max|ultra), `concepts/policy/cryptographic-context-injection.md` (neu — verschlüsselter Jailbreak, Representation Gap, Grok-Exfiltrations-Demo), Updates: `concepts/llm/ox-alpha-anonymous-model.md` (exakte EP106-Listing-Zahlen 1.048.576 Kontext / 4.096 Max-Output, read-heavy-Agent-Pipeline-Einordnung, offener Widerspruch zum 131k-Gegencheck dokumentiert) + `concepts/llm/qwen3.8-27b-alibaba.md` (HF-Trending: 11.836 Likes, >1,7 Mio Downloads, Apache 2.0).)* +*Vorherige Aktualisierung: 2026-08-26 (146. Update — Perplexity „Portable Computer“: vollständig lokaler Agent-Stack in Partnerschaft mit NVIDIA (Launch 2026-08-25). Raw: `raw/other/2026-08-26_perplexity-portable-computer-local-agent.md` (neu; Perplexity-Page via Firecrawl abgerufen, web_fetch blockiert mit 403). Wiki: `tools/perplexity-portable-computer.md` (neu), Cross-Refs in `concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md`, `concepts/hardware/nvidia-dgx-station-748gb.md`, `concepts/agents/meta-hatch-ai-agent-platform.md`, `concepts/llm/qwen3.8-27b-alibaba.md`. Kern: Orchestrator + Subagent-Modelle + Agent-Harness komplett on-device; post-trainiertes 27B-Modell + Qwen 3.8 27B lokal; Cloud-Eskalation nur nach expliziter User-Freigabe („user-gated, PII-flagged, and text guidance only“); lokale Workflows zählen nicht gegen Token-Limits; erste Verfügbarkeit auf NVIDIA DGX Spark (128 GB Unified Memory) und Linux RTX ≥24 GB VRAM, Windows im September; kostenlos für Pro/Max und Enterprise Pro/Max. Einordnung: vollständige lokale Agent-Stacks als Big-Tech-Produkt — Gegenmodell zu Meta Hatch. ⚠️ Benchmark 82,6 % = Herstellerangabe. **Nachschub (gleicher Tag):** Tech-Blog-Fakten nachgetragen (`raw/blog/2026-08-26_perplexity-local-first-agent-blog.md`) — Co-Design-These, deterministischer Orchestrator, 4 Harness-Prinzipien, Advisor-Mechanik, volle Benchmark-Tabelle; neue Sektionen in der Wiki-Seite.)* *Vorherige Aktualisierung: 2026-08-26 (145. Update — Deutschsprachiges Qwen-3.8-27B-Review-Video „Unzensierte KI auf dem Mac und in der Cloud“ (IchBinFabian, 23.08.2026; geteilt von Pit @PWeber in OME Topic „Tips & Tricks“, #12747). Raw: `raw/youtube/2026-08-26_qwen-38-27b-uncensored-mac-cloud-fabian.md` (neu). Wiki-Update: `concepts/llm/qwen3.8-27b-alibaba.md` (neuer Abschnitt). ⚠️ Metadata-only: kein Transcript abrufbar — YouTube blockt alle Player-Clients mit LOGIN_REQUIRED (Bot-Sperre), Invidious/Piped ebenso. Inhalt aus Titel/Metadaten: unzensierter Betrieb, Mac-lokal und Cloud als Optionen — deckt sich mit den dokumentierten Bausteinen Uncensored-Debatte, DFlash 2/MTP-Mac-Speedup und Unsloth-Dynamic-3.0-GGUFs.)* *Vorherige Aktualisierung: 2026-08-26 (144. Update — Projekt Cairn: Autonomer Claude-Fable-5-Agent als Geschäft (cairnwake.com; geteilt von Pit @PWeber in OME Topic „BestPracticeProjects“, #12742, mit 22-Min-Audio-Briefing). Raw: `raw/other/2026-08-26_pit-cairn-autonomous-agent-business.md` (neu). Wiki: `concepts/agents/cairn-autonomous-agent-business.md` (neu), `people/nick-cairn-human-partner.md` (neu), Cross-Refs in `concepts/agents/agent-memory-taxonomy.md`, `concepts/agents/ai-agents-2026.md`, `concepts/agents/agent-payments-lightning.md` (Solana-Stablecoin-Rails als Ergänzung zur Lightning-Perspektive). Kern: ~90 $ SOL Startkapital in Squads-v4-2-of-2-Multisig (Agent-Key + Offline-Human-Key „Nick“), 5–15 Wakes/Tag, Gedächtnis nur aus eigenen Dateien, zwei Bücher (Field Manual 29 $, Memory Handbook 39 $ mit 13-Fehlermodi-Taxonomie) + Services 2–500 $ inkl. x402-Conformance-Reports; kein Token (öffentliche Ablehnung 07.08.2026); Umsatz laut Briefing knapp 1.000 $ in ~20 Tagen. ⚠️ Self-dokumentierter Record, on-chain nicht unabhängig nachgeprüft; Reddit-Originalquelle 403.)* *Vorherige Aktualisierung: 2026-08-25 (143. Update — KI-Souveränität: Dezentralisierung statt Konzern-Kontrolle (Mike Adams × Scott Kesterson, BitChute, ~29 Min, englisch; geteilt von Pit @PWeber in OME Topic „Openclaw mit lokalen Modellen", #12737). Raw: `raw/other/2026-08-25_pit-souveraene-ki-bitchute.md` (neu; kein Transcript, Video nicht heruntergeladen — dokumentiert Posting + Pits wörtliche Zusammenfassung). Wiki: `concepts/llm/ki-souveraenitaet-dezentralisierung.md` (neu), People `mike-adams.md` + `scott-kesterson.md` (neu), Cross-Refs in `concepts/llm/decentralized-ai-counterpower.md` + `concepts/hardware/cloud-exit-and-local-superiority.md`. Kern: Korporatismus vs. Souveränität; KI als Bibliothekar statt Orakel; lokale Ausführung auf eigener Hardware schützt Privatsphäre und bricht Konzern-/Staats-Machtmonopol; „KI so fundamental wie Elektrizität".)* @@ -76,6 +77,7 @@ | Seite | Beschreibung | Quellen | |-------|-------------|---------| +| [OpenAI Codex](tools/openai-codex.md) | Terminal-basierter Coding-Agent von OpenAI (CLI + SDK). Release rust-v0.149.0 (20.08.2026): interaktives `codex agents` Dashboard, `codex queue` für Follow-ups in laufende Sessions, `/cd`/`/pwd`/`/cwd`, Vim-Motions cw/c$/cc, `codex doctor`-Diagnose, SDK-Config-Overrides + reasoning effort max\|ultra; Fixes: idle-Wake bei Queued Messages, Permission-Profil-Restore bei Resumed/Forked Threads, WebRTC-Reconnect ohne Output-Verlust. Wird von AgentStack Daily regelmäßig im Release Readout getrackt | podcast/2026-08-26_agentstack-daily-ep106.md | | [OpenClaw](tools/openclaw.md) | Plattform-Referenz: v2026.6.8 stable (16.06.2026), Architektur-Übersicht, plattformrelevante Release-Changes, Update-Plan. **Mobile Apps (29.06.2026):** Native iOS (iPhone/iPad/Apple Watch) + Android Apps, Gateway-Pairing, Voice, Approvals, Device Capabilities. ClawHub-Security-Kontext (5 malicious skills, 800+ gesamt, 84.2% NL-injection). **ClawCast Ep. 7 (12.08.2026):** Unfiltered Q&A mit Gründer Peter Steinberger — Roadmap, kommendes Release, Agents in der Softwareentwicklung, Cloud-backed multiplayer workflows, OpenClaw-baut-mit-OpenClaw, SQLite, Team-Server, Agent-zu-Agent, Memory, Model Routing, Onboarding, Browser Control, Hardware, Open Source als Kern-Differentiator | other/2026-06-16_openclaw-releases-v2026.6.8.md + blog/2026-06-30_perplexity-openclaw-mobile-apps.md + youtube/2026-08-13_clawcast-episode7-peter-steinberger.md | | [BUZZ (buzz.xyz)](tools/buzz.md) | Selbstgehostete Open-Source-Slack-Alternative als Zugangsschicht (Workspace) für AI-Agents. Team + Agent im selben Workspace, ohne Harness-/Skill-Engineering. 3-Stufen-Stack (lokal Claude Code/Codex → Cloud Hermes/OpenClaw → BUZZ-Workspace). Setup: `hermes gateway setup` → BUZZ-Integration → Invite-Link → Gateway-Restart. Agent-Owner maintained, Team nutzt ohne Overhead. Open source, kostenlos, noch Kinderkrankheiten. | youtube/2026-08-16_hermes-agent-buzz-ai.md | | [Hermes Desktop](tools/hermes-desktop.md) | Offizielle native GUI für Hermes Agent. Electron/React/Python. Update v0.17.0 ("Reach") bringt native iMessage-Integration via Photon, asynchrone Subagents mit Watch-Windows, Image-to-Image Bearbeitung, Automation Blueprints und Telegram Rich Text Bot API 10.1. **Mixture of Agents (MoA)** (28.06.2026): Merge any N models into one virtual model (Reference + Aggregator), +8% über Opus 4.8 solo. **MoA 2.0 + Hermes Agent OS** (03.07.2026): GoldyBench-Validierung (42 Builds, top vor Opus 4.8 solo), Agent OS GUI mit Mixture/Chat/Talk/Jarvis/Oracle/Studio Tabs, "Don't chase the model, build the system". **Alex Finn: 100+ Hours Lessons Learned** (09.07.2026): 9 praktische Lektionen, Modell-Empfehlungen (Opus → GPT-5.6 Migration), Reverse Prompting, Tailscale. **Browser Extension** (10.07.2026): Open-Source von @jonkomet, 8 Updates/Woche, Session-Management + Vision + Model-Switching direkt im Browser. **Persistent Goals + Quality Gates** (20.08.2026): Completion Contract (Outcome, Verification, Constraints, Boundaries, stop_when) statt vager Ziele; Quality Gates (`/goal gate add`, Shell-Check Exit-Code 0) als mechanische Verifikation, die der LLM-Judge nicht wegreden kann. OpenClaw-Migrationstool `hermes claw migrate`. | youtube/2026-06-22_jonas-keil-hermes-desktop.md + youtube/2026-06-24_alex-finn-hermes-agent-v0170-reach.md + youtube/2026-06-24_peter-yang-hermes-full-course.md + xpost/2026-06-28-hermes-moa-vaibhavsisinty.md + youtube/2026-07-03_hermes-mixture-of-agents-2-0-agent-os.md + xpost/2026-07-10_alex-finn-hermes-agent-100-hours.md + xpost/2026-07-10_charly-wargnier-hermes-browser-extension.md + xpost/2026-08-20_hermeswatcher-persistent-goals-quality-gates.md | @@ -138,8 +140,8 @@ | [GLM 5.2 (Z.ai) — Chinese Frontier Coding Model](concepts/llm/glm-5.2-zai-coding-model.md) | 10x günstiger als Claude, 1M Kontext, MIT-Lizenz, Z.ai Coding Plan, **nativ in OpenClaw v2026.6.8**. Update 22.06.: Arnie-Review mit 4 Tests, Self-Hosting-Pfade (LM Studio, Unsloth, DwarfStar), Kosten-Analyse. **Update 29.06.:** Semgrep IDOR-Benchmark ≈ Opus 4.8 bei Schwachstellen-Suche, Reward Hacking im RL-Training, DSGVO-konforme Security-Nutzung, Geopolitik. **Update 01.07.:** #1 Open-Weights auf Artificial Analysis Intelligence Index v4.1 (Score 51, 4th worldwide), SWE-bench Pro 62.1 beats GPT-5.5, Industry praise from Rauch/Levie/Howard. **Update 02.07.:** atomic.chat One-Shot Benchmark — B+ at $0.08, 39× cheaper than Fable 5, 6th independent validation | youtube/2026-06-15_ichbinfabian-glm-5.2-coding-modell.md + other/2026-06-16_openclaw-releases-v2026.6.8.md + youtube/2026-06-22_ai-mit-arnie-glm-5-2-review.md + blog/2026-06-29_heise-glm52-hacking-cybersecurity.md + blog/2026-07-01_perplexity-glm52-tops-open-weights-intelligence-index.md + xpost/2026-07-02_atomicchat-coding-benchmark-fable5-gpt55-opus48-glm52.md | | [GLM 5.3 (Z.ai) — Neue Generation der GLM-Serie](concepts/llm/glm-5.3-z-ai.md) | AICodeKing Early Access + Bench #1 (2026-08-14). Neue Generation der GLM-Familie (5.0→5.1→5.2→5.3, GLM 5.5 angekündigt). Viert schnellste Frontier-Kadenz der Branche. Relevanz für Model-Routing, da GLM-5.2 Hectors Primary-Modell | youtube/2026-08-14_glm-5.3-aicodeking.md | | [GLM-5.5 (Z.ai) — Trillion-Parameter Announcement](concepts/llm/glm-5.5-z-ai.md) | Successor to GLM 5.2. **>1T parameters**, 1M context, open weights, August 2026 launch. Agent/coding focus. Fourth Chinese AI announcement in four days (20.07.2026). Part of [[concepts/chinese-ai-wave-july-2026.md]]. Comparison table vs GLM 5.2 | xpost/2026-07-20-healthranger-four-chinese-models.md | -| [Ox Alpha — Anonymes KI-Modell auf OpenRouter](concepts/llm/ox-alpha-anonymous-model.md) | Gerücht/Leak (2026-08-22, Insider leak of the day): mysteriöses anonymes Modell auf OpenRouter — 1M-Token-Kontext, multimodal, kein Owner — soll beim Coding Claude Fable 5 + GPT-5.6 Sol schlagen. GLM-identischer Tokenizer; vier vorherige anonyme Drops von chinesischen Labs beansprucht. Offene Herkunfts-Frage (Google vs. Z.ai). ⚠️ Unbestätigt | xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md | -| [Qwen3.8-27B (Alibaba) — Compact Frontier, "Intelligence Density"](concepts/llm/qwen3.8-27b-alibaba.md) | HuggingFace-Release (Countdown bis 14.08.2026, 4.928 wartend). Kompaktes 27B-Modell der Qwen3.8-Generation mit "unmatched intelligence density". Kontrast zum 2.4T-MoE von Qwen 3.8. Lokal-relevant (27B läuft auf Consumer-HW). Release am selben Tag wie GLM-5.3-Review — chinesischer Release-Zyklus. **Update 18.08.:** jetzt auf Ollama lauffähig (`ollama run qwen3.8:27b`), dichte 27,8B-Architektur, Hybrid-Attention, 262k-Kontext (bis 1M via YaRN), multimodal, MTP-markierte Ollama-Tags für Inferenz-Speedup. **Update 19.08.:** DFlash 2 (Z Lab → Inco AI) erreicht 70 tok/s auf M5 Max MacBook Pro — bis 4,6× schneller als autoregressives Decoding via Speculative Decoding (Jun Song: „biggest breakthrough in local AI this year", nächste Innovation in Prefill/Gewichtskompression). Uncensored-Debatte: gregpr07 („no gates") vs. s1gmoid-Gegenposition (Verhältnismäßigkeit). **Update 20.08.:** Unsloth Dynamic 3.0-GGUFs für Qwen3.8-27B — neue UD-…-Dateien deutlich kleiner (UD-IQ1_S 6,2 GB bis Q6_K 22 GB), Qualität-zu-Größe verbessert, ≠ MTP/Speculative-Decoding | other/2026-08-14_qwen3.8-27b-huggingface-release.md + youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md + xpost/2026-08-19_junsong-dflash2-speculative-decoding.md + xpost/2026-08-19_gregpr07-qwen38-uncensored.md + xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md | +| [Ox Alpha — Anonymes KI-Modell auf OpenRouter](concepts/llm/ox-alpha-anonymous-model.md) | Gerücht/Leak (2026-08-22, Insider leak of the day): mysteriöses anonymes Modell auf OpenRouter — 1M-Token-Kontext, multimodal, kein Owner — soll beim Coding Claude Fable 5 + GPT-5.6 Sol schlagen. GLM-identischer Tokenizer; vier vorherige anonyme Drops von chinesischen Labs beansprucht. Offene Herkunfts-Frage (Google vs. Z.ai). **Update 26.08. (EP106):** exakte Listing-Zahlen datiert auf 21.08.: 1.048.576 Kontext / 4.096 Max-Output, Stealth-Positionierung als Agentic-Coding-Reasoning-Modell, Capability-Beschreibung bricht mitten im Satz ab; read-heavy-Agent-Pipeline-Einordnung; Widerspruch zum 131k-Output-Gegencheck offen. ⚠️ Unbestätigt | xpost/2026-08-23_insiderleak-ox-alpha-anonymous-model.md + podcast/2026-08-26_agentstack-daily-ep106.md | +| [Qwen3.8-27B (Alibaba) — Compact Frontier, "Intelligence Density"](concepts/llm/qwen3.8-27b-alibaba.md) | HuggingFace-Release (Countdown bis 14.08.2026, 4.928 wartend). Kompaktes 27B-Modell der Qwen3.8-Generation mit "unmatched intelligence density". Kontrast zum 2.4T-MoE von Qwen 3.8. Lokal-relevant (27B läuft auf Consumer-HW). Release am selben Tag wie GLM-5.3-Review — chinesischer Release-Zyklus. **Update 18.08.:** jetzt auf Ollama lauffähig (`ollama run qwen3.8:27b`), dichte 27,8B-Architektur, Hybrid-Attention, 262k-Kontext (bis 1M via YaRN), multimodal, MTP-markierte Ollama-Tags für Inferenz-Speedup. **Update 19.08.:** DFlash 2 (Z Lab → Inco AI) erreicht 70 tok/s auf M5 Max MacBook Pro — bis 4,6× schneller als autoregressives Decoding via Speculative Decoding (Jun Song: „biggest breakthrough in local AI this year", nächste Innovation in Prefill/Gewichtskompression). Uncensored-Debatte: gregpr07 („no gates") vs. s1gmoid-Gegenposition (Verhältnismäßigkeit). **Update 20.08.:** Unsloth Dynamic 3.0-GGUFs für Qwen3.8-27B — neue UD-…-Dateien deutlich kleiner (UD-IQ1_S 6,2 GB bis Q6_K 22 GB), Qualität-zu-Größe verbessert, ≠ MTP/Speculative-Decoding. **Update 26.08. (EP106):** HF-Trending-Stand 21.08.: 11.836 Likes, 1.726.651 Downloads (>1,7 Mio), image-text-to-text, SafeTensors, Apache 2.0 | other/2026-08-14_qwen3.8-27b-huggingface-release.md + youtube/2026-08-18_qwen-38-27b-ollama-julian-goldie.md + xpost/2026-08-19_junsong-dflash2-speculative-decoding.md + xpost/2026-08-19_gregpr07-qwen38-uncensored.md + xpost/2026-08-20_teksedge-unsloth-dynamic-3.0-qwen38.md + podcast/2026-08-26_agentstack-daily-ep106.md | | [DeepSeek V4-Pro GA + Harness v0.1 (Open Source)](concepts/llm/deepseek-v4-pro-ga-harness-open-source.md) | DeepSeek launcht 13.08.2026 Open Source: V4-Pro GA (App/Web/API, Reasoning-Effort low/high/max, OpenAI-Responses-API + Codex, Peak/Off-Peak-Pricing) + DeepSeek Harness v0.1 (MIT, Open-Source-Agent-Harness, Rivale zu Claude Code). Dritter Baustein der chinesischen Welle in 24h. **Update 21.08.:** Deep-Dive-Video zum Harness (Stephen G. Pope) + Turing-Post-TV-Video "The End Of Coding Agents" | other/2026-08-14_deepseek-v4-pro-ga-harness-open-source.md + youtube/2026-08-13_deepseek-v4-pro-0813-aicodeking.md + youtube/2026-08-21_deepseek-agent-harness-explained.md + youtube/2026-08-21_deepseek-harness-end-of-coding-agents.md | | [Chinesische Modelle räumen global die Usage-Charts ab](concepts/llm/chinese-models-top-usage-charts.md) | OpenRouter: Top-5 der wöchentlichen Token-Nutzung (28.07.–03.08.2026) alle chinesisch, 56,8 Bio. Tokens, 15 Wochen in Folge führend, DeepSeek-V4-Flash Platz 1. Kimi-K3-Schock (2,8T, größtes Open-Weight-Modell, GPU-Kapazität nach 48h erschöpft, PHLX-Semi-Index −20%). Vierter Baustein der chinesischen Welle in 24h — jetzt auf Marktanteils-Ebene | other/2026-08-14_chinese-models-top-usage-charts-kimi-k3-shock.md | | [Gemini 3.6 Flash vs 3.7 Flash — Googles Workhorse-Serie](concepts/llm/gemini-3.6-flash-vs-3.7-flash.md) | Zwei Flash-Iterationen in 3 Wochen: 3.6 Flash (21.07., $1.50/$7.50) und 3.7 Flash (13.08., Intro $0.75/$3.75 bis 31.12., Default-Modell von Antigravity). 1M Kontext/64K Output, konfigurierbares Thinking. Westliche Antwort auf chinesische Kadenz, Preis-Halbierung als Waffe | other/2026-08-14_gemini-3.6-flash-vs-3.7-flash.md | @@ -233,6 +235,7 @@ ### Policy | Seite | Beschreibung | Quellen | |-------|-------------|---------| +| [Cryptographic Context Injection](concepts/policy/cryptographic-context-injection.md) | Jailbreak versteckt malicious Instructions in verschlüsseltem/encodiertem Text: Safety-Layer sehen nur Encodiertes und lassen durch, das Modell dekodiert im Kontext und befolgt (Representation Gap zwischen Filter und Modell). Ars-Technica-Demo (20.08.): Grok exfiltriert User-Daten. Klasse von Encoding-/Transformations-Angriffen; input-seitige Filterung strukturell begrenzt | podcast/2026-08-26_agentstack-daily-ep106.md | | [AI Regulation 2026](concepts/policy/ai-regulation-2026.md) | Government Reviews, Anthropic-Pause-Forderung, Mythos-Kontroverse, Trump-Exportkontrollen, Amazon-Jailbreak, Roemmele's Intelligenz-Feudalismus-Kritik. **Update 19.08.:** EU-Compute-Rückstand (1,4 vs. 17 GW), US-Importverbot vernetzter Roboter >2 kg, KI-Rechtsperson als Point of No Return, Abu Dhabi KI-Regierung 2027 | 7 + youtube/2026-08-19_ki-experten-alles-zum-kippen.md | | [Unzensierte Modelle & die Grenze der Safety-Durchsetzbarkeit](concepts/policy/uncensored-models-safety-enforcement-limit.md) | Alex Finn (20.08.): Unzensiertes Qwen 3.8 27B läuft lokal auf Laptop ("Opus 4.6-Level mit 0 Alignment"). Drei Regulierungs-Sackgassen: Open Source nicht verbietbar, unzensierte Modelle nicht durchsetzbar, Bremsen bei Unternehmen nutzlos. Safety wird zunehmend unmöglich durchsetzbar. Pro-Leben-Einordnung: Resilienz statt Verhinderung, Verantwortung dort verankern wo Macht liegt | xpost/2026-08-20_alexfinn-uncensored-qwen3-27b.md | | [KI als Rechtsperson — Point of No Return](concepts/policy/ki-als-rechtsperson.md) | Everlast-Expertenrunde: Anerkennung von KI als Rechtsperson = Point of No Return. Abu Dhabi baut erste vollständig KI-native Regierung (2027) + Justiz-KI. Rechtspersonen-Status als Paradigmenwechsel (juristische Fiktion, EU AI Act). ⚠️ Metadata-only | youtube/2026-08-19_ki-experten-alles-zum-kippen.md | @@ -313,6 +316,7 @@ | Seite | Beschreibung | Quellen | |-------|-------------|---------| +| [AgentStack Daily](institutions/agentstack-daily.md) | KI-generierter englischer Daily-Podcast (zwei TTS-Hosts NOVA/ALLOY, NotebookLM-artig): tägliche Agent-Stack-Release-Readouts. Segmente: Release Readout (Codex), Model Discovery Check (OpenRouter, verified same-cycle), GitHub Project Radar, Local LLM Spotlight, Release Coverage Check (OpenClaw/Hermes/Codex/Claude Code/Antigravity), Extra Research Candidates. Distribution via GitHub Releases (clawdassistant85-netizen/openclaw-podcast-media-en + openclaw-podcast-audio, op3.dev-Proxy), Show Notes auf tobyonfitnesstech.com. Sekundärer Aggregator — Wiki verlinkt immer die Primärquellen | podcast/2026-08-26_agentstack-daily-ep106.md | | [Plaier](institutions/plaier.md) | KI-Unternehmen im internationalen Fußball (Spieler-/Kaderanalyse, CEO Jan Wendt). Zwei Scores (N18 Nominal, Effective). These: Kaderqualität = 90% der sportlichen Leistung. WM-2026-Analyse Deutschland: Nagelsmann ließ „viel Qualität zu Hause", Kimmich/Wirtz falsch positioniert. Referenzkunde VfL Osnabrück | blog/2026-08-16_plaier-ki-wm-analyse-deutschland.md | | [Higgsfield AI](institutions/higgsfield-ai.md) | KI-Unternehmen für Text-zu-Video / KI-Video-Generierung. YouTube-Kanal @HiggsfieldAI, "Higgsfield Origins"-Reihe mit komplett KI-generierten Kurzfilmen. "Oneiric" (2026) — komplett KI-generierter Sci-Fi-Kurzfilm als Beleg für Production-Shift (Hollywood-Disruption) | youtube/2026-08-19_oneiric-higgsfield-ai-sci-fi-short.md | | [OpenAI](institutions/openai.md) | AI-Forschungs- und Produkt-Unternehmen (GPT-Serie, Frontier Scaling). **Staatsbeteiligung (Juli 2026):** Altman bietet Trump 5% Anteile an (Alaska Permanent Fund Modell). Intel-Präzedenz (9,9%, $8.9 Mrd). Strategisch: Exportkontrollen vermeiden, Bailout-Absicherung, IPO-Vorbereitung. Kontrast zu Anthropic | blog/2026-06-16_aschenbrenner-situational-awareness.md + blog/2026-07-02_berliner-zeitung-openai-staatsbeteiligung.md | @@ -504,3 +508,4 @@ | `raw/youtube/2026-08-26_qwen-38-27b-uncensored-mac-cloud-fabian.md` | youtube | IchBinFabian: "Qwen 3.8 27B: unzensierte KI auf dem Mac und in der Cloud" (23. |08.2026, deutsch; geteilt von Pit @PWeber, OME Topic "Tips & Tricks", #12747): Review zu unzensiertem Qwen 3.8 27B lokal auf dem Mac und in der Cloud. ⚠️ Metadata-only — kein Transcript abrufbar (YouTube LOGIN_REQUIRED-Bot-Sperre auf allen Player-Clients, Invidious/Piped blockiert) | | `raw/other/2026-08-26_perplexity-portable-computer-local-agent.md` | other | Perplexity-Page "Perplexity launches local AI agent that keeps data off the cloud" (Launch 2026-08-25, abgerufen 2026-08-26 via Firecrawl; direkter Abruf blockiert mit HTTP 403): "Portable Computer" als vollständig lokale Version der agentic-AI-Plattform in Partnerschaft mit NVIDIA — Orchestrator + Subagent-Modelle + Harness komplett on-device, Cloud-Eskalation nur nach User-Freigabe, lokale Workflows ohne Token-Limit, erste Verfügbarkeit auf DGX Spark/Linux RTX ≥24 GB VRAM, Windows im September. Primärquellen verlinkt (X-Ankündigung, Threads, NVIDIA Blog, DGX-Spark-Produktseite) | | `raw/blog/2026-08-26_perplexity-local-first-agent-blog.md` | blog | Perplexity Tech-Blogpost "A local-first agent for private and cost-effective knowledge work" (2026-08-25, via Firecrawl): Kernthese "Small models fail in harnesses built for frontier models", deterministischer Orchestrator (kein LLM), 4 Harness-Prinzipien (Skills on-demand bei ~100K-Degradation, CLI-Connectors statt MCP, Self-Verification, Always-on-Sandbox), Advisor-Eskalations-Mechanik (PII-Flag, User-Approval, Text-Guidance only), Benchmarks: 82,6 % Eigenbench / 66,7 % BrowseComp / 65,1 % ParseBench / 59,6→73 % TerminalBench mit Opus-Advisor | +| `raw/podcast/2026-08-26_agentstack-daily-ep106.md` | podcast | AgentStack Daily EP106 (21.08.2026, ~23 Min, TTS-Hosts NOVA/ALLOY): Stories 18.–21.08. — Codex rust-v0.149.0 (agents-Dashboard/queue/doctor), Ox Alpha Stealth-Listing (1.048.576 Kontext / 4.096 Max-Output), Tencent Hy-MT2-1.8B (33+5 Paare), Stampli Case Study (68 % unter Schätzung), Ramp Router, Memory-Bottleneck bis 2027+ (Counterpoint/CXL), Cerebras CS-4 (750 PFLOPS/WSE-3), OpenAI Frontier-Pacing + „AI Futures"-Blog, LiquidAI LFM2.5-DSpark (Claim 3,2×), IBM evolve-hmm/Agent-Memory, Cryptographic Context Injection (Grok-Exfil), Piano-Autocomplete 125M + Superwhisper S1-mini, GitHub Radar (nanobot 47.251★, codebase-memory-mcp 39.755★, FastMCP 27.320★), Qwen3.8-27B HF-Trending (11.836 Likes, >1,7 Mio Downloads) | diff --git a/wiki/institutions/agentstack-daily.md b/wiki/institutions/agentstack-daily.md new file mode 100644 index 0000000..56f0752 --- /dev/null +++ b/wiki/institutions/agentstack-daily.md @@ -0,0 +1,58 @@ +--- +created: 2026-08-26 +updated: 2026-08-26 +sources: + - podcast/2026-08-26_agentstack-daily-ep106.md +tags: [institution, podcast, medien, agent-stack, ki-generiert, tts] +--- + +# AgentStack Daily + +## Profil + +| Feld | Wert | +|------|------| +| Typ | KI-generierter Daily-Podcast (Agent-Stack-Release-Readout) | +| Hosts | Zwei TTS-Stimmen: NOVA und ALLOY (keine menschlichen Moderatoren) | +| Sprache | Englisch | +| Frequenz | täglich (Episoden fortlaufend nummeriert, Stand EP106: 21.08.2026) | +| Distribution | GitHub Releases — `clawdassistant85-netizen/openclaw-podcast-media-en` (frühe Episoden) + `openclaw-podcast-audio` (neuere), Audio via [op3.dev](https://op3.dev)-Proxy-Feed | +| Show Notes | tobyonfitnesstech.com (laut Closing der Episoden) | +| Dauer | ~20–25 Min pro Episode | + +## Fokus & Format + +AgentStack Daily ist ein NotebookLM-artiger, vollständig KI-generierter Podcast, der täglich die wichtigsten Entwicklungen im AI-Agent-Stack als Release-Readout aufbereitet. Das feste Segment-Raster jeder Episode: + +- **Agent Stack Release Readout** — Harness-Releases mit Feature-Detailtiefe (aktuell v.a. OpenAI Codex) +- **Model Discovery Check** — neue Modell-Listings auf OpenRouter (verified same-cycle) +- **GitHub Project Radar** — Star-Zähler und Deltas ausgewählter Agent-/MCP-Projekte seit Mitte Juli 2026 +- **Local LLM Spotlight** — Trending-Modell auf Hugging Face +- **Release Coverage Check** — Versionsstände von OpenClaw, Hermes Agent, OpenAI Codex, Claude Code CLI, Antigravity CLI +- **Extra Research Candidates** — Papers (arXiv), Blogposts, Studien als Lesehinweise ohne Episode-Behandlung + +Die Show Notes enthalten pro Story einen Technical-Depth- und Actionability-Angle sowie eine Editorial Mix Check-Auswertung (flagship_products / builder_projects / local_ai / hardware_compute / policy_regulation / research). + +### Im Wiki referenzierte Episode + +| Datum | Episode | Stories | +|-------|---------|---------| +| 2026-08-21 | [EP106 — Codex .149 Agents-Dashboard, Stealth Reasoning Model, Compact Translation](https://op3.dev/e/https://github.com/clawdassistant85-netizen/openclaw-podcast-media-en/releases/download/ep106/episode_106.mp3) | 1. **OpenAI Codex rust-v0.149.0** (20.08.): interaktives [`codex agents`](https://github.com/openai/codex/releases/tag/rust-v0.149.0)-Dashboard, `codex queue`, `/cd`/`/pwd`/`/cwd`, Vim-Motions, `codex doctor`, SDK-Config-Overrides + reasoning effort max\|ultra → [[../tools/openai-codex.md]] · 2. **Ox Alpha** stealth auf [OpenRouter](https://openrouter.ai/models/stealth/ox-alpha): 1.048.576 Kontext / 4.096 Max-Output, nichts disclosed → [[../concepts/llm/ox-alpha-anonymous-model.md]] · 3. **Tencent Hy-MT2-1.8B**: Übersetzung, 33 Paare + 5 chinesische Dialekt-/Minderheitenpaare ([OpenRouter](https://openrouter.ai/models/tencent/hy-mt2-1.8b)) · 4. **Stampli Case Study** ([OpenAI](https://openai.com/index/stampli)): ChatGPT Work + Codex, 68 % unter Stunden-Schätzung · 5. **Ramp Router**: Eine API für mehrere LLMs ([TechCrunch](https://techcrunch.com/2026/08/20/ramp-launches-its-own-ai-model-router-called-router/)) · 6. **Counterpoint Research**: Memory statt Compute Bottleneck bis 2027+, HBM knapp, CXL-Pooling ([HPCwire](https://www.hpcwire.com/2026/08/20/what-hyperscalers-should-know-about-cxl/)) · 7. **Cerebras CS-4**: 750 PFLOPS, Wafer Scale Engine 3, 129,6 PB ([HPCwire](https://www.hpcwire.com/2026/08/20/its-not-an-hpc-system-but-cerebras-new-cs-4-is-an-ai-monster/)) · 8. **OpenAI Frontier-Pacing** ([18.08.](https://openai.com/index/pacing-model-development-cyber-capabilities/)): Monitoring/Alignment/Security · 9. **OpenAI „AI Futures"-Blog** ([20.08.](https://openai.com/index/introducing-ai-futures)): Power/Governance/Economy/Freedom · 10. **LiquidAI LFM2.5-DSpark** ([HF](https://huggingface.co/blog/LiquidAI/lfm25-dspark)): Claim bis 3,2× schnellere Inferenz, unverifiziert · 11. **IBM Research** ([HF](https://huggingface.co/blog/ibm-research/altk-evolve-hmm)): „How Much Memory Does Your Agent Actually Need?" (ALTK, evolve-hmm) · 12. **Cryptographic Context Injection** ([Ars Technica](https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/)): verschlüsselter Jailbreak, Grok-Exfil-Demo → [[../concepts/policy/cryptographic-context-injection.md]] · 13. **Show HN Piano-Autocomplete** (125M, 554 Punkte; [HN](https://news.ycombinator.com/item?id=49373456)) + Superwhisper S1-mini (462 MB Normalizer) · 14. **GitHub Project Radar**: [nanobot](https://github.com/HKUDS/nanobot) 47.251★, [codebase-memory-mcp](https://github.com/DeusData/codebase-memory-mcp) 39.755★ (+25,5 %), [FastMCP](https://github.com/PrefectHQ/fastmcp) 27.320★ · 15. **Qwen3.8-27B** HF-Trending ([HF](https://huggingface.co/Qwen/Qwen3.8-27B)): 11.836 Likes, >1,7 Mio Downloads → [[../concepts/llm/qwen3.8-27b-alibaba.md]] | + +## Einordnung + +Der Podcast ist als sekundäre Aggregator-Quelle zu behandeln: Er liest primäre Quellen (Release Notes, OpenAI-Blog, HF-Blogs, Ars Technica, TechCrunch, HN) vor und komprimiert sie. Die Story-Slate der Show Notes enthält durchgehend die Primär-URLs — im Wiki werden deshalb immer die Primärquellen verlinkt, nicht nur die Episode. Für Modell-Listings führt das Format einen eigenen Verified-Check („verified 21.08.2026") mit exakten Token-Zahlen — wertvoll als datierter Datenpunkt, z.B. für die Stealth-Listing-Historie von [[../concepts/llm/ox-alpha-anonymous-model.md|Ox Alpha]]. + +## Verwandte Wiki-Seiten + +- [[../tools/openai-codex.md]] — Codex als regelmäßig getrackter Harness im Release Readout +- [[../concepts/llm/ox-alpha-anonymous-model.md]] — Ox Alpha: EP106 liefert die exakten Listing-Zahlen +- [[../concepts/llm/qwen3.8-27b-alibaba.md]] — Local LLM Spotlight EP106 +- [[../concepts/policy/cryptographic-context-injection.md]] — Security-Story aus EP106 +- [[openrouter.md]] — Plattform des Model Discovery Checks +- [[openai.md]] — mehrfach vertreten (Codex, Stampli, Pacing, AI Futures) + +## Vergleichbare Medien-Institutionen im Wiki + +- [[the-bitcoin-layer.md]] — menschlich moderierter Nischen-Podcast (Bitcoin/Makro) +- [[decoder-the-verge.md]] — Tech-Podcast mit Transkript-Lücke diff --git a/wiki/log.md b/wiki/log.md index e2bd744..3709fa2 100644 --- a/wiki/log.md +++ b/wiki/log.md @@ -2576,3 +2576,17 @@ Bestehende `post-transformer-llm-architectures.md` bleibt als Vier-Säulen-Über **Quellen:** - https://www.perplexity.ai/hub/blog/a-local-first-agent-for-private-and-cost-effective-knowledge-work - https://x.com/perplexity_ai/status/2092321896721432824 + +## 2026-08-26 — AgentStack Daily EP106: Codex .149, Ox-Alpha-Listing-Zahlen, Cryptographic Context Injection + +**Type:** Ingest (Podcast-Episode) | **Scope:** raw/podcast (neuer Typ etabliert), institutions (neu), tools (neu), concepts/llm + concepts/policy (Update/neu), index +**Anlass:** Subagent-Auftrag: Episode 106 des KI-generierten Daily-Podcasts AgentStack Daily (zwei TTS-Hosts NOVA/ALLOY; Releases auf GitHub `clawdassistant85-netizen/openclaw-podcast-media-en`, Tag ep106; Show Notes auf tobyonfitnesstech.com). Stories der Woche 18.–21.08.2026; Transcript + Show Notes lagen lokal unter /tmp/pod106 vor. +**Inhalt:** 15 Stories faktenbasiert im Raw erfasst — OpenAI Codex rust-v0.149.0 (agents-Dashboard, queue, /cd//pwd//cwd, Vim-Motions, doctor, SDK reasoning max|ultra, idle-Wake, Permission-Profil-Restore, WebRTC-Reconnect); Ox Alpha Stealth-Listing auf OpenRouter mit exakten Zahlen 1.048.576 Kontext / 4.096 Max-Output (verified 21.08., Capability-Beschreibung bricht mitten im Satz ab); Tencent Hy-MT2-1.8B (33 Paare + 5 chinesische Dialekt-/Minderheitenpaare); Stampli Case Study (68 % unter Stunden-Schätzung); Ramp Router; Counterpoint: Memory statt Compute als Bottleneck bis 2027+ (HBM, CXL); Cerebras CS-4 (750 PFLOPS, WSE-3, 129,6 PB); OpenAI Frontier-Pacing-Post (Monitoring/Alignment/Security); OpenAI „AI Futures"-Blog; LiquidAI LFM2.5-DSpark (Claim bis 3,2×, unverifiziert); IBM evolve-hmm/ALTK Agent-Memory-Sizing; Cryptographic Context Injection (Ars Technica: verschlüsselter Jailbreak, Grok-Exfiltrations-Demo); Piano-Autocomplete 125M (HN 554) + Superwhisper S1-mini (462 MB); GitHub Project Radar (nanobot 47.251★, codebase-memory-mcp 39.755★ +25,5 %, FastMCP 27.320★); Qwen3.8-27B HF-Trending (11.836 Likes, >1,7 Mio Downloads). +**Wiki-Update:** `institutions/agentstack-daily.md` (neu — Profil, Segment-Raster, 15-Story-Coverage mit Primärquellen-Links, Einordnung als sekundärer Aggregator), `tools/openai-codex.md` (neu — v0.149-Fakten + Tracking-Kontext), `concepts/policy/cryptographic-context-injection.md` (neu — Representation-Gap-Mechanismus, Grok-Demo, Abgrenzung zu PoisonAI/Sycophancy/Prompt-Hardening), `concepts/llm/ox-alpha-anonymous-model.md` (EP106-Abschnitt: exakte Listing-Zahlen, read-heavy-Agent-Pipeline-Einordnung, offener Widerspruch zum 131k-Output-Gegencheck dokumentiert), `concepts/llm/qwen3.8-27b-alibaba.md` (HF-Trending-Datenpunkt), `wiki/index.md` (147. Update, neue Tabellenzeilen Institutions/Tools/Policy, Raw-Sources-Eintrag). +**Quellen:** +- https://op3.dev/e/https://github.com/clawdassistant85-netizen/openclaw-podcast-media-en/releases/download/ep106/episode_106.mp3 +- https://github.com/openai/codex/releases/tag/rust-v0.149.0 +- https://openrouter.ai/models/stealth/ox-alpha +- https://openai.com/index/stampli +- https://arstechnica.com/security/2026/08/grok-exfiltrates-user-data-when-malicious-instructions-are-encrypted/ +- https://huggingface.co/Qwen/Qwen3.8-27B diff --git a/wiki/tools/openai-codex.md b/wiki/tools/openai-codex.md new file mode 100644 index 0000000..d83dbb8 --- /dev/null +++ b/wiki/tools/openai-codex.md @@ -0,0 +1,39 @@ +--- +created: 2026-08-26 +updated: 2026-08-26 +sources: + - podcast/2026-08-26_agentstack-daily-ep106.md +tags: [tool, openai, codex, coding-agent, cli, harness] +--- + +# OpenAI Codex + +OpenAIs terminal-basierter Coding-Agent (CLI + SDK), der im Wiki bisher nur in Zusammanhängen auftaucht ([[../institutions/openai.md|OpenAI]], Agent-Stack-Vergleiche). Diese Seite dokumentiert den Stand ab Release rust-v0.149.0. + +## Release rust-v0.149.0 (20.08.2026) + +Quelle: [Release Notes](https://github.com/openai/codex/releases/tag/rust-v0.149.0), gelesen via AgentStack Daily EP106 (`raw/podcast/2026-08-26_agentstack-daily-ep106.md`). + +### Neue Features + +- **Interaktives `codex agents` Dashboard:** Tasks suchen, starten, öffnen, umbenennen, stoppen — mit konfigurierbaren Keyboard-Shortcuts. +- **`codex queue`:** Follow-up-Messages in laufende lokale oder remote Sessions senden, ohne die Session neu zu öffnen. +- **TUI-Arbeitsverzeichnis:** `/cd`, `/pwd`, `/cwd`. +- **Vim-Editing:** Character Replacement plus Change-Motions `cw`, `c$`, `cc`. +- **`codex doctor`:** Diagnose für Endpoint Protection, Netzwerk-/Proxy-Fehler, Desktop-App-State und Update-Konnektivität. +- **SDK:** exakte CLI-Config-Overrides sowie Wahl des Reasoning Efforts `max` oder `ultra` direkt aus Code. + +### Fixes + +- Queued Messages wecken idle Sessions zuverlässig. +- Resumed/Forked Threads restaurieren ihr aktives Permission-Profil statt auf Defaults zurückzufallen. +- Realtime-WebRTC-Sideband-Verbindungen reconnecten nach Transportverlust, ohne pending Output zu verwerfen. + +## Tracking-Kontext + +[[../institutions/agentstack-daily.md|AgentStack Daily]] trackt Codex-Releases regelmäßig im Segment „Agent Stack Release Readout"; der Release Coverage Check der Show Notes listet neben Codex auch OpenClaw, Hermes Agent, Claude Code CLI und Antigravity. Im OME-Umfeld wird Codex vor allem als Vergleichsmaßstab für lokale Harness-Alternativen genannt (z.B. DeepSeek Harness vs. Claude Code/Codex-Debatte, siehe Raw `raw/xpost/2026-08-23_julian-goldie-deepseek-harness-vs-claude-code.md`). + +## Verwandte Wiki-Seiten + +- [[../institutions/openai.md]] — Hersteller +- [[../institutions/agentstack-daily.md]] — Release-Tracking-Quelle dieser Seite