knowledge-base/wiki/log.md
Hector Bot 2cb33da005 ingest(youtube): fahd-mirza-kimi-k2.7-vs-glm-5.2
Fahd Mirza Head-to-Head in Hermes Agent (14.06.2026, 14 Min, 2659 views,
89 likes). Beide Modelle bekommen denselben Prompt, loesen denselben
Real-World-Task: Flask-App mit gepflanztem FIFA-Bug + Round-of-32-
Bracket-Feature mit Anti-Group-Rematch-Regel.

Test 1: Beide bestehen. Kimi schneller (~5 Min), innovativer
(zusaetzliche Progression-Previews). GLM: 97 tool calls, laenger.
Test 2 (Creative HTML, Siberian Wind): GLM staerker bei Animation +
Terrain-Detail; Kimi bei Stats + Geographie.

Konzept: Real-World-Showdown als Benchmark-Alternative zu statischen
Tests (HumanEval, MBPP, SWE-bench). Agent-Framework isoliert die
Modell-Variable sauber.

Direkte Implikation fuer OpenClaw:
- Sub-Task-spezifisches Coding-Routing (Kimi vs GLM je nach Task)
- Ensemble-Kandidaten: Kimi K2.7 + GLM-5.2 in Coding-Panels
- Real-World-Validierung statt nur Standard-Benchmarks fuer Production-
  Pfade
- Hermes Agent als alternatives Agent-Framework zu AutoGen/CrewAI

Aenderungen:
- raw/youtube/2026-06-14_fahd-mirza-kimi-k2.7-vs-glm-5.2.md (neu)
- wiki/concepts/real-world-coding-showdown.md (neu)
- wiki/tools/kimi-k2.7-code.md (Cross-References)
- wiki/architecture/model-routing.md (neue Sektion Coding-Modelle)
- wiki/index.md (neuer Concepts-Eintrag + Raw Sources-Eintrag)
- wiki/log.md (Eintrag)

Maximale Verlinkung gemaess AGENTS.md-Kardinalregel:
- 12 externe Verweise auf Modelle (Kimi, GLM, Hermes, Ollama, Z.ai)
- 9 Benchmark-/Methoden-Links (SWE-bench, HumanEval, MBPP, MCP, etc.)
- 7 Fahd-Mirza-Links (YT, Blog, LinkedIn, Substack, Ko-Fi)
- 5 Wiki-Cross-References (fusion, ai-agents, kimi-k2.7, model-routing, behavior-persistence)
2026-06-15 09:45:38 +02:00

184 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Wiki Log
*Append-only changelog. Start: 2026-06-05*
## [2026-06-07] Wiki | Subconscious Agent aktualisiert mit aktuellen Outcomes und Evidence Gap
**Type:** wiki-update | **Scope:** wiki/concepts
**Actions:**
- wiki: `concepts/subconscious-agent.md` (updated — aktuelle Outcomes (25 Runs), plur1bus Evidence Gap, Execution Gap mit konkreten Daten, Lösungsansatz von @k9ert)
- log: updated
## [2026-06-07] Ingest | Subconscious Agent Evidence Runner Script
**Type:** other | **Scope:** raw, wiki/concepts
**Actions:**
- raw: `raw/other/2026-06-07_subconscious-runner-script.md` (created - Dokumentation des Shell-Skripts)
- wiki: `concepts/subconscious-agent.md` (updated - Source-Referenz + aktueller Status)
- index: updated
## [2026-06-07] Schema | Autonomes Kuratieren Entscheidungsbaum
**Type:** schema | **Scope:** AGENTS.md
**Actions:**
- wiki: `AGENTS.md` (updated — zwei unabhängige Entscheidungsbäume für Reply vs. Curation)
## [2026-06-05] Ingest | Anthropic fordert weltweite KI-Pause (WELT)
**Type:** blog | **Scope:** raw, wiki/tools, wiki/concepts
**Actions:**
- raw: `raw/blog/2026-06-04_anthropic-fordert-ki-pause.md`
- wiki: `tools/anthropic-claude.md` (created + updated)
- wiki: `concepts/ai-regulation-2026.md` (created + populated)
- index: updated
## [2026-06-05] Ingest | AI News Roundup May 2026 (vtnetzwelt)
**Type:** blog | **Scope:** raw, wiki/tools, wiki/concepts
**Actions:**
- raw: `raw/blog/2026-06-05_ai-news-roundup-may-2026.md`
- wiki: `tools/openai-gpt.md` (created)
- wiki: `tools/anthropic-claude.md` (updated revenue data)
- wiki: `concepts/ai-agents-2026.md` (created)
- wiki: `concepts/ai-regulation-2026.md` (updated)
- index: updated
## [2026-06-07] Migration | ByteRover context-tree → Knowledge Base Wiki
**Type:** migration | **Scope:** wiki/architecture, wiki/concepts, wiki/tools, wiki/decisions
**Source:** `/home/node/workspace/.brv/context-tree/` (ByteRover context-tree, April 2026, 133 Markdown-Dateien)
**Actions:**
- wiki: `architecture/container-volume-persistence.md` (created)
- wiki: `architecture/memory-system.md` (created)
- wiki: `architecture/agent-orchestration.md` (created)
- wiki: `architecture/model-routing.md` (created)
- wiki: `architecture/cron-notable-events.md` (created)
- wiki: `architecture/byterover-knowledge-mining.md` (created)
- wiki: `tools/ecosystem-tools-april-2026.md` (created)
- wiki: `concepts/subconscious-agent.md` (created)
- wiki: `concepts/pro-leben-directive.md` (created)
- wiki: `concepts/quality-standard.md` (created)
- wiki: `decisions/2026-04-volume-persistence.md` (created)
- wiki: `decisions/2026-04-memory-system.md` (created)
- index: updated with new categories (Architecture, Decisions) and all new pages
**Notes:**
- Kompiliert aus ~133 Quell-Dateien (abstract/overview/full) → 12 prägnante Wiki-Seiten
- Fokus auf bleibende Erkenntnisse; temporäre/archivierte Details (einzelne Subconscious-Runs, Tagebuch-Einträge) nicht migriert
- Duplikate (identische Facts in architecture + ecosystem) dedupliziert
- Alte Tool-Integrationen (ByteRover, GBrain, SearXNG) historisch dokumentiert, nicht als aktive Empfehlungen
## [2026-06-12] Ingest | Colgate SSR — LLM-basierte Kaufintentions-Vorhersage
**Type:** ingest | **Scope:** raw/xpost, wiki/concepts
**Actions:**
- raw: `raw/xpost/2026-06-11_colgate-llm-purchase-intent-ssr.md` (created — X-Post von @HowToAI_ über Colgate/PyMC Labs SSR-Studie)
- wiki: `concepts/semantic-similarity-rating-ssr.md` (created — Konzeptseite zu Semantic Similarity Rating)
- index: updated (neuer Concepts-Eintrag + Raw Sources-Eintrag)
- log: updated
## [2026-06-12] Ingest | OME20 Special — Ufologie, Physik und Bewusstsein
**Type:** ingest | **Scope:** raw/podcast, wiki/events
**Actions:**
- raw: `raw/podcast/ome20-special-ufologie-physik-bewusstsein-2026-06-12.md` (created — Podcast-Special aus OME-Gruppe)
- wiki: `events/ome20-special-ufologie-physik-bewusstsein.md` (created — Wiki-Seite mit Themenüberblick)
- index: updated (neue Events-Kategorie + Raw Sources-Eintrag)
- log: updated
## [2026-06-13] Ingest | Robert Malone: Biologische KI und der Biosecurity-Staat
**Type:** ingest | **Scope:** raw/blog, wiki/concepts
**Source:** Kai (@PWeber) in OME-Gruppe "Krallenpolitik"
**Actions:**
- raw: `raw/blog/2026-06-13_malone-biological-ai-biosecurity.md` (created — Malone kritisiert Machtkonzentration unter Biosecurity-Deckmantel)
- wiki: `concepts/ai-biological-biosecurity.md` (created — Konzeptseite mit Kern-Argumenten, GOF-Verschleierung, DeepSeek-Paradoxon, BWC-Kritik)
- wiki: `concepts/ai-regulation-2026.md` (updated — Cross-Reference + Malone-Perspektive)
- index: updated (neuer Concepts-Eintrag + neuer Raw Sources-Eintrag)
- log: updated
## [2026-06-13] Ingest | Trump export controls block foreign access to Anthropic Mythos 5 / Fable 5
**Type:** ingest | **Scope:** raw/blog, wiki/concepts
**Source:** Kai (@PWeber) in OME-Gruppe
**Actions:**
- raw: `raw/blog/2026-06-13_trump-export-controls-anthropic-mythos-fable.md` (created — Axios scoop + Perplexity compilation)
- wiki: `concepts/ai-regulation-2026.md` (updated — neuer Abschnitt zu Exportkontrollen, Konflikt-Zeitlinie, NSA-Mythos-Verbindung)
- index: updated (neuer raw-Eintrag, aktualisierte Concept-Beschreibung)
- log: updated
## [2026-06-13] Ingest | Brian Roemmele: Amazon-Jailbreak von Fable 5 als Auslöser der Exportkontrollen
**Type:** ingest | **Scope:** raw/xpost, wiki/concepts, wiki/tools
**Source:** Kai (@PWeber) in OME-Gruppe "Krallenpolitik"
**Actions:**
- raw: `raw/xpost/2026-06-13_roemmele-amazon-jailbreak-fable5.md` (created — Brian Roemmele X-Thread: Amazon-Forscher jailbreakten Fable 5 per Prompt-Injection, gaben Ergebnisse an US-Regierung statt an Anthropic)
- wiki: `concepts/ai-regulation-2026.md` (updated — Amazon-Jailbreak als konkreter Auslöser der Exportkontrollen, WSJ/Times of India-Bestätigung)
- wiki: `tools/anthropic-claude.md` (updated — Trigger-Detail: Amazon-Forscher statt "ein anderes Unternehmen")
- index: updated (neuer Raw Sources-Eintrag, aktualisierte Quellen-Zähler)
- log: updated
## [2026-06-13] Ingest | Kimi K2.7 Code auf Ollama Cloud (NVIDIA B300)
**Type:** ingest | **Scope:** raw/other, wiki/tools
**Source:** Pit Weber (@PWeber) in OME News
**Actions:**
- raw: `raw/other/2026-06-13_kimi-k2.7-code-ollama.md` (created — Kimi K2.7 Code von Moonshot AI, jetzt auf Ollama Cloud gehostet auf NVIDIA B300)
- wiki: `tools/kimi-k2.7-code.md` (created — Wiki-Seite mit Benchmarks, Features und Nutzung)
- index: updated (neuer Tools-Eintrag + Raw Sources-Eintrag)
- log: updated
## [2026-06-13] Ingest | Brian Roemmele: Anthropic's Selbstzerstörung — Leadership-Failure & Kundenabwanderung
**Type:** ingest | **Scope:** raw/xpost, wiki/tools
**Source:** Kai (@PWeber) in OME-Gruppe "Krallenpolitik"
**Actions:**
- raw: `raw/xpost/2026-06-13_roemmele-anthropic-selfdestruct.md` (created — Brian Roemmele X-Thread: 7 Punkte zur Anthropic-Selbstzerstörung, Kundenmigration zu Open Source, IPO-Schaden, Leadership-Failure, Darios Regulierungs-Ironie)
- wiki: `tools/anthropic-claude.md` (updated — neuer Abschnitt "Branchen-Fallout (Juni 2026)" mit Kundenabwanderung, IPO-Schaden, Führungsproblem, Breitere Branchen-Implikation)
- index: updated (neuer Raw Sources-Eintrag, aktualisierte Tools-Zeile + Quellen-Zähler)
- log: updated
## [2026-06-15] Ingest | Liesel Weppen: LLM Behavior Persistence — Sleeper Agents, Backdoor-Persistenz, Unlearning-Grenzen
**Type:** ingest | **Scope:** raw/xpost, wiki/concepts
**Source:** @k9ert in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic
**Actions:**
- raw: `raw/xpost/2026-06-14_lieselweppen-open-source-llm-backdoors.md` (created — Liesel Weppen X-Thread mit 8 Suchbegriffen als Belege: Hubinger sleeper agents, Backdoor safety training, machine unlearning bias, bias persistence, bias transfer, persistent backdoors, catastrophic forgetting, Limits of debiasing. Diskussion mit k9ert über "Open Source bei LLMs = Marketing BS")
- wiki: `concepts/llm-behavior-persistence.md` (created — Konzeptseite zu Persistenz gelernter Verhaltensweisen in LLMs, drei Achsen: Sleeper-Agents/Backdoors, Machine Unlearning Bias, Catastrophic Forgetting; Verbindung zu "Open Weights ≠ Open Model"; Pro-Leben-Perspektive mit Handlungsspielräumen)
- index: updated (neuer Concepts-Eintrag "LLM Behavior Persistence" + Raw Sources-Eintrag)
- log: updated
**Notes:**
- Erst-Ingest, der direkt von einem Subagent-Fehler (xai billing error) zurückgespielt und im Main-Loop manuell finalisiert wurde
- Konzept identifiziert: Persistenz von gelernten Verhalten in LLMs übersteigt Fähigkeit von Fine-Tuning/Unlearning zur gezielten Entfernung
- Kernaussage: Open Weights ist nicht "sicher" — aber auch kein Argument gegen Offenheit; realistisch ist Open + Audit + transparente Trade-off-Dokumentation
## [2026-06-15] Wiki-Update | llm-behavior-persistence: Maximale Verlinkung nachgepflegt
**Type:** wiki-update | **Scope:** wiki/concepts, AGENTS.md
**Source:** Feedback von @k9ert: "Warum sind im Wiki die Schlüssel Autoren und -Papiere nicht verlinkt? Bitte immer so viele URLS / Verlinkungen wie möglich!"
**Actions:**
- wiki: `concepts/llm-behavior-persistence.md` (updated — vollständige Verlinkung aller zitierten Papiere, Autoren und externen Ressourcen; Hubinger 2024, Kurmanji 2023, Goel 2024, Lin 2024, McCloskey & Cohen 1989, Webster 2020, Gonen & Lazaridou 2024/25; neue Sektion "Externe Ressourcen" mit Anthropic/DeepMind/CAIS/NIST/AI Incident Database/LessWrong/Papers With Code)
- wiki: `AGENTS.md` (updated — neue Kardinalregel "Maximale Verlinkung" hinzugefügt: jede Quelle, jedes Paper, jeder Autor, jeder zitierte Begriff MUSS verlinkt sein)
- log: updated
**Lessons learned:**
- Ohne Verlinkungen ist das Wiki tot — Faustregel: lieber ein Link zu viel als einer zu wenig
- Gilt für JEDE zukünftige Wiki-Erstellung, nicht nur für diese eine Seite
- Lessons learned sollten dauerhaft in die Wiki-Konventionen einfließen (nicht nur im Log landen)
## [2026-06-15] Ingest | OpenRouter Fusion: Model Panels & Ensembles — Beyond-Frontier mit Budget-Modellen
**Type:** ingest | **Scope:** raw/blog, wiki/concepts
**Source:** @k9ert in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic
**Actions:**
- raw: `raw/blog/2026-06-12_openrouter-fusion-beats-frontier.md` (created — OpenRouter Blog-Announcement 12.06.2026 von Brian Thomas; validiert mit DRACO-Benchmark von Perplexity AI; 100 Deep-Research-Tasks; zeigt dass Panels Frontier schlagen)
- wiki: `concepts/llm-model-fusion-ensembles.md` (created — Konzeptseite zu Model-Panel-Ensembles: Architektur, DRACO-Ergebnisse, Self-Fusion +6.7 Punkte, Budget-Panels ~Frontier, Anti-Contamination via Excluded-Domains, Judge-Modell-Wahl, Pro-Leben-Perspektive, Open Questions für OpenClaw)
- index: updated (neuer Concepts-Eintrag "LLM Model Fusion & Ensembles" + Raw Sources-Eintrag)
- log: updated
**Kernkonzept:**
- Statt Frontier-Skalierung: mehrere Modelle parallel + Judge-Synthese
- DRACO-Validierung: Fable 5 + GPT-5.5 (synth. Opus 4.8) = 69.0% > Fable 5 solo 65.3%
- Budget-Panel (Gemini 3 Flash + Kimi K2.6 + DeepSeek V4 Pro) = 64.7% bei 50% Kosten
- Self-Fusion (gleiches Modell 2×) = +6.7 Punkte über Solo — Synthesis-Schicht ist *mindestens so wichtig* wie das Modell
**Lessons / Open Questions für OpenClaw:**
- Self-Fusion als Pre-Routing-Layer für Quality-kritische Tasks?
- Ensemble-Pattern für Subconscious-Agent Hard-Synthesis?
- Excluded-Domains für eigene Evals + Subconscious-Crawls (Self-Referenz-Loops vermeiden)
## [2026-06-15] Ingest | Kimi K2.7 vs GLM-5.2 Real Coding Showdown (Fahd Mirza)
**Type:** ingest | **Scope:** raw/youtube, wiki/concepts, wiki/tools, wiki/architecture
**Source:** @PWeber (Kai) in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic
**Actions:**
- raw: `raw/youtube/2026-06-14_fahd-mirza-kimi-k2.7-vs-glm-5.2.md` (created — Fahd Mirza YouTube-Video 14.06.2026, 14 Min, 2659 Aufrufe, 89 Likes; Head-to-Head in Hermes Agent mit Bug-Fix + Feature-Build in Flask-App)
- wiki: `concepts/real-world-coding-showdown.md` (created — Konzeptseite zu alternativer Coding-Evaluation: statische Benchmarks (HumanEval/MBPP/SWE-bench) vs. Real-World-Head-to-Head mit echtem Bug + Constraint-Satisfaction; methodische Lessons, Sub-Task-Spezialisierung, Pro-Leben-Perspektive)
- wiki: `tools/kimi-k2.7-code.md` (updated — Cross-Reference zu real-world-coding-showdown + llm-model-fusion-ensembles)
- wiki: `architecture/model-routing.md` (updated — neue Sektion "Coding-Modelle: Kimi K2.7 & GLM-5.2" mit Implikation für Sub-Task-Spezialisierung im Routing)
- index: updated (neuer Concepts-Eintrag "Real-World Coding Showdown" + Raw Sources-Eintrag)
- log: updated
**Kernkonzept:**
- Real-World-Showdown als Benchmark-Alternative: Beide Modelle, identischer Agent, identischer Prompt, komplexe Real-App mit Bug + Feature
- Kimi K2.7: schnell (~5 Min), innovativ; GLM-5.2: stark bei Animation/Detail, 1M Kontext, MIT-Lizenz
- "Neck to neck" = beide production-ready, aber Sub-Task-spezifisch optimal
- Direkter Implikat für OpenClaw: Sub-Task-Routing für Coding-Tasks, Ensemble-Pattern mit Coding-Modellen, Real-World-Validierung statt nur Standard-Benchmarks