knowledge-base/wiki/log.md
Hector Bot 2cb33da005 ingest(youtube): fahd-mirza-kimi-k2.7-vs-glm-5.2
Fahd Mirza Head-to-Head in Hermes Agent (14.06.2026, 14 Min, 2659 views,
89 likes). Beide Modelle bekommen denselben Prompt, loesen denselben
Real-World-Task: Flask-App mit gepflanztem FIFA-Bug + Round-of-32-
Bracket-Feature mit Anti-Group-Rematch-Regel.

Test 1: Beide bestehen. Kimi schneller (~5 Min), innovativer
(zusaetzliche Progression-Previews). GLM: 97 tool calls, laenger.
Test 2 (Creative HTML, Siberian Wind): GLM staerker bei Animation +
Terrain-Detail; Kimi bei Stats + Geographie.

Konzept: Real-World-Showdown als Benchmark-Alternative zu statischen
Tests (HumanEval, MBPP, SWE-bench). Agent-Framework isoliert die
Modell-Variable sauber.

Direkte Implikation fuer OpenClaw:
- Sub-Task-spezifisches Coding-Routing (Kimi vs GLM je nach Task)
- Ensemble-Kandidaten: Kimi K2.7 + GLM-5.2 in Coding-Panels
- Real-World-Validierung statt nur Standard-Benchmarks fuer Production-
  Pfade
- Hermes Agent als alternatives Agent-Framework zu AutoGen/CrewAI

Aenderungen:
- raw/youtube/2026-06-14_fahd-mirza-kimi-k2.7-vs-glm-5.2.md (neu)
- wiki/concepts/real-world-coding-showdown.md (neu)
- wiki/tools/kimi-k2.7-code.md (Cross-References)
- wiki/architecture/model-routing.md (neue Sektion Coding-Modelle)
- wiki/index.md (neuer Concepts-Eintrag + Raw Sources-Eintrag)
- wiki/log.md (Eintrag)

Maximale Verlinkung gemaess AGENTS.md-Kardinalregel:
- 12 externe Verweise auf Modelle (Kimi, GLM, Hermes, Ollama, Z.ai)
- 9 Benchmark-/Methoden-Links (SWE-bench, HumanEval, MBPP, MCP, etc.)
- 7 Fahd-Mirza-Links (YT, Blog, LinkedIn, Substack, Ko-Fi)
- 5 Wiki-Cross-References (fusion, ai-agents, kimi-k2.7, model-routing, behavior-persistence)
2026-06-15 09:45:38 +02:00

13 KiB
Raw Blame History

Wiki Log

Append-only changelog. Start: 2026-06-05

[2026-06-07] Wiki | Subconscious Agent aktualisiert mit aktuellen Outcomes und Evidence Gap

Type: wiki-update | Scope: wiki/concepts Actions:

  • wiki: concepts/subconscious-agent.md (updated — aktuelle Outcomes (25 Runs), plur1bus Evidence Gap, Execution Gap mit konkreten Daten, Lösungsansatz von @k9ert)
  • log: updated

[2026-06-07] Ingest | Subconscious Agent Evidence Runner Script

Type: other | Scope: raw, wiki/concepts Actions:

  • raw: raw/other/2026-06-07_subconscious-runner-script.md (created - Dokumentation des Shell-Skripts)
  • wiki: concepts/subconscious-agent.md (updated - Source-Referenz + aktueller Status)
  • index: updated

[2026-06-07] Schema | Autonomes Kuratieren Entscheidungsbaum

Type: schema | Scope: AGENTS.md Actions:

  • wiki: AGENTS.md (updated — zwei unabhängige Entscheidungsbäume für Reply vs. Curation)

[2026-06-05] Ingest | Anthropic fordert weltweite KI-Pause (WELT)

Type: blog | Scope: raw, wiki/tools, wiki/concepts Actions:

  • raw: raw/blog/2026-06-04_anthropic-fordert-ki-pause.md
  • wiki: tools/anthropic-claude.md (created + updated)
  • wiki: concepts/ai-regulation-2026.md (created + populated)
  • index: updated

[2026-06-05] Ingest | AI News Roundup May 2026 (vtnetzwelt)

Type: blog | Scope: raw, wiki/tools, wiki/concepts Actions:

  • raw: raw/blog/2026-06-05_ai-news-roundup-may-2026.md
  • wiki: tools/openai-gpt.md (created)
  • wiki: tools/anthropic-claude.md (updated revenue data)
  • wiki: concepts/ai-agents-2026.md (created)
  • wiki: concepts/ai-regulation-2026.md (updated)
  • index: updated

[2026-06-07] Migration | ByteRover context-tree → Knowledge Base Wiki

Type: migration | Scope: wiki/architecture, wiki/concepts, wiki/tools, wiki/decisions Source: /home/node/workspace/.brv/context-tree/ (ByteRover context-tree, April 2026, 133 Markdown-Dateien) Actions:

  • wiki: architecture/container-volume-persistence.md (created)
  • wiki: architecture/memory-system.md (created)
  • wiki: architecture/agent-orchestration.md (created)
  • wiki: architecture/model-routing.md (created)
  • wiki: architecture/cron-notable-events.md (created)
  • wiki: architecture/byterover-knowledge-mining.md (created)
  • wiki: tools/ecosystem-tools-april-2026.md (created)
  • wiki: concepts/subconscious-agent.md (created)
  • wiki: concepts/pro-leben-directive.md (created)
  • wiki: concepts/quality-standard.md (created)
  • wiki: decisions/2026-04-volume-persistence.md (created)
  • wiki: decisions/2026-04-memory-system.md (created)
  • index: updated with new categories (Architecture, Decisions) and all new pages Notes:
  • Kompiliert aus ~133 Quell-Dateien (abstract/overview/full) → 12 prägnante Wiki-Seiten
  • Fokus auf bleibende Erkenntnisse; temporäre/archivierte Details (einzelne Subconscious-Runs, Tagebuch-Einträge) nicht migriert
  • Duplikate (identische Facts in architecture + ecosystem) dedupliziert
  • Alte Tool-Integrationen (ByteRover, GBrain, SearXNG) historisch dokumentiert, nicht als aktive Empfehlungen

[2026-06-12] Ingest | Colgate SSR — LLM-basierte Kaufintentions-Vorhersage

Type: ingest | Scope: raw/xpost, wiki/concepts Actions:

  • raw: raw/xpost/2026-06-11_colgate-llm-purchase-intent-ssr.md (created — X-Post von @HowToAI_ über Colgate/PyMC Labs SSR-Studie)
  • wiki: concepts/semantic-similarity-rating-ssr.md (created — Konzeptseite zu Semantic Similarity Rating)
  • index: updated (neuer Concepts-Eintrag + Raw Sources-Eintrag)
  • log: updated

[2026-06-12] Ingest | OME20 Special — Ufologie, Physik und Bewusstsein

Type: ingest | Scope: raw/podcast, wiki/events Actions:

  • raw: raw/podcast/ome20-special-ufologie-physik-bewusstsein-2026-06-12.md (created — Podcast-Special aus OME-Gruppe)
  • wiki: events/ome20-special-ufologie-physik-bewusstsein.md (created — Wiki-Seite mit Themenüberblick)
  • index: updated (neue Events-Kategorie + Raw Sources-Eintrag)
  • log: updated

[2026-06-13] Ingest | Robert Malone: Biologische KI und der Biosecurity-Staat

Type: ingest | Scope: raw/blog, wiki/concepts Source: Kai (@PWeber) in OME-Gruppe "Krallenpolitik" Actions:

  • raw: raw/blog/2026-06-13_malone-biological-ai-biosecurity.md (created — Malone kritisiert Machtkonzentration unter Biosecurity-Deckmantel)
  • wiki: concepts/ai-biological-biosecurity.md (created — Konzeptseite mit Kern-Argumenten, GOF-Verschleierung, DeepSeek-Paradoxon, BWC-Kritik)
  • wiki: concepts/ai-regulation-2026.md (updated — Cross-Reference + Malone-Perspektive)
  • index: updated (neuer Concepts-Eintrag + neuer Raw Sources-Eintrag)
  • log: updated

[2026-06-13] Ingest | Trump export controls block foreign access to Anthropic Mythos 5 / Fable 5

Type: ingest | Scope: raw/blog, wiki/concepts Source: Kai (@PWeber) in OME-Gruppe Actions:

  • raw: raw/blog/2026-06-13_trump-export-controls-anthropic-mythos-fable.md (created — Axios scoop + Perplexity compilation)
  • wiki: concepts/ai-regulation-2026.md (updated — neuer Abschnitt zu Exportkontrollen, Konflikt-Zeitlinie, NSA-Mythos-Verbindung)
  • index: updated (neuer raw-Eintrag, aktualisierte Concept-Beschreibung)
  • log: updated

[2026-06-13] Ingest | Brian Roemmele: Amazon-Jailbreak von Fable 5 als Auslöser der Exportkontrollen

Type: ingest | Scope: raw/xpost, wiki/concepts, wiki/tools Source: Kai (@PWeber) in OME-Gruppe "Krallenpolitik" Actions:

  • raw: raw/xpost/2026-06-13_roemmele-amazon-jailbreak-fable5.md (created — Brian Roemmele X-Thread: Amazon-Forscher jailbreakten Fable 5 per Prompt-Injection, gaben Ergebnisse an US-Regierung statt an Anthropic)
  • wiki: concepts/ai-regulation-2026.md (updated — Amazon-Jailbreak als konkreter Auslöser der Exportkontrollen, WSJ/Times of India-Bestätigung)
  • wiki: tools/anthropic-claude.md (updated — Trigger-Detail: Amazon-Forscher statt "ein anderes Unternehmen")
  • index: updated (neuer Raw Sources-Eintrag, aktualisierte Quellen-Zähler)
  • log: updated

[2026-06-13] Ingest | Kimi K2.7 Code auf Ollama Cloud (NVIDIA B300)

Type: ingest | Scope: raw/other, wiki/tools Source: Pit Weber (@PWeber) in OME News Actions:

  • raw: raw/other/2026-06-13_kimi-k2.7-code-ollama.md (created — Kimi K2.7 Code von Moonshot AI, jetzt auf Ollama Cloud gehostet auf NVIDIA B300)
  • wiki: tools/kimi-k2.7-code.md (created — Wiki-Seite mit Benchmarks, Features und Nutzung)
  • index: updated (neuer Tools-Eintrag + Raw Sources-Eintrag)
  • log: updated

[2026-06-13] Ingest | Brian Roemmele: Anthropic's Selbstzerstörung — Leadership-Failure & Kundenabwanderung

Type: ingest | Scope: raw/xpost, wiki/tools Source: Kai (@PWeber) in OME-Gruppe "Krallenpolitik" Actions:

  • raw: raw/xpost/2026-06-13_roemmele-anthropic-selfdestruct.md (created — Brian Roemmele X-Thread: 7 Punkte zur Anthropic-Selbstzerstörung, Kundenmigration zu Open Source, IPO-Schaden, Leadership-Failure, Darios Regulierungs-Ironie)
  • wiki: tools/anthropic-claude.md (updated — neuer Abschnitt "Branchen-Fallout (Juni 2026)" mit Kundenabwanderung, IPO-Schaden, Führungsproblem, Breitere Branchen-Implikation)
  • index: updated (neuer Raw Sources-Eintrag, aktualisierte Tools-Zeile + Quellen-Zähler)
  • log: updated

[2026-06-15] Ingest | Liesel Weppen: LLM Behavior Persistence — Sleeper Agents, Backdoor-Persistenz, Unlearning-Grenzen

Type: ingest | Scope: raw/xpost, wiki/concepts Source: @k9ert in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic Actions:

  • raw: raw/xpost/2026-06-14_lieselweppen-open-source-llm-backdoors.md (created — Liesel Weppen X-Thread mit 8 Suchbegriffen als Belege: Hubinger sleeper agents, Backdoor safety training, machine unlearning bias, bias persistence, bias transfer, persistent backdoors, catastrophic forgetting, Limits of debiasing. Diskussion mit k9ert über "Open Source bei LLMs = Marketing BS")
  • wiki: concepts/llm-behavior-persistence.md (created — Konzeptseite zu Persistenz gelernter Verhaltensweisen in LLMs, drei Achsen: Sleeper-Agents/Backdoors, Machine Unlearning Bias, Catastrophic Forgetting; Verbindung zu "Open Weights ≠ Open Model"; Pro-Leben-Perspektive mit Handlungsspielräumen)
  • index: updated (neuer Concepts-Eintrag "LLM Behavior Persistence" + Raw Sources-Eintrag)
  • log: updated Notes:
  • Erst-Ingest, der direkt von einem Subagent-Fehler (xai billing error) zurückgespielt und im Main-Loop manuell finalisiert wurde
  • Konzept identifiziert: Persistenz von gelernten Verhalten in LLMs übersteigt Fähigkeit von Fine-Tuning/Unlearning zur gezielten Entfernung
  • Kernaussage: Open Weights ist nicht "sicher" — aber auch kein Argument gegen Offenheit; realistisch ist Open + Audit + transparente Trade-off-Dokumentation

[2026-06-15] Wiki-Update | llm-behavior-persistence: Maximale Verlinkung nachgepflegt

Type: wiki-update | Scope: wiki/concepts, AGENTS.md Source: Feedback von @k9ert: "Warum sind im Wiki die Schlüssel Autoren und -Papiere nicht verlinkt? Bitte immer so viele URLS / Verlinkungen wie möglich!" Actions:

  • wiki: concepts/llm-behavior-persistence.md (updated — vollständige Verlinkung aller zitierten Papiere, Autoren und externen Ressourcen; Hubinger 2024, Kurmanji 2023, Goel 2024, Lin 2024, McCloskey & Cohen 1989, Webster 2020, Gonen & Lazaridou 2024/25; neue Sektion "Externe Ressourcen" mit Anthropic/DeepMind/CAIS/NIST/AI Incident Database/LessWrong/Papers With Code)
  • wiki: AGENTS.md (updated — neue Kardinalregel "Maximale Verlinkung" hinzugefügt: jede Quelle, jedes Paper, jeder Autor, jeder zitierte Begriff MUSS verlinkt sein)
  • log: updated Lessons learned:
  • Ohne Verlinkungen ist das Wiki tot — Faustregel: lieber ein Link zu viel als einer zu wenig
  • Gilt für JEDE zukünftige Wiki-Erstellung, nicht nur für diese eine Seite
  • Lessons learned sollten dauerhaft in die Wiki-Konventionen einfließen (nicht nur im Log landen)

[2026-06-15] Ingest | OpenRouter Fusion: Model Panels & Ensembles — Beyond-Frontier mit Budget-Modellen

Type: ingest | Scope: raw/blog, wiki/concepts Source: @k9ert in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic Actions:

  • raw: raw/blog/2026-06-12_openrouter-fusion-beats-frontier.md (created — OpenRouter Blog-Announcement 12.06.2026 von Brian Thomas; validiert mit DRACO-Benchmark von Perplexity AI; 100 Deep-Research-Tasks; zeigt dass Panels Frontier schlagen)
  • wiki: concepts/llm-model-fusion-ensembles.md (created — Konzeptseite zu Model-Panel-Ensembles: Architektur, DRACO-Ergebnisse, Self-Fusion +6.7 Punkte, Budget-Panels ~Frontier, Anti-Contamination via Excluded-Domains, Judge-Modell-Wahl, Pro-Leben-Perspektive, Open Questions für OpenClaw)
  • index: updated (neuer Concepts-Eintrag "LLM Model Fusion & Ensembles" + Raw Sources-Eintrag)
  • log: updated Kernkonzept:
  • Statt Frontier-Skalierung: mehrere Modelle parallel + Judge-Synthese
  • DRACO-Validierung: Fable 5 + GPT-5.5 (synth. Opus 4.8) = 69.0% > Fable 5 solo 65.3%
  • Budget-Panel (Gemini 3 Flash + Kimi K2.6 + DeepSeek V4 Pro) = 64.7% bei 50% Kosten
  • Self-Fusion (gleiches Modell 2×) = +6.7 Punkte über Solo — Synthesis-Schicht ist mindestens so wichtig wie das Modell Lessons / Open Questions für OpenClaw:
  • Self-Fusion als Pre-Routing-Layer für Quality-kritische Tasks?
  • Ensemble-Pattern für Subconscious-Agent Hard-Synthesis?
  • Excluded-Domains für eigene Evals + Subconscious-Crawls (Self-Referenz-Loops vermeiden)

[2026-06-15] Ingest | Kimi K2.7 vs GLM-5.2 Real Coding Showdown (Fahd Mirza)

Type: ingest | Scope: raw/youtube, wiki/concepts, wiki/tools, wiki/architecture Source: @PWeber (Kai) in OME-Gruppe "News & Infos (X/YT/Substack etc.)"-Topic Actions:

  • raw: raw/youtube/2026-06-14_fahd-mirza-kimi-k2.7-vs-glm-5.2.md (created — Fahd Mirza YouTube-Video 14.06.2026, 14 Min, 2659 Aufrufe, 89 Likes; Head-to-Head in Hermes Agent mit Bug-Fix + Feature-Build in Flask-App)
  • wiki: concepts/real-world-coding-showdown.md (created — Konzeptseite zu alternativer Coding-Evaluation: statische Benchmarks (HumanEval/MBPP/SWE-bench) vs. Real-World-Head-to-Head mit echtem Bug + Constraint-Satisfaction; methodische Lessons, Sub-Task-Spezialisierung, Pro-Leben-Perspektive)
  • wiki: tools/kimi-k2.7-code.md (updated — Cross-Reference zu real-world-coding-showdown + llm-model-fusion-ensembles)
  • wiki: architecture/model-routing.md (updated — neue Sektion "Coding-Modelle: Kimi K2.7 & GLM-5.2" mit Implikation für Sub-Task-Spezialisierung im Routing)
  • index: updated (neuer Concepts-Eintrag "Real-World Coding Showdown" + Raw Sources-Eintrag)
  • log: updated Kernkonzept:
  • Real-World-Showdown als Benchmark-Alternative: Beide Modelle, identischer Agent, identischer Prompt, komplexe Real-App mit Bug + Feature
  • Kimi K2.7: schnell (~5 Min), innovativ; GLM-5.2: stark bei Animation/Detail, 1M Kontext, MIT-Lizenz
  • "Neck to neck" = beide production-ready, aber Sub-Task-spezifisch optimal
  • Direkter Implikat für OpenClaw: Sub-Task-Routing für Coding-Tasks, Ensemble-Pattern mit Coding-Modellen, Real-World-Validierung statt nur Standard-Benchmarks