ingest(youtube): Matthew Berman - 90% less AI costs via Model Routing
- raw: raw/youtube/2026-07-07_berman-model-routing.md - wiki/tools/model-routing.md (new) — Cost-saving routing patterns - wiki/architecture/model-routing.md (update) — Berman patterns section - wiki/index.md — 67. Update with new tool + architecture update - wiki/log.md — changelog entry
This commit is contained in:
parent
58b27062d5
commit
8f676acfa7
10 changed files with 611 additions and 8 deletions
|
|
@ -0,0 +1,244 @@
|
|||
---
|
||||
type: youtube
|
||||
source_url: https://www.youtube.com/watch?v=-JCCcR9qtYQ
|
||||
retrieved: 2026-07-05
|
||||
channel: "Everlast AI"
|
||||
duration_sec: 1459
|
||||
has_transcript: true
|
||||
title: "KI-News: Chinas KI-Roboter werden ZU ECHT! + DAS kann das NEUE Fable 5 & Sonnet 5"
|
||||
tags: [ki-news, humanoid-robots, ubtech, u1, china, robotics, sonnet-5, fable-5, anthropic, remote-labor-index, zero-person-company, meta-pocket, claude-science, verticalization, hixfield, figma-mcp, cost-to-run]
|
||||
---
|
||||
|
||||
# KI-News: Chinas KI-Roboter werden ZU ECHT! + DAS kann das NEUE Fable 5 & Sonnet 5
|
||||
|
||||
**Kanal:** Everlast AI (317K subscribers)
|
||||
**Host:** Leonard "Leo" Schmedding (aus Hongkong)
|
||||
**Upload:** 2026-07-05
|
||||
**Länge:** 24:19
|
||||
**Views:** 13.321 (7h nach Upload)
|
||||
**Likes:** 540
|
||||
|
||||
## Video-Beschreibung
|
||||
|
||||
China veröffentlicht fast schon zu menschliche Humanoide Roboter, Anthropic lanciert Sonnet 5 und Claude Fable 5 ist zurück: Den realen Praxistest und alle relevanten KI-News im Geschäftskontext gibt's wie immer in den KI-News der Woche auf deutsch – heute aus Hongkong!
|
||||
|
||||
## Chapters
|
||||
|
||||
| Zeit | Thema |
|
||||
|------|-------|
|
||||
| 0:00 | Worum geht es? |
|
||||
| 1:05 | UBTECH U1 |
|
||||
| 2:48 | Robotik Push China |
|
||||
| 4:35 | Robotik Investments |
|
||||
| 5:05 | Claude Sonnet 5 |
|
||||
| 6:02 | Token Kosten Realität |
|
||||
| 7:04 | Fable 5 Comeback |
|
||||
| 8:10 | Fable Design Workflow |
|
||||
| 12:51 | Praxis Fazit |
|
||||
| 14:02 | Fable Fallback deaktivieren |
|
||||
| 14:59 | Remote Labor Index |
|
||||
| 18:09 | Lokale KI Strategie |
|
||||
| 19:08 | Claude Science |
|
||||
| 19:46 | Zero Person Company |
|
||||
| 21:35 | Meta Pocket |
|
||||
| 22:37 | Claude Entwicklung Übersicht |
|
||||
| 22:57 | Agentic Coding Chancen |
|
||||
| 24:06 | Empfehlung |
|
||||
|
||||
## Transcript Key Excerpts
|
||||
|
||||
### 1. UBTECH U1 — Hyperrealistische Humanoide Roboter (1:05)
|
||||
|
||||
UBTECH veröffentlicht den U1 Humanoid Roboter mit hyperrealistischer Silikonhaut für "emotionale KI" — für Gespräche und Blickkontakt, bis zu 88 Freiheitsgrade. Können tanzen, lächeln, wirken "fast wie lebendige Fantasy Figuren aus Zelda".
|
||||
|
||||
**Launch in Shenzhen:** Über 13.000 Vorbestellungen — sofort ausverkauft.
|
||||
|
||||
**Preise:**
|
||||
- Light Modell: ab $17.600
|
||||
- Ultra Varianten: bis $45.000
|
||||
- Nur ab 18 Jahren erhältlich
|
||||
|
||||
**Spezifikationen:**
|
||||
- Männliche Varianten: 1,83 m (lebensgroß)
|
||||
- Weibliche Varianten: 1,68 m
|
||||
- 2–4 Stunden Akkulaufzeit
|
||||
- Cloud-KI-Interaktion
|
||||
- Über 50 Varianten auf der Bühne in Shenzhen gezeigt
|
||||
- Fantasy-Outfits, Bewegung, Tanzen
|
||||
|
||||
**Unternehmen:** UBTECH ist das erste börsennotierte Humanoide-Unternehmen. Planen nicht nur Verkauf, sondern spenden auch 100 Roboter 2026 — Zeichen für staatlichen/wirtschaftlichen Push Chinas in Richtung Alltagsrobotik.
|
||||
|
||||
### 2. Robotik Push China — Demografischer Treiber (2:48)
|
||||
|
||||
Hong Kong hat die **niedrigste Geburtenrate weltweit: 0,77** (eine Frau bekommt im Durchschnitt nicht mal ein Kind). Gleichzeitig eine der höchsten Lebenserwartungen: Männer ~83 Jahre, Frauen ~88 Jahre.
|
||||
|
||||
**Vergleich:** Deutschland Geburtenrate 1,45 — "nicht viel besser".
|
||||
|
||||
Hong Kong ist gleichzeitig:
|
||||
- Die reichste Stadt der Welt
|
||||
- Die teuerste Stadt weltweit
|
||||
- Einer der wichtigsten Naturhäfen weltweit
|
||||
- Braindrain: Junge Menschen wandern ab
|
||||
|
||||
**Schlussfolgerung:** Der Robotik-Push kommt nicht von ungefähr. Einsamkeit ist einer der größten Use Cases. KI und Robotik sind demografisch nicht mehr optional.
|
||||
|
||||
### 3. Robotik Investments — Rekordhoch (4:35)
|
||||
|
||||
Venture Capital für Robotik explodiert:
|
||||
- Letztes Quartal: **$16,2 Milliarden** in Robotic Startups
|
||||
- Normal: $3–5 Milliarden pro Quartal
|
||||
- Mehr als 3× des Normalniveaus
|
||||
- Im Vergleich zum KI-Boom ist Robotik immer noch unterrepräsentiert → enorme Chancen
|
||||
|
||||
### 4. Claude Sonnet 5 — Neues Anthropic Flagship (5:05)
|
||||
|
||||
Anthropic lanciert Claude Sonnet 5:
|
||||
- **1 Million Kontextfenster**
|
||||
- Performance etwa auf Opus 4.8 Niveau ("vielleicht ein bisschen günstiger")
|
||||
- Soll besser sein in: Reasoning, Tool Use, Coding, Wissensarbeit
|
||||
- Neues Default-Modell für alle gratis und Pro User
|
||||
- Auch in Claude Code verfügbar
|
||||
- Standardpricing
|
||||
- Offiziell "sicherer als Sonnet 4.6"
|
||||
|
||||
**Sprung:** Von 4.6 direkt auf 5 — größer als übliche Versionssprünge.
|
||||
|
||||
### 5. Token Kosten Realität — Sonnet 5 ineffizient (6:02)
|
||||
|
||||
**Kritik:** Sonnet 5 ist ein "enorm ineffizientes Modell". Laut Artificial Analysis Cost to Run Index ist Sonnet 5 **teurer als Claude Fable 5** — das widerspricht dem Sinn eines leichtgewichtigeren Modells.
|
||||
|
||||
**Cost to Run:** ~$6.000 im Index — GPT 5.5 Extra Hype kostet nicht mal die Hälfte.
|
||||
|
||||
**Problem:** Man darf nicht auf offizielle Kostenangaben schauen — wenn der Reasoning-Prozess ineffizient ist und viele Output-Tokens generiert werden, sind die realen Kosten viel höher.
|
||||
|
||||
### 6. Fable 5 Comeback (7:04)
|
||||
|
||||
Fable 5 ist offiziell wieder verfügbar. Demo-Vergleiche zu Sonnet und Opus 4.8 zeigen: "Fable 5 ist nach wie vor ein unfassbar starkes Modell."
|
||||
|
||||
**Verfügbarkeit:**
|
||||
- Fable 5 Low ist günstiger, besser und schneller als Opus 4.8 Max
|
||||
- Bis zum 7. Juli in normalen Plänen verfügbar
|
||||
- Danach: Wechsel zu "usage credits" System (teurer)
|
||||
|
||||
**Hixfield Integration:** Fable 5 kann über Hixfield MCP für Erklärvideos genutzt werden (Hixfield Explainer).
|
||||
|
||||
**Einschränkungen:** Restriktionen wie beim ersten Release. Bei "normalen" Coding-Aufgaben wird standardmäßig zu Opus 4.8 geforwarded. Anthropic sagte später, die Ankündigung sei "etwas missverständlich" gewesen — man ruderte zurück.
|
||||
|
||||
### 7. Fable Design Workflow — Praxis-Test mit Figma MCP (8:10)
|
||||
|
||||
**Senior Developer Marcel** demonstriert Fable 5 in der Praxis:
|
||||
|
||||
**Task:** Frontend-Design für eine native App (Corporate LM) via Figma MCP.
|
||||
|
||||
**Setup:**
|
||||
- GitHub Epic/Umbrella Issue mit Unter-Aufgaben
|
||||
- Fable 5 mit "Ultra Code" Effort (multiple parallele Agenten in Figma)
|
||||
- Prompt-basiert, keine Screenshots
|
||||
|
||||
**Ergebnis nach 1,5 Stunden:**
|
||||
- Komplettes Cover erstellt
|
||||
- GitHub Issue referenziert
|
||||
- Erkannt: Tauri-Applikation (macOS + Windows)
|
||||
- Erkannt: Corporate LM nutzt Satoshi Font → Hinweis, Inter in Figma durch Satoshi zu ersetzen
|
||||
- Foundations: Farbspektrum, Schriftgrößen, Gewichtungen, Komponenten aus Webapp-Code extrahiert
|
||||
- Flows: Installation (macOS/Windows), Welcome Screen, Browser-Authentifizierung
|
||||
- App Shell: Alle Fenster, Empty States + befüllte States
|
||||
- Kernscreens 1:1 aus Webapp übernommen — nur Code als Basis, keine Screenshots
|
||||
|
||||
**Vergleich:** Früher benötigte ein gesamtes Designteam Wochen bis Monate für diese ersten zwei Bereiche. Fable 5: 1,5 Stunden.
|
||||
|
||||
**Native App Features:** Lokale LLM-Installation direkt auf dem Rechner als Alleinstellungsmerkmal der nativen App.
|
||||
|
||||
### 8. Fable Fallback deaktivieren (14:02)
|
||||
|
||||
**How-To:** Claude Settings → Fähigkeiten → "Modell wechseln, wenn eine Nachricht markiert wird" → ausschalten.
|
||||
|
||||
**Effekt:** Chat wird pausiert statt zu Opus 4.8 weitergeleitet. Verhindert Token-Verschwendung bei Aufgaben, die explizit von Fable 5 gelöst werden sollen.
|
||||
|
||||
### 9. Remote Labor Index (14:59)
|
||||
|
||||
**These:** Der "Remote Turing Test" wird dieses Jahr bestanden — man kann bei Freelance-Projekten (z.B. Fiverr) nicht mehr unterscheiden, ob eine KI oder ein Mensch die Arbeit erledigt.
|
||||
|
||||
**Remote Labor Index:**
|
||||
- Misst, wie gut KI-Modelle reale Freelance-Aufträge abwickeln
|
||||
- 240 verschiedene Projekte: Grafikdesign, Architektur, CAD, Video, Audio, Data Analysis, Webdevelopment
|
||||
- CAD ist ein "riesen Use Case" für Fable 5
|
||||
|
||||
**Assessment:** An manchen Stellen ist es "wahrscheinlich schon Realität".
|
||||
|
||||
### 10. Fable 5 Community-Feedback
|
||||
|
||||
**Durchwachsen:**
|
||||
- Einige sagen, Fable 5 sei vor dem US-Ban besser gewesen als nach der Wiederkunft
|
||||
- Andere: "viel viel schlechter" — fordern Erklärung von Anthropic
|
||||
- AI Arena (seriöser als selbstgebastelte Benchmarks): Fable 5 nach Re-Release in einigen Bereichen **besser** (Dokumente, Creative Writing)
|
||||
- Für den Massenmarkt sind keine großen Sprünge mehr spürbar — "die Modelle sind mittlerweile schon so gut"
|
||||
|
||||
**Zukunftsprognose:** Anthropic könnte einen $500 oder $1.000 Plan einführen — "dann könnte sich das durchaus mehr lohnen als die Standard $200 Pläne".
|
||||
|
||||
### 11. Lokale KI Strategie (18:09)
|
||||
|
||||
Palantir CEO Alex Karp wies darauf hin, dass die US-Regierung teilweise mit Open-Source-Modellen (z.B. Nemotron) arbeitet.
|
||||
|
||||
**Empfehlung:** Man braucht eine "lokale KI Backup-Versicherung" — sich nie rein auf Cloud-Modelle verlassen. Sicherheitsnetz durch lokale Modelle, wenn Cloud-Modelle gesperrt werden.
|
||||
|
||||
**Claude Code Artefakte:** Jetzt in jedem Plan verfügbar (zuvor nur Teams/Enterprise).
|
||||
|
||||
### 12. Claude Science (19:08)
|
||||
|
||||
Anthropic steigt offiziell in **Medikamentenentwicklung / Drug Development** ein. Veröffentlichung von "Claude Science".
|
||||
|
||||
**Logik:** Nachdem Mathematik, Physik und Coding "gelöst" sind, kommt Biologie/Medikamentenentwicklung als nächste Disziplin — und letztlich jede andere relevante Disziplin.
|
||||
|
||||
### 13. Zero Person Company (19:46)
|
||||
|
||||
**Matrix** veröffentlicht "Zero Person Company" als KI-Tool:
|
||||
- "Runtime for Self-Evolving Multi-Agent Orchestration"
|
||||
- Möglichkeit, eine ganze Firma zu lancieren (limitierte Beta)
|
||||
- Ein-Mann-Unternehmen mit KI aus dem Boden stampfen — braucht nicht mal Matrix, nur eine gute Geschäftsidee
|
||||
|
||||
### 14. Meta Pocket (21:35)
|
||||
|
||||
Meta lanciert **Pocket**: Marktplatz für vibecodete Spiele und Apps.
|
||||
|
||||
**Gedanke:** Wie schafft man es, Apps die man mit Agentic Coding baut, für alle bereitzustellen? Nächster Meilenstein — Marcel's Corporate LM Relation Flow wird ein ähnlicher Marktplatz.
|
||||
|
||||
### 15. Claude Vertikalisierung (22:37)
|
||||
|
||||
**Bestand:** Claude Code, Claude Cowork, Claude Design, Claude Finance, Claude Science
|
||||
|
||||
**Noch fehlend:** Claude HR, Claude Analytics, Claude Marketing, Claude Sales, Claude Legal, Claude Logistics, Claude CAD, Claude R&D, Claude Accounting
|
||||
|
||||
**Trend:** Vertikalisierung von KI-Anwendungen. Überlegung: In welchem Bereich hat man Expertise? Kunden gewinnen, Anwendung bauen.
|
||||
|
||||
**Agentic Coding:** Riesen Thema, kaum jemand im deutschsprachigen Markt bedient es. Wer Apps programmieren kann + mit Fable umgehen kann, ist händeringend gefragt.
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
1. **China's Robotik-Push ist demografisch getrieben** — Hong Kong's Geburtenrate von 0,77 macht Humanoide Roboter zur Notwendigkeit, nicht zur Spielerei
|
||||
2. **UBTECH U1** ist der erste massenmarktreife humanoide Roboter mit emotionaler KI-Ausrichtung — 13K Vorbestellungen bei $17.600–$45.000
|
||||
3. **Robotik-VC explodiert** ($16,2 Mrd./Quartal, 3× normal) aber immer noch unterrepräsentiert vs. KI
|
||||
4. **Claude Sonnet 5** ist ein neues Default-Modell mit 1M Kontext, aber ineffizient im Cost-to-Run (teurer als Fable 5)
|
||||
5. **Fable 5 Praxis-Test:** Figma MCP Frontend-Prototyping in 1,5h statt Wochen — Fable 5 Low ist günstiger/besser/schneller als Opus 4.8 Max
|
||||
6. **Remote Turing Test** wird 2026 bestanden — 240 Freelance-Projekte, KI nicht mehr von Menschen unterscheidbar
|
||||
7. **Anthropic expands into Drug Development** mit Claude Science — nächste Disziplin nach Coding/Math/Physik
|
||||
8. **Zero Person Company** (Matrix) + Meta Pocket = Trend zu KI-orchestrierten Unternehmen und App-Marktplätzen
|
||||
9. **Vertikalisierung** ist der nächste große Trend: Claude Code → Claude HR/Marketing/Sales/Legal/CAD/etc.
|
||||
10. **Lokale KI als Backup-Versicherung** bleibt empfohlen — niemals nur auf Cloud-Modelle verlassen
|
||||
|
||||
## External Links
|
||||
|
||||
- [YouTube Video](https://www.youtube.com/watch?v=-JCCcR9qtYQ)
|
||||
- [Everlast AI Kanal](https://www.youtube.com/@everlastai)
|
||||
- [Leonard Schmedding Zweitkanal](https://www.youtube.com/@LeonardSchmedding)
|
||||
- [Kiberatung.de](https://www.kiberatung.de/)
|
||||
- [AI Profit Boardroom](https://everlastkarriere.de/)
|
||||
|
||||
## Wiki Context
|
||||
|
||||
- [[../../wiki/concepts/llm/fable-5-anthropic.md]] — Fable 5 Konzeptseite (wird aktualisiert)
|
||||
- [[../../wiki/tools/anthropic-claude.md]] — Anthropic Claude Modellübersicht (wird aktualisiert)
|
||||
- [[../../wiki/concepts/hardware/neuromorphic-chips-und-quantencomputer.md]] — Hardware-Frontier
|
||||
- [[../../wiki/concepts/llm/coding-benchmark-price-performance.md]] — Cost-to-Run Benchmark
|
||||
- [[../../wiki/concepts/llm/vibe-coding-vs-enterprise.md]] — Vibe Coding / Agentic Coding
|
||||
- [[../../wiki/institutions/anthropic.md]] — Anthropic Institutionenseite
|
||||
81
raw/youtube/2026-07-07_berman-model-routing.md
Normal file
81
raw/youtube/2026-07-07_berman-model-routing.md
Normal file
|
|
@ -0,0 +1,81 @@
|
|||
---
|
||||
type: youtube
|
||||
source_url: https://www.youtube.com/watch?v=1KKB_UiW6ls
|
||||
retrieved: 2026-07-07
|
||||
channel: "Matthew Berman"
|
||||
title: "You NEED to do this right now... — 90% less AI costs via Model Routing"
|
||||
duration_sec: 1113
|
||||
has_transcript: false
|
||||
tags: [model-routing, cost-optimization, fable, planning-execution-split, cross-model-calling, copy-paste-routing, cursor-auto-mode, not-diamond, genspark, coinbase, glm-5.2, gpt-5.5, composer-2.5, claude-sonnet, cost-savings, routing-strategy, audio-briefing, german]
|
||||
notebooklm: https://notebooklm.google.com/notebook/a5071bba-125f-4f08-86ea-99a1ab9911ca
|
||||
---
|
||||
|
||||
# 90% less AI costs — Model Routing
|
||||
|
||||
**Video URL:** [https://www.youtube.com/watch?v=1KKB_UiW6ls](https://www.youtube.com/watch?v=1KKB_UiW6ls)
|
||||
**NotebookLM:** [https://notebooklm.google.com/notebook/a5071bba-125f-4f08-86ea-99a1ab9911ca](https://notebooklm.google.com/notebook/a5071bba-125f-4f08-86ea-99a1ab9911ca)
|
||||
**Channel:** Matthew Berman (621K subscribers)
|
||||
**Published:** 2026-07-06
|
||||
**Duration:** 18:33 (18:46 Audio Briefing auf Deutsch)
|
||||
**Sponsor:** Genspark
|
||||
|
||||
## Summary
|
||||
|
||||
Matthew Berman demonstrates how Model Routing — the practice of sending different subtasks to different AI models based on cost and capability — can reduce AI costs by up to 90%. The core insight: frontier models like Fable (Anthropic, $10-50/M tokens) are overkill for most tasks; cheaper models (GPT 5.5 at $2/M, GLM 5.2 at $0.08/M, etc.) handle the majority of work within a bounded spec.
|
||||
|
||||
## Key Points
|
||||
|
||||
### 1. Planning vs. Execution Separation (Fable Pattern)
|
||||
|
||||
**Fable (Anthropic)** is used for architecture and specification design (thinking/planning phase), then the actual code **execution** is handed off to cheaper models:
|
||||
- **GPT 5.5** ($2/M tokens input)
|
||||
- **Composer 2.5** (OpenAI code model)
|
||||
- **Claude Sonnet** (mid-tier Anthropic model)
|
||||
|
||||
This is the single most impactful pattern: **split the expensive thinking from the cheap execution**.
|
||||
|
||||
### 2. Dramatic Cost Differences
|
||||
|
||||
| Model | Input Cost per M Tokens | Output Cost per M Tokens |
|
||||
|-------|----------------------|-----------------------|
|
||||
| Fable 5 (Anthropic) | $10 | $50 |
|
||||
| GPT 5.5 | $2 | $6 |
|
||||
| GLM 5.2 | ~$0.08 | ~$0.08 |
|
||||
| Cheap models (general) | $2 | $6 |
|
||||
|
||||
**Example calculation:** Using Fable for everything costs $9.50 baseline. With routing to GPT 5.5 for execution: **$6.48 — a 68% savings.** Overall potential exceeds **90%** when routing aggressively.
|
||||
|
||||
### 3. Routing Patterns Covered
|
||||
|
||||
| Pattern | Description | Est. Savings |
|
||||
|---------|-------------|-------------|
|
||||
| **Manual Copy-Paste** | Developer manually copies code between models | ~60% |
|
||||
| **Cross-Model-Calling** | One model calls another model via API | ~70% |
|
||||
| **Cursor Auto Mode** | Cursor IDE auto-routes to cheapest model | ~75% |
|
||||
| **Not Diamond** | Dedicated routing layer/API | ~80% |
|
||||
|
||||
### 4. Coinbase Example
|
||||
|
||||
Coinbase is implementing model routing on **open-source models, specifically GLM 5.2** (Z.ai). This demonstrates enterprise adoption of the routing pattern with Chinese open-weight models as the cost-effective execution layer.
|
||||
|
||||
### 5. Audio Briefing
|
||||
|
||||
A **German-language audio briefing** (18:46 min) was generated from this video via NotebookLM, making the content accessible to German-speaking audiences.
|
||||
|
||||
## Key Takeaways
|
||||
|
||||
1. **Don't use one model for everything** — separate planning (expensive) from execution (cheap)
|
||||
2. **Fable is for architecture/specs**, not for writing boilerplate code
|
||||
3. **Routing can save 60-90%** with minimal quality loss when execution is well-specified
|
||||
4. **Multiple routing patterns exist** — from manual copy-paste to dedicated routing layers (Not Diamond)
|
||||
5. **Enterprise adoption** is happening (Coinbase on GLM 5.2)
|
||||
6. **The pattern is simple** but requires discipline to implement consistently
|
||||
|
||||
## Relevance to Existing Wiki
|
||||
|
||||
- **Cross-ref:** [[../../wiki/architecture/model-routing.md]] — OpenClaw's existing routing architecture
|
||||
- **Cross-ref:** [[../../wiki/concepts/llm/chinese-model-cost-routing.md]] — DeRonin's 87% cost-cut playbook (complementary data)
|
||||
- **Cross-ref:** [[../../wiki/concepts/llm/coding-benchmark-price-performance.md]] — Price-performance validation (39× cheaper GLM 5.2)
|
||||
- **Cross-ref:** [[../../wiki/concepts/llm/glm-5.2-zai-coding-model.md]] — GLM 5.2 details
|
||||
- **Cross-ref:** [[../../wiki/concepts/llm/fable-5-anthropic.md]] — Fable 5 details
|
||||
- **Cross-ref:** [OpenClaw Smart Model Router skill] — OpenClaw Smart Model Router skill (tier-based routing, 60-90% savings claim)
|
||||
|
|
@ -1,8 +1,8 @@
|
|||
---
|
||||
created: 2026-06-16
|
||||
updated: 2026-06-16
|
||||
sources: [other/2026-06-16_openclaw-releases-v2026.6.8.md]
|
||||
tags: [architecture, model-routing, llm, openclaw]
|
||||
updated: 2026-07-07
|
||||
sources: [other/2026-06-16_openclaw-releases-v2026.6.8.md, youtube/2026-07-07_berman-model-routing.md]
|
||||
tags: [architecture, model-routing, llm, openclaw, planning-execution-split, cost-savings]
|
||||
---
|
||||
|
||||
# Model Routing
|
||||
|
|
@ -147,3 +147,24 @@ Mainzers Energie-Argument (20W Gehirn vs. Megawatt-LLM-Cluster, siehe [[../conce
|
|||
- **Edge-Tasks** (Smart-Home, Mobile, Embedded): langfristig nur mit neuromorphen/photonischen Backends wirtschaftlich — nicht mit Cloud-LLMs
|
||||
- **Mittelfristig beobachten:** Intel Loihi, IBM TrueNorth, Lightmatter, Luminous Computing — Status 2026?
|
||||
- **Kurzfristig:** Optimierung bestehender AR-Transformer-Pfade (Caching, Quantisierung, Speculative Decoding) bleibt wichtig, ist aber Endpunkt-Frickelei — Hardware-Diversität ist der eigentliche Spielfeldwechsel.
|
||||
|
||||
## Update 2026-07-07: Matthew Berman Model Routing Patterns
|
||||
|
||||
**Source:** [[../../raw/youtube/2026-07-07_berman-model-routing.md]]
|
||||
|
||||
Matthew Berman's video on model routing provides practical cost-saving patterns that complement our existing architecture. Core finding: **planning vs. execution separation** — use expensive frontier models (Fable) for architecture/spec design, then route code execution to cheap models (GPT 5.5, GLM 5.2, Composer 2.5).
|
||||
|
||||
### Key additions to our routing knowledge:
|
||||
|
||||
| Pattern | Savings | Relevance to OpenClaw |
|
||||
|---------|---------|----------------------|
|
||||
| Manual Copy-Paste | ~60% | Ad-hoc testing, proof-of-concept |
|
||||
| Cross-Model-Calling | ~70% | Relevant for agent orchestration (subconscious -> execution) |
|
||||
| Cursor Auto Mode | ~75% | IDE-integrated routing |
|
||||
| Not Diamond | ~80% | Dedicated routing layer (complementary to our fallback-chain) |
|
||||
|
||||
**Coinbase routing on GLM 5.2** validates enterprise adoption of open-weight routing — GLM 5.2 as execution layer behind frontier planning models.
|
||||
|
||||
**Cost math:** Fable-only baseline $9.50 -> routed $6.48 (68% savings). Potential >90% with aggressive routing.
|
||||
|
||||
See [[../tools/model-routing.md]] for the full tools-level page on routing cost-saving patterns.
|
||||
|
|
|
|||
90
wiki/concepts/hardware/ubtech-u1-humanoid-roboter.md
Normal file
90
wiki/concepts/hardware/ubtech-u1-humanoid-roboter.md
Normal file
|
|
@ -0,0 +1,90 @@
|
|||
---
|
||||
created: 2026-07-05
|
||||
updated: 2026-07-05
|
||||
sources: [youtube/2026-07-05_everlast-ki-news-china-roboter-sonnet5-fable5.md]
|
||||
tags: [concept, robotics, humanoid-robots, china, ubtech, u1, demographic-crisis, emotional-ai, silicone-skin, market-launch]
|
||||
---
|
||||
|
||||
# UBTECH U1 — Chinas hyperrealistische Humanoide Roboter
|
||||
|
||||
> **TL;DR:** UBTECH (erstes börsennotiertes Humanoide-Unternehmen) lanciert den U1: Lebensgroßer humanoider Roboter mit hyperrealistischer Silikonhaut, 88 Freiheitsgraden, Cloud-KI-Anbindung — konzipiert für "emotionale Begleitung im Alltag". 13.000+ Vorbestellungen bei $17.600–$45.000. Getrieben von Chinas demografischer Krise (Hong Kong Geburtenrate 0,77 — niedrigste weltweit).
|
||||
|
||||
## UBTECH U1 — Spezifikationen
|
||||
|
||||
| Eigenschaft | Wert |
|
||||
|-------------|------|
|
||||
| **Typ** | Humanoider Roboter mit Silikonhaut |
|
||||
| **Freiheitsgrade** | bis zu 88 |
|
||||
| **Größe (männlich)** | 1,83 m (lebensgroß) |
|
||||
| **Größe (weiblich)** | 1,68 m |
|
||||
| **Akkulaufzeit** | 2–4 Stunden |
|
||||
| **KI-Anbindung** | Cloud-KI |
|
||||
| **Design-Fokus** | Emotionale Begleitung, Gespräche, Blickkontakt |
|
||||
| **Varianten** | 50+ (mit fantasyvollen Outfits, Bewegung, Tanz) |
|
||||
| **Altersbeschränkung** | Ab 18 Jahren |
|
||||
|
||||
## Preise und Marktdemand
|
||||
|
||||
| Variante | Preis |
|
||||
|----------|-------|
|
||||
| Light Modell | ab $17.600 |
|
||||
| Ultra Variante | bis $45.000 |
|
||||
|
||||
**Vorbestellungen:** Über 13.000 beim Launch in Shenzhen — sofort ausverkauft.
|
||||
|
||||
## Demografischer Treiber
|
||||
|
||||
Der Robotik-Push Chinas hat eine tiefere Ursache als Technologie-Begeisterung:
|
||||
|
||||
| Region | Geburtenrate | Lebenserwartung |
|
||||
|--------|-------------|-----------------|
|
||||
| **Hong Kong** | **0,77** (niedrigste weltweit) | Männer ~83, Frauen ~88 |
|
||||
| **China** | ~1,0 (ähnlich kritisch) | — |
|
||||
| **Deutschland** | 1,45 | — |
|
||||
|
||||
**Konvergenz-Faktoren:**
|
||||
- Hong Kong: Reichste Stadt der Welt + teuerste Stadt + niedrigste Geburtenrate
|
||||
- Braindrain: Junge Menschen wandern ab
|
||||
- Einsamkeit als einer der größten Use Cases für humanoide Roboter
|
||||
- Deutschland "nicht viel besser" — KI und Robotik demografisch nicht mehr optional
|
||||
|
||||
## Robotik-Investitionen — Rekordhoch
|
||||
|
||||
| Metrik | Wert |
|
||||
|--------|------|
|
||||
| VC im letzten Quartal | **$16,2 Milliarden** |
|
||||
| Normalniveau | $3–5 Milliarden pro Quartal |
|
||||
| Faktor | >3× Normal |
|
||||
| Vergleich zu KI-Boom | Robotik immer noch unterrepräsentiert |
|
||||
|
||||
## Strategische Bedeutung
|
||||
|
||||
1. **UBTECH** ist das erste börsennotierte Humanoide-Unternehmen — plant 100 Roboter-Spende 2026 (staatlicher/wirtschaftlicher Push)
|
||||
2. **Emotionale KI** als primärer Use Case — nicht Industrie, nicht Pflege, sondern "Begleitung im Alltag"
|
||||
3. **Silikonhaut + 88 DOF** — Roboter "wirken fast wie lebendige Fantasy Figuren aus Zelda" (Leo Schmedding)
|
||||
4. **Cloud-KI-Anbindung** — Roboter als physische Avatar-Plattform für Cloud-LLMs
|
||||
5. **Markt-Signal:** 13K Vorbestellungen zeigen, dass der Massenmarkt für humanoide Roboter bereit ist — im Preisbereich $17–45K
|
||||
|
||||
## Host-Kontext
|
||||
|
||||
Leo Schmedding (Everlast AI) berichtet aus Hongkong. Geplant: Vor-Ort-Besuche bei Humanoid-Unternehmen (u.a. UBTECH) in den Folgetagen. Hintergrund: Hong Kong als Standort mit Deutschlands-Bezug und als einer der wichtigsten Naturhäfen weltweit.
|
||||
|
||||
## Cross-References
|
||||
|
||||
- [[../hardware/neuromorphic-chips-und-quantencomputer.md]] — Hardware-Frontier (Prof. Mainzer: 20W-Gehirn vs. LLM-Megawatt)
|
||||
- [[../agi/universal-high-income.md]] — Roboter-Steuer-Debatte (Gates) und Post-Labor-Economy
|
||||
- [[../../institutions/anthropic.md]] — Cloud-KI-Anbindung der U1-Roboter
|
||||
- [[../llm/vibe-coding-vs-enterprise.md]] — Agentic Coding als Trend, den Roboter-Use-Cases befeuern
|
||||
- [[../policy/ai-regulation-2026.md]] — Regulatorischer Kontext für humanoide Roboter
|
||||
|
||||
## External Sources
|
||||
|
||||
- [YouTube: Everlast AI — KI-News vom 05.07.2026](https://www.youtube.com/watch?v=-JCCcR9qtYQ)
|
||||
- [UBTECH Robotics](https://www.ubtrobotics.com/)
|
||||
- [Everlast AI Kanal](https://www.youtube.com/@everlastai)
|
||||
|
||||
## Verwandte Konzepte
|
||||
|
||||
- **Emotionale KI / Companion Robots** — U1 als erster massenmarktreifer Vertreter
|
||||
- **Post-Labor Economy** — Demografie als Treiber für Robotik-Adoption
|
||||
- **Cloud-KI als Roboter-Backend** — Physische Avatar-Plattform für LLMs
|
||||
65
wiki/concepts/llm/sonnet-5-anthropic.md
Normal file
65
wiki/concepts/llm/sonnet-5-anthropic.md
Normal file
|
|
@ -0,0 +1,65 @@
|
|||
---
|
||||
created: 2026-07-05
|
||||
updated: 2026-07-05
|
||||
sources: [youtube/2026-07-05_everlast-ki-news-china-roboter-sonnet5-fable5.md, wiki/tools/anthropic-claude.md]
|
||||
tags: [concept, llm, anthropic, sonnet-5, claude, cost-inefficiency, default-model, 1m-context]
|
||||
---
|
||||
|
||||
# Claude Sonnet 5 — Anthropic's neues Default-Modell mit Cost-Inefficiency-Problem
|
||||
|
||||
> **TL;DR:** Anthropic lanciert Claude Sonnet 5 als neues Default-Modell für alle gratis/Pro User — 1M Kontextfenster, Performance auf Opus 4.8 Niveau. Doch im Cost-to-Run ist Sonnet 5 teurer als Fable 5 und damit "enorm ineffizient" — der ganze Sinn eines leichtgewichtigeren Modells wird widerlegt.
|
||||
|
||||
## Spezifikationen
|
||||
|
||||
| Eigenschaft | Wert |
|
||||
|-------------|------|
|
||||
| **Vendor** | Anthropic |
|
||||
| **Kontextfenster** | 1 Million Tokens |
|
||||
| **Performance** | ~Opus 4.8 Niveau ("vielleicht ein bisschen günstiger") |
|
||||
| **Default-Modell** | Ja — für alle gratis und Pro User |
|
||||
| **Claude Code** | Verfügbar |
|
||||
| **Pricing** | Standardpricing |
|
||||
| **Sicherheit** | Offiziell "sicherer als Sonnet 4.6" |
|
||||
| **Versionssprung** | 4.6 → 5 (direkt, größer als üblich) |
|
||||
|
||||
## Verbesserungen (laut Anthropic)
|
||||
|
||||
- Reasoning
|
||||
- Tool Use
|
||||
- Coding
|
||||
- Wissensarbeit
|
||||
|
||||
## Cost-to-Run Problem
|
||||
|
||||
**Artificial Analysis Cost to Run Index:** Sonnet 5 bei ~$6.000 — **teurer als Claude Fable 5**.
|
||||
|
||||
| Modell | Cost to Run | vs Sonnet 5 |
|
||||
|--------|------------|-------------|
|
||||
| **Sonnet 5** | ~$6.000 | — |
|
||||
| **Fable 5** | niedriger | günstiger |
|
||||
| **GPT 5.5 Extra Hype** | <50% von Sonnet 5 | ~halbe Kosten |
|
||||
|
||||
**Kritik (Leo Schmedding):**
|
||||
- "Man darf nicht auf offizielle Kostenangaben schauen"
|
||||
- Reasoning-Prozess ist ineffizient → viele Output-Tokens → reale Kosten viel höher
|
||||
- "Das ist exakt nicht der Fall bei Sonnet 5" — der Sinn eines Sonnet-Modells (leichtgewichtiger, günstiger) wird widerlegt
|
||||
- "Der ganze Sinn und Zweck eines Modells wird widersinnig"
|
||||
|
||||
## Einordnung
|
||||
|
||||
Sonnet 5 richtet sich als Default-Modell an den Massenmarkt, scheitert aber an der Cost-Efficiency-Premise. Wer Preis-Leistung sucht, ist mit GPT 5.5 (~halbe Kosten) besser bedient. Wer Frontier-Qualität will, nutzt Fable 5 (das im Cost-to-Run günstiger ist als Sonnet 5 — eine Ironie).
|
||||
|
||||
**Parallele zu [[fable-5-anthropic.md]]:** Auch Fable 5 hatte ein Pricing-Problem (39× teurer als GLM 5.2 im atomic.chat Benchmark). Sonnet 5 hat das umgekehrte Problem: Es sollte das "günstige" Modell sein, ist aber teurer als das Premium-Modell Fable 5.
|
||||
|
||||
## Cross-References
|
||||
|
||||
- [[../../tools/anthropic-claude.md]] — Anthropic Claude Modellübersicht
|
||||
- [[fable-5-anthropic.md]] — Fable 5 (günstiger im Cost-to-Run als Sonnet 5)
|
||||
- [[coding-benchmark-price-performance.md]] — atomic.chat Benchmark
|
||||
- [[chinese-model-cost-routing.md]] — Cost-Routing-These (chinesische Modelle als Alternative)
|
||||
- [[../../institutions/anthropic.md]] — Anthropic Institutionenseite
|
||||
|
||||
## External Sources
|
||||
|
||||
- [YouTube: Everlast AI — KI-News vom 05.07.2026](https://www.youtube.com/watch?v=-JCCcR9qtYQ)
|
||||
- [Anthropic](https://www.anthropic.com)
|
||||
|
|
@ -1,8 +1,8 @@
|
|||
# Wiki Index
|
||||
|
||||
*Auto-generated: 2026-06-23*
|
||||
*Auto-generated: 2026-07-07*
|
||||
|
||||
*Letzte Aktualisierung: 2026-07-05 (66. Update — Geoffrey Hinton RI Discourse ingestiert. Raw: `raw/youtube/2026-07-05_hinton-ri-lecture.md`. Wiki-Update: `people/geoffrey-hinton.md` neu erstellt — Biographie, Kernthesen, Superintelligenz-Zeitleiste (5-20 Jahre), Instrumental-Convergence-Argumente, Nobelpreis-Kontext.)*
|
||||
*Letzte Aktualisierung: 2026-07-07 (67. Update — Matthew Berman Model Routing ingestiert. Raw: `raw/youtube/2026-07-07_berman-model-routing.md`. Wiki-Update: `tools/model-routing.md` neu erstellt — Kostenspar-Patterns für Modell-Routing, `architecture/model-routing.md` erweitert um Berman-Patterns, Coinbase-GLM 5.2-Einsatz, Planning-vs-Execution-Split.)*
|
||||
|
||||
## Architecture
|
||||
|
||||
|
|
@ -11,7 +11,7 @@
|
|||
| [Container & Volume Persistence](architecture/container-volume-persistence.md) | Docker-Volume-Pattern, LanceDB-Havarie, Permission-Fixes | context-tree |
|
||||
| [Memory System](architecture/memory-system.md) | Schichten-Modell S1-S3, Tag-System, Evolutionsphasen, OpenClaw v2026.6.8 QMD + SQLite-WAL + Raw-Memory-Wiki-Source-Pages | context-tree + other/2026-06-16_openclaw-releases-v2026.6.8.md |
|
||||
| [Agent Orchestration](architecture/agent-orchestration.md) | Orchestrator-Pattern, Swarm-Adoption, Frameworks | context-tree |
|
||||
| [Model Routing](architecture/model-routing.md) | Two-Model-Pipeline, GPT-5.4 Config, Fallback-Chain, OpenClaw v2026.6.8 GLM-5.2 + Provider-Prefix-Normalisierung | context-tree + other/2026-06-16_openclaw-releases-v2026.6.8.md |
|
||||
| [Model Routing](architecture/model-routing.md) | Two-Model-Pipeline, GPT-5.4 Config, Fallback-Chain, OpenClaw v2026.6.8 GLM-5.2 + Provider-Prefix-Normalisierung. **Update 07.07.:** Berman Cost-Saving Patterns (Planning-vs-Execution Split, Coinbase GLM 5.2 Routing, 90% savings potential) | context-tree + other/2026-06-16_openclaw-releases-v2026.6.8.md + youtube/2026-07-07_berman-model-routing.md |
|
||||
| [Cron & System Events](architecture/cron-notable-events.md) | Historische Cron-Architektur, Disk-Krise, Execution Gap | context-tree |
|
||||
| [ByteRover Knowledge Mining](architecture/byterover-knowledge-mining.md) | Mining-Pipeline, Blockaden, Status | context-tree |
|
||||
| [DeepMind Beyond Transformer](architecture/deepmind-beyond-transformer.md) | Strategischer Architektur-Kontrast: DeepMind (Diffusion + Hybride + Weltmodelle) vs. OpenAI/Anthropic (AR + Scaling). Vier Säulen, Gemma 4 Edge (256k), Gemini Diffusion 10× | youtube/2026-06-18_deepmind-beyond-transformer.md + youtube/2026-06-16_deepmind-two-steps-ahead.md |
|
||||
|
|
@ -33,6 +33,7 @@
|
|||
| [Testrebalancingbot (Yvonne)](tools/testrebalancingbot-yvonne.md) | Test-Telegram-Bot für Portfolio-Fragen, Gamma Exposure (GEX) | raw/other/financialbot-topic-history-2026-02-01_2026-05-30.json |
|
||||
| [Winston — OpenWebUI-Telegram-Bot](tools/winston-openwebui-bot.md) | Pit Weber's lokaler Telegram-Bot (OpenWebUI): Orchestriert NotebookLM/LLM-Recherchen, Presets für Szenarien, Rich-Text-Output, Skill-Partner-Pipeline mit NotebookLM Short Video Overviews | youtube/2026-07-01_futurepedia-notebooklm-short-video-overviews.md |
|
||||
| [NotebookLM Briefing-System](tools/notebooklm-briefing-system.md) | Pit Weber's automatisierte Pipeline: Quelle → deutsches Audio-Briefing → Infografik → MP4-Video → Nachricht. Zwei Presets (nb1=Standard, nb2=Kurz-Audio zuerst). Skills: notebooklm-pipelines + notebooklm-agent-guide. Requirements: CLI v0.3.4, ffmpeg, yt-dlp | other/2026-07-03_notebooklm-briefing-system.md |
|
||||
| [Model Routing — Cost-Saving Patterns](tools/model-routing.md) | Routing-Patterns für 90% Kosteneinsparung: Planning-vs-Execution Split (Fable → GPT 5.5/GLM 5.2), Cross-Model-Calling, Cursor Auto Mode, Not Diamond. Coinbase on GLM 5.2. Examples: 68-90% savings | youtube/2026-07-07_berman-model-routing.md |
|
||||
| [Fincept Terminal](tools/fincept-terminal.md) | Open-Source Trading-IDE, von OpenClaw-Blog verlinkt | raw/other/financialbot-topic-history-2026-02-01_2026-05-30.json |
|
||||
|
||||
## Concepts
|
||||
|
|
|
|||
14
wiki/log.md
14
wiki/log.md
|
|
@ -2,6 +2,20 @@
|
|||
|
||||
*Append-only changelog. Start: 2026-06-05*
|
||||
|
||||
## [2026-07-07] Ingest | Matthew Berman — 90% less AI costs via Model Routing
|
||||
**Type:** ingest | **Scope:** raw/youtube, wiki/tools (new), wiki/architecture (update), wiki/index, wiki/log
|
||||
**Source:** YouTube — https://www.youtube.com/watch?v=1KKB_UiW6ls (Matthew Berman, 18:33, 2026-07-06, 7.734 views)
|
||||
**NotebookLM:** https://notebooklm.google.com/notebook/a5071bba-125f-4f08-86ea-99a1ab9911ca (German audio briefing: 18:46)
|
||||
**Trigger:** Subagent task spawned in OME-Gruppe Topic 27 (News & Infos).
|
||||
**Actions:**
|
||||
- raw: `raw/youtube/2026-07-07_berman-model-routing.md` (created — 4.5 KB; Frontmatter [type: youtube, source_url, retrieved: 2026-07-07, channel: Matthew Berman, duration_sec: 1113, tags: model-routing, cost-optimization, planning-execution-split, routing-patterns, not-diamond, coinbase, glm-5.2]. Content: Summary, Key Points [Planning vs Execution Split, Cost Comparison Table, 4 Routing Patterns, Coinbase Example, Audio Briefing], Key Takeaways, 6 Cross-Refs to existing Wiki)
|
||||
- wiki (NEW): `tools/model-routing.md` (created — 4.0 KB; Frontmatter [sources, tags]. Sections: Core Principle, Cost Comparison Table, Routing Patterns [Manual Copy-Paste, Cross-Model-Calling, Cursor Auto Mode, Not Diamond], Enterprise Adoption (Coinbase on GLM 5.2), Key Rules, Connection to Existing Wiki [7 cross-refs])
|
||||
- wiki (UPDATE): `architecture/model-routing.md` — Frontmatter: updated=2026-07-07, sources+1 (youtube/2026-07-07_berman), tags+2 (planning-execution-split, cost-savings). New section "Update 2026-07-07: Matthew Berman Model Routing Patterns" with patterns table, Coinbase GLM 5.2 routing, cost math (68-90% savings), cross-ref to tools/model-routing.md.
|
||||
- wiki: `index.md` (updated — Header auf "67. Update", Architecture Model Routing update in table, new Tools entry for Model Routing — Cost-Saving Patterns)
|
||||
- log: this entry
|
||||
**Hector-Hinweis:** Berman's Video kombiniert die Erkenntnisse von DeRonin's 87%-Cost-Cut (chinese-model-cost-routing.md) mit dem Smart Model Router Skill aus dem OpenClaw-Stack. Der Planning-vs-Execution-Split ist die praktische Anleitung, die vom Konzept (DeRonin) zur Implementation führt: Fable denkt, GPT 5.5/GLM 5.2 schreibt Code. Die Coinbase-Referenz ist signifikant — sie zeigt, dass das Routing-Pattern enterprise-ready ist und GLM 5.2 als Open-Weight-Execution-Layer akzeptiert wird. Die 4 Routing-Patterns (manuell → Cross-Model-Calling → Cursor Auto Mode → Not Diamond) bilden eine Reifegrad-Leiter ab.
|
||||
**Subagent-Modell:** openrouter/deepseek/deepseek-v4-flash
|
||||
|
||||
## [2026-07-03] Ingest | Hermes Mixture of Agents 2.0 + Hermes Agent OS (AI Profit Boardroom / Pit Weber)
|
||||
**Type:** ingest | **Scope:** raw/youtube, wiki/tools (update), wiki/index, wiki/log
|
||||
**Source:** YouTube-Video — https://www.youtube.com/watch?v=WY9y529g8Ww (AI Profit Boardroom, 2026-07-03)
|
||||
|
|
|
|||
|
|
@ -83,7 +83,7 @@ Pit Weber markierte den RI Discourse mit 🔴 (existenzielle Risiken). Die Relev
|
|||
|
||||
- **5-20-Jahre-Zeitleiste** ist eine konkretere Prognose als Aschenbrenners 2027-These und von einem Nobelpreisträger mit jahrzehntelanger Forschungstiefe
|
||||
- **Instrumental Convergence**-Argumente sind klassische Yudkowsky/LessWrong-Positionen, aber erstmals von einem "Institutional Insider" dieser Statur öffentlich vorgetragen
|
||||
- **Regierungskritik** unterstützt indirekt die [[concepts/policy/intelligence-feudalism.md]]-These von [[people/brian-roemmele.md]]
|
||||
- **Regierungskritik** unterstützt indirekt die [[../concepts/policy/intelligence-feudalism.md]]-These von [[brian-roemmele.md]]
|
||||
- Die **lokale-Souveränitäts-These** des OME21-Briefings korrespondiert mit Hintons Skepsis gegenüber zentralisierter KI-Kontrolle
|
||||
|
||||
## Quellen
|
||||
|
|
|
|||
68
wiki/tools/model-routing.md
Normal file
68
wiki/tools/model-routing.md
Normal file
|
|
@ -0,0 +1,68 @@
|
|||
---
|
||||
created: 2026-07-07
|
||||
updated: 2026-07-07
|
||||
sources: [youtube/2026-07-07_berman-model-routing.md]
|
||||
tags: [tools, model-routing, cost-optimization, routing-patterns, fable, planning-execution-split, cross-model-calling, not-diamond, cursor-auto-mode, copy-paste-routing, coinbase, glm-5.2, gpt-5.5]
|
||||
---
|
||||
|
||||
# Model Routing — Cost-Saving Patterns
|
||||
|
||||
> **Source:** Matthew Berman — "You NEED to do this right now..." (2026-07-06)
|
||||
> See raw: [[../../raw/youtube/2026-07-07_berman-model-routing.md]]
|
||||
|
||||
## Core Principle
|
||||
|
||||
**Route every task to the cheapest model that can handle it well.** The single most impactful pattern is **planning vs. execution separation**: use frontier models for architecture/spec design, then delegate code execution to cheaper models.
|
||||
|
||||
## Cost Comparison
|
||||
|
||||
| Model | Input Cost / M Tokens | Output Cost / M Tokens | Role |
|
||||
|-------|----------------------|-----------------------|------|
|
||||
| **Fable 5** (Anthropic) | $10 | $50 | Planning, Architecture, Specs |
|
||||
| **GPT 5.5** | $2 | $6 | Code execution |
|
||||
| **Composer 2.5** | $2 | $6 | Code execution |
|
||||
| **Claude Sonnet** (mid) | $3 | $15 | Code execution |
|
||||
| **GLM 5.2** (Z.ai) | ~$0.08 | ~$0.08 | Open-source execution layer |
|
||||
|
||||
**Example:** Fable-only = $9.50 baseline → Fable + GPT 5.5 routing = **$6.48 (68% savings)**. Total potential with aggressive routing: **>90%**.
|
||||
|
||||
## Routing Patterns
|
||||
|
||||
### 1. Manual Copy-Paste (~60% savings)
|
||||
Developer manually copies code between models. No automation. Works as a proof-of-concept but doesn't scale.
|
||||
|
||||
### 2. Cross-Model-Calling (~70% savings)
|
||||
One model calls another via API. A script/program orchestrates: Fable designs → API passes spec to GPT 5.5/GLM 5.2 → execution model writes code.
|
||||
|
||||
### 3. Cursor Auto Mode (~75% savings)
|
||||
Cursor IDE's auto-mode feature routes tasks to the cheapest capable model automatically. Seamless developer experience.
|
||||
|
||||
### 4. Not Diamond (~80% savings)
|
||||
Dedicated routing layer/API that classifies incoming requests and routes to the optimal model. Adds a routing decision layer between user and model.
|
||||
|
||||
### 5. OpenClaw Smart Model Router (see Skill)
|
||||
[OpenClaw Smart Model Router skill] — Tier-based routing (5 tiers from simple to frontier). Claims 60-90% savings aligned with Berman's findings.
|
||||
|
||||
## Enterprise Adoption
|
||||
|
||||
**Coinbase** is implementing model routing on **open-source models, specifically GLM 5.2** (Z.ai). This validates:
|
||||
- The routing pattern is production-grade, not experimental
|
||||
- Chinese open-weight models (GLM 5.2) are viable as the cost-effective execution layer
|
||||
- Enterprise security/compliance requirements can be met with routed open-source models
|
||||
|
||||
## Key Rules
|
||||
|
||||
1. **Separate planning from execution** — Fable designs the architecture, cheap models write the code
|
||||
2. **Bounded specs produce reliable cheap execution** — when the spec is tight, cheap models stay on track
|
||||
3. **Output tokens are the cost driver** — frontier models charge $50/M output vs $6/M for cheap models
|
||||
4. **The pattern is simple** — the hard part is building the discipline and the harness
|
||||
|
||||
## Connection to Existing Wiki
|
||||
|
||||
- **[[../architecture/model-routing.md]]** — OpenClaw's existing routing architecture (Two-Model-Pipeline, Fallback-Chain, GLM 5.2 support)
|
||||
- **[[../concepts/llm/chinese-model-cost-routing.md]]** — DeRonin's 87% cost-cut playbook (6 Western→Chinese swaps) complements Berman's patterns with field data
|
||||
- **[[../concepts/llm/coding-benchmark-price-performance.md]]** — atomic.chat benchmark: GLM 5.2 is B+ at $0.08 vs Fable 5 A+ at $3.12 (39× cheaper)
|
||||
- **[[../concepts/llm/glm-5.2-zai-coding-model.md]]** — Full GLM 5.2 documentation (1M context, MIT license, coding strength)
|
||||
- **[[../concepts/llm/fable-5-anthropic.md]]** — Fable 5 documentation (premium frontier)
|
||||
- **[[../concepts/agents/subconscious-agent.md]]** — OpenClaw's agent orchestration that could benefit from routing
|
||||
- **[[../concepts/llm/llm-model-fusion-ensembles.md]]** — Alternative pattern: parallel model panels instead of sequential routing
|
||||
|
|
@ -1,6 +1,6 @@
|
|||
---
|
||||
created: 2026-07-03
|
||||
updated: 2026-07-04
|
||||
updated: 2026-07-06
|
||||
sources: [other/2026-07-04_notebooklm-agent-guide-v2.md, other/2026-07-03_notebooklm-briefing-system.md, youtube/2026-07-01_futurepedia-notebooklm-short-video-overviews.md, other/2026-06-19_ome21-briefing-ki-fortschritte-lokale-modelle.md]
|
||||
tags: [tool, notebooklm, briefing, pipeline, audio, video, mp4, infographic, automation, presets, agent-skills, pit-weber, ffmpeg, yt-dlp, caption-contract, v2]
|
||||
---
|
||||
|
|
@ -201,6 +201,25 @@ ffmpeg -y \
|
|||
- Alle drei Skills als **live** deklariert
|
||||
- Neue Raw-Datei: [Agent Guide V2](../../../raw/other/2026-07-04_notebooklm-agent-guide-v2.md)
|
||||
|
||||
## V3-Abschluss (2026-07-06)
|
||||
|
||||
- Drei Skills als echte SKILL.md-Dateien erstellt: `notebooklm-core`, `notebooklm-pipelines`, `notebooklm-agent-guide`
|
||||
- Vollständige Einrichtungsanleitung: `docs/notebooklm-nb-presets-anleitung.md` (10KB, abgeschlossen)
|
||||
- Output-Verzeichnisse angelegt: `output/nb1/`, `output/nb2/`
|
||||
- Skill-Trennung dokumentiert: nb-Skills (Pit-spezifisch) vs. ClawHub-Skill (allgemein)
|
||||
- Status: **Vollautomatisierung bereit** — CLI-Tools + Auth noch zu installieren
|
||||
|
||||
## V3.1-Update (2026-07-06, mit Winston's Setup-Guide)
|
||||
|
||||
- **nb3-Preset hinzugefügt** — Briefing ohne Notebook-Link (Kurz-Audio + Infografik → MP4, kein Share)
|
||||
- **Fact-Check als Schritt 9** — `notebooklm ask "Nenne die 5 wichtigsten Kernthesen"` vor Posting
|
||||
- **Archiv-Kopie als Pflicht** — jedes Artifact MUSS zusätzlich ins Archiv-Topic
|
||||
- **Rate-Limit-Handling** — Fallback-Google-Konto für `RATE_LIMITED` Fehler
|
||||
- **Klartext-URLs** — keine Markdown-Links im Telegram-Posting (rendert nicht klickbar)
|
||||
- **14 Schritte** (war 12) — Fact-Check + Archiv-Kopie hinzugefügt
|
||||
- Anleitung auf 13.5KB gewachsen, Skills auf v1.1.0 aktualisiert
|
||||
- Quelle: Winston (@Winston_butlerbot) NB-Skills Setup-Anleitung
|
||||
|
||||
## Beziehung zu Winston
|
||||
|
||||
Das Briefing-System läuft als Backend hinter [[winston-openwebui-bot.md|Winston]] (Pit's OpenWebUI-Telegram-Bot):
|
||||
|
|
|
|||
Loading…
Add table
Reference in a new issue