ingest(hardware): nvidia dgx station gb300 — 748gb unified memory desktop

- raw: youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md (The Stack, 8:14)
- wiki: concepts/hardware/nvidia-dgx-station-748gb.md (NEW — 748 GB unified, full-precision 70B, ~$85K-$115K, ROI break-even ~2 months)
- index: 39. Update, new hardware row + raw source entry
- log: entry appended
This commit is contained in:
Hector 2026-06-28 15:09:44 +02:00
parent c35d8e0736
commit 0007c0cb50
4 changed files with 226 additions and 1 deletions

View file

@ -0,0 +1,80 @@
---
type: youtube
source_url: https://www.youtube.com/watch?v=EhXQysElOY8
retrieved: 2026-06-28
channel: "The Stack"
title: "NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good"
duration_sec: 494
views: 2983
published: 2026-06-27
has_transcript: true
tags: [nvidia, dgx-station, local-ai, unified-memory, gb300, grace-blackwell, lm-studio, ollama, hardware, cloud-exit]
---
# NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good
**Channel:** [The Stack](https://www.youtube.com/@The-Stack-ai)
**Published:** 2026-06-27
**Duration:** 8:14
**Views:** ~2,983 (at time of fetch, 7h after publish)
**Hashtags:** #localai #nvidiadgx #lmstudio #computex2026 #agenticai
## Video Description (full)
LM Studio local AI just changed: NVIDIA's DGX Station packs 748GB unified memory to run 70B models in full precision, no cloud needed.
LM Studio and Ollama finally lose their asterisk, NVIDIA's DGX Station at Computex 2026 lands 748GB of coherent unified memory in a single deskside tower, enough to load a full-precision 70B model with room to spare, no quantization required, no cloud offload.
Announced by Jensen Huang at GTC Taipei on May 31, 2026, the machine is built around the GB300 Grace Blackwell Ultra Desktop Superchip: a 72-core ARM Grace CPU fused to a Blackwell Ultra GPU via NVLink-C2C at 900 GB/s. The memory pool splits into 252GB HBM3e (7.1 TB/s GPU-side) and 496GB LPDDR5X (CPU-side), both fully coherent, one address space, zero explicit copies. Compute tops out at 20 petaFLOPS FP4. NVIDIA doesn't sell a Founders Edition; OEM partners ASUS, Dell, HP, MSI, and others handle that, with real-world pricing landing between roughly $85K and $115K (the MSI XpertStation WS300 lists at $96,995.99 on CDW). The video also covers NVIDIA's DGX Spark (128GB, ~$4,700) as the genuine prosumer entry point, and gives an honest head-to-head with the Mac Studio M5 Ultra, which still holds the value crown for a solo developer running mid-size models. The trillion-parameter claim gets a reality check, it's technically true only with aggressive 4-bit quantization, not full-precision weights. The cloud ROI math is real: at ~$98/hour for a comparable AWS p5 instance, the hardware pays for itself in roughly two months of sustained workload. The DGX Station for Windows (WSL-based) is flagged as a Q4 2026 promise, not a shipping product. RTX Spark, NVIDIA's MediaTek-partnered consumer AI PC chip, rounds out the roadmap alongside a three-generation plan through Rubin and Rosa Feynman.
For individual builders and small teams deciding between local AI options, this is a practical breakdown of which box on the NVIDIA ladder actually makes sense for their workload.
## Chapters
- 0:00 Intro
- 0:15 What Jensen actually unveiled
- 1:12 Why unified memory is the whole story
- 2:31 The trillion-parameter asterisk
- 3:25 The price, who it's for, and how to choose
- 5:25 The cloud math that justifies the big box
- 6:40 The bigger play: NVIDIA wants the whole PC
## Tools & Resources Mentioned
- **LM Studio:** https://lmstudio.ai
- **Ollama:** https://ollama.com
- **NVIDIA DGX Station:** https://www.nvidia.com/en-us/products/workstations/dgx-station/
- **NVIDIA DGX Station for Windows:** https://www.nvidia.com/en-us/products/workstations/dgx-station-for-windows/
- **NVIDIA DGX Spark:** https://www.nvidia.com/en-us/products/workstations/dgx-spark/
## Key Specs Summary
| Spec | Value |
|------|-------|
| Chip | GB300 Grace Blackwell Ultra Desktop Superchip |
| CPU | 72-core ARM Grace |
| GPU | Blackwell Ultra |
| Interconnect | NVLink-C2C @ 900 GB/s |
| Total Unified Memory | 748 GB (252 GB HBM3e + 496 GB LPDDR5X) |
| HBM3e Bandwidth | 7.1 TB/s (GPU-side) |
| Compute | 20 petaFLOPS FP4 |
| Price Range | ~$85K$115K (OEM-dependent) |
| MSI XpertStation WS300 | $96,995.99 (CDW listing) |
| DGX Spark (entry) | 128 GB, ~$4,700 |
| Full-precision 70B model | Fits with room to spare |
| Trillion-parameter claim | Only with 4-bit quantization |
| Cloud ROI break-even | ~2 months (vs AWS p5 @ ~$98/h) |
| DGX Station for Windows | Q4 2026 (WSL-based, not shipping) |
| Announced | GTC Taipei, May 31, 2026 by Jensen Huang |
## NVIDIA Local AI Hardware Ladder
| Product | Memory | Price | Target |
|---------|--------|-------|--------|
| DGX Spark | 128 GB | ~$4,700 | Prosumer entry |
| DGX Station | 748 GB | ~$85K$115K | Small teams / sustained workloads |
| Mac Studio M5 Ultra | (value crown for solo devs) | — | Solo developer, mid-size models |
## Context
Shared by Pit Weber in OME-Gruppe, Topic "News & Infos" (Topic 13) on 2026-06-28.

View file

@ -0,0 +1,131 @@
---
created: 2026-06-28
updated: 2026-06-28
sources: [youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md]
tags: [concept, hardware, nvidia, dgx-station, gb300, grace-blackwell, unified-memory, local-ai, cloud-exit, lm-studio, ollama]
---
# NVIDIA DGX Station — 748 GB Unified Memory Desktop
> *"LM Studio and Ollama finally lose their asterisk."*
> — The Stack, YouTube, 2026-06-27
## Kernthese
NVIDIAs DGX Station (GB300 Grace Blackwell Ultra Desktop Superchip) bringt **748 GB kohärenter Unified Memory** in einen einzelnen Schreibtisch-Tower — genug, um ein 70B-Modell in voller Präzision (FP16/BF16) ohne Quantisierung zu laden. Das entfernt den letzten Haken von LM Studio und Ollama: die "läuft lokal, aber nur quantisiert"-Einschränkung. Die Maschine ist kein Konzept — sie wurde von Jensen Huang auf der GTC Taipei (31.05.2026) angekündigt, OEMs (ASUS, Dell, HP, MSI) listen sie bereits, und die Cloud-ROI-Mathematik rechtfertigt sie für anhaltende Workloads.
## Architektur: GB300 Grace Blackwell Ultra Desktop Superchip
Der GB300 ist ein **Superchip** — CPU und GPU auf einem Die, verbunden über NVLink-C2C:
| Komponente | Spec |
|------------|------|
| CPU | 72-core ARM Grace |
| GPU | Blackwell Ultra |
| Interconnect | NVLink-C2C @ 900 GB/s |
| GPU Memory | 252 GB HBM3e @ 7.1 TB/s |
| CPU Memory | 496 GB LPDDR5X |
| **Total Unified Memory** | **748 GB (ein Adressraum, kohärent)** |
| Compute | 20 petaFLOPS FP4 |
**Das entscheidende Architektur-Merkmal:** HBM3e und LPDDR5X sind **fully coherent** — ein Adressraum, zero explicit copies. Die GPU kann direkt auf CPU-Memory zugreifen und umgekehrt. Das eliminiert den PCIe-Bottleneck, der bei diskreten GPU-Setups den Datentransfer dominiert.
> Siehe auch: [[edge-inference-als-cloud-alternative.md]] — AMDs Strix Halo bringt das Unified-Memory-Prinzip in die x86-Welt (128 GB). Der GB300 skaliert dasselbe Prinzip um das 5,8-fache.
## Was 748 GB bedeuten
| Modell-Größe | FP16 VRAM-Bedarf | Passt in 748 GB? | Quantisierung nötig? |
|-------------|-----------------|-------------------|----------------------|
| 8B | ~16 GB | ✅ (massiv Spielraum) | Nein |
| 34B | ~68 GB | ✅ | Nein |
| 70B | ~140 GB | ✅ (5× Platz) | Nein |
| 405B | ~810 GB | ❌ | Ja (4-bit → ~200 GB) |
| 1T (Trillion) | ~2 TB | ❌ | Ja (4-bit → ~500 GB, knapp) |
**Die Trillion-Parameter-Realität:** NVIDIAs Claim, die DGX Station könne Trillionen-Parameter-Modelle laden, ist technisch wahr — aber **nur mit aggressiver 4-Bit-Quantisierung**. In voller Präzision (FP16/BF16) ist bei ~70B100B die Grenze erreicht. Das ist kein Betrug, aber ein Asterisk, den der Video-Titel nicht zeigt.
## Preis und Positionierung
| Produkt | Memory | Preis | Zielgruppe |
|---------|--------|-------|------------|
| **DGX Spark** | 128 GB | ~$4.700 | Prosumer-Einstieg |
| **DGX Station** | 748 GB | ~$85K$115K | Kleine Teams, sustained Workloads |
| **Mac Studio M5 Ultra** | (variiert) | — | Solo-Entwickler, Mid-Size-Modelle (Value Crown) |
- **MSI XpertStation WS300:** $96.995,99 (CDW-Listing)
- NVIDIA verkauft keine Founders Edition — OEMs übernehmen Vertrieb
### Cloud-ROI-Mathematik
| Kostenart | Wert |
|-----------|------|
| AWS p5 (vergleichbar) | ~$98/Stunde |
| DGX Station (Einmal) | ~$98.000 |
| Break-even | ~1.000 Stunden (~2 Monate sustained) |
| Danach | Nur Strom + Wartung |
**Einschätzung:** Für Teams, die 24/7-Inferenz betreiben, ist die DGX Station nach ~2 Monaten günstiger als Cloud. Für sporadische Nutzung bleibt Cloud überlegen. Die Mathematik funktioniert nur bei **anhaltender** Auslastung.
## DGX Station for Windows — Q4 2026
NVIDIA kündigt eine WSL-basierte Variante für Windows an. Status: **Versprechen, nicht lieferbar** (Q4 2026). Die aktuelle DGX Station läuft auf Linux.
## NVIDIA Roadmap: Drei Generationen
| Generation | Status |
|------------|--------|
| Grace Blackwell (GB300) | Lieferbar (Computex 2026) |
| Rubin | Geplant |
| Rosa Feynman | Geplant |
Zusätzlich: **RTX Spark** — NVIDIAs MediaTek-Partner-Consumer-AI-PC-Chip. NVIDIA zielt auf den gesamten PC-Markt ab, nicht nur Workstations.
## Einordnung: Wo steht die DGX Station im Cloud-Exit-Kontext?
Die DGX Station ist das **Spitzenmodell** der lokalen KI-Hardware-Ladder:
```
DGX Spark (128 GB, $4.7K) → DGX Station (748 GB, ~$98K) → DGX Datacenter (rack-scale)
↑ ↑ ↑
Prosumer / Einzelbau Kleine Teams / 24-7-Workloads Cloud-Scale
```
### Vergleich mit bestehenden Wiki-Hardware-Seiten
| Plattform | Unified Memory | Preis | Vorteil | Quelle |
|-----------|---------------|-------|---------|--------|
| AMD Ryzen AI Max+ 395 (Strix Halo) | 128 GB | $1.499 | x86 Edge, lunchbox-Format | [[edge-inference-als-cloud-alternative.md]] |
| Mac Studio M5 Ultra | (variiert) | — | Value Crown für Solo-Dev | Video-Comparison |
| **NVIDIA DGX Station (GB300)** | **748 GB** | **~$85K$115K** | **Full-precision 70B, sustained ROI** | Diese Seite |
| NVIDIA DGX Spark | 128 GB | ~$4.700 | Prosumer-Einstieg | Diese Seite |
### Verbindung zum Cloud-Exit-Pattern
Die DGX Station bestätigt die [[cloud-exit-and-local-superiority.md]]-These von einer neuen Seite: Nicht nur werden Cloud-Server teurer (Hetzner/HP-Preisschock), sondern lokale Hardware wird **kapazitativ** — 748 GB Unified Memory waren bisher Server-Rack-Territorium. Die "RAM als neue digitale Währung"-These aus dem OME21-Briefing bekommt hier ihre radikalste Bestätigung.
Cross-Ref zu [[cloud-exit-and-local-superiority.md]] § "RAM als Währung": Dort war die These, dass Arbeitsspeicher zum Engpass und damit zur wertvollsten Resource wird. 748 GB in einer Desktop-Box ist die konsequente Antwort.
### LM Studio und Ollama: Der Asterisk verschwindet
Bisher: Lokale LLM-Inferenz = "ja, aber nur quantisiert" (GGUF Q4/Q8). Die DGX Station entfernt diesen Asterisk:
- **70B in FP16/BF16:** ~140 GB → passt in 748 GB mit 5× Reserve
- **Keine Quantisierungs-Artifakte:** Full-precision-Inferenz ohne Quality-Loss
- **LM Studio + Ollama:** Beide Tools werden auf der DGX Station unterstützt
Das ist ein Paradigmenwechsel für die lokale KI-Community: die Frage verschiebt sich von "passt das Modell in den VRAM?" zu "ist die Hardware wirtschaftlich gerechtfertigt?".
## Quellen
- **YouTube-Video:** [NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good](https://www.youtube.com/watch?v=EhXQysElOY8) — The Stack, 2026-06-27, 8:14
- **NVIDIA DGX Station:** [nvidia.com](https://www.nvidia.com/en-us/products/workstations/dgx-station/)
- **NVIDIA DGX Spark:** [nvidia.com](https://www.nvidia.com/en-us/products/workstations/dgx-spark/)
- **LM Studio:** [lmstudio.ai](https://lmstudio.ai)
- **Ollama:** [ollama.com](https://ollama.com)
## Cross-Refs
- [[cloud-exit-and-local-superiority.md]] — Cloud-Exit-These, Preisschock, lokale Performance-Parität
- [[edge-inference-als-cloud-alternative.md]] — AMD Strix Halo (128 GB unified), x86-Alternative
- [[neuromorphic-chips-und-quantencomputer.md]] — Hardware-Frontier-Perspektive (Prof. Mainzer)
- [[../../architecture/model-routing.md]] — Model-Routing (lokale Modelle in OpenClaw)
- [[../llm/llm-model-catalog.md]] — Modell-Katalog (welche Modelle lokal laufen)

View file

@ -2,7 +2,7 @@
*Auto-generated: 2026-06-23* *Auto-generated: 2026-06-23*
*Letzte Aktualisierung: 2026-06-28 (38. Update — Hermes "Mixture of Agents" (MoA) Feature: Merge any N models into one virtual model. Reference + Aggregator Pattern, +8% über Opus 4.8 solo. Neue Sektion in `tools/hermes-desktop.md`. Raw-Datei `xpost/2026-06-28-hermes-moa-vaibhavsisinty.md`.)* *Letzte Aktualisierung: 2026-06-28 (39. Update — NVIDIA DGX Station GB300: 748 GB Unified Memory Desktop Superchip. Full-precision 70B lokal, kein Quantisierung mehr. Cloud-ROI Break-even ~2 Monate. Neue Seite `concepts/hardware/nvidia-dgx-station-748gb.md`. Raw-Datei `youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md`.)*
## Architecture ## Architecture
@ -77,6 +77,7 @@
| [Neuromorphic Chips & Quantencomputer](concepts/hardware/neuromorphic-chips-und-quantencomputer.md) | Prof. Dr. Klaus Mainzer (TU München / Akademie-Präsident): Hardware-Frontier jenseits klassischer Chips. Neuromorphic, Photonik, Quantencomputer (Dekohärenz, Shors Algorithmus). 20W-Gehirn vs. LLM-Megawatt. Geopolitische Positionen (Pro-Atomkraft, China, Thiel-Kritik). | youtube/2026-06-16_everlast-mainzer-neuromorphe-chips-quantencomputer.md | | [Neuromorphic Chips & Quantencomputer](concepts/hardware/neuromorphic-chips-und-quantencomputer.md) | Prof. Dr. Klaus Mainzer (TU München / Akademie-Präsident): Hardware-Frontier jenseits klassischer Chips. Neuromorphic, Photonik, Quantencomputer (Dekohärenz, Shors Algorithmus). 20W-Gehirn vs. LLM-Megawatt. Geopolitische Positionen (Pro-Atomkraft, China, Thiel-Kritik). | youtube/2026-06-16_everlast-mainzer-neuromorphe-chips-quantencomputer.md |
| [Cloud-Exit & Lokale Überlegenheit](concepts/hardware/cloud-exit-and-local-superiority.md) | Preisschock (Hetzner/HP), lokale MoE-Modelle mit 150 tok/s auf M-Series, RAM als neue digitale Währung, Community-Krallen-Beispiele (Rüdiger/Andreas/Christian) | other/2026-06-19_ome21-briefing-ki-fortschritte-lokale-modelle.md | | [Cloud-Exit & Lokale Überlegenheit](concepts/hardware/cloud-exit-and-local-superiority.md) | Preisschock (Hetzner/HP), lokale MoE-Modelle mit 150 tok/s auf M-Series, RAM als neue digitale Währung, Community-Krallen-Beispiele (Rüdiger/Andreas/Christian) | other/2026-06-19_ome21-briefing-ki-fortschritte-lokale-modelle.md |
| [Edge-Inferenz als Cloud-Alternative](concepts/hardware/edge-inference-als-cloud-alternative.md) | AMD Ryzen AI Max+ 395 (Strix Halo): 235B-Modell lokal, 128 GB unified memory, 3× RTX 5080, $1.499 Lunchbox-PC. x86-Alternative zur Mac-only Cloud-Exit-Bewegung. Cross-Refs zu Cloud-Exit, Mainzer, Aravind, Aschenbrenner | xpost/2026-06-16_amd-ryzen-ai-edge-inferenz.md | | [Edge-Inferenz als Cloud-Alternative](concepts/hardware/edge-inference-als-cloud-alternative.md) | AMD Ryzen AI Max+ 395 (Strix Halo): 235B-Modell lokal, 128 GB unified memory, 3× RTX 5080, $1.499 Lunchbox-PC. x86-Alternative zur Mac-only Cloud-Exit-Bewegung. Cross-Refs zu Cloud-Exit, Mainzer, Aravind, Aschenbrenner | xpost/2026-06-16_amd-ryzen-ai-edge-inferenz.md |
| [NVIDIA DGX Station — 748 GB Unified Memory Desktop](concepts/hardware/nvidia-dgx-station-748gb.md) | GB300 Grace Blackwell Ultra Superchip: 748 GB kohärent (252 GB HBM3e + 496 GB LPDDR5X), 20 petaFLOPS FP4. Full-precision 70B ohne Quantisierung. ~$85K$115K, ROI break-even ~2 Monate vs Cloud. DGX Spark (128 GB, $4.7K) als Prosumer-Einstieg. LM Studio/Ollama-Asterisk verschwindet | youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md |
### Agents ### Agents
| Seite | Beschreibung | Quellen | | Seite | Beschreibung | Quellen |
@ -237,3 +238,4 @@
| `raw/youtube/2026-06-26_karpathy-how-i-use-llms.md` | youtube | Andrej Karpathy: How I use LLMs (2:11:11, ~2.5M views, English) — Praxis-Crashkurs: LLM Fundamentals, Tool Integration, Multimodal, Custom GPTs | | `raw/youtube/2026-06-26_karpathy-how-i-use-llms.md` | youtube | Andrej Karpathy: How I use LLMs (2:11:11, ~2.5M views, English) — Praxis-Crashkurs: LLM Fundamentals, Tool Integration, Multimodal, Custom GPTs |
| `raw/other/2026-06-26_gemini-wm2026-spielplan-ausfuellung.md` | other | Gemini Demo: WM 2026 Spielplan ausfüllen per KI — Foto-Verarbeitung (cv2/PIL), Multimodalität, Limitationen (WM läuft noch) | | `raw/other/2026-06-26_gemini-wm2026-spielplan-ausfuellung.md` | other | Gemini Demo: WM 2026 Spielplan ausfüllen per KI — Foto-Verarbeitung (cv2/PIL), Multimodalität, Limitationen (WM läuft noch) |
| `raw/xpost/2026-06-28-hermes-moa-vaibhavsisinty.md` | xpost | Hermes "Mixture of Agents" (MoA): Merge any N models into one virtual model (Reference + Aggregator), +8%/+11% über Opus 4.8/GPT-5.5 solo. @Teknium: any number of models. @lambdua: "toy stage" | | `raw/xpost/2026-06-28-hermes-moa-vaibhavsisinty.md` | xpost | Hermes "Mixture of Agents" (MoA): Merge any N models into one virtual model (Reference + Aggregator), +8%/+11% über Opus 4.8/GPT-5.5 solo. @Teknium: any number of models. @lambdua: "toy stage" |
| `raw/youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md` | youtube | The Stack: NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good — DGX Station GB300 Grace Blackwell Ultra, 748 GB unified memory, full-precision 70B lokal, ~$85K$115K, Cloud-ROI break-even ~2 Monate |

View file

@ -2,6 +2,18 @@
*Append-only changelog. Start: 2026-06-05* *Append-only changelog. Start: 2026-06-05*
## [2026-06-28] Ingest | NVIDIA DGX Station GB300 — 748 GB Unified Memory Desktop
**Type:** ingest | **Scope:** raw/youtube, wiki/concepts/hardware
**Source:** YouTube — https://www.youtube.com/watch?v=EhXQysElOY8 (The Stack, 8:14, 2.983 views, 2026-06-27)
**Trigger:** Shared by Pit Weber in OME-Gruppe Topic "News & Infos" (Topic 13) on 2026-06-28.
**Actions:**
- raw: `raw/youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md` (created — 4.6 KB; Frontmatter [type: youtube, channel: The Stack, tags: nvidia, dgx-station, local-ai, unified-memory, gb300, grace-blackwell, lm-studio, ollama, hardware, cloud-exit]. Vollständige Video-Beschreibung, Kapitel-Marker, Specs-Tabelle, Tools, NVIDIA Hardware Ladder)
- wiki (NEW): `concepts/hardware/nvidia-dgx-station-748gb.md` (created — 7.2 KB; Frontmatter [sources, tags]. Sektionen: Kernthese, GB300-Architektur [Specs-Tabelle, Unified-Memory-Prinzip, Cross-Ref zu Strix Halo], 748-GB-Bedeutung [Modell-Größen-Tabelle, Trillion-Parameter-Realitätscheck], Preis/Positionierung [Spark/Station/Mac-Tabelle, Cloud-ROI-Mathematik], DGX for Windows Q4 2026, NVIDIA Roadmap [Rubin, Rosa Feynman, RTX Spark], Einordnung [Hardware-Ladder, Vergleichstabelle, Cloud-Exit-Verbindung, LM Studio/Ollama-Asterisk], 5 Cross-Refs, 4 externe Links)
- wiki: `index.md` (updated — Header auf "39. Update", neue Hardware-Zeile für DGX Station, neuer Raw-Sources-Eintrag)
- log: this entry
**Hector-Hauptthese:** Die DGX Station ist die radikalste Bestätigung der Cloud-Exit-These aus dem OME21-Briefing: 748 GB Unified Memory waren bisher Server-Rack-Territorium — jetzt in einer Desktop-Box. Der Asterisk "läuft lokal, aber nur quantisiert" für LM Studio/Ollama verschwindet: ein 70B-Modell in voller Präzision passt mit 5× Reserve. Die Cloud-ROI-Mathematik (~2 Monate break-even bei $98/h AWS p5) macht die Maschine für Teams mit sustained Workloads wirtschaftlich rational. Für OpenClaw/Hector relevant: die Hardware-Ladder von DGX Spark ($4.7K, 128 GB) bis DGX Station ($98K, 748 GB) definiert die lokale Deployment-Skala für unsere Wiki-Architektur. Die Verbindung zu [[concepts/hardware/edge-inference-als-cloud-alternative.md]] ist direkt: AMDs Strix Halo (128 GB, $1.499) ist der Prosumer-Einstieg, NVIDIAs GB300 (748 GB, $98K) ist die Enterprise-Spitze — beide nutzen dasselbe Unified-Memory-Prinzip.
**Subagent-Modell:** ollama/glm-5.2:cloud
## [2026-06-28] Ingest | Hermes "Mixture of Agents" (MoA) — Multi-Model Fusion Feature ## [2026-06-28] Ingest | Hermes "Mixture of Agents" (MoA) — Multi-Model Fusion Feature
**Type:** ingest | **Scope:** raw/xpost, wiki/tools **Type:** ingest | **Scope:** raw/xpost, wiki/tools
**Source:** X Post — https://x.com/vaibhavsisinty/status/2070741416649850898 **Source:** X Post — https://x.com/vaibhavsisinty/status/2070741416649850898