- raw/blog/2026-09-17_nvidia-ai-agents-3d-scenes-simulation.md (neu) - wiki/concepts/agents/nvidia-agenten-3d-scene-prep.md (neu) - wiki/institutions/nvidia.md (neu) - Cross-Refs: harness-loop-graph-engineering, hermes-bot-mode, harness-vs-standalone-chat-llm, robotics-model-benchmark, tools/openclaw, tools/nvidia-dgx-spark - index.md (205. Update), log.md
6.8 KiB
6.8 KiB
| created | updated | sources | tags | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 2026-08-27 | 2026-09-17 |
|
|
NVIDIA DGX Spark
NVIDIAs Desktop-AI-Supercomputer (Projekt „DIGITS", Launch Oktober 2025): ein vollständiges GB10-Grace-Blackwell-System im Desktop-Format, ausgerichtet auf lokale Inferenz und Fine-Tuning kleinerer bis mittlerer Modelle. Erste Zielplattform für Perplexitys vollständig lokalen Agent-Stack („Portable Computer", 2026-08-25).
Spezifikationen (Stand 08/2026)
| Komponente | Wert |
|---|---|
| Chip | NVIDIA GB10 Grace Blackwell Superchip |
| CPU | 20 Arm-Core (10× Cortex-X925 + 10× Cortex-A725, LMSYS-Teardown) |
| GPU | Blackwell-GPU mit 5. Gen Tensor-Cores, ~1 PFLOP FP4 (Herstellerangabe) |
| Memory | 128 GB LPDDR5x unified (CPU+GPU geteilt) |
| Bandbreite | ~273 GB/s (LPDDR5x — zentraler Bottleneck, LMSYS) |
| Storage | 1–4 TB NVMe (je nach SKU) |
| Netzwerk | 2× QSFP ConnectX-7 (200 Gb/s), 10 GbE RJ-45, WLAN |
| Anschlüsse | 4× USB-C, HDMI, 240-W-USB-C-Netzteil |
| OS | DGX OS (Ubuntu-basiert); DGX-OS-Update 2026: bis zu 1,9× Inference-Speedups (Herstellerangabe) |
| Verbindbarkeit | Bis zu 4 Systeme clusterbar (NVIDIA-Claim: bis zu 700B-Parameter-Modelle, ursprünglich 2× für ~405B FP4) |
Preis (US + EU, Stand 08/2026)
| Markt | Preis | Notiz |
|---|---|---|
| US (Launch, Okt 2025) | $3.999 | Ursprünglicher Listenpreis |
| US (Feb 2026) | $4.699 | Preiserhöhung (intuitionlabs.ai-Review) |
| US (Amazon) | $4.899,99 | Effektiver Straßenpreis |
| EU | NVIDIA-Forum, Feb 2026; Verkauf in Europa u. a. über buyzero.de (pi3g) |
Single-Stream-Benchmarks (reale Messungen)
Ollama-Benchmark-Suite (Ollama v0.12.6, FW 580.95.05, 10 Runs, temp 0, 500 Token — ollama.com/blog/nvidia-spark-performance)
| Modell | Quant | Prefill (tok/s) | Decode (tok/s) |
|---|---|---|---|
| gpt-oss 20B | MXFP4 | 3.224 | 58,27 |
| llama3.1 8B | q4 | 1.905 | 38,02 |
| gpt-oss 120B | MXFP4 (MoE) | 1.169 | 41,14 |
| deepseek-r1 14B | q4 | 1.438 | 19,99 |
| gemma3 12B | q4 | 1.323 | 24,25 |
| gemma3 27B | q4 | 766 | 10,83 |
| qwen3 32B | q4 | 598 | 9,41 |
| llama3.1 70B | q4 | 276 | 4,42 |
LMSYS In-Depth Review (13.10.2025, SGLang/Ollama)
- Llama 3.1 8B FP8 (SGLang, batch 1): 7.991 prefill / 20,5 decode — batch 32: 7.949 / 368
- GPT-OSS 20B MXFP4 (Ollama): 2.053 prefill / 49,7 decode
- Llama 3.1 70B FP8 (SGLang): 803 prefill / 2,7 decode — dichte 70B-Modelle sind auf der Box praktisch unbrauchbar (<3 tok/s)
- EAGLE3 speculative decoding: bis ~2× Speedup; kein Thermal-Throttling über längere Läufe
Community / weitere Messungen
- llama.cpp-Diskussion (#16578): gpt-oss-120B ~35 tok/s — nahe am Ollama-Wert; Community-Kritik am Preis-Leistungs-Verhältnis vs. RTX 6000 Pro oder Mac Studio
- Kern-Insight (auch Devsplainers, 26.08.): Memory-Bandbreite statt Petaflops entscheidet Box-Tauglichkeit — 1-PFLOP-Headline, aber 273 GB/s shared LPDDR5x limitieren dichte Modelle; MoE-Modelle (gpt-oss, DeepSeek) profitieren vom Aktivierungsverhältnis
2er-Kopplung
Zwei Sparks über ConnectX-7 (200 Gb/s RDMA) ergeben einen 256-GB-unified-Memory-Pool; NVIDIA demonstrierte Llama 3.1 405B (FP4) lokal auf zwei Geräten. Die aktuelle Produktseite nennt „bis zu 4 Systeme → 700B Parameter" — reale Multi-Node-Benchmarks dazu sind (Stand 08/2026) nicht unabhängig verifiziert.
Einordnung im Hardware-Stack
| Gerät | Memory | Bandbreite | Preis (ca.) | Position |
|---|---|---|---|---|
| DGX Spark | 128 GB | 273 GB/s | $4.699–4.900 / €4.800 | Prosumer-Einstieg, MoE-freundlich |
| Mac Studio M5 Ultra | 512 GB | 1,2 TB/s | ab 6.599 € | Mehr Memory + Bandbreite fürs Geld |
| DGX Station (GB300) | 748 GB | — | ~$85K–115K | Full-precision 70B+, Workstation |
- Reale Stärke: MoE-Modelle (gpt-oss-120B @ 41 t/s, DeepSeek-R1-14B), 8B–27B-Klasse für Agent-Workloads, Perplexity „Portable Computer" als erste vollständig lokale Agent-Plattform auf Spark — siehe perplexity-portable-computer.md.
- Reale Schwäche: dichte 70B+ (2,7 t/s) und der Preis nach der Erhöhung — LMSYS und llama.cpp-Community kommen zu einem kritischen Fazit; DRAM-Teuerung 2026 verstärkt das Problem (siehe ../concepts/hardware/china-local-ai-box.md).
- Verwandte lokale Infrastruktur: ../concepts/hardware/edge-inference-als-cloud-alternative.md (AMD Strix Halo, 128 GB als x86-Alternative), ../concepts/hardware/cloud-exit-and-local-superiority.md, ../concepts/hardware/nvidia-dgx-station-748gb.md (DGX-Ladder Spark → Station → Datacenter), ../concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md (Apple-Gegenseite).
- Software: Ollama, llama.cpp, NIM, LM Studio, DGX OS; „NemoClaw"-Referenz für die OpenClaw-Community auf der NVIDIA-Produktseite — siehe openclaw.md.
Quellen
- NVIDIA Produktseite: https://www.nvidia.com/en-us/products/workstations/dgx-spark/ (Specs, 4-Systeme/700B-Claim, DGX-OS-Update, NemoClaw)
- Ollama-Benchmark-Suite: https://ollama.com/blog/nvidia-spark-performance
- LMSYS In-Depth Review: https://lmsys.org/blog/2025-10-13-nvidia-dgx-spark/
- llama.cpp-Diskussion: https://github.com/ggml-org/llama.cpp/discussions/16578
- Preis US: intuitionlabs.ai-Review (Feb 2026, $4.699), Amazon ($4.899,99); Preis EU: NVIDIA-Developer-Forum (€4.800, Feb 2026), buyzero.de/pi3g
- Kontext im Wiki:
concepts/hardware/nvidia-dgx-station-748gb.md(Spark-Zeile),concepts/hardware/china-local-ai-box.md(DGX-Spark-Lektion: 1 PFLOP Headline, <3 tok/s dichter 70B),tools/perplexity-portable-computer.md(erste Spark-Agent-Plattform)
Verwandte Wiki-Seiten
- NVIDIA-Agenten-Workflow für SimReady-Szenen (17.09.2026): Der DGX Spark wird dort explizit als Prototyping-System für NemoClaw-Subagents, Blender-MCP-Anbindung, USD-Authoring und ovrtx-/ovphysx-Loops empfohlen — die Kategorie „Agenten-Harness auf dem Desktop" neben „lokale Inferenz".
- ../concepts/hardware/nvidia-dgx-station-748gb.md — DGX-Station-Oberklasse (748 GB)
- perplexity-portable-computer.md — erster vollständig lokaler Agent-Stack auf Spark
- ../concepts/hardware/china-local-ai-box.md — chinesische Konkurrenz-Boxen, Memory-Bandbreite-These
- ../concepts/hardware/edge-inference-als-cloud-alternative.md — Strix-Halo-Alternative
- ../concepts/hardware/apple-m6-m5-ultra-mac-mini-studio-update.md — Apple-TB5/RDMA-Cluster-Gegenposition
- ../concepts/hardware/cloud-exit-and-local-superiority.md — Cloud-Exit-Kontext