knowledge-base/raw/youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md
Hector 0007c0cb50 ingest(hardware): nvidia dgx station gb300 — 748gb unified memory desktop
- raw: youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md (The Stack, 8:14)
- wiki: concepts/hardware/nvidia-dgx-station-748gb.md (NEW — 748 GB unified, full-precision 70B, ~$85K-$115K, ROI break-even ~2 months)
- index: 39. Update, new hardware row + raw source entry
- log: entry appended
2026-06-28 15:09:44 +02:00

80 lines
No EOL
4.5 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: youtube
source_url: https://www.youtube.com/watch?v=EhXQysElOY8
retrieved: 2026-06-28
channel: "The Stack"
title: "NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good"
duration_sec: 494
views: 2983
published: 2026-06-27
has_transcript: true
tags: [nvidia, dgx-station, local-ai, unified-memory, gb300, grace-blackwell, lm-studio, ollama, hardware, cloud-exit]
---
# NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good
**Channel:** [The Stack](https://www.youtube.com/@The-Stack-ai)
**Published:** 2026-06-27
**Duration:** 8:14
**Views:** ~2,983 (at time of fetch, 7h after publish)
**Hashtags:** #localai #nvidiadgx #lmstudio #computex2026 #agenticai
## Video Description (full)
LM Studio local AI just changed: NVIDIA's DGX Station packs 748GB unified memory to run 70B models in full precision, no cloud needed.
LM Studio and Ollama finally lose their asterisk, NVIDIA's DGX Station at Computex 2026 lands 748GB of coherent unified memory in a single deskside tower, enough to load a full-precision 70B model with room to spare, no quantization required, no cloud offload.
Announced by Jensen Huang at GTC Taipei on May 31, 2026, the machine is built around the GB300 Grace Blackwell Ultra Desktop Superchip: a 72-core ARM Grace CPU fused to a Blackwell Ultra GPU via NVLink-C2C at 900 GB/s. The memory pool splits into 252GB HBM3e (7.1 TB/s GPU-side) and 496GB LPDDR5X (CPU-side), both fully coherent, one address space, zero explicit copies. Compute tops out at 20 petaFLOPS FP4. NVIDIA doesn't sell a Founders Edition; OEM partners ASUS, Dell, HP, MSI, and others handle that, with real-world pricing landing between roughly $85K and $115K (the MSI XpertStation WS300 lists at $96,995.99 on CDW). The video also covers NVIDIA's DGX Spark (128GB, ~$4,700) as the genuine prosumer entry point, and gives an honest head-to-head with the Mac Studio M5 Ultra, which still holds the value crown for a solo developer running mid-size models. The trillion-parameter claim gets a reality check, it's technically true only with aggressive 4-bit quantization, not full-precision weights. The cloud ROI math is real: at ~$98/hour for a comparable AWS p5 instance, the hardware pays for itself in roughly two months of sustained workload. The DGX Station for Windows (WSL-based) is flagged as a Q4 2026 promise, not a shipping product. RTX Spark, NVIDIA's MediaTek-partnered consumer AI PC chip, rounds out the roadmap alongside a three-generation plan through Rubin and Rosa Feynman.
For individual builders and small teams deciding between local AI options, this is a practical breakdown of which box on the NVIDIA ladder actually makes sense for their workload.
## Chapters
- 0:00 Intro
- 0:15 What Jensen actually unveiled
- 1:12 Why unified memory is the whole story
- 2:31 The trillion-parameter asterisk
- 3:25 The price, who it's for, and how to choose
- 5:25 The cloud math that justifies the big box
- 6:40 The bigger play: NVIDIA wants the whole PC
## Tools & Resources Mentioned
- **LM Studio:** https://lmstudio.ai
- **Ollama:** https://ollama.com
- **NVIDIA DGX Station:** https://www.nvidia.com/en-us/products/workstations/dgx-station/
- **NVIDIA DGX Station for Windows:** https://www.nvidia.com/en-us/products/workstations/dgx-station-for-windows/
- **NVIDIA DGX Spark:** https://www.nvidia.com/en-us/products/workstations/dgx-spark/
## Key Specs Summary
| Spec | Value |
|------|-------|
| Chip | GB300 Grace Blackwell Ultra Desktop Superchip |
| CPU | 72-core ARM Grace |
| GPU | Blackwell Ultra |
| Interconnect | NVLink-C2C @ 900 GB/s |
| Total Unified Memory | 748 GB (252 GB HBM3e + 496 GB LPDDR5X) |
| HBM3e Bandwidth | 7.1 TB/s (GPU-side) |
| Compute | 20 petaFLOPS FP4 |
| Price Range | ~$85K$115K (OEM-dependent) |
| MSI XpertStation WS300 | $96,995.99 (CDW listing) |
| DGX Spark (entry) | 128 GB, ~$4,700 |
| Full-precision 70B model | Fits with room to spare |
| Trillion-parameter claim | Only with 4-bit quantization |
| Cloud ROI break-even | ~2 months (vs AWS p5 @ ~$98/h) |
| DGX Station for Windows | Q4 2026 (WSL-based, not shipping) |
| Announced | GTC Taipei, May 31, 2026 by Jensen Huang |
## NVIDIA Local AI Hardware Ladder
| Product | Memory | Price | Target |
|---------|--------|-------|--------|
| DGX Spark | 128 GB | ~$4,700 | Prosumer entry |
| DGX Station | 748 GB | ~$85K$115K | Small teams / sustained workloads |
| Mac Studio M5 Ultra | (value crown for solo devs) | — | Solo developer, mid-size models |
## Context
Shared by Pit Weber in OME-Gruppe, Topic "News & Infos" (Topic 13) on 2026-06-28.