- raw: youtube/2026-06-28_nvidia-748gb-ram-desktop-local-ai.md (The Stack, 8:14) - wiki: concepts/hardware/nvidia-dgx-station-748gb.md (NEW — 748 GB unified, full-precision 70B, ~$85K-$115K, ROI break-even ~2 months) - index: 39. Update, new hardware row + raw source entry - log: entry appended
80 lines
No EOL
4.5 KiB
Markdown
80 lines
No EOL
4.5 KiB
Markdown
---
|
||
type: youtube
|
||
source_url: https://www.youtube.com/watch?v=EhXQysElOY8
|
||
retrieved: 2026-06-28
|
||
channel: "The Stack"
|
||
title: "NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good"
|
||
duration_sec: 494
|
||
views: 2983
|
||
published: 2026-06-27
|
||
has_transcript: true
|
||
tags: [nvidia, dgx-station, local-ai, unified-memory, gb300, grace-blackwell, lm-studio, ollama, hardware, cloud-exit]
|
||
---
|
||
|
||
# NVIDIA'S 748GB Ram Desktop Makes Local AI INSANELY Good
|
||
|
||
**Channel:** [The Stack](https://www.youtube.com/@The-Stack-ai)
|
||
**Published:** 2026-06-27
|
||
**Duration:** 8:14
|
||
**Views:** ~2,983 (at time of fetch, 7h after publish)
|
||
**Hashtags:** #localai #nvidiadgx #lmstudio #computex2026 #agenticai
|
||
|
||
## Video Description (full)
|
||
|
||
LM Studio local AI just changed: NVIDIA's DGX Station packs 748GB unified memory to run 70B models in full precision, no cloud needed.
|
||
|
||
LM Studio and Ollama finally lose their asterisk, NVIDIA's DGX Station at Computex 2026 lands 748GB of coherent unified memory in a single deskside tower, enough to load a full-precision 70B model with room to spare, no quantization required, no cloud offload.
|
||
|
||
Announced by Jensen Huang at GTC Taipei on May 31, 2026, the machine is built around the GB300 Grace Blackwell Ultra Desktop Superchip: a 72-core ARM Grace CPU fused to a Blackwell Ultra GPU via NVLink-C2C at 900 GB/s. The memory pool splits into 252GB HBM3e (7.1 TB/s GPU-side) and 496GB LPDDR5X (CPU-side), both fully coherent, one address space, zero explicit copies. Compute tops out at 20 petaFLOPS FP4. NVIDIA doesn't sell a Founders Edition; OEM partners ASUS, Dell, HP, MSI, and others handle that, with real-world pricing landing between roughly $85K and $115K (the MSI XpertStation WS300 lists at $96,995.99 on CDW). The video also covers NVIDIA's DGX Spark (128GB, ~$4,700) as the genuine prosumer entry point, and gives an honest head-to-head with the Mac Studio M5 Ultra, which still holds the value crown for a solo developer running mid-size models. The trillion-parameter claim gets a reality check, it's technically true only with aggressive 4-bit quantization, not full-precision weights. The cloud ROI math is real: at ~$98/hour for a comparable AWS p5 instance, the hardware pays for itself in roughly two months of sustained workload. The DGX Station for Windows (WSL-based) is flagged as a Q4 2026 promise, not a shipping product. RTX Spark, NVIDIA's MediaTek-partnered consumer AI PC chip, rounds out the roadmap alongside a three-generation plan through Rubin and Rosa Feynman.
|
||
|
||
For individual builders and small teams deciding between local AI options, this is a practical breakdown of which box on the NVIDIA ladder actually makes sense for their workload.
|
||
|
||
## Chapters
|
||
|
||
- 0:00 Intro
|
||
- 0:15 What Jensen actually unveiled
|
||
- 1:12 Why unified memory is the whole story
|
||
- 2:31 The trillion-parameter asterisk
|
||
- 3:25 The price, who it's for, and how to choose
|
||
- 5:25 The cloud math that justifies the big box
|
||
- 6:40 The bigger play: NVIDIA wants the whole PC
|
||
|
||
## Tools & Resources Mentioned
|
||
|
||
- **LM Studio:** https://lmstudio.ai
|
||
- **Ollama:** https://ollama.com
|
||
- **NVIDIA DGX Station:** https://www.nvidia.com/en-us/products/workstations/dgx-station/
|
||
- **NVIDIA DGX Station for Windows:** https://www.nvidia.com/en-us/products/workstations/dgx-station-for-windows/
|
||
- **NVIDIA DGX Spark:** https://www.nvidia.com/en-us/products/workstations/dgx-spark/
|
||
|
||
## Key Specs Summary
|
||
|
||
| Spec | Value |
|
||
|------|-------|
|
||
| Chip | GB300 Grace Blackwell Ultra Desktop Superchip |
|
||
| CPU | 72-core ARM Grace |
|
||
| GPU | Blackwell Ultra |
|
||
| Interconnect | NVLink-C2C @ 900 GB/s |
|
||
| Total Unified Memory | 748 GB (252 GB HBM3e + 496 GB LPDDR5X) |
|
||
| HBM3e Bandwidth | 7.1 TB/s (GPU-side) |
|
||
| Compute | 20 petaFLOPS FP4 |
|
||
| Price Range | ~$85K–$115K (OEM-dependent) |
|
||
| MSI XpertStation WS300 | $96,995.99 (CDW listing) |
|
||
| DGX Spark (entry) | 128 GB, ~$4,700 |
|
||
| Full-precision 70B model | Fits with room to spare |
|
||
| Trillion-parameter claim | Only with 4-bit quantization |
|
||
| Cloud ROI break-even | ~2 months (vs AWS p5 @ ~$98/h) |
|
||
| DGX Station for Windows | Q4 2026 (WSL-based, not shipping) |
|
||
| Announced | GTC Taipei, May 31, 2026 by Jensen Huang |
|
||
|
||
## NVIDIA Local AI Hardware Ladder
|
||
|
||
| Product | Memory | Price | Target |
|
||
|---------|--------|-------|--------|
|
||
| DGX Spark | 128 GB | ~$4,700 | Prosumer entry |
|
||
| DGX Station | 748 GB | ~$85K–$115K | Small teams / sustained workloads |
|
||
| Mac Studio M5 Ultra | (value crown for solo devs) | — | Solo developer, mid-size models |
|
||
|
||
## Context
|
||
|
||
Shared by Pit Weber in OME-Gruppe, Topic "News & Infos" (Topic 13) on 2026-06-28. |