knowledge-base/raw/youtube/2026-09-16_marfil-draws-deepseek-v41-flash-any-hardware.md

5.7 KiB
Raw Permalink Blame History

type source_url retrieved channel duration_sec has_transcript title uploaded tags
youtube https://www.youtube.com/watch?v=Z0lkcQK2Oj8 2026-09-17 Marfil Draws 806 false Run DeepSeek V4.1 Flash on ANY hardware (16GB to 512GB): can it work? 2026-09-16
deepseek
deepseek-v4.1-flash
local-inference
hardware
quantization
dgx-spark
apple-silicon
moe

Run DeepSeek V4.1 Flash on ANY hardware (16GB to 512GB): can it work?

Video von Marfil Draws (@themarfildraws, 332 Abonnenten), hochgeladen 16.09.2026, Länge 13:26 (806 s), ca. 2.983 Views / 40 Likes zum Abrufzeitpunkt. Kein Transcript abrufbar; Inhalt aus Beschreibung, Kapitelmarken und Quellenliste (Firecrawl-Scrape, 17.09.2026).

Charakter des Videos: Aggregation fremder Messwerte, keine eigene Messung. Der Autor bezieht sich auf 25 namentlich genannte Quellen. Eigenangabe des Autors: Zahlen aus Reddit, HN, X und Forenposts sind individuelle Nutzerberichte, keine Mittelwerte.

Kapitelmarken

Zeit Kapitel
00:00 Why a Flash model is so big
01:12 Apps, programs and the Q2 or Q4 builds
03:02 How to judge speed
03:42 The computer you already own
05:24 128 GB machines
07:41 Settings that change the result
08:34 Two machines and 256 GB
10:31 512 GB and Spark clusters
11:49 Pro graphics cards
12:34 The API and why people buy anyway

Architektur (Kapitel 00:00)

  • Mixture-of-Experts: 384 Experten pro Layer, 6 pro Token aktiv.
  • 552B Hauptparameter + 196B in zwei Engram-Tabellen.
  • Pro Token ~4,51 GB Experten, aber nur ~12,4 KB Tabellen → Tabellen können auf der SSD liegen, Experten nicht.
  • 1M-Token-Kontext belegt < 1 GB.

Builds und Quantisierung (Kapitel 01:12)

  • Download gesamt ~510 GB.
  • DwarfStar Q2: 366-GB-Datei, ~163 GB RAM-Bedarf.
  • DwarfStar Q4: 519 GB, ~316 GB RAM-Bedarf.
  • llama.cpp unterstützt nur Cloud; auf NVIDIA-PCs nur inoffizielle Mods — einer davon mit altem Chat-Template → schwächeres Tool-Calling.
  • Q2 gegen Q4: 90,08 % vs. 96,7 % Übereinstimmung mit der DeepSeek-API beim wahrscheinlich nächsten Token.

Geschwindigkeitsmaßstab (Kapitel 03:02)

  • 5 tok/s zu langsam für Agenten (nachts ok), ~15 Minimum, Coding ≥ 20.
  • Zweite Bremse ist der Prefill: 256-GB-M3-Ultra mit 15.688-Token-Prompt ≈ 99 s pro Call, 6,6 s bei Wiederverwendung.

Hardware-Stufen (Kapitel 03:4207:41)

Konfiguration Preis Durchsatz (tok/s) Notiz
16-GB-Mac-mini ~23 s pro Token 108 s bis erstes Wort, ~13 h für 2.048 Token
RTX 5090 + 128 GB RAM 5,12 54 % jedes Tokens wartet auf die SSD
Strix-Halo-Boxen 5,19,7
1× DGX Spark $4.699 710
M5 Max Mac Studio (128 GB) $5.399 Q2 1516,4; Q4 10,410,6 stock, 14,38 mit Mod Ein früher Q4-Load startete den Mac neu

Einstellungen (Kapitel 07:41)

  • Fester Harness-Prompt: Claude Code ~2025k, opencode ~10k Token plus wechselnde Attribution-Zeile → lokal ~90 % langsamer, bis der Header auf null steht.
  • Reasoning-Effort hoch/max → Tool-Call-Schleifen; DeepSeek selbst empfiehlt 6080.
  • Ein Test: 788k statt 141k Output-Token.

Zwei Maschinen, 256 GB, Cluster (Kapitel 08:3410:31)

Konfiguration Preis Durchsatz (tok/s) Notiz
2× M5 Max $10.867 2530
2× DGX Spark 21,9 (2,9-Bit 31,6) Ein OOM fror beide ein
M3 Ultra 256 GB 29,85 auf Code Q2 voll im RAM
M3 Ultra 512 GB 18,118,7 Q4 im RAM
3× Spark im Dreieck ~$14.532 37,9 Erstmals komplett im RAM, aber nur ~32k Kontext
4× Spark ~$20.091 4574 1M Kontext — laut Video „paying twice for at best 50 % more speed"

Pro-Grafikkarten und Pruning (Kapitel 11:49)

  • 2× RTX PRO 6000 (~$32.000): 113,3 auf Q2.
  • Ein geprunter Build: 141 tok/s, aber eine Experten-Auswahl scored 45 % in Chemie — laut Video „pruning can quietly delete whole subjects".
  • 4 Karten ($77k98k): 200260; Speculative Decoding crasht bei ≥ ~4.000 Token.

API-Ökonomie (Kapitel 12:34)

  • API: 15 ¢/M Input off-peak, 60 ¢ Output, doppelt zu Peak, ~200 tok/s.
  • Beliebtester Gegenfall: $3 Overnight-API-Lauf; $9.499 Mac gegen $200/Monat → 47,5 Monate Payback (grobe Arithmetik, ohne Wiederverkauf/Speed/Qualität).
  • Fazit im Video: „On money alone, no" — gekauft wird für Privacy und Ownership.

Quellen (vom Autor genannt)

  • NVIDIA Developer Forum: „DeepSeek v4.1 Flash" (108 Posts zum Abruf)
  • r/LocalLLaMA: „DeepSeek V4.1 Flash: Stronger, Faster, More Accessible" (219 Punkte, 57 Kommentare)
  • Kai: „What It Actually Costs to Run DeepSeek V4.1 Flash Locally?" (256.687 Views, 174 Kommentare heruntergeladen)
  • NVIDIA Marketplace US: DGX-Spark-Produktseite
  • DwarfStar QA_BEFORE_RELEASES.md (Kopie c1_ds4_qa)
  • drowzeys README: V4.1 Flash auf einer 256-GB-M3-Ultra mit oMLX
  • B&H Photo: RTX PRO 6000 Blackwell
  • MikroTik: CRS804 DDQ
  • Petronella Technology: DGX Spark 0.5 m QSFP112 400G DAC
  • Hacker News: „Show HN: Sunk Cost How long until a local LLM rig pays for itself?" (46 Punkte, 98 Kommentare)
  • r/LocalLLM: „OK guys, let's be honest 1 minute about local LLM" (273 Punkte, 528 Kommentare)
  • antirez/ds4: QA_BEFORE_RELEASES.md (Sektion 17, V4.1-Messungen)
  • DeepSeek: offizielle Modelcard, config.json, V4.1-Flash Technical Report (PDF)
  • LibertAI: NVFP4-Repack + FP4-Engram
  • antirez/ds4 (DwarfStar): DGX-Spark-Guide
  • Manolo Remiddi: „Apple Just Made the Best Local AI Machine. Do Not Buy It Yet." (235.693 Views, 437 Kommentare)