120 lines
8 KiB
Markdown
120 lines
8 KiB
Markdown
|
|
---
|
|||
|
|
created: 2026-09-17
|
|||
|
|
updated: 2026-09-17
|
|||
|
|
sources: [youtube/2026-09-16_marfil-draws-deepseek-v41-flash-any-hardware.md]
|
|||
|
|
tags: [concept, hardware, deepseek-v4.1-flash, local-inference, quantization, dgx-spark, apple-silicon, strix-halo, moe, engram, cost-analysis]
|
|||
|
|
---
|
|||
|
|
|
|||
|
|
# DeepSeek V4.1 Flash auf lokaler Hardware
|
|||
|
|
|
|||
|
|
> **TL;DR:** [DeepSeek V4.1 Flash](../llm/deepseek-v4.1-flash.md) ist als 552B-MoE mit ~510 GB Download ein Extremfall für lokale Inferenz. Das Video „Run DeepSeek V4.1 Flash on ANY hardware" ([Marfil Draws](../../people/marfil-draws.md), 16.09.2026) sortiert die Hardware-Stufen von 16 GB bis 512 GB und kommt zum Schluss: Unterhalb von ~15 tok/s ist das Modell für Agenten unbrauchbar, und **rein monetär lohnt sich lokaler Betrieb nicht** ($9.499-Mac vs. $200/Monat → 47,5 Monate Payback). Gekauft wird für Privacy und Ownership, nicht für die Rechnung. ⚠️ Alle Durchsatz- und Preiszahlen sind **aggregierte Einzelberichte** aus Foren, Reddit und HN — keine eigene Messung, keine Mittelwerte.
|
|||
|
|
|
|||
|
|
## Quelle
|
|||
|
|
|
|||
|
|
| Quelle | Kanal | Datum | Länge | Reach |
|
|||
|
|
|--------|-------|-------|-------|-------|
|
|||
|
|
| [Run DeepSeek V4.1 Flash on ANY hardware (16GB to 512GB)](https://www.youtube.com/watch?v=Z0lkcQK2Oj8) | Marfil Draws | 2026-09-16 | 13:26 | ~2.983 Views, 40 Likes |
|
|||
|
|
|
|||
|
|
**Geteilt von:** Kai (@PWeber) im OME-Topic „Tips & Tricks". Kein Transcript abrufbar (YouTube-Bot-Sperre); Grundlage sind Beschreibung, Kapitelmarken und Quellenliste.
|
|||
|
|
|
|||
|
|
## Warum das Modell so groß ist
|
|||
|
|
|
|||
|
|
| Größe | Wert |
|
|||
|
|
|-------|------|
|
|||
|
|
| Experten | 384 pro Layer, 6 aktiv pro Token |
|
|||
|
|
| Parameter | 552B Haupt + 196B in zwei Engram-Tabellen |
|
|||
|
|
| Pro Token gelesen | ~4,51 GB Experten, ~12,4 KB Tabellen |
|
|||
|
|
| Engram-Tabellen | können auf der SSD liegen — die Experten nicht |
|
|||
|
|
| 1M-Token-Kontext | < 1 GB |
|
|||
|
|
|
|||
|
|
**Primärquellen-Check:** Die Werte decken sich mit der HF-`config.json` (abgerufen 17.09.2026): `n_routed_experts: 384`, `num_experts_per_tok: 6`, `engram_layer_ids: [1, 14]`, `max_position_embeddings: 1048576`, `expert_dtype: "fp4"`. Die Architektur-Beschreibung des Videos ist belegt; nur die abgeleiteten Durchsatz- und Preiszahlen bleiben Einzelberichte.
|
|||
|
|
|
|||
|
|
## Quantisierung: Q2 gegen Q4
|
|||
|
|
|
|||
|
|
| Build | Dateigröße | RAM-Bedarf | API-Übereinstimmung |
|
|||
|
|
|-------|-----------|------------|---------------------|
|
|||
|
|
| DwarfStar **Q2** | 366 GB | ~163 GB | 90,08 % |
|
|||
|
|
| DwarfStar **Q4** | 519 GB | ~316 GB | 96,7 % |
|
|||
|
|
|
|||
|
|
Der Download liegt gesamt bei ~510 GB. Laut Video unterstützt **llama.cpp nur Cloud**; auf NVIDIA-PCs laufen nur inoffizielle Mods — einer mit veraltetem Chat-Template, was Tool-Calling verschlechtert. Die Quantisierungsfrage ist damit kein Detail: Der Sprung Q2 → Q4 kostet knapp die doppelte RAM-Menge für 6,6 Prozentpunkte API-Nähe.
|
|||
|
|
|
|||
|
|
## Geschwindigkeit beurteilen
|
|||
|
|
|
|||
|
|
- **5 tok/s** = zu langsam für Agenten (nachts akzeptabel), **~15 tok/s** Minimum, **Coding ≥ 20 tok/s**.
|
|||
|
|
- Die zweite Bremse neben dem Durchsatz ist der **Prefill**: 256-GB-M3-Ultra mit 15.688-Token-Prompt ≈ **99 s pro Call**, 6,6 s bei Wiederverwendung. Kontext-Caching ist damit wichtiger als der Token-Durchsatz.
|
|||
|
|
|
|||
|
|
## Hardware-Stufen
|
|||
|
|
|
|||
|
|
| Stufe | Konfiguration | Preis | tok/s | Notiz |
|
|||
|
|
|-------|---------------|-------|-------|-------|
|
|||
|
|
| Bestand | 16-GB-Mac-mini | — | ~23 s **pro Token** | 108 s bis erstes Wort, ~13 h für 2.048 Token |
|
|||
|
|
| Bestand | RTX 5090 + 128 GB RAM | — | 5,12 | 54 % jedes Tokens warten auf die SSD |
|
|||
|
|
| 128 GB | Strix-Halo-Box | — | 5,1–9,7 | |
|
|||
|
|
| 128 GB | 1× DGX Spark | $4.699 | 7–10 | vgl. [[../../tools/nvidia-dgx-spark.md]] |
|
|||
|
|
| 128 GB | M5 Max Mac Studio | $5.399 | Q2 15–16,4 / Q4 10,4–10,6 | mit Mod 14,38; ein früher Q4-Load startete den Mac neu |
|
|||
|
|
| 256 GB | 2× M5 Max | $10.867 | 25–30 | |
|
|||
|
|
| 256 GB | 2× DGX Spark | — | 21,9 (2,9-Bit 31,6) | ein OOM fror beide ein |
|
|||
|
|
| 256 GB | M3 Ultra | — | 29,85 auf Code | Q2 voll im RAM |
|
|||
|
|
| 512 GB | M3 Ultra | — | 18,1–18,7 | Q4 im RAM |
|
|||
|
|
| Cluster | 3× Spark im Dreieck | ~$14.532 | 37,9 | erstmals komplett im RAM, aber nur ~32k Kontext |
|
|||
|
|
| Cluster | 4× Spark | ~$20.091 | 45–74 | 1M Kontext — laut Video „paying twice for at best 50 % more speed" |
|
|||
|
|
|
|||
|
|
## Einstellungen, die das Ergebnis verändern
|
|||
|
|
|
|||
|
|
- **Harness-Prompt-Größe:** Claude Code ~20–25k feste Token, opencode ~10k, plus wechselnde Attribution-Zeile — laut Video macht das lokal ~**90 % langsamer**, bis der Header auf null steht.
|
|||
|
|
- **Reasoning-Effort:** hohe/max-Werte erzeugen Tool-Call-Schleifen; DeepSeek empfiehlt selbst **60–80**.
|
|||
|
|
- Ein dokumentierter Testfall: **788k statt 141k** Output-Token.
|
|||
|
|
|
|||
|
|
## Pro-Karten und die Pruning-Falle
|
|||
|
|
|
|||
|
|
| Konfiguration | Preis | tok/s | Notiz |
|
|||
|
|
|---------------|-------|-------|-------|
|
|||
|
|
| 2× RTX PRO 6000 | ~$32.000 | 113,3 auf Q2 | |
|
|||
|
|
| geprunter Build | — | 141 | eine Experten-Auswahl scored **45 % in Chemie** |
|
|||
|
|
| 4 Karten | $77k–98k | 200–260 | Speculative Decoding crasht bei ≥ ~4.000 Token |
|
|||
|
|
|
|||
|
|
Die Pruning-Anekdote ist der lehrreichste Einzelbefund: Geschwindigkeit durch Experten-Reduktion kann still ganze Fachgebiete löschen. Schneller heißt nicht besser.
|
|||
|
|
|
|||
|
|
## API gegen lokalen Betrieb
|
|||
|
|
|
|||
|
|
| Position | Wert |
|
|||
|
|
|----------|------|
|
|||
|
|
| API Input off-peak | 15 ¢/M Token |
|
|||
|
|
| API Output | 60 ¢/M Token |
|
|||
|
|
| Peak | doppelte Rate |
|
|||
|
|
| API-Durchsatz | ~200 tok/s |
|
|||
|
|
| Payback $9.499-Mac gegen $200/Monat | 47,5 Monate |
|
|||
|
|
|
|||
|
|
Das Video zieht das Fazit „On money alone, no" — die Rechnung spricht für die API. Die 47,5 Monate sind eine grobe Arithmetik ohne Wiederverkauf, Speed- und Qualitätsvergleich.
|
|||
|
|
|
|||
|
|
## Einordnung
|
|||
|
|
|
|||
|
|
Das Video ist bemerkenswert sauber für die Gattung: Es ist ausdrücklich eine **Sammlung fremder Messwerte** mit namentlicher Quellenliste und eigener Fehlerangabe („individual user reports, not averages"). Die Architektur-Angaben sind an der HF-`config.json` verifizierbar; die Durchsatz- und Preiszahlen sind es nicht — sie stammen aus Foren, Reddit und HN und wurden nicht nachgefahren.
|
|||
|
|
|
|||
|
|
Zwei Dinge bleiben hängen. Erstens: Die **Bottleneck-Diagnose** (Memory-Bandbreite statt Petaflops) ist konsistent mit dem, was dieses Wiki an anderer Stelle zu lokaler Hardware dokumentiert ([[../../tools/nvidia-dgx-spark.md]], [[../llm/local-llm-laptop-guide.md]], [[lokale-ki-coding.md]]). Zweitens: Die Payback-Rechnung spricht für genau den Weg, den dieses Setup schon geht — V4.1-Flash läuft hier über die Ollama-Cloud ([[../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md]]) statt über eigene Hardware.
|
|||
|
|
|
|||
|
|
Für die eigene Praxis bleibt der nüchterne Befund: Unter ~15 tok/s ist ein Agent unbrauchbar, und darunter fällt praktisch jede bezahlbare Einzelmaschine. Lokale V4.1-Flash-Inferenz ist heute eine Ownership- und Privacy-Entscheidung, keine Kostenentscheidung.
|
|||
|
|
|
|||
|
|
## Cross-References
|
|||
|
|
|
|||
|
|
- [[../llm/deepseek-v4.1-flash.md]] — Modell-Seite (Architektur, API, Benchmarks)
|
|||
|
|
- [[../../tools/nvidia-dgx-spark.md]] — DGX Spark (128-GB-Stufe, Cluster)
|
|||
|
|
- [[../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md]] — der Cloud-Weg, den dieses Setup nutzt
|
|||
|
|
- [[lokale-ki-coding.md]] — lokale Coding-Kosten und Hardware-Abhängigkeit
|
|||
|
|
- [[../llm/local-llm-laptop-guide.md]] — Laptop-Klasse statt 512-GB-Klasse
|
|||
|
|
- [[../llm/mlx-moe-local-ai-optimization.md]] — MLX/MoE auf Apple Silicon
|
|||
|
|
- [[../llm/chinese-model-cost-routing.md]] — Kosten-Routing zwischen Cloud und lokal
|
|||
|
|
- [[nvidia-dgx-station-748gb.md]] — Enterprise-Stufe
|
|||
|
|
- [[cloud-exit-and-local-superiority.md]] — Cloud-Exit-These
|
|||
|
|
- [[../../institutions/nvidia.md]] — NVIDIA als Institution
|
|||
|
|
|
|||
|
|
## Externe Links
|
|||
|
|
|
|||
|
|
- [Video (YouTube)](https://www.youtube.com/watch?v=Z0lkcQK2Oj8)
|
|||
|
|
- [DeepSeek V4.1 Flash — HF-Modelcard](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash)
|
|||
|
|
- [DeepSeek V4.1 Flash — config.json](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/raw/main/config.json)
|
|||
|
|
- [DeepSeek V4.1 Flash — Technical Report (PDF)](https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf)
|
|||
|
|
- [antirez/ds4 (DwarfStar) — DGX-Spark-Guide](https://github.com/antirez/ds4/blob/main/docs/DGX_SPARK.md)
|
|||
|
|
- [NVIDIA Developer Forum — DeepSeek v4.1 Flash](https://forums.developer.nvidia.com/t/382725)
|
|||
|
|
- [Hacker News — Show HN: Sunk Cost](https://news.ycombinator.com/item?id=49706656)
|