--- type: other source_url: https://api-docs.deepseek.com/news/news260910 retrieved: 2026-09-10 title: "DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (offizielle Ankündigung)" author: "DeepSeek" tags: [deepseek, deepseek-v4.1-flash, release, architecture, moe, causal-encoder-decoder, kv-cache, multimodal, api, pricing, open-weights] --- # DeepSeek-V4.1-Flash — Offizielle Ankündigung (API-Docs, 10.09.2026) > **Quelle:** Offizielle DeepSeek-API-Docs-Ankündigung (10.09.2026): https://api-docs.deepseek.com/news/news260910 — Primärquelle zum V4.1-Flash-Release, verlinkt aus dem X-Post von @deepseek_ai (https://x.com/deepseek_ai/status/2097930608790167907). ## Kernaussagen (wörtlich aus der Ankündigung) > 🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. > > 🔹 Introducing the smallest model in our new architecture family, with native visual understanding. > > 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models. ## Asymmetrische Architektur - **552B-Parameter-MoE** (Mixture-of-Experts). - **Neue Causal-Encoder–Decoder-Architektur (CED):** nur **8B aktive Parameter für Input**, **16B für Output** (asymmetrisch, Input/Output nicht symmetrisch). - Neue Pre-Training-Methoden + größer skaliertes RL-Post-Training liefern Benchmark-Ergebnisse vor Flaggschiff-Modellen, inkl. **DeepSeek-V4-Pro**. ## Kleinere KV-Cache - Gegenüber der Vorgängergeneration braucht V4.1-Flash nur **1/4 des HBM** und **1/8 des SSD-Speichers** für den KV-Cache. - Cache-Hit-Gebühren machen oft einen großen Anteil der Agent-Kosten aus — Cache-Kompression senkt diese Kosten deutlich. ## API-Verfügbarkeit - **Live auf der DeepSeek-API mit nativem Multimodal-Support** — Modellname: `deepseek-flash`. - **V4-Flash & V4-Flash-Vision-Exp sind retired.** Für Kompatibilität routen `deepseek-v4-flash` und `deepseek-v4-flash-vision-exp` temporär auf V4.1-Flash. - Tests mehrerer Parteien sehen V4.1-Flash **vor V4-Pro** bei Performance, Kosten, Speed und Gesamtlaufzeit. V4-Pro wird ausphasen. - **Ab 04:00 UTC, 14.09.2026:** alle `deepseek-v4-pro`-Requests routen auf V4.1-Flash zu V4.1-Flash-Raten — bis V4.1-Pro launcht. - **Offizielle Partner:** WorkBuddy (inkl. CodeBuddy) und OpenCode unterstützen V4.1-Flash vollständig. ## Pricing - Effizientere Architektur → niedrigere API-Preise. - Peak/Off-Peak-Pricing: **Off-Peak = 50 % der Peak-Raten**. - **Neue Preise ab 04:00 UTC, 10.09.2026.** ## Open Source & Deployment - Zusammenarbeit mit der Open-Source-Community für V4.1-Flash-Inference-Support; weitere Deployment-Optionen geplant. - Großskalige Deployments (2.000 GPUs + Storage-Cluster) ansprechbar. - **Modell (Hugging Face):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash - **Paper (Tech Report):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf ## Einordnung - Primärquelle (offiziell) zum V4.1-Flash-Release; X-Post von @deepseek_ai als Ankündigung: `raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md`. - V4.1 Flash ist eine neue Modell-Version gegenüber dem V4-Flash im aktiven Stack ([[../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md]]).