knowledge-base/raw/other/2026-09-10_deepseek-v41-flash-official-announcement.md

58 lines
3.2 KiB
Markdown
Raw Normal View History

---
type: other
source_url: https://api-docs.deepseek.com/news/news260910
retrieved: 2026-09-10
title: "DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (offizielle Ankündigung)"
author: "DeepSeek"
tags: [deepseek, deepseek-v4.1-flash, release, architecture, moe, causal-encoder-decoder, kv-cache, multimodal, api, pricing, open-weights]
---
# DeepSeek-V4.1-Flash — Offizielle Ankündigung (API-Docs, 10.09.2026)
> **Quelle:** Offizielle DeepSeek-API-Docs-Ankündigung (10.09.2026): https://api-docs.deepseek.com/news/news260910 — Primärquelle zum V4.1-Flash-Release, verlinkt aus dem X-Post von @deepseek_ai (https://x.com/deepseek_ai/status/2097930608790167907).
## Kernaussagen (wörtlich aus der Ankündigung)
> 🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
>
> 🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
>
> 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
## Asymmetrische Architektur
- **552B-Parameter-MoE** (Mixture-of-Experts).
- **Neue Causal-EncoderDecoder-Architektur (CED):** nur **8B aktive Parameter für Input**, **16B für Output** (asymmetrisch, Input/Output nicht symmetrisch).
- Neue Pre-Training-Methoden + größer skaliertes RL-Post-Training liefern Benchmark-Ergebnisse vor Flaggschiff-Modellen, inkl. **DeepSeek-V4-Pro**.
## Kleinere KV-Cache
- Gegenüber der Vorgängergeneration braucht V4.1-Flash nur **1/4 des HBM** und **1/8 des SSD-Speichers** für den KV-Cache.
- Cache-Hit-Gebühren machen oft einen großen Anteil der Agent-Kosten aus — Cache-Kompression senkt diese Kosten deutlich.
## API-Verfügbarkeit
- **Live auf der DeepSeek-API mit nativem Multimodal-Support** — Modellname: `deepseek-flash`.
- **V4-Flash & V4-Flash-Vision-Exp sind retired.** Für Kompatibilität routen `deepseek-v4-flash` und `deepseek-v4-flash-vision-exp` temporär auf V4.1-Flash.
- Tests mehrerer Parteien sehen V4.1-Flash **vor V4-Pro** bei Performance, Kosten, Speed und Gesamtlaufzeit. V4-Pro wird ausphasen.
- **Ab 04:00 UTC, 14.09.2026:** alle `deepseek-v4-pro`-Requests routen auf V4.1-Flash zu V4.1-Flash-Raten — bis V4.1-Pro launcht.
- **Offizielle Partner:** WorkBuddy (inkl. CodeBuddy) und OpenCode unterstützen V4.1-Flash vollständig.
## Pricing
- Effizientere Architektur → niedrigere API-Preise.
- Peak/Off-Peak-Pricing: **Off-Peak = 50 % der Peak-Raten**.
- **Neue Preise ab 04:00 UTC, 10.09.2026.**
## Open Source & Deployment
- Zusammenarbeit mit der Open-Source-Community für V4.1-Flash-Inference-Support; weitere Deployment-Optionen geplant.
- Großskalige Deployments (2.000 GPUs + Storage-Cluster) ansprechbar.
- **Modell (Hugging Face):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- **Paper (Tech Report):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
## Einordnung
- Primärquelle (offiziell) zum V4.1-Flash-Release; X-Post von @deepseek_ai als Ankündigung: `raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md`.
- V4.1 Flash ist eine neue Modell-Version gegenüber dem V4-Flash im aktiven Stack ([[../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md]]).