- raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md (neu) - raw/other/2026-09-10_deepseek-v41-flash-official-announcement.md (neu) - wiki/concepts/llm/deepseek-v4.1-flash.md (neu) - wiki/concepts/llm/deepseek-v4-pro-ga-harness-open-source.md (V4.1-Flash-Nachtrag + Einordnung) - wiki/people/aicodeking.md (Review durch offizielles Release bestätigt) - wiki/index.md (199. Update), wiki/log.md
3.2 KiB
3.2 KiB
| type | source_url | retrieved | title | author | tags | |||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| other | https://api-docs.deepseek.com/news/news260910 | 2026-09-10 | DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (offizielle Ankündigung) | DeepSeek |
|
DeepSeek-V4.1-Flash — Offizielle Ankündigung (API-Docs, 10.09.2026)
Quelle: Offizielle DeepSeek-API-Docs-Ankündigung (10.09.2026): https://api-docs.deepseek.com/news/news260910 — Primärquelle zum V4.1-Flash-Release, verlinkt aus dem X-Post von @deepseek_ai (https://x.com/deepseek_ai/status/2097930608790167907).
Kernaussagen (wörtlich aus der Ankündigung)
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
Asymmetrische Architektur
- 552B-Parameter-MoE (Mixture-of-Experts).
- Neue Causal-Encoder–Decoder-Architektur (CED): nur 8B aktive Parameter für Input, 16B für Output (asymmetrisch, Input/Output nicht symmetrisch).
- Neue Pre-Training-Methoden + größer skaliertes RL-Post-Training liefern Benchmark-Ergebnisse vor Flaggschiff-Modellen, inkl. DeepSeek-V4-Pro.
Kleinere KV-Cache
- Gegenüber der Vorgängergeneration braucht V4.1-Flash nur 1/4 des HBM und 1/8 des SSD-Speichers für den KV-Cache.
- Cache-Hit-Gebühren machen oft einen großen Anteil der Agent-Kosten aus — Cache-Kompression senkt diese Kosten deutlich.
API-Verfügbarkeit
- Live auf der DeepSeek-API mit nativem Multimodal-Support — Modellname:
deepseek-flash. - V4-Flash & V4-Flash-Vision-Exp sind retired. Für Kompatibilität routen
deepseek-v4-flashunddeepseek-v4-flash-vision-exptemporär auf V4.1-Flash. - Tests mehrerer Parteien sehen V4.1-Flash vor V4-Pro bei Performance, Kosten, Speed und Gesamtlaufzeit. V4-Pro wird ausphasen.
- Ab 04:00 UTC, 14.09.2026: alle
deepseek-v4-pro-Requests routen auf V4.1-Flash zu V4.1-Flash-Raten — bis V4.1-Pro launcht. - Offizielle Partner: WorkBuddy (inkl. CodeBuddy) und OpenCode unterstützen V4.1-Flash vollständig.
Pricing
- Effizientere Architektur → niedrigere API-Preise.
- Peak/Off-Peak-Pricing: Off-Peak = 50 % der Peak-Raten.
- Neue Preise ab 04:00 UTC, 10.09.2026.
Open Source & Deployment
- Zusammenarbeit mit der Open-Source-Community für V4.1-Flash-Inference-Support; weitere Deployment-Optionen geplant.
- Großskalige Deployments (2.000 GPUs + Storage-Cluster) ansprechbar.
- Modell (Hugging Face): https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- Paper (Tech Report): https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
Einordnung
- Primärquelle (offiziell) zum V4.1-Flash-Release; X-Post von @deepseek_ai als Ankündigung:
raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md. - V4.1 Flash ist eine neue Modell-Version gegenüber dem V4-Flash im aktiven Stack (../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md).