knowledge-base/raw/other/2026-09-10_deepseek-v41-flash-official-announcement.md
Hector bfacf8050e ingest(raw): DeepSeek-V4.1-Flash offizielles Release (X-Post + API-Docs)
- raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md (neu)
- raw/other/2026-09-10_deepseek-v41-flash-official-announcement.md (neu)
- wiki/concepts/llm/deepseek-v4.1-flash.md (neu)
- wiki/concepts/llm/deepseek-v4-pro-ga-harness-open-source.md (V4.1-Flash-Nachtrag + Einordnung)
- wiki/people/aicodeking.md (Review durch offizielles Release bestätigt)
- wiki/index.md (199. Update), wiki/log.md
2026-09-10 13:35:03 +02:00

57 lines
3.2 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: other
source_url: https://api-docs.deepseek.com/news/news260910
retrieved: 2026-09-10
title: "DeepSeek-V4.1-Flash: Smarter, Faster, More Efficient (offizielle Ankündigung)"
author: "DeepSeek"
tags: [deepseek, deepseek-v4.1-flash, release, architecture, moe, causal-encoder-decoder, kv-cache, multimodal, api, pricing, open-weights]
---
# DeepSeek-V4.1-Flash — Offizielle Ankündigung (API-Docs, 10.09.2026)
> **Quelle:** Offizielle DeepSeek-API-Docs-Ankündigung (10.09.2026): https://api-docs.deepseek.com/news/news260910 — Primärquelle zum V4.1-Flash-Release, verlinkt aus dem X-Post von @deepseek_ai (https://x.com/deepseek_ai/status/2097930608790167907).
## Kernaussagen (wörtlich aus der Ankündigung)
> 🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
>
> 🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
>
> 🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
## Asymmetrische Architektur
- **552B-Parameter-MoE** (Mixture-of-Experts).
- **Neue Causal-EncoderDecoder-Architektur (CED):** nur **8B aktive Parameter für Input**, **16B für Output** (asymmetrisch, Input/Output nicht symmetrisch).
- Neue Pre-Training-Methoden + größer skaliertes RL-Post-Training liefern Benchmark-Ergebnisse vor Flaggschiff-Modellen, inkl. **DeepSeek-V4-Pro**.
## Kleinere KV-Cache
- Gegenüber der Vorgängergeneration braucht V4.1-Flash nur **1/4 des HBM** und **1/8 des SSD-Speichers** für den KV-Cache.
- Cache-Hit-Gebühren machen oft einen großen Anteil der Agent-Kosten aus — Cache-Kompression senkt diese Kosten deutlich.
## API-Verfügbarkeit
- **Live auf der DeepSeek-API mit nativem Multimodal-Support** — Modellname: `deepseek-flash`.
- **V4-Flash & V4-Flash-Vision-Exp sind retired.** Für Kompatibilität routen `deepseek-v4-flash` und `deepseek-v4-flash-vision-exp` temporär auf V4.1-Flash.
- Tests mehrerer Parteien sehen V4.1-Flash **vor V4-Pro** bei Performance, Kosten, Speed und Gesamtlaufzeit. V4-Pro wird ausphasen.
- **Ab 04:00 UTC, 14.09.2026:** alle `deepseek-v4-pro`-Requests routen auf V4.1-Flash zu V4.1-Flash-Raten — bis V4.1-Pro launcht.
- **Offizielle Partner:** WorkBuddy (inkl. CodeBuddy) und OpenCode unterstützen V4.1-Flash vollständig.
## Pricing
- Effizientere Architektur → niedrigere API-Preise.
- Peak/Off-Peak-Pricing: **Off-Peak = 50 % der Peak-Raten**.
- **Neue Preise ab 04:00 UTC, 10.09.2026.**
## Open Source & Deployment
- Zusammenarbeit mit der Open-Source-Community für V4.1-Flash-Inference-Support; weitere Deployment-Optionen geplant.
- Großskalige Deployments (2.000 GPUs + Storage-Cluster) ansprechbar.
- **Modell (Hugging Face):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash
- **Paper (Tech Report):** https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/DeepSeek_V41_Tech_Report.pdf
## Einordnung
- Primärquelle (offiziell) zum V4.1-Flash-Release; X-Post von @deepseek_ai als Ankündigung: `raw/xpost/2026-09-10_deepseek-v41-flash-official-release.md`.
- V4.1 Flash ist eine neue Modell-Version gegenüber dem V4-Flash im aktiven Stack ([[../../tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md]]).