knowledge-base/wiki/tools/ollama-cloud-deepseek-v4-flash-200tps-zdr.md

78 lines
5.6 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
title: Ollama Cloud — DeepSeek-V4-Flash 200+ tps & Zero Data Retention
type: tool
category: tools
tags: [ollama, deepseek-v4-flash, cloud, performance, privacy, zdr]
source: https://x.com/ollama/status/2085975426321858799
date: 2026-08-08
---
# Ollama Cloud: DeepSeek-V4-Flash mit 200+ tps und ZDR
## Kernaussage (offizieller Ollama-Post, 2026-08-08)
> Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR).
## Rollout-Kontext (vorheriger Post)
> DeepSeek-V4-Flash-0731 is now fully rolled out as the new default for deepseek-v4-flash on Ollama's cloud.
> - Fast: 120+ output tps on Ollama's cloud
> - Private: zero data retention hosting in US & Europe
## Fakten
- **DeepSeek-V4-Flash-0731** ist neuer Default für `deepseek-v4-flash` auf Ollama Cloud
- **200+ tps** Output-Speed (Folge-Post; Rollout-Post nennt 120+ tps)
- **Zero Data Retention (ZDR)** — Hosting in US & Europa
- Positionierung: Speed + Effizienz + Frontier-Level-Performance
## Relevanz für Hector-Setup
- `ollama/deepseek-v4-flash:0731:cloud` ist unser **Tier-0-Modell** (Heartbeat/Cron-Jobs, siehe AGENTS.md Barbell-Routing)
- 200+ tps Output bestätigt Eignung für schnelle, kostengünstige Status-/Script-Jobs
- ZDR (US/EU) relevant für Datenschutz-Betrachtung bei Cloud-Inferenz
- Ergänzt die lokale dSpark-Optimierung (siehe `unsloth-dspark.md`) — Cloud vs. lokal beide auf Flash-0731
## DeepSeek V4 Pro 0813 — Third-Party Review (2026-08-13)
Am 13.08.2026 veröffentlichte der KI-Coding-Kanal [[../people/aicodeking.md|AICodeKing]] ein **"Fully Tested"-Review** zu **DeepSeek V4 Pro 0813** (Build-Datum 13.08.2026, Titel: "Okay, this is ACTUALLY CRAZY!").
- **Source:** https://www.youtube.com/watch?v=W_e-YR4QRvg
- **Raw:** `raw/youtube/2026-08-13_deepseek-v4-pro-0813-aicodeking.md`
- **Einordnung:** Während dieser Page der **V4-Flash-0731**-Default (Tier-0) im Fokus steht, ist **DeepSeek V4 Pro** das größere Schwestermodell (`ollama/deepseek-v4-pro:cloud`, DRACO-Solo 60.3%, siehe [[../concepts/llm/llm-model-catalog.md]]). Dieses Review liefert unabhängige Third-Party-Benchmark-Sicht auf die V4-Pro-0813-Version.
- ⚠️ Kein Transkript verfügbar (yt-dlp scheitert ohne Cookies) — nur Titel-Metadaten verifiziert. Konkrete Benchmark-Details bei Gelegenheit nachziehen.
- Kontext: chinesische Open-Weight-Welle, siehe [[../concepts/chinese-ai-wave-july-2026.md]] und [[../concepts/llm/chinese-model-cost-routing.md]].
## DeepSeek API-Preiserhöhung (V4 Pro + V4 Flash) — ab 17.08.2026
Am 13.08.2026 meldete Pit Weber (OME Topic 13) eine **DeepSeek-API-Preiserhöhung** für **V4-Pro und V4-Flash**, wirksam ab **17. August 2026**.
- **Erhöhung um 50% bis 1.100%** über den aktuellen Preisen, je nach Modell, Token-Typ und Nutzungszeit.
- **Neue Peak/Off-Peak-Preise.** Peak-Zeiten täglich **9:0012:00 und 14:0018:00 Peking-Zeit**.
- **V4-Pro uncached Input Peak:** 3 → **9 Yuan/Mio Tokens**; **Output Peak:** 6 → **27 Yuan/Mio**.
- **V4-Pro Off-Peak:** 4,5 Yuan Input / 13,5 Yuan Output pro Mio Tokens.
- **V4-Flash:** 1,5 Yuan Input / 4,5 Yuan Output pro Mio Tokens.
- **Kontext:** DeepSeek hatte die V4-Pro-Standardpreise am **31. Mai 2026 um ~75% gesenkt**.
- **Begründung:** stark wachsendes Nutzungsvolumen.
- **Einordnung:** DeepSeek war der **Preisdrücker an der Frontier** (V4 Pro 0813: ~$0,43/$0,87 pro Mio, ~57× billiger als Fable 5). Diese Erhöhung ist eine **bemerkenswerte Wende** — der Preisdrücker erhöht selbst. Passt zur Marktdisziplinierungs-Diskussion (Grok 4.6 + DeepSeek V4 Pro drücken Frontier-Preise).
- ⚠️ Betrifft die **DeepSeek-API direkt**; Ollama-Cloud-Preise können abweichen. Für Cost-Routing relevant: Peak/Off-Peak-Zeiten (Peking-Zeit) als neuer Kostenfaktor.
- **Source:** https://www.perplexity.ai/discover/you/deepseek-to-raise-api-prices-u-bo6gHsYdTuy3WR3nKpo3_w (Perplexity-Link liefert per web_fetch 403/Cloudflare — Fakten aus verifizierter Websuche: investing.com, thenews.com.pk, kelo.com, deepseek.ai/pricing, roic.ai)
- **Raw:** `raw/other/2026-08-13_deepseek-api-preiserhoehung.md`
## Datenhaltung & Cache-Isolation (Update 2026-09-23)
Netbits (@NetLightning) hat die ZDR-Zusage im OME-Topic hinterfragt; eigene Prüfung der Primärquellen bestätigt die Zitate dieser Seite und ordnet sie ein:
- **Bestätigt:** „we require no logging, no training, and zero data retention policies in place" und „primarily in the United States … may route to Europe and Singapore" (Pricing-FAQ, abgerufen 23.09.2026). Partner sind **NVIDIA Cloud Providers (NCPs)**, gehostet werden **native Gewichte** der Modellgeber.
- **Offene Flanke:** Die ZDR-Zusage ist eine **Anforderung an Partner**, keine eigene Messung. Die Privacy Policy (§5) nennt „model inference providers" selbst als Dritte. Ob unser `deepseek-v4.1-flash` bei welchem NCP landet und wie dessen Retention real aussieht, ist **nicht unabhängig geprüft**.
- **Prompt-Cache:** Ollamas Cache-Isolation ist ungetestet. Der Stanford-Audit (ICML 2025) hat **DeepSeek** als per-user isoliert bestätigt — das betrifft die DeepSeek-API, nicht Ollamas Serving-Schicht.
Vollständige Analyse, Paper-Zitate und Verifikationsprotokoll: [[../concepts/policy/inference-provider-data-retention.md]].
## Quellen
- X-Post: https://x.com/ollama/status/2085975426321858799
- Datenhaltung/Cache-Verifikation: `raw/other/2026-09-23_netbits-ollama-cloud-privacy-prompt-caching.md`
- Raw: `raw/xpost/2026-08-08_ollama-deepseek-v4-flash-200tps-zdr.md`
- YouTube: `raw/youtube/2026-08-13_deepseek-v4-pro-0813-aicodeking.md`
- Raw: `raw/other/2026-08-13_deepseek-api-preiserhoehung.md`