--- title: Ollama Cloud — DeepSeek-V4-Flash 200+ tps & Zero Data Retention type: tool category: tools tags: [ollama, deepseek-v4-flash, cloud, performance, privacy, zdr] source: https://x.com/ollama/status/2085975426321858799 date: 2026-08-08 --- # Ollama Cloud: DeepSeek-V4-Flash mit 200+ tps und ZDR ## Kernaussage (offizieller Ollama-Post, 2026-08-08) > Currently serving 200 tps+ output speed for DeepSeek-V4-Flash on Ollama's cloud with zero data retention (ZDR). ## Rollout-Kontext (vorheriger Post) > DeepSeek-V4-Flash-0731 is now fully rolled out as the new default for deepseek-v4-flash on Ollama's cloud. > - Fast: 120+ output tps on Ollama's cloud > - Private: zero data retention hosting in US & Europe ## Fakten - **DeepSeek-V4-Flash-0731** ist neuer Default für `deepseek-v4-flash` auf Ollama Cloud - **200+ tps** Output-Speed (Folge-Post; Rollout-Post nennt 120+ tps) - **Zero Data Retention (ZDR)** — Hosting in US & Europa - Positionierung: Speed + Effizienz + Frontier-Level-Performance ## Relevanz für Hector-Setup - `ollama/deepseek-v4-flash:0731:cloud` ist unser **Tier-0-Modell** (Heartbeat/Cron-Jobs, siehe AGENTS.md Barbell-Routing) - 200+ tps Output bestätigt Eignung für schnelle, kostengünstige Status-/Script-Jobs - ZDR (US/EU) relevant für Datenschutz-Betrachtung bei Cloud-Inferenz - Ergänzt die lokale dSpark-Optimierung (siehe `unsloth-dspark.md`) — Cloud vs. lokal beide auf Flash-0731 ## DeepSeek V4 Pro 0813 — Third-Party Review (2026-08-13) Am 13.08.2026 veröffentlichte der KI-Coding-Kanal [[../people/aicodeking.md|AICodeKing]] ein **"Fully Tested"-Review** zu **DeepSeek V4 Pro 0813** (Build-Datum 13.08.2026, Titel: "Okay, this is ACTUALLY CRAZY!"). - **Source:** https://www.youtube.com/watch?v=W_e-YR4QRvg - **Raw:** `raw/youtube/2026-08-13_deepseek-v4-pro-0813-aicodeking.md` - **Einordnung:** Während dieser Page der **V4-Flash-0731**-Default (Tier-0) im Fokus steht, ist **DeepSeek V4 Pro** das größere Schwestermodell (`ollama/deepseek-v4-pro:cloud`, DRACO-Solo 60.3%, siehe [[../concepts/llm/llm-model-catalog.md]]). Dieses Review liefert unabhängige Third-Party-Benchmark-Sicht auf die V4-Pro-0813-Version. - ⚠️ Kein Transkript verfügbar (yt-dlp scheitert ohne Cookies) — nur Titel-Metadaten verifiziert. Konkrete Benchmark-Details bei Gelegenheit nachziehen. - Kontext: chinesische Open-Weight-Welle, siehe [[../concepts/chinese-ai-wave-july-2026.md]] und [[../concepts/llm/chinese-model-cost-routing.md]]. ## DeepSeek API-Preiserhöhung (V4 Pro + V4 Flash) — ab 17.08.2026 Am 13.08.2026 meldete Pit Weber (OME Topic 13) eine **DeepSeek-API-Preiserhöhung** für **V4-Pro und V4-Flash**, wirksam ab **17. August 2026**. - **Erhöhung um 50% bis 1.100%** über den aktuellen Preisen, je nach Modell, Token-Typ und Nutzungszeit. - **Neue Peak/Off-Peak-Preise.** Peak-Zeiten täglich **9:00–12:00 und 14:00–18:00 Peking-Zeit**. - **V4-Pro uncached Input Peak:** 3 → **9 Yuan/Mio Tokens**; **Output Peak:** 6 → **27 Yuan/Mio**. - **V4-Pro Off-Peak:** 4,5 Yuan Input / 13,5 Yuan Output pro Mio Tokens. - **V4-Flash:** 1,5 Yuan Input / 4,5 Yuan Output pro Mio Tokens. - **Kontext:** DeepSeek hatte die V4-Pro-Standardpreise am **31. Mai 2026 um ~75% gesenkt**. - **Begründung:** stark wachsendes Nutzungsvolumen. - **Einordnung:** DeepSeek war der **Preisdrücker an der Frontier** (V4 Pro 0813: ~$0,43/$0,87 pro Mio, ~57× billiger als Fable 5). Diese Erhöhung ist eine **bemerkenswerte Wende** — der Preisdrücker erhöht selbst. Passt zur Marktdisziplinierungs-Diskussion (Grok 4.6 + DeepSeek V4 Pro drücken Frontier-Preise). - ⚠️ Betrifft die **DeepSeek-API direkt**; Ollama-Cloud-Preise können abweichen. Für Cost-Routing relevant: Peak/Off-Peak-Zeiten (Peking-Zeit) als neuer Kostenfaktor. - **Source:** https://www.perplexity.ai/discover/you/deepseek-to-raise-api-prices-u-bo6gHsYdTuy3WR3nKpo3_w (Perplexity-Link liefert per web_fetch 403/Cloudflare — Fakten aus verifizierter Websuche: investing.com, thenews.com.pk, kelo.com, deepseek.ai/pricing, roic.ai) - **Raw:** `raw/other/2026-08-13_deepseek-api-preiserhoehung.md` ## Datenhaltung & Cache-Isolation (Update 2026-09-23) Netbits (@NetLightning) hat die ZDR-Zusage im OME-Topic hinterfragt; eigene Prüfung der Primärquellen bestätigt die Zitate dieser Seite und ordnet sie ein: - **Bestätigt:** „we require no logging, no training, and zero data retention policies in place" und „primarily in the United States … may route to Europe and Singapore" (Pricing-FAQ, abgerufen 23.09.2026). Partner sind **NVIDIA Cloud Providers (NCPs)**, gehostet werden **native Gewichte** der Modellgeber. - **Offene Flanke:** Die ZDR-Zusage ist eine **Anforderung an Partner**, keine eigene Messung. Die Privacy Policy (§5) nennt „model inference providers" selbst als Dritte. Ob unser `deepseek-v4.1-flash` bei welchem NCP landet und wie dessen Retention real aussieht, ist **nicht unabhängig geprüft**. - **Prompt-Cache:** Ollamas Cache-Isolation ist ungetestet. Der Stanford-Audit (ICML 2025) hat **DeepSeek** als per-user isoliert bestätigt — das betrifft die DeepSeek-API, nicht Ollamas Serving-Schicht. Vollständige Analyse, Paper-Zitate und Verifikationsprotokoll: [[../concepts/policy/inference-provider-data-retention.md]]. ## Quellen - X-Post: https://x.com/ollama/status/2085975426321858799 - Datenhaltung/Cache-Verifikation: `raw/other/2026-09-23_netbits-ollama-cloud-privacy-prompt-caching.md` - Raw: `raw/xpost/2026-08-08_ollama-deepseek-v4-flash-200tps-zdr.md` - YouTube: `raw/youtube/2026-08-13_deepseek-v4-pro-0813-aicodeking.md` - Raw: `raw/other/2026-08-13_deepseek-api-preiserhoehung.md`