knowledge-base/raw/xpost/2026-09-10_pawel-huryn-deepseek-v41-flash-real-task-benchmark.md

2.1 KiB

type source_url retrieved author is_thread post_date engagement tags
xpost https://x.com/PawelHuryn/status/2098002428054397185 2026-09-10 @PawelHuryn true 2026-09-10
likes
6098
deepseek
deepseek-v4.1-flash
benchmark
coding-benchmark
real-world-tasks
bug-fixing
price-performance
pareto-frontier
cost-routing

Paweł Huryn: DeepSeek V4.1 Flash im Real-Task-Test (X-Post, 10.09.2026)

Quelle: X-Post von Paweł Huryn (@PawelHuryn, 10.09.2026): https://x.com/PawelHuryn/status/2098002428054397185 — Thread, geteilt in OME Topic 13. 6.098 Likes; Autor: AI PM, Newsletter „Product Compass" (productcompass.pm), 131.218 Follower.

Test-Setup

  • Aufgabe: „2 repos, 105 hidden bugs, find and fix what you can."
  • Metrik: Anzahl gefundener und gefixter versteckter Bugs (von 105).
  • Verglichene Modelle / Effort-Level: Opus 5 (max), Grok 4.6 (max), DeepSeek V4.1 Flash (max), GPT-5.6 Luna (xhigh), Opus 5 (high).

Ergebnisse — Bugs gefixt (von 105)

Modell Effort Bugs gefixt
Opus 5 max 27
Grok 4.6 max 27
DeepSeek V4.1 Flash max 24
GPT-5.6 Luna xhigh 23
Opus 5 high 21

Ergebnisse — Kosten

Modell Effort Kosten (USD)
Opus 5 max $51.33
Opus 5 high $38.77
Grok 4.6 max $16.96
GPT-5.6 Luna xhigh $2.50
DeepSeek V4.1 Flash max $1.80

Claim des Autors (wörtlich, auszugsweise)

„A really strong model for everyday tasks. And look at the cost: […] It's at the Pareto frontier."

„Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵"

Anmerkungen zur Quellenlage

  • Einzel-Review eines unabhängigen Testers (AI-PM-Newsletter). Keine Peer-Review, keine offengelegte Test-Suite, keine Reproduktionsdaten im Post.
  • Bewertungsmetrik ist die reine Anzahl gefixter Bugs; Schwere/Relevanz der Bugs und Fehlerquote (falsche Fixes) sind nicht ausgewiesen.
  • Thread kündigt weitere Ergebnisse an (weitere Effort-Level, zusätzliche Tests zu DeepSeek V4 Pro).