| type |
source_url |
retrieved |
author |
is_thread |
post_date |
engagement |
tags |
| xpost |
https://x.com/PawelHuryn/status/2098002428054397185 |
2026-09-10 |
@PawelHuryn |
true |
2026-09-10 |
|
| deepseek |
| deepseek-v4.1-flash |
| benchmark |
| coding-benchmark |
| real-world-tasks |
| bug-fixing |
| price-performance |
| pareto-frontier |
| cost-routing |
|
Paweł Huryn: DeepSeek V4.1 Flash im Real-Task-Test (X-Post, 10.09.2026)
Quelle: X-Post von Paweł Huryn (@PawelHuryn, 10.09.2026): https://x.com/PawelHuryn/status/2098002428054397185 — Thread, geteilt in OME Topic 13. 6.098 Likes; Autor: AI PM, Newsletter „Product Compass" (productcompass.pm), 131.218 Follower.
Test-Setup
- Aufgabe: „2 repos, 105 hidden bugs, find and fix what you can."
- Metrik: Anzahl gefundener und gefixter versteckter Bugs (von 105).
- Verglichene Modelle / Effort-Level: Opus 5 (max), Grok 4.6 (max), DeepSeek V4.1 Flash (max), GPT-5.6 Luna (xhigh), Opus 5 (high).
Ergebnisse — Bugs gefixt (von 105)
| Modell |
Effort |
Bugs gefixt |
| Opus 5 |
max |
27 |
| Grok 4.6 |
max |
27 |
| DeepSeek V4.1 Flash |
max |
24 |
| GPT-5.6 Luna |
xhigh |
23 |
| Opus 5 |
high |
21 |
Ergebnisse — Kosten
| Modell |
Effort |
Kosten (USD) |
| Opus 5 |
max |
$51.33 |
| Opus 5 |
high |
$38.77 |
| Grok 4.6 |
max |
$16.96 |
| GPT-5.6 Luna |
xhigh |
$2.50 |
| DeepSeek V4.1 Flash |
max |
$1.80 |
Claim des Autors (wörtlich, auszugsweise)
„A really strong model for everyday tasks. And look at the cost: […] It's at the Pareto frontier."
„Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵"
Anmerkungen zur Quellenlage
- Einzel-Review eines unabhängigen Testers (AI-PM-Newsletter). Keine Peer-Review, keine offengelegte Test-Suite, keine Reproduktionsdaten im Post.
- Bewertungsmetrik ist die reine Anzahl gefixter Bugs; Schwere/Relevanz der Bugs und Fehlerquote (falsche Fixes) sind nicht ausgewiesen.
- Thread kündigt weitere Ergebnisse an (weitere Effort-Level, zusätzliche Tests zu DeepSeek V4 Pro).