knowledge-base/raw/xpost/2026-09-10_pawel-huryn-deepseek-v41-flash-real-task-benchmark.md

53 lines
2.1 KiB
Markdown
Raw Permalink Normal View History

---
type: xpost
source_url: https://x.com/PawelHuryn/status/2098002428054397185
retrieved: 2026-09-10
author: "@PawelHuryn"
is_thread: true
post_date: 2026-09-10
engagement: {likes: 6098}
tags: [deepseek, deepseek-v4.1-flash, benchmark, coding-benchmark, real-world-tasks, bug-fixing, price-performance, pareto-frontier, cost-routing]
---
# Paweł Huryn: DeepSeek V4.1 Flash im Real-Task-Test (X-Post, 10.09.2026)
> **Quelle:** X-Post von Paweł Huryn (@PawelHuryn, 10.09.2026): https://x.com/PawelHuryn/status/2098002428054397185 — Thread, geteilt in OME Topic 13. 6.098 Likes; Autor: AI PM, Newsletter „Product Compass" (productcompass.pm), 131.218 Follower.
## Test-Setup
- **Aufgabe:** „2 repos, 105 hidden bugs, find and fix what you can."
- **Metrik:** Anzahl gefundener und gefixter versteckter Bugs (von 105).
- **Verglichene Modelle / Effort-Level:** Opus 5 (max), Grok 4.6 (max), DeepSeek V4.1 Flash (max), GPT-5.6 Luna (xhigh), Opus 5 (high).
## Ergebnisse — Bugs gefixt (von 105)
| Modell | Effort | Bugs gefixt |
|---|---|---|
| Opus 5 | max | 27 |
| Grok 4.6 | max | 27 |
| **DeepSeek V4.1 Flash** | **max** | **24** |
| GPT-5.6 Luna | xhigh | 23 |
| Opus 5 | high | 21 |
## Ergebnisse — Kosten
| Modell | Effort | Kosten (USD) |
|---|---|---|
| Opus 5 | max | $51.33 |
| Opus 5 | high | $38.77 |
| Grok 4.6 | max | $16.96 |
| GPT-5.6 Luna | xhigh | $2.50 |
| **DeepSeek V4.1 Flash** | **max** | **$1.80** |
## Claim des Autors (wörtlich, auszugsweise)
> „A really strong model for everyday tasks. And look at the cost: […] It's at the Pareto frontier."
> „Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵"
## Anmerkungen zur Quellenlage
- Einzel-Review eines unabhängigen Testers (AI-PM-Newsletter). Keine Peer-Review, keine offengelegte Test-Suite, keine Reproduktionsdaten im Post.
- Bewertungsmetrik ist die reine Anzahl gefixter Bugs; Schwere/Relevanz der Bugs und Fehlerquote (falsche Fixes) sind nicht ausgewiesen.
- Thread kündigt weitere Ergebnisse an (weitere Effort-Level, zusätzliche Tests zu DeepSeek V4 Pro).