53 lines
2.1 KiB
Markdown
53 lines
2.1 KiB
Markdown
|
|
---
|
||
|
|
type: xpost
|
||
|
|
source_url: https://x.com/PawelHuryn/status/2098002428054397185
|
||
|
|
retrieved: 2026-09-10
|
||
|
|
author: "@PawelHuryn"
|
||
|
|
is_thread: true
|
||
|
|
post_date: 2026-09-10
|
||
|
|
engagement: {likes: 6098}
|
||
|
|
tags: [deepseek, deepseek-v4.1-flash, benchmark, coding-benchmark, real-world-tasks, bug-fixing, price-performance, pareto-frontier, cost-routing]
|
||
|
|
---
|
||
|
|
|
||
|
|
# Paweł Huryn: DeepSeek V4.1 Flash im Real-Task-Test (X-Post, 10.09.2026)
|
||
|
|
|
||
|
|
> **Quelle:** X-Post von Paweł Huryn (@PawelHuryn, 10.09.2026): https://x.com/PawelHuryn/status/2098002428054397185 — Thread, geteilt in OME Topic 13. 6.098 Likes; Autor: AI PM, Newsletter „Product Compass" (productcompass.pm), 131.218 Follower.
|
||
|
|
|
||
|
|
## Test-Setup
|
||
|
|
|
||
|
|
- **Aufgabe:** „2 repos, 105 hidden bugs, find and fix what you can."
|
||
|
|
- **Metrik:** Anzahl gefundener und gefixter versteckter Bugs (von 105).
|
||
|
|
- **Verglichene Modelle / Effort-Level:** Opus 5 (max), Grok 4.6 (max), DeepSeek V4.1 Flash (max), GPT-5.6 Luna (xhigh), Opus 5 (high).
|
||
|
|
|
||
|
|
## Ergebnisse — Bugs gefixt (von 105)
|
||
|
|
|
||
|
|
| Modell | Effort | Bugs gefixt |
|
||
|
|
|---|---|---|
|
||
|
|
| Opus 5 | max | 27 |
|
||
|
|
| Grok 4.6 | max | 27 |
|
||
|
|
| **DeepSeek V4.1 Flash** | **max** | **24** |
|
||
|
|
| GPT-5.6 Luna | xhigh | 23 |
|
||
|
|
| Opus 5 | high | 21 |
|
||
|
|
|
||
|
|
## Ergebnisse — Kosten
|
||
|
|
|
||
|
|
| Modell | Effort | Kosten (USD) |
|
||
|
|
|---|---|---|
|
||
|
|
| Opus 5 | max | $51.33 |
|
||
|
|
| Opus 5 | high | $38.77 |
|
||
|
|
| Grok 4.6 | max | $16.96 |
|
||
|
|
| GPT-5.6 Luna | xhigh | $2.50 |
|
||
|
|
| **DeepSeek V4.1 Flash** | **max** | **$1.80** |
|
||
|
|
|
||
|
|
## Claim des Autors (wörtlich, auszugsweise)
|
||
|
|
|
||
|
|
> „A really strong model for everyday tasks. And look at the cost: […] It's at the Pareto frontier."
|
||
|
|
|
||
|
|
> „Other effort levels (high) and extra tests for DeepSeek V4 Pro dropping every hour in this thread 🧵"
|
||
|
|
|
||
|
|
## Anmerkungen zur Quellenlage
|
||
|
|
|
||
|
|
- Einzel-Review eines unabhängigen Testers (AI-PM-Newsletter). Keine Peer-Review, keine offengelegte Test-Suite, keine Reproduktionsdaten im Post.
|
||
|
|
- Bewertungsmetrik ist die reine Anzahl gefixter Bugs; Schwere/Relevanz der Bugs und Fehlerquote (falsche Fixes) sind nicht ausgewiesen.
|
||
|
|
- Thread kündigt weitere Ergebnisse an (weitere Effort-Level, zusätzliche Tests zu DeepSeek V4 Pro).
|