knowledge-base/raw/other/2026-09-17_ollama-benchmarks-glm-53-flash-vs-deepseek-v41-flash.md

57 lines
2.6 KiB
Markdown
Raw Permalink Normal View History

---
type: other
source_url: "https://ollama.com/library/glm-5.3-flash + https://ollama.com/library/deepseek-v4.1-flash"
retrieved: 2026-09-17
title: "Ollama-Benchmark-Tabellen: GLM-5.3-Flash vs. DeepSeek V4.1 Flash"
tags: [benchmarks, ollama, glm-5.3-flash, deepseek-v4.1-flash, model-comparison]
---
# Ollama-Benchmark-Tabellen: GLM-5.3-Flash vs. DeepSeek V4.1 Flash
Abgerufen am 17.09.2026 von den offiziellen Ollama-Modellseiten. Beide Tabellen sind Hersteller-Selbstauskunft (Z.ai bzw. DeepSeek), nicht unabhängig nachgemessen.
## Befund zur Quellenlage
- `ollama.com/library/glm-5.3-flash` enthält eine **vollständige Benchmark-Tabelle** (Coding, Agentic, Vision). Die in einem Gruppenbeitrag aufgestellte Behauptung, für den Flash-Ableger lägen keine Zahlen vor, ist damit widerlegt.
- `ollama.com/library/deepseek-v4.1-flash` enthält ebenfalls eine vollständige Tabelle. Deren GLM-Spalte ist im HTML mit **„GLM-5.3"** überschrieben (nicht „GLM-5.3-Flash") — der dortige Vergleichswert bezieht sich also auf das große Geschwistermodell.
- Die GLM-5.3-Flash-Tabelle listet **keine** Terminal-Bench-3.0- oder -4.0-Zeile. Der kursierende TB-4.0-Wert 37,9 stammt aus der GLM-5.3-Spalte der DeepSeek-Seite.
## GLM-5.3-Flash (eigene Tabelle, Auszug)
| Benchmark | Wert |
|---|---|
| Terminal Bench 2.1 | 84,3 |
| DeepSWE v1.1 | 63,4 |
| NL2Repo | 56,3 |
| Toolathlon Verified | 78,4 |
| AutomationBench v1.0.6 | 48,8 |
| Agents' Last Exam | 26,3 |
| HLE w/ Tools | 55,3 |
| GDPval-AA v2 | 1773 |
| BabyVision | 53,4 |
| MVBench | 77,8 |
| MMVU | 80,5 |
Architektur laut Seite: 320B total / 18B aktiv, 45 Layer, 1M Kontext, MIT-Lizenz, nativ multimodal (Text/Bild/Video). Angaben zu Attention-Compute (3,0×) und KV-Cache (4,4×) gegenüber GLM-5.3.
## DeepSeek-V4.1-Flash (eigene Tabelle, Auszug)
| Benchmark | Wert |
|---|---|
| Terminal-Bench 2.1 | 90,6 |
| Terminal-Bench 3.0 | 30,0 |
| Terminal-Bench 4.0 | 31,2 |
| DeepSWE v1.1 | 74,2 |
| NL2Repo-Bench | 64,0 |
| CyberGym | 88,1 |
| HLE w/ tools | 63,9 |
| AutomationBench | 54,8 |
| Agent's Last Exam | 31,8 |
| BabyVision w/ tools | 89,6 |
Architektur laut Seite: 552B Backbone, Causal Encoder-Decoder (20+20 Layer), 8B aktiv im Prefill / 16B im Decode, KV-Cache 890 Byte/Token. Vergleichsspalten: Opus-5.0, GPT-5.6 Sol, K3, GLM-5.3, DS-V4-Pro, DS-V4-Flash.
## Unabhängige Gegenprobe
Die Seite `artificialanalysis.ai/models/comparisons/deepseek-v4-1-flash-vs-glm-5-3-flash` existiert und vergleicht die beiden Flash-Modelle (Intelligence Index v4.3, zehn Evaluations). Die Einzelwerte werden clientseitig gerendert und wurden bei diesem Abruf nicht ausgelesen.