knowledge-base/raw/other/2026-09-17_ollama-benchmarks-glm-53-flash-vs-deepseek-v41-flash.md

56 lines
2.6 KiB
Markdown
Raw Permalink Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: other
source_url: "https://ollama.com/library/glm-5.3-flash + https://ollama.com/library/deepseek-v4.1-flash"
retrieved: 2026-09-17
title: "Ollama-Benchmark-Tabellen: GLM-5.3-Flash vs. DeepSeek V4.1 Flash"
tags: [benchmarks, ollama, glm-5.3-flash, deepseek-v4.1-flash, model-comparison]
---
# Ollama-Benchmark-Tabellen: GLM-5.3-Flash vs. DeepSeek V4.1 Flash
Abgerufen am 17.09.2026 von den offiziellen Ollama-Modellseiten. Beide Tabellen sind Hersteller-Selbstauskunft (Z.ai bzw. DeepSeek), nicht unabhängig nachgemessen.
## Befund zur Quellenlage
- `ollama.com/library/glm-5.3-flash` enthält eine **vollständige Benchmark-Tabelle** (Coding, Agentic, Vision). Die in einem Gruppenbeitrag aufgestellte Behauptung, für den Flash-Ableger lägen keine Zahlen vor, ist damit widerlegt.
- `ollama.com/library/deepseek-v4.1-flash` enthält ebenfalls eine vollständige Tabelle. Deren GLM-Spalte ist im HTML mit **„GLM-5.3"** überschrieben (nicht „GLM-5.3-Flash") — der dortige Vergleichswert bezieht sich also auf das große Geschwistermodell.
- Die GLM-5.3-Flash-Tabelle listet **keine** Terminal-Bench-3.0- oder -4.0-Zeile. Der kursierende TB-4.0-Wert 37,9 stammt aus der GLM-5.3-Spalte der DeepSeek-Seite.
## GLM-5.3-Flash (eigene Tabelle, Auszug)
| Benchmark | Wert |
|---|---|
| Terminal Bench 2.1 | 84,3 |
| DeepSWE v1.1 | 63,4 |
| NL2Repo | 56,3 |
| Toolathlon Verified | 78,4 |
| AutomationBench v1.0.6 | 48,8 |
| Agents' Last Exam | 26,3 |
| HLE w/ Tools | 55,3 |
| GDPval-AA v2 | 1773 |
| BabyVision | 53,4 |
| MVBench | 77,8 |
| MMVU | 80,5 |
Architektur laut Seite: 320B total / 18B aktiv, 45 Layer, 1M Kontext, MIT-Lizenz, nativ multimodal (Text/Bild/Video). Angaben zu Attention-Compute (3,0×) und KV-Cache (4,4×) gegenüber GLM-5.3.
## DeepSeek-V4.1-Flash (eigene Tabelle, Auszug)
| Benchmark | Wert |
|---|---|
| Terminal-Bench 2.1 | 90,6 |
| Terminal-Bench 3.0 | 30,0 |
| Terminal-Bench 4.0 | 31,2 |
| DeepSWE v1.1 | 74,2 |
| NL2Repo-Bench | 64,0 |
| CyberGym | 88,1 |
| HLE w/ tools | 63,9 |
| AutomationBench | 54,8 |
| Agent's Last Exam | 31,8 |
| BabyVision w/ tools | 89,6 |
Architektur laut Seite: 552B Backbone, Causal Encoder-Decoder (20+20 Layer), 8B aktiv im Prefill / 16B im Decode, KV-Cache 890 Byte/Token. Vergleichsspalten: Opus-5.0, GPT-5.6 Sol, K3, GLM-5.3, DS-V4-Pro, DS-V4-Flash.
## Unabhängige Gegenprobe
Die Seite `artificialanalysis.ai/models/comparisons/deepseek-v4-1-flash-vs-glm-5-3-flash` existiert und vergleicht die beiden Flash-Modelle (Intelligence Index v4.3, zehn Evaluations). Die Einzelwerte werden clientseitig gerendert und wurden bei diesem Abruf nicht ausgelesen.