90 lines
4.7 KiB
Markdown
90 lines
4.7 KiB
Markdown
---
|
||
type: other
|
||
source_url: https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/
|
||
retrieved: 2026-09-22
|
||
title: "N8 Programs: Jev vs. GPT-5.6 Terra auf System-1-Benchmarks"
|
||
author: "N8 Programs (@N8Programs); Reddit-Repost von /u/stealthispost"
|
||
tags: [reddit, xpost, typesafe-ai, jev, gpt-5-6-terra, benchmark, system-one, calibration, ece, cost]
|
||
---
|
||
|
||
# N8 Programs — Jev vs. GPT-5.6 Terra (Reddit-Repost, 17.09.2026)
|
||
|
||
## Quellenkette
|
||
|
||
- **Geteilter Reddit-Post:** https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/
|
||
- **Reddit-Autor:** `/u/stealthispost`
|
||
- **Reddit-Zeitpunkt laut RSS:** 17.09.2026, 04:44 UTC
|
||
- **Reddit-Zugriff:** normale Seite/JSON blockiert; `.rss` liefert HTTP 200 und den vollständigen Posttext.
|
||
- **Upstream-Thread:** N8 Programs ([@N8Programs](https://x.com/N8Programs)), beginnend https://x.com/N8Programs/status/2100088523403432357, 16.09.2026 05:04 UTC.
|
||
- **Upstream-Autorprofil:** „Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.“; GitHub https://github.com/N8python.
|
||
- **Abruf 22.09.2026:** Ausgangspost 179.846 Views, 907 Likes, 70 Reposts, 29 Quotes; keine Community Note (FxTwitter).
|
||
|
||
## Wortlaut des Ausgangsposts
|
||
|
||
> Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.
|
||
|
||
Der Reddit-Repost zitiert außerdem den Folgetext:
|
||
|
||
> Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. […] Jev is ~18x cheaper than Terra! […] expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst.
|
||
|
||
## Methode laut Thread und Bildern
|
||
|
||
- Multiple-Choice-Benchmarks auf gemeinsamer 0–100-%-Skala.
|
||
- Terra: `reasoning=none`, „letter-only output“; beim Schach ausdrückliche `a/b/c/d`-Instruktion.
|
||
- Benchmarks: MMLU, GPQA Diamond, ARC-Easy, ARC-Challenge, WinoGrande, HellaSwag, GSM8K mit 4 bzw. 10 Auswahlmöglichkeiten, Schach mit 4 legalen Zügen.
|
||
- Die Radar-Grafik trägt den Untertitel „Official benchmark comparison“ und den Hinweis: „Each spoke is one benchmark condition. Polygon area is not an aggregate score.“
|
||
- Kein öffentliches Repository, keine Prompts, Beispieldateien oder Run-Logs im Thread verlinkt.
|
||
|
||
## Ergebnistabelle (aus Thread-Bild abgelesen)
|
||
|
||
| Benchmark | Jev | Terra |
|
||
|---|---:|---:|
|
||
| MMLU | 90,91 % | 87,89 % |
|
||
| GPQA Diamond | 68,18 % | 56,06 % |
|
||
| ARC-Easy | 99,33 % | 98,82 % |
|
||
| ARC-Challenge | 97,61 % | 96,76 % |
|
||
| WinoGrande | 91,63 % | 78,69 % |
|
||
| HellaSwag | 94,86 % | **95,32 %** |
|
||
| GSM8K · 4 choices | 79,68 % | **87,72 %** |
|
||
| GSM8K · 10 choices | 54,59 % | **69,83 %** |
|
||
| Chess · 4 legal moves | 49,45 % | **50,25 %** |
|
||
|
||
Hinweis: Der Text sagt, Jev übertreffe Terra bei „WinoGrande **or HellaSwag**“. Die veröffentlichte Tabelle zeigt bei HellaSwag **Terra 95,32 % vor Jev 94,86 %**.
|
||
|
||
## Kosten (aus Thread-Bild abgelesen)
|
||
|
||
Titel: „Evaluation cost: Jev vs. Terra — Canonical benchmark runs · USD“.
|
||
|
||
| Benchmark | Jev (estimated) | Terra (reported) |
|
||
|---|---:|---:|
|
||
| MMLU | $0,2749 | $4,6387 |
|
||
| GPQA Diamond | $0,0048 | $0,1106 |
|
||
| ARC-Easy | $0,0411 | $0,5427 |
|
||
| ARC-Challenge | $0,0207 | $0,2897 |
|
||
| WinoGrande | $0,0197 | $0,2652 |
|
||
| HellaSwag | $0,2315 | $4,9176 |
|
||
| GSM8K · 4 choices | $0,0247 | $0,3815 |
|
||
| GSM8K · 10 choices | $0,0321 | $0,4663 |
|
||
| Chess · 4 legal moves | $0,0504 | $1,3741 |
|
||
| **Gesamt** | **$0,6999** | **$12,9865** |
|
||
|
||
Fußnoten im Bild:
|
||
|
||
- Jev: „recorded input tokens × $0.042/million; output free.“
|
||
- Terra: „summed API-reported costs. Auxiliary runs excluded.“
|
||
|
||
## Kalibrierung (aus Thread-Bild abgelesen)
|
||
|
||
Reliability-Diagramm „Jev reliability diagram“:
|
||
|
||
- **N = 33.735**
|
||
- **ECE = 1,74 Prozentpunkte**
|
||
- 10 gleich breite Bins
|
||
- Wilson-95-%-Konfidenzintervalle
|
||
- Textangabe des Autors: beste Einzel-Bench **0,26 pp**, schlechteste **6,96 pp**.
|
||
|
||
## Weitere Thread-Aussagen
|
||
|
||
- N8 Programs beschreibt Jev als günstigen Kandidaten für ungefähr „one Terra forward pass over a context“.
|
||
- Frontier-Modelle wie Astra hätten vermutlich bessere No-CoT-Fähigkeiten, seien für Massenklassifikation aber zu teuer.
|
||
- Schlussurteil: Jev sei „exactly what it says on the tin“, unter der Annahme, dass öffentliche Benchmarks in Alltagsaufgaben übertragen und nicht gezielt trainiert wurden.
|