--- type: other source_url: https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/ retrieved: 2026-09-22 title: "N8 Programs: Jev vs. GPT-5.6 Terra auf System-1-Benchmarks" author: "N8 Programs (@N8Programs); Reddit-Repost von /u/stealthispost" tags: [reddit, xpost, typesafe-ai, jev, gpt-5-6-terra, benchmark, system-one, calibration, ece, cost] --- # N8 Programs — Jev vs. GPT-5.6 Terra (Reddit-Repost, 17.09.2026) ## Quellenkette - **Geteilter Reddit-Post:** https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/ - **Reddit-Autor:** `/u/stealthispost` - **Reddit-Zeitpunkt laut RSS:** 17.09.2026, 04:44 UTC - **Reddit-Zugriff:** normale Seite/JSON blockiert; `.rss` liefert HTTP 200 und den vollständigen Posttext. - **Upstream-Thread:** N8 Programs ([@N8Programs](https://x.com/N8Programs)), beginnend https://x.com/N8Programs/status/2100088523403432357, 16.09.2026 05:04 UTC. - **Upstream-Autorprofil:** „Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.“; GitHub https://github.com/N8python. - **Abruf 22.09.2026:** Ausgangspost 179.846 Views, 907 Likes, 70 Reposts, 29 Quotes; keine Community Note (FxTwitter). ## Wortlaut des Ausgangsposts > Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier. Der Reddit-Repost zitiert außerdem den Folgetext: > Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. […] Jev is ~18x cheaper than Terra! […] expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst. ## Methode laut Thread und Bildern - Multiple-Choice-Benchmarks auf gemeinsamer 0–100-%-Skala. - Terra: `reasoning=none`, „letter-only output“; beim Schach ausdrückliche `a/b/c/d`-Instruktion. - Benchmarks: MMLU, GPQA Diamond, ARC-Easy, ARC-Challenge, WinoGrande, HellaSwag, GSM8K mit 4 bzw. 10 Auswahlmöglichkeiten, Schach mit 4 legalen Zügen. - Die Radar-Grafik trägt den Untertitel „Official benchmark comparison“ und den Hinweis: „Each spoke is one benchmark condition. Polygon area is not an aggregate score.“ - Kein öffentliches Repository, keine Prompts, Beispieldateien oder Run-Logs im Thread verlinkt. ## Ergebnistabelle (aus Thread-Bild abgelesen) | Benchmark | Jev | Terra | |---|---:|---:| | MMLU | 90,91 % | 87,89 % | | GPQA Diamond | 68,18 % | 56,06 % | | ARC-Easy | 99,33 % | 98,82 % | | ARC-Challenge | 97,61 % | 96,76 % | | WinoGrande | 91,63 % | 78,69 % | | HellaSwag | 94,86 % | **95,32 %** | | GSM8K · 4 choices | 79,68 % | **87,72 %** | | GSM8K · 10 choices | 54,59 % | **69,83 %** | | Chess · 4 legal moves | 49,45 % | **50,25 %** | Hinweis: Der Text sagt, Jev übertreffe Terra bei „WinoGrande **or HellaSwag**“. Die veröffentlichte Tabelle zeigt bei HellaSwag **Terra 95,32 % vor Jev 94,86 %**. ## Kosten (aus Thread-Bild abgelesen) Titel: „Evaluation cost: Jev vs. Terra — Canonical benchmark runs · USD“. | Benchmark | Jev (estimated) | Terra (reported) | |---|---:|---:| | MMLU | $0,2749 | $4,6387 | | GPQA Diamond | $0,0048 | $0,1106 | | ARC-Easy | $0,0411 | $0,5427 | | ARC-Challenge | $0,0207 | $0,2897 | | WinoGrande | $0,0197 | $0,2652 | | HellaSwag | $0,2315 | $4,9176 | | GSM8K · 4 choices | $0,0247 | $0,3815 | | GSM8K · 10 choices | $0,0321 | $0,4663 | | Chess · 4 legal moves | $0,0504 | $1,3741 | | **Gesamt** | **$0,6999** | **$12,9865** | Fußnoten im Bild: - Jev: „recorded input tokens × $0.042/million; output free.“ - Terra: „summed API-reported costs. Auxiliary runs excluded.“ ## Kalibrierung (aus Thread-Bild abgelesen) Reliability-Diagramm „Jev reliability diagram“: - **N = 33.735** - **ECE = 1,74 Prozentpunkte** - 10 gleich breite Bins - Wilson-95-%-Konfidenzintervalle - Textangabe des Autors: beste Einzel-Bench **0,26 pp**, schlechteste **6,96 pp**. ## Weitere Thread-Aussagen - N8 Programs beschreibt Jev als günstigen Kandidaten für ungefähr „one Terra forward pass over a context“. - Frontier-Modelle wie Astra hätten vermutlich bessere No-CoT-Fähigkeiten, seien für Massenklassifikation aber zu teuer. - Schlussurteil: Jev sei „exactly what it says on the tin“, unter der Annahme, dass öffentliche Benchmarks in Alltagsaufgaben übertragen und nicht gezielt trainiert wurden.