knowledge-base/raw/other/2026-09-17_reddit-n8-jev-terra-benchmark.md

4.7 KiB
Raw Permalink Blame History

type source_url retrieved title author tags
other https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/ 2026-09-22 N8 Programs: Jev vs. GPT-5.6 Terra auf System-1-Benchmarks N8 Programs (@N8Programs); Reddit-Repost von /u/stealthispost
reddit
xpost
typesafe-ai
jev
gpt-5-6-terra
benchmark
system-one
calibration
ece
cost

N8 Programs — Jev vs. GPT-5.6 Terra (Reddit-Repost, 17.09.2026)

Quellenkette

Wortlaut des Ausgangsposts

Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.

Der Reddit-Repost zitiert außerdem den Folgetext:

Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. […] Jev is ~18x cheaper than Terra! […] expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst.

Methode laut Thread und Bildern

  • Multiple-Choice-Benchmarks auf gemeinsamer 0100-%-Skala.
  • Terra: reasoning=none, „letter-only output“; beim Schach ausdrückliche a/b/c/d-Instruktion.
  • Benchmarks: MMLU, GPQA Diamond, ARC-Easy, ARC-Challenge, WinoGrande, HellaSwag, GSM8K mit 4 bzw. 10 Auswahlmöglichkeiten, Schach mit 4 legalen Zügen.
  • Die Radar-Grafik trägt den Untertitel „Official benchmark comparison“ und den Hinweis: „Each spoke is one benchmark condition. Polygon area is not an aggregate score.“
  • Kein öffentliches Repository, keine Prompts, Beispieldateien oder Run-Logs im Thread verlinkt.

Ergebnistabelle (aus Thread-Bild abgelesen)

Benchmark Jev Terra
MMLU 90,91 % 87,89 %
GPQA Diamond 68,18 % 56,06 %
ARC-Easy 99,33 % 98,82 %
ARC-Challenge 97,61 % 96,76 %
WinoGrande 91,63 % 78,69 %
HellaSwag 94,86 % 95,32 %
GSM8K · 4 choices 79,68 % 87,72 %
GSM8K · 10 choices 54,59 % 69,83 %
Chess · 4 legal moves 49,45 % 50,25 %

Hinweis: Der Text sagt, Jev übertreffe Terra bei „WinoGrande or HellaSwag“. Die veröffentlichte Tabelle zeigt bei HellaSwag Terra 95,32 % vor Jev 94,86 %.

Kosten (aus Thread-Bild abgelesen)

Titel: „Evaluation cost: Jev vs. Terra — Canonical benchmark runs · USD“.

Benchmark Jev (estimated) Terra (reported)
MMLU $0,2749 $4,6387
GPQA Diamond $0,0048 $0,1106
ARC-Easy $0,0411 $0,5427
ARC-Challenge $0,0207 $0,2897
WinoGrande $0,0197 $0,2652
HellaSwag $0,2315 $4,9176
GSM8K · 4 choices $0,0247 $0,3815
GSM8K · 10 choices $0,0321 $0,4663
Chess · 4 legal moves $0,0504 $1,3741
Gesamt $0,6999 $12,9865

Fußnoten im Bild:

  • Jev: „recorded input tokens × $0.042/million; output free.“
  • Terra: „summed API-reported costs. Auxiliary runs excluded.“

Kalibrierung (aus Thread-Bild abgelesen)

Reliability-Diagramm „Jev reliability diagram“:

  • N = 33.735
  • ECE = 1,74 Prozentpunkte
  • 10 gleich breite Bins
  • Wilson-95-%-Konfidenzintervalle
  • Textangabe des Autors: beste Einzel-Bench 0,26 pp, schlechteste 6,96 pp.

Weitere Thread-Aussagen

  • N8 Programs beschreibt Jev als günstigen Kandidaten für ungefähr „one Terra forward pass over a context“.
  • Frontier-Modelle wie Astra hätten vermutlich bessere No-CoT-Fähigkeiten, seien für Massenklassifikation aber zu teuer.
  • Schlussurteil: Jev sei „exactly what it says on the tin“, unter der Annahme, dass öffentliche Benchmarks in Alltagsaufgaben übertragen und nicht gezielt trainiert wurden.