4.7 KiB
| type | source_url | retrieved | title | author | tags | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| other | https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/ | 2026-09-22 | N8 Programs: Jev vs. GPT-5.6 Terra auf System-1-Benchmarks | N8 Programs (@N8Programs); Reddit-Repost von /u/stealthispost |
|
N8 Programs — Jev vs. GPT-5.6 Terra (Reddit-Repost, 17.09.2026)
Quellenkette
- Geteilter Reddit-Post: https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/
- Reddit-Autor:
/u/stealthispost - Reddit-Zeitpunkt laut RSS: 17.09.2026, 04:44 UTC
- Reddit-Zugriff: normale Seite/JSON blockiert;
.rssliefert HTTP 200 und den vollständigen Posttext. - Upstream-Thread: N8 Programs (@N8Programs), beginnend https://x.com/N8Programs/status/2100088523403432357, 16.09.2026 05:04 UTC.
- Upstream-Autorprofil: „Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.“; GitHub https://github.com/N8python.
- Abruf 22.09.2026: Ausgangspost 179.846 Views, 907 Likes, 70 Reposts, 29 Quotes; keine Community Note (FxTwitter).
Wortlaut des Ausgangsposts
Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.
Der Reddit-Repost zitiert außerdem den Folgetext:
Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. […] Jev is ~18x cheaper than Terra! […] expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst.
Methode laut Thread und Bildern
- Multiple-Choice-Benchmarks auf gemeinsamer 0–100-%-Skala.
- Terra:
reasoning=none, „letter-only output“; beim Schach ausdrücklichea/b/c/d-Instruktion. - Benchmarks: MMLU, GPQA Diamond, ARC-Easy, ARC-Challenge, WinoGrande, HellaSwag, GSM8K mit 4 bzw. 10 Auswahlmöglichkeiten, Schach mit 4 legalen Zügen.
- Die Radar-Grafik trägt den Untertitel „Official benchmark comparison“ und den Hinweis: „Each spoke is one benchmark condition. Polygon area is not an aggregate score.“
- Kein öffentliches Repository, keine Prompts, Beispieldateien oder Run-Logs im Thread verlinkt.
Ergebnistabelle (aus Thread-Bild abgelesen)
| Benchmark | Jev | Terra |
|---|---|---|
| MMLU | 90,91 % | 87,89 % |
| GPQA Diamond | 68,18 % | 56,06 % |
| ARC-Easy | 99,33 % | 98,82 % |
| ARC-Challenge | 97,61 % | 96,76 % |
| WinoGrande | 91,63 % | 78,69 % |
| HellaSwag | 94,86 % | 95,32 % |
| GSM8K · 4 choices | 79,68 % | 87,72 % |
| GSM8K · 10 choices | 54,59 % | 69,83 % |
| Chess · 4 legal moves | 49,45 % | 50,25 % |
Hinweis: Der Text sagt, Jev übertreffe Terra bei „WinoGrande or HellaSwag“. Die veröffentlichte Tabelle zeigt bei HellaSwag Terra 95,32 % vor Jev 94,86 %.
Kosten (aus Thread-Bild abgelesen)
Titel: „Evaluation cost: Jev vs. Terra — Canonical benchmark runs · USD“.
| Benchmark | Jev (estimated) | Terra (reported) |
|---|---|---|
| MMLU | $0,2749 | $4,6387 |
| GPQA Diamond | $0,0048 | $0,1106 |
| ARC-Easy | $0,0411 | $0,5427 |
| ARC-Challenge | $0,0207 | $0,2897 |
| WinoGrande | $0,0197 | $0,2652 |
| HellaSwag | $0,2315 | $4,9176 |
| GSM8K · 4 choices | $0,0247 | $0,3815 |
| GSM8K · 10 choices | $0,0321 | $0,4663 |
| Chess · 4 legal moves | $0,0504 | $1,3741 |
| Gesamt | $0,6999 | $12,9865 |
Fußnoten im Bild:
- Jev: „recorded input tokens × $0.042/million; output free.“
- Terra: „summed API-reported costs. Auxiliary runs excluded.“
Kalibrierung (aus Thread-Bild abgelesen)
Reliability-Diagramm „Jev reliability diagram“:
- N = 33.735
- ECE = 1,74 Prozentpunkte
- 10 gleich breite Bins
- Wilson-95-%-Konfidenzintervalle
- Textangabe des Autors: beste Einzel-Bench 0,26 pp, schlechteste 6,96 pp.
Weitere Thread-Aussagen
- N8 Programs beschreibt Jev als günstigen Kandidaten für ungefähr „one Terra forward pass over a context“.
- Frontier-Modelle wie Astra hätten vermutlich bessere No-CoT-Fähigkeiten, seien für Massenklassifikation aber zu teuer.
- Schlussurteil: Jev sei „exactly what it says on the tin“, unter der Annahme, dass öffentliche Benchmarks in Alltagsaufgaben übertragen und nicht gezielt trainiert wurden.