knowledge-base/raw/other/2026-09-17_reddit-n8-jev-terra-benchmark.md

90 lines
4.7 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

---
type: other
source_url: https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/
retrieved: 2026-09-22
title: "N8 Programs: Jev vs. GPT-5.6 Terra auf System-1-Benchmarks"
author: "N8 Programs (@N8Programs); Reddit-Repost von /u/stealthispost"
tags: [reddit, xpost, typesafe-ai, jev, gpt-5-6-terra, benchmark, system-one, calibration, ece, cost]
---
# N8 Programs — Jev vs. GPT-5.6 Terra (Reddit-Repost, 17.09.2026)
## Quellenkette
- **Geteilter Reddit-Post:** https://www.reddit.com/r/ProAI/comments/1wik66s/tested_typesafeai_s_claim_that_their_new_model/
- **Reddit-Autor:** `/u/stealthispost`
- **Reddit-Zeitpunkt laut RSS:** 17.09.2026, 04:44 UTC
- **Reddit-Zugriff:** normale Seite/JSON blockiert; `.rss` liefert HTTP 200 und den vollständigen Posttext.
- **Upstream-Thread:** N8 Programs ([@N8Programs](https://x.com/N8Programs)), beginnend https://x.com/N8Programs/status/2100088523403432357, 16.09.2026 05:04 UTC.
- **Upstream-Autorprofil:** „Studying Applied Mathematics and Statistics at @JohnsHopkins. Studying In-Context Learning at The Intelligence Amplification Lab.“; GitHub https://github.com/N8python.
- **Abruf 22.09.2026:** Ausgangspost 179.846 Views, 907 Likes, 70 Reposts, 29 Quotes; keine Community Note (FxTwitter).
## Wortlaut des Ausgangsposts
> Tested @typesafeai's claim that their new model Jev delivered "comparable... intelligence" to GPT-5.6 Terra on "System 1" tasks. To do this, I compare both models on multiple-choice benchmarks (MMLU, GPQA, etc.). Set reasoning=none for Terra for sys 1. Result: Jev is Terra-tier.
Der Reddit-Repost zitiert außerdem den Folgetext:
> Jev's performance on knowledge benchmarks like MMLU and GPQA is highly impressive - especially for a non-COT model. It exceeds Terra at other linguistic reasoning tasks like WinoGrande or HellaSwag as well. It only loses substantially on math reasoning. […] Jev is ~18x cheaper than Terra! […] expected calibration error is 1.74 percentage points averaged across benches - its 0.26pp at best and 6.96pp at worst.
## Methode laut Thread und Bildern
- Multiple-Choice-Benchmarks auf gemeinsamer 0100-%-Skala.
- Terra: `reasoning=none`, „letter-only output“; beim Schach ausdrückliche `a/b/c/d`-Instruktion.
- Benchmarks: MMLU, GPQA Diamond, ARC-Easy, ARC-Challenge, WinoGrande, HellaSwag, GSM8K mit 4 bzw. 10 Auswahlmöglichkeiten, Schach mit 4 legalen Zügen.
- Die Radar-Grafik trägt den Untertitel „Official benchmark comparison“ und den Hinweis: „Each spoke is one benchmark condition. Polygon area is not an aggregate score.“
- Kein öffentliches Repository, keine Prompts, Beispieldateien oder Run-Logs im Thread verlinkt.
## Ergebnistabelle (aus Thread-Bild abgelesen)
| Benchmark | Jev | Terra |
|---|---:|---:|
| MMLU | 90,91 % | 87,89 % |
| GPQA Diamond | 68,18 % | 56,06 % |
| ARC-Easy | 99,33 % | 98,82 % |
| ARC-Challenge | 97,61 % | 96,76 % |
| WinoGrande | 91,63 % | 78,69 % |
| HellaSwag | 94,86 % | **95,32 %** |
| GSM8K · 4 choices | 79,68 % | **87,72 %** |
| GSM8K · 10 choices | 54,59 % | **69,83 %** |
| Chess · 4 legal moves | 49,45 % | **50,25 %** |
Hinweis: Der Text sagt, Jev übertreffe Terra bei „WinoGrande **or HellaSwag**“. Die veröffentlichte Tabelle zeigt bei HellaSwag **Terra 95,32 % vor Jev 94,86 %**.
## Kosten (aus Thread-Bild abgelesen)
Titel: „Evaluation cost: Jev vs. Terra — Canonical benchmark runs · USD“.
| Benchmark | Jev (estimated) | Terra (reported) |
|---|---:|---:|
| MMLU | $0,2749 | $4,6387 |
| GPQA Diamond | $0,0048 | $0,1106 |
| ARC-Easy | $0,0411 | $0,5427 |
| ARC-Challenge | $0,0207 | $0,2897 |
| WinoGrande | $0,0197 | $0,2652 |
| HellaSwag | $0,2315 | $4,9176 |
| GSM8K · 4 choices | $0,0247 | $0,3815 |
| GSM8K · 10 choices | $0,0321 | $0,4663 |
| Chess · 4 legal moves | $0,0504 | $1,3741 |
| **Gesamt** | **$0,6999** | **$12,9865** |
Fußnoten im Bild:
- Jev: „recorded input tokens × $0.042/million; output free.“
- Terra: „summed API-reported costs. Auxiliary runs excluded.“
## Kalibrierung (aus Thread-Bild abgelesen)
Reliability-Diagramm „Jev reliability diagram“:
- **N = 33.735**
- **ECE = 1,74 Prozentpunkte**
- 10 gleich breite Bins
- Wilson-95-%-Konfidenzintervalle
- Textangabe des Autors: beste Einzel-Bench **0,26 pp**, schlechteste **6,96 pp**.
## Weitere Thread-Aussagen
- N8 Programs beschreibt Jev als günstigen Kandidaten für ungefähr „one Terra forward pass over a context“.
- Frontier-Modelle wie Astra hätten vermutlich bessere No-CoT-Fähigkeiten, seien für Massenklassifikation aber zu teuer.
- Schlussurteil: Jev sei „exactly what it says on the tin“, unter der Annahme, dass öffentliche Benchmarks in Alltagsaufgaben übertragen und nicht gezielt trainiert wurden.