knowledge-base/raw/xpost/2026-09-18_cjzafir-jev-3-40-dollar-erfahrung.md

84 lines
5.8 KiB
Markdown
Raw Permalink Normal View History

---
type: xpost
source_url: https://x.com/cjzafir/status/2100991512020725788
retrieved: 2026-09-18
author: "@cjzafir"
is_thread: false
tags: [jev, typesafe-ai, erfahrungsbericht, kosten, kontextfenster, rag, system-1-controller]
---
# CJ Zafir: „I burned $3.40 on Jev in 24 hours" (18.09.2026)
## Post (Wortlaut)
> I burned $3.40 on Jev in 24 hours.
>
> It's not Jev vs current LLMs. It'll be Jev + LLMs.
>
> Here are all the goods and bads of Jev.
>
> 1. Jev or a Jev-like system will be integrated into all LLMs for automatic model routing, small decisions, if/else queries, retrieval, tool selection, scoring, etc. For all these things, LLMs lose a lot of tokens on thinking and planning and use smarter instead of faster models. That can be fixed with Jev.
>
> 2. Jev can do things that require fine-tuning with LLMs. It can learn what to do from just 1 prompt.
>
> 3. It's crazy fast. It scans, analyzes, and gives its decisions instantly without a thinking blocker.
>
> 4. It doesn't waste tokens on thinking in plain English and planning. It scores probabilities and picks the best one.
>
> 5. Output cost of $0. Which makes sense because it's not built for free-form text generation like LLMs. It just outputs decisions.
>
> 6. It can fix RAG. The current 32k context is too low, but in the near future, we'll see systems like Jev with a 1M context window, and that'll fix the retrieval problem.
>
> 7. Jev has 3 decision types: multiple-choice questions, scoring the probability, and Noul.
> Choice → Which one?
> Score → How much?
> Noul → Yes or no?
> That's how it doesn't need reasoning.
>
> 8. Jev works on parallel sampling. LLMs generate sequentially: Token 1 ↓ Token 2 ↓ Token 3 ↓ Token 4. Jev can evaluate multiple questions against the same state in parallel. That's one of the reasons it has great latency.
>
> 9. Instead of RLHF (Reinforcement Learning from Human Feedback), it works on RLCD (Reinforcement Learning for Calibrated Decisions). With RLCD, the model's goal is not just to predict the answer; it also shows how much confidence it has. A model saying: true is different from: true, 99% confidence.
>
> 10. Jev response times are roughly: 70500 ms, which is 193.6× faster and 444.6× cheaper than frontier LLM workflows. The biggest issue with LLMs today is that they take too much time to execute tasks. With a Jev-like system, that'll be at least 50% faster.
## Metadaten
- **URL:** https://x.com/cjzafir/status/2100991512020725788 (von Pit Weber in OME Topic „JEV" geteilt, 18.09.2026 22:48 UTC)
- **Autor:** CJ Zafir ([@cjzafir](https://x.com/cjzafir), **64.375 Follower**, verifiziert; Bio: „I fine-tune small language models (SLMs) and beat large language models (LLMs)")
- **Erstellt:** 18.09.2026 16:53:05 UTC (also ~5,9 h vor dem Teilen)
- **Engagement (Abruf 18.09.2026 ~22:48 UTC):** 33.238 Views · 517 Likes · 38 Retweets · 32 Replies
- **Typ:** Einzelpost, Liste in 10 Punkten
## Neuheitswert
Der erste Beitrag der Jev-Akte mit einer **Ausgabenangabe aus eigener Nutzung**: 3,40 US-Dollar in 24 Stunden. Alle bisherigen Kostenaussagen waren Anbieterpreise oder Demo-Rechnungen Dritter.
## Gegenrechnung zur Ausgabe
Bei $0,042 pro Million Input-Token entsprechen $3,40 rund **81 Millionen Input-Token**. Zum Vergleich: der bekannte Demo-Durchlauf „1.700 E-Mails" verbrauchte 4,2 Mio. Input-Token — $3,40 sind also etwa **19 solcher Postfach-Durchläufe**. Die Zahl ist plausibel für einen Entwickler, der 24 Stunden lang systematisch experimentiert.
## Prüfung der Einzelaussagen
| # | Aussage | Prüfung |
|---|---------|---------|
| 1 | Jev als System-1-Controller neben LLMs (Routing, Tool-Wahl, Scoring) | ✅ deckt sich mit der Anbieter-Dokumentation und den genannten Anwendungsfällen |
| 2 | „Can learn what to do from just 1 prompt" | ⚠️ **Überzeichnet.** Belegt ist nicht Lernen, sondern **Konfiguration**: Beispiel + Schema definieren die Entscheidung. TypeSafe dokumentiert keine Gewichtsanpassung zur Laufzeit |
| 3 | „crazy fast … without a thinking blocker" | ✅ plausibel, deckt sich mit dem dokumentierten parallelen Sampling |
| 4 | Keine Planungs-Token, nur Wahrscheinlichkeiten | ✅ korrekt für Choice/Score/Noul |
| 5 | Output $0 | ✅ dokumentiert ($42/Btok bzw. $0,042/Mtok Input; Output kostenlos) |
| 6 | „current 32k context is too low; 1M will fix RAG" | ⚠️ **teilweise falsch.** Dokumentiert sind **64k pro Request** und **32k für State + längste Frage** — die genannte 32k sind die engere der beiden Grenzen, nicht die einzige. Das angekündigte 1M-Fenster existiert **nicht** in der Dokumentation und würde Retrieval-Probleme (Qualität, Aktualität, Berechtigungen, „lost in the middle") ohnehin nicht lösen |
| 7 | Drei Entscheidungstypen Choice/Score/Noul | ✅ korrekt (Anbieter-Primitive) |
| 8 | Paralleles Sampling statt sequenzieller Token | ✅ korrekt, Kern der Architektur |
| 9 | RLCD statt RLHF, Ziel ist Kalibrierung | ✅ korrekt (Anbieterangabe) |
| 10 | 70500 ms, 193,6× / 444,6× | ✅ Anbieterwerte, korrekt wiedergegeben — es sind **Spannen-Obergrenzen** aus Workflow-Evals, keine Agenten-End-to-End-Werte |
## Wichtige nicht genannte Einschränkung
Die Dokumentation nennt als Eingabe **ausschließlich Text** (String, JSON-Objekt oder Array) — „No image, audio, or video input". Für Computer-Use-Anwendungen heißt das: Bildschirminhalte können **nicht** direkt an Jev gehen; ein Agent müsste den Zustand erst in Text übersetzen. Das relativiert die Desktop-Demos, die den Eindruck erwecken, Jev „sehe" den Bildschirm.
## Offene Punkte
- ⚠️ Erfahrungsbericht eines einzelnen Entwicklers; keine Angabe, wofür die $3,40 ausgegeben wurden (Verteilung über Aufgabentypen fehlt).
- ⚠️ Der Autor hat ein Eigeninteresse an kleinen Modellen (Bio) — seine These „Jev + LLMs" passt zu dieser Position.
- ⚠️ Keine Messung der eigenen Latenz oder Genauigkeit.