Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model
![]()
Introduction: The Jev Announcement
On September 15, 2026, Diogo Almeida, a co-inventor of ChatGPT, announced a new type of AI model called “Jev” on X, after two years of stealth development. The announcement describes Jev like this:
• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions
The post got a big reaction, and other claims are also spreading on social media:
Claims spreading on social media (not in the announcement post itself)
- Input costs $0.042 per million tokens
- Returns a decision from up to 255 options
- No hallucinations, by design
- 67.8% agreement with the average judgment of frontier models
- “About 5,000 requests cost me only $2”
In this post I do not take these numbers and claims as given. Instead, I look at the direction shown in the announcement, a model optimized for decisions, and explain technically why that leads to speed and low cost.
This direction, “do not generate text; read the decision and its probability directly from inside the model”, shares the same idea as NerveReflex, an architecture I have been researching and publishing since May 2026. In the second half, I share results from running the same kind of setup (I call it “Open Jev” here) with an open model on my own workstation.
1. What a Decision-Focused Model Is
LLMs like ChatGPT write text for humans, one token at a time.
A model “optimized for decisions”, as Jev is described, can be understood as an AI that returns only a decision, for programs.
- Input: text or data, plus a question like “Is this A or B?”
- Output: which of the predefined options applies (the decision), and its probability (confidence)
flowchart LR
subgraph LLM["Typical LLM"]
A1["Customer message"] --> B1["Understand text"] --> C1["'Thank you. This is about billing...'<br/>(writes text token by token)"]
end
subgraph DEC["Decision-focused model"]
A2["Customer message"] --> B2["Understand text"] --> D2["{ category: billing, confidence: 0.94 }<br/>(returns decision + probability only)"]
end
Note: the outputs in the diagram are examples for explanation.
2. Why It Is Fast and Cheap: How It Works
Jev’s internal design is not public. Here I use how LLMs generally compute to explain why focusing on decisions leads to speed and low cost.
Most LLM cost comes from writing out text
LLM inference has two main stages.
- Prefill (reading the input): the model reads the whole input in parallel at once. This is efficient on a GPU and does not take much time or cost.
- Decode (writing text): the model predicts the next token, then uses it to predict the next one, and repeats this tens or hundreds of times. Each step needs the previous token, so it is hard to parallelize. This is the part that uses the most GPU memory and time.
flowchart LR
IN["Input text<br/>(hundreds to thousands of tokens)"] --> PF["Prefill<br/>one parallel pass"]
PF --> HS["hidden state<br/>(what the model understood)"]
HS --> D1["Decode 1"] --> D2["Decode 2"] --> D3["..."] --> DN["Decode N"]
DN --> OUT1["Text output"]
HS -.->|"Decision-focused / NerveReflex"| OUT2["Read decision + probability directly<br/>(no decode)"]
This is why cloud LLM pricing often sets the output token price higher than the input token price.
If no text is generated, there is almost no output-side cost
After reading the question, the model’s internals (the hidden state, a set of numbers) already hold what it understood. If you get the probability for each option from there in one pass, there is no need to repeat decode.
In my view, “output tokens free” in the announcement can be explained naturally by this structure: if the design does not generate text, there is almost no compute cost on the output side.
3. How to Read the Claims
Here I look at the claims spreading on social media from a technical point of view. I have not verified the numbers myself, so I only check whether each claim makes sense given how it works.
| Claim | Verdict (from how it works) | Note |
|---|---|---|
| Much faster | Fair | Without decode, response time gets much shorter |
| Much cheaper | Fair | Output-side compute is almost gone, so the cost structure changes |
| Output format never breaks | Fair | With no path to output anything outside the options, no format errors |
| No hallucinations | Too far | ”No format errors” and “no wrong decisions” are different things |
About “no hallucinations”
A design that returns only decisions cannot output a string outside the allowed options. So it avoids problems like a normal LLM breaking the format when you ask it to “answer in JSON”.
But that does not mean the model never picks the wrong option between A and B. If the 67.8% agreement figure quoted above is correct, the decisions differ from frontier models in about 30% of cases. A decision-focused model can still pick a wrong option with high confidence. The Open Jev test in the second half also shows this.
4. My Work on NerveReflex
“Do not output text; read the decision directly from the internal representation” is also the basic policy of NerveReflex, a research project I published in May 2026.
AI does not need to build a sentence every time it sorts or classifies business items. This is where NerveReflex started. In May, I compared two ways of doing the same contract-clause risk classification.
| Item | Default Gemma (text output) | NerveReflex (classify from internal representation) |
|---|---|---|
| How output is produced | Generates JSON token by token | Reads the hidden state directly |
| Generated tokens | 37 tokens | 0 tokens |
| Model runs | 38 (1 read + 37 writes) | 1 (read only) |
| Time | about 5 s | about 0.07 s (about 70x faster) |
| Output format | Free text (needs parsing, can break) | Fixed structure (cannot break) |
5. Testing “Open Jev” with an Open Model
Jev’s model and source code are not public. But if you understand the principle, you can reproduce the same approach with an open model on your own workstation.
I ran the open model Gemma 4 26B with vLLM on a single workstation (ASUS Ascent GX10). Using vLLM’s constrained output, I built a setup that does not write text and only reads the probability of each option from a single token step (Open Jev). Then I tested it on legal contract data.
flowchart LR
C["Contract clause<br/>(about 540 Japanese characters)"] --> G["Gemma 4 26B<br/>on vLLM"]
G -->|Prefill| H["hidden state"]
H --> L["Probabilities for 1 token<br/>constrained to 4 options"]
L --> P["Risk decision + confidence"]
P --> T["Temperature calibration"]
Note: the May test (Gemma 3 4B, classifier heads) and this test (Gemma 4 26B, vLLM constrained output) use different models and implementations, so their timing numbers cannot be compared directly.
Result 1: Speed (same model, same clause of about 540 characters)
The model and the input are the same. The only difference is whether the model wrote out text token by token.
Result 2: Accuracy, speed, and cost vs. frontier models
Measured on risk classification of 239 contract clauses.
| Model | Setup | Accuracy | Latency | Cost for 239 items |
|---|---|---|---|---|
| Gemma 4 26B (Open Jev) | Local (no thinking, 1-token decision) | 74.1% | 0.11 s | $0 (no API cost) |
| Claude Sonnet 5 | Cloud API (no thinking, 1-token decision) | 77.0% | 2.17 s | $0.47 |
| Claude Fable 5.1 | Cloud API (with thinking) | 80.0% | 6.06 s | about $3.8 |
- No-thinking vs. no-thinking: The accuracy gap between cloud Claude Sonnet 5 and the open model on my workstation was 2.9 points (77.0% vs. 74.1%). The local open model was about 20x faster.
- Cost: Jev is also said to have free output tokens, but Open Jev on local hardware has no API cost at all (electricity and hardware costs are separate).
The key to operation: auto-process only confident decisions, using temperature calibration
AI confidence scores tend to be higher than the real accuracy. A model may say “99%” even when the evidence is weak.
So I applied a statistical correction called temperature calibration, to get closer to a state where “90% confidence” really means about 90% correct. With that, a practical operating rule became possible.
flowchart TD
IN["Contract clause"] --> OJ["Open Jev (local)<br/>about 0.1 s, no API cost"]
OJ --> Q{"Calibrated confidence<br/>≥ 0.9?"}
Q -->|"Yes (about 43%)<br/>accuracy 91.7%"| AUTO["Auto-process"]
Q -->|"No (about 57%)"| ESC["Frontier model<br/>or human expert reviews"]
- The confident 40-plus percent of items are handled by the local model in about 0.1 s with no API cost.
- Only the less confident remaining items (just under 60%) go to a frontier model or a human expert.
With this setup, I expect total waiting time and API cost to drop a lot compared with sending every item to a frontier model.
6. Discussion: A Different Role from Frontier Models
In my view, the core of Jev and NerveReflex is a simple and reasonable design change: switching how an existing model gives output, from “generating text” to “reading internal probabilities and representations”.
| AI that thinks deeply (System 2) | AI that reacts quickly (System 1) | |
|---|---|---|
| Examples | Frontier models (with thinking) | Jev, NerveReflex, Open Jev |
| Output | Thousands of reasoning tokens and text | Typed decision and confidence |
| Latency | Seconds to minutes | Milliseconds |
| Good for | Complex reasoning, design, writing | Sorting, detection, routing, pre-approval checks |
Most business systems do not need long text. They need typed decisions like “Can this request be approved?” or “Which team should get this email?”. I think this text-free approach can give good cost-effectiveness for automating that kind of work.
Closing
With Jev’s release, the idea of “letting AI decide from its internal representation, fast and cheap, without generating unneeded text” is getting wide attention.
Paying for output tokens on frontier model APIs is not the only way to use AI.
By running open models on your own hardware and combining them with text-free decisions, I believe individuals and small teams without large capital can also build practical AI systems with little waste. I will keep researching and building this non-verbal AI architecture.
- Test environment: ASUS Ascent GX10 (single workstation)
- Model: Gemma-4-26B-A4B (FP8), vLLM
Join the conversation on LinkedIn — share your thoughts and comments.
Discuss on LinkedIn