Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model

Introduction: The Jev Announcement

On September 15, 2026, Diogo Almeida, a co-inventor of ChatGPT, announced a new type of AI model called “Jev” on X, after two years of stealth development. The announcement describes Jev like this:

• 20-200x faster • 40-400x cheaper (w/ output tokens free) • Frontier composable intelligence optimized for decisions

— Diogo Almeida (@CompleteSkeptic), September 15, 2026

The post got a big reaction, and other claims are also spreading on social media:

Claims spreading on social media (not in the announcement post itself)

  • Input costs $0.042 per million tokens
  • Returns a decision from up to 255 options
  • No hallucinations, by design
  • 67.8% agreement with the average judgment of frontier models
  • “About 5,000 requests cost me only $2”

In this post I do not take these numbers and claims as given. Instead, I look at the direction shown in the announcement, a model optimized for decisions, and explain technically why that leads to speed and low cost.

This direction, “do not generate text; read the decision and its probability directly from inside the model”, shares the same idea as NerveReflex, an architecture I have been researching and publishing since May 2026. In the second half, I share results from running the same kind of setup (I call it “Open Jev” here) with an open model on my own workstation.

1. What a Decision-Focused Model Is

LLMs like ChatGPT write text for humans, one token at a time.

A model “optimized for decisions”, as Jev is described, can be understood as an AI that returns only a decision, for programs.

  • Input: text or data, plus a question like “Is this A or B?”
  • Output: which of the predefined options applies (the decision), and its probability (confidence)
flowchart LR
    subgraph LLM["Typical LLM"]
        A1["Customer message"] --> B1["Understand text"] --> C1["'Thank you. This is about billing...'<br/>(writes text token by token)"]
    end
    subgraph DEC["Decision-focused model"]
        A2["Customer message"] --> B2["Understand text"] --> D2["{ category: billing, confidence: 0.94 }<br/>(returns decision + probability only)"]
    end

Note: the outputs in the diagram are examples for explanation.

2. Why It Is Fast and Cheap: How It Works

Jev’s internal design is not public. Here I use how LLMs generally compute to explain why focusing on decisions leads to speed and low cost.

Most LLM cost comes from writing out text

LLM inference has two main stages.

  1. Prefill (reading the input): the model reads the whole input in parallel at once. This is efficient on a GPU and does not take much time or cost.
  2. Decode (writing text): the model predicts the next token, then uses it to predict the next one, and repeats this tens or hundreds of times. Each step needs the previous token, so it is hard to parallelize. This is the part that uses the most GPU memory and time.
flowchart LR
    IN["Input text<br/>(hundreds to thousands of tokens)"] --> PF["Prefill<br/>one parallel pass"]
    PF --> HS["hidden state<br/>(what the model understood)"]
    HS --> D1["Decode 1"] --> D2["Decode 2"] --> D3["..."] --> DN["Decode N"]
    DN --> OUT1["Text output"]
    HS -.->|"Decision-focused / NerveReflex"| OUT2["Read decision + probability directly<br/>(no decode)"]

This is why cloud LLM pricing often sets the output token price higher than the input token price.

If no text is generated, there is almost no output-side cost

After reading the question, the model’s internals (the hidden state, a set of numbers) already hold what it understood. If you get the probability for each option from there in one pass, there is no need to repeat decode.

In my view, “output tokens free” in the announcement can be explained naturally by this structure: if the design does not generate text, there is almost no compute cost on the output side.

3. How to Read the Claims

Here I look at the claims spreading on social media from a technical point of view. I have not verified the numbers myself, so I only check whether each claim makes sense given how it works.

ClaimVerdict (from how it works)Note
Much fasterFairWithout decode, response time gets much shorter
Much cheaperFairOutput-side compute is almost gone, so the cost structure changes
Output format never breaksFairWith no path to output anything outside the options, no format errors
No hallucinationsToo far”No format errors” and “no wrong decisions” are different things

About “no hallucinations”

A design that returns only decisions cannot output a string outside the allowed options. So it avoids problems like a normal LLM breaking the format when you ask it to “answer in JSON”.

But that does not mean the model never picks the wrong option between A and B. If the 67.8% agreement figure quoted above is correct, the decisions differ from frontier models in about 30% of cases. A decision-focused model can still pick a wrong option with high confidence. The Open Jev test in the second half also shows this.

4. My Work on NerveReflex

“Do not output text; read the decision directly from the internal representation” is also the basic policy of NerveReflex, a research project I published in May 2026.

AI does not need to build a sentence every time it sorts or classifies business items. This is where NerveReflex started. In May, I compared two ways of doing the same contract-clause risk classification.

The same risk classification, written out as tokens vs. read directly from the internal representation (May 2026, Gemma 3 4B).
ItemDefault Gemma (text output)NerveReflex (classify from internal representation)
How output is producedGenerates JSON token by tokenReads the hidden state directly
Generated tokens37 tokens0 tokens
Model runs38 (1 read + 37 writes)1 (read only)
Timeabout 5 sabout 0.07 s (about 70x faster)
Output formatFree text (needs parsing, can break)Fixed structure (cannot break)

5. Testing “Open Jev” with an Open Model

Jev’s model and source code are not public. But if you understand the principle, you can reproduce the same approach with an open model on your own workstation.

I ran the open model Gemma 4 26B with vLLM on a single workstation (ASUS Ascent GX10). Using vLLM’s constrained output, I built a setup that does not write text and only reads the probability of each option from a single token step (Open Jev). Then I tested it on legal contract data.

flowchart LR
    C["Contract clause<br/>(about 540 Japanese characters)"] --> G["Gemma 4 26B<br/>on vLLM"]
    G -->|Prefill| H["hidden state"]
    H --> L["Probabilities for 1 token<br/>constrained to 4 options"]
    L --> P["Risk decision + confidence"]
    P --> T["Temperature calibration"]

Note: the May test (Gemma 3 4B, classifier heads) and this test (Gemma 4 26B, vLLM constrained output) use different models and implementations, so their timing numbers cannot be compared directly.

Result 1: Speed (same model, same clause of about 540 characters)

Time to classify the same contract clause (seconds, shorter is faster)
Generate JSON text
1.774 s
Read option probabilities
0.074 s
Scaled with 1.774 s = 100%. About a 24x difference.

The model and the input are the same. The only difference is whether the model wrote out text token by token.

Result 2: Accuracy, speed, and cost vs. frontier models

Measured on risk classification of 239 contract clauses.

ModelSetupAccuracyLatencyCost for 239 items
Gemma 4 26B (Open Jev)Local (no thinking, 1-token decision)74.1%0.11 s$0 (no API cost)
Claude Sonnet 5Cloud API (no thinking, 1-token decision)77.0%2.17 s$0.47
Claude Fable 5.1Cloud API (with thinking)80.0%6.06 sabout $3.8
Accuracy (%, higher is better)
Gemma 4 26B (Open Jev)
74.1%
Claude Sonnet 5
77.0%
Claude Fable 5.1 (thinking)
80.0%
Latency per item (seconds, shorter is faster)
Gemma 4 26B (Open Jev)
0.11 s
Claude Sonnet 5
2.17 s
Claude Fable 5.1 (thinking)
6.06 s
Latency scaled with 6.06 s = 100%.
  • No-thinking vs. no-thinking: The accuracy gap between cloud Claude Sonnet 5 and the open model on my workstation was 2.9 points (77.0% vs. 74.1%). The local open model was about 20x faster.
  • Cost: Jev is also said to have free output tokens, but Open Jev on local hardware has no API cost at all (electricity and hardware costs are separate).

The key to operation: auto-process only confident decisions, using temperature calibration

AI confidence scores tend to be higher than the real accuracy. A model may say “99%” even when the evidence is weak.

So I applied a statistical correction called temperature calibration, to get closer to a state where “90% confidence” really means about 90% correct. With that, a practical operating rule became possible.

After calibration: decisions with confidence ≥ 0.9 (share of all items, and their accuracy)
Share of all items
about 43%
Accuracy in that range
91.7%
(Ref.) Accuracy on all
74.1%
flowchart TD
    IN["Contract clause"] --> OJ["Open Jev (local)<br/>about 0.1 s, no API cost"]
    OJ --> Q{"Calibrated confidence<br/>≥ 0.9?"}
    Q -->|"Yes (about 43%)<br/>accuracy 91.7%"| AUTO["Auto-process"]
    Q -->|"No (about 57%)"| ESC["Frontier model<br/>or human expert reviews"]
  • The confident 40-plus percent of items are handled by the local model in about 0.1 s with no API cost.
  • Only the less confident remaining items (just under 60%) go to a frontier model or a human expert.

With this setup, I expect total waiting time and API cost to drop a lot compared with sending every item to a frontier model.

6. Discussion: A Different Role from Frontier Models

In my view, the core of Jev and NerveReflex is a simple and reasonable design change: switching how an existing model gives output, from “generating text” to “reading internal probabilities and representations”.

AI that thinks deeply (System 2)AI that reacts quickly (System 1)
ExamplesFrontier models (with thinking)Jev, NerveReflex, Open Jev
OutputThousands of reasoning tokens and textTyped decision and confidence
LatencySeconds to minutesMilliseconds
Good forComplex reasoning, design, writingSorting, detection, routing, pre-approval checks

Most business systems do not need long text. They need typed decisions like “Can this request be approved?” or “Which team should get this email?”. I think this text-free approach can give good cost-effectiveness for automating that kind of work.

Closing

With Jev’s release, the idea of “letting AI decide from its internal representation, fast and cheap, without generating unneeded text” is getting wide attention.

Paying for output tokens on frontier model APIs is not the only way to use AI.

By running open models on your own hardware and combining them with text-free decisions, I believe individuals and small teams without large capital can also build practical AI systems with little waste. I will keep researching and building this non-verbal AI architecture.

  • Test environment: ASUS Ascent GX10 (single workstation)
  • Model: Gemma-4-26B-A4B (FP8), vLLM

Share this article

Join the conversation on LinkedIn — share your thoughts and comments.

Discuss on LinkedIn

Related Posts