Reflex-1: Frontier-LLM Accuracy at Jev Speed with a 4B Open Model

View on GitHub

Challenge

When AI is used for business classification and judgment, frontier LLMs such as GPT-5.6 and Gemini 3.8 are accurate but take 1–3 seconds per question and are billed per use in the cloud. The dedicated judgment API Jev is fast at 0.3 seconds, but breaks down on decisions that require reasoning and does not accept images. Neither publishes its weights, so neither can be trained on a company's own judgments.

Solution

Built on Google Gemma 4 E4B, Reflex-1 is a 4B open model that combines classification with latent reasoning (reasoning inside the model instead of writing the reasoning out as text). Requests use a Jev-compatible shape, and training per judgment (LoRA plus small latent-reasoning modules) lets it learn judgments specific to a company and handle them at the same speed. Each reasoning step can be decoded back into words afterwards.

Result

On classification, Reflex-1 stays in the same accuracy band as frontier LLMs with a median latency of 0.10 s per decision: 12× faster than GPT-5.6 and faster than Jev (0.30 s). On a trained reasoning decision (a 10-hop logic puzzle), it reaches 100% in 0.8 s, where frontier LLMs take 2.7 s. The weights are 9.8 GB at 4-bit and fit one GPU, and the model reads images. It was built on a personal setup (one GX10) in 4 days, with about USD 2.5 in external API costs for the comparisons.

Reflex-1: frontier-LLM accuracy at Jev speed. Latency per decision: Reflex-1 0.10 s, Jev 0.30 s, GPT-5.6 1.24 s, Gemini 3.8 Flash 2.28 s

Reflex-1 is a derived result of the research project NerveReflex (an AI architecture in which models cooperate through internal representations rather than language), turning its ideas into an open model that can be used in practice.

On classification it returns answers in the same accuracy band as frontier LLMs such as GPT-5.6 and Gemini 3.8, faster than the dedicated judgment API Jev, and once trained on your own judgments it reaches frontier-level accuracy at Jev-level speed even on decisions that require reasoning. It runs on a local GPU and reads images. Built on a personal setup, in a few days, for a few dollars.

12×faster than GPT-5.6 on classification (same accuracy band)
0.10 smedian latency per decision. Jev: 0.30 s
100% · 0.8 strained reasoning decision (10-hop logic puzzle, same 50 questions). Frontier LLM: 100% in 2.7 s. Jev: 52% in 0.3 s
4B · 9.8 GBopen weights, 4-bit, fits one GPU. Reads images. No per-call fee

Classification: same accuracy, a different speed class

Business classification tasks (routing inquiries, emotion of posts, moderation) were compared with identical inputs and options. Accuracy lands in the same band for all four; only the time to answer differs, by an order of magnitude.

Latency per decision (5 text classification tasks)
Latency per decision (5 text classification tasks)
Text classification: accuracy × speed
mean of intent, utterance intent, emotion, sentiment, moderation
Text classification: accuracy × speed
Image classification: accuracy × speed
mean of food photos and e-commerce products. Jev has no image input
Image classification: accuracy × speed
Upper-left is fast and accurate. Reflex-1 runs sequentially on one GPU (GX10); Jev and frontier models via API, network included. Details under Evaluations.
Instant decisions vs Jev — the exact options shown to the models; a tag appears the moment each responds.
Image classification (Jev: text only) — 30 food and product photos, about 0.3 s each (CC0 / public-domain photos from Wikimedia Commons).

Reasoning too: train it, and get frontier-LLM accuracy at Jev speed

Jev is fast but breaks down on decisions that require reasoning. Frontier LLMs are accurate with reasoning on, but take seconds per question. Reflex-1 gets both through latent reasoning: it reasons inside the model instead of writing its reasoning out as text. Train it on a judgment and it returns frontier-LLM accuracy even on decisions that require reasoning, in about a second, like Jev. Each step can be decoded back into words afterwards. The chart shows the result of training on five logic tasks that require reasoning.

Decisions that require reasoning: accuracy × speed (mean of 5 trained logic tasks)
Decisions that require reasoning: accuracy × speed (mean of 5 trained logic tasks)
Tasks: BIG-Bench Hard logic tasks deepened with the same rules (Web of Lies at 10 and 20 hops, object tracking with 7 people and 14 swaps, boolean expressions, navigation). Trained per task.
Reasoning vs frontier and Jev — starts with how to read the puzzle; tags on the left are decoded latent steps.

Judgments specific to your company can be learned, too

This reasoning was acquired by per-task training; it does not transfer to arbitrary reasoning as-is. The flip side: judgments specific to your company — matching invoices against orders, deciding approval routes, checking policy compliance — can be trained the same way, and Reflex-1 then reasons through them at the same speed. Closed models like Jev and frontier LLMs cannot be trained on your own judgments like this.

How Reflex-1 differs from frontier LLMs and Jev

frontier LLM
GPT-5.6 / Gemini 3.8
Jev
typesafe/jev-1.13
Reflex-1
Instant decisions▲accurate, but 1–3 s●same band, 0.3 s●same band, 0.1–0.2 s
Decisions that require reasoning▲accurate with thinking, seconds×breaks down●~1 s (trained judgments)
Train on your own judgments×not possible×not possible●trainable
Inspect the reasoning●read thinking tokens×—●decode latent steps
Images●supported×not supported●~0.2 s per image
Runs locally, footprint×cloud only×cloud only●one GPU, 9.8 GB at 4-bit
Open weights×closed×closed●open
Per-call cost×pay per token×pay per call●none
Request format▲chat●Decisions API●same as Jev
● yes / strong ▲ with conditions × no

Evaluations: frontier LLMs vs Jev vs Reflex-1

Same question sets and options. Jev and Reflex-1: full sets (200 per classification task, 100 per reasoning task); frontier LLMs: reference values on subsets drawn from them (100 and 50). Numbers are accuracy %, latency in parentheses.

Instant decisions (text)

Give a text and options; pick one. Which desk should handle an inquiry / intent of a Japanese voice-assistant utterance / emotion (4 classes) and positive / negative / neutral (3 classes) of a social post / whether a comment insults someone (moderation).

Taskfrontier LLMJevReflex-1
GPT-5.6 SolGemini 3.8 Flash
Intent (CLINC150)979696.593.5
Japanese utterance intent (MASSIVE)969692.590.0
Emotion (tweet_eval)847981.082.0
Sentiment (tweet_eval)787774.574.5
Moderation: does it insult someone? (civil_comments)737673.580.0
median latency (s)1.1–1.32.1–2.60.300.09–0.21

Instant decisions (images)

Give one photo; pick its category. Food photos (10 dishes) / e-commerce product photos (10 categories) / scanned document types (16 classes).

Taskfrontier LLMJevReflex-1
GPT-5.6 SolGemini 3.8 Flash
Food photos (Food-101, 10 classes)9898—98.5
E-commerce products (10 classes)9096—83.0
Document type (RVL-CDIP, 16 classes)———58.0
median latency (s)1.8–2.02.9–3.1—0.21–0.26

Reasoning (trained logic tasks)

Logic puzzles that require reasoning (Reflex-1 trained per task). Web of Lies: 10 or 20 people each say whether the previous one lies; is the last one truthful? / Object tracking: after 7 people swap balls 14 times, which ball does a given person hold? / Boolean expressions: evaluate a not/and/or expression / Navigation: after following the steps, are you back at the start?

Taskfrontier LLMJevReflex-1
GPT-5.6 Sol
reasoning off
Gemini 3.8 Flash
reasoning on
Web of Lies, 10 hops46100 (2.7 s)59.0100 (0.8 s)
Web of Lies, 20 hops50100 (3.1 s)58.072 (1.5 s)
Object tracking, 7 people / 14 swaps10100 (3.5 s)26.270 (1.2 s)
Boolean expressions (BBH)92100 (2.4 s)98.9100 (0.4 s)
Navigation (BBH)78100 (2.5 s)98.280.8 (0.5 s)
median latency (s)1.2–1.32.4–3.50.300.4–1.5

Try it

Requests use a Jev-compatible shape (state, questions, instructions, criteria; choice type). mode, skill, trace, image and the text type are Reflex-1 extensions.

POST /v1/decisions
{"state": "I want to know my credit rating",
 "questions": {"desk": {"type": "choice",
   "instructions": "Which desk should handle this inquiry?",
   "criteria": {"credit_score": "", "transfer": "", "none": "none of these"}}}}
→ {"answers": {"desk": {"choice": "credit_score", "confidence": 0.97,
     "probs": {"credit_score": 0.97, "transfer": 0.01, "none": 0.02}}}, "latency_ms": 90}

{"state": "Question: Ryan tells the truth. Michael says Ryan lies. ...",
 "mode": "deep", "skill": "web_of_lies", "trace": true, "questions": {...}}
→ {"answers": {"a": {"choice": "No", "confidence": 1.0,
     "trace": ["start -> Ryan tells the truth", "Michael says Ryan lies -> Michael lies", "..."]}},
   "latency_ms": 820}

Cite: Suzuki, G. (2026). Reflex-1: frontier-LLM accuracy at Jev speed with a 4B open model. Zenodo. doi:10.5281/zenodo.23005201

How it was built

No large compute cluster was used; Reflex-1 was developed on a single local GPU machine (ASUS Ascent GX10). Training, evaluation and the videos were all done on this machine.

ItemDetails
Machine1 × ASUS Ascent GX10 (NVIDIA GB10 Grace Blackwell, 20-core CPU, 128 GB unified CPU/GPU memory)
OS / softwareUbuntu 24.04 (aarch64) · PyTorch 2.14 (CUDA 13) · Transformers 5.17 · PEFT 0.21 (LoRA) · bitsandbytes 0.50 · vLLM 0.19
Base modelGoogle Gemma 4 E4B (Apache 2.0)
TrainingLoRA plus small latent-reasoning modules. 1.4–2.2 hours per task on one GPU
Development time4 days, including the comparison measurements
External costAbout $2.5 in API calls for the comparisons (Jev, GPT-5.6, Gemini 3.8). No cloud compute for training

Notes

  • Decisions that require reasoning need per-judgment training. Reflex-1 is strong on the kinds and depths of judgment it was trained on, and does not transfer as-is to untrained kinds of reasoning. Classification works without any training.
  • Reasoning mode currently runs on NVIDIA GPUs (CUDA). Support for common runtimes such as llama.cpp is coming shortly.

The ideas behind Reflex-1 and the earlier experiments are summarized on the NerveReflex project page:

Measured in September 2026 on one GX10 (NVIDIA GB10). Jev (typesafe/jev-1.13) and frontier models via OpenRouter; total API cost about $2.5. Reflex-1 is fine-tuned from Google Gemma 4 E4B (Apache 2.0); not affiliated with or endorsed by Google. Jev is a product of TypeSafe AI; results were measured by the author via OpenRouter and are not endorsed by TypeSafe AI. Data: CLINC150 (Larson et al. 2019, CC BY 3.0), MASSIVE (FitzGerald et al. 2022, CC BY 4.0), Civil Comments (Jigsaw, CC0), BIG-Bench Hard (Suzgun et al. 2022, MIT), tweet_eval (Barbieri et al. 2020), Food-101 (Bossard et al. 2014), Fashion Product Images, RVL-CDIP (Harley et al. 2015). Photos in the videos: CC0 / public domain from Wikimedia Commons. Concept: NerveReflex.

Share this article

Join the conversation on LinkedIn — share your thoughts and comments.

Discuss on LinkedIn