Reflex-1: Frontier-LLM Accuracy at Jev Speed with a 4B Open Model
View on GitHubChallenge
When AI is used for business classification and judgment, frontier LLMs such as GPT-5.6 and Gemini 3.8 are accurate but take 1–3 seconds per question and are billed per use in the cloud. The dedicated judgment API Jev is fast at 0.3 seconds, but breaks down on decisions that require reasoning and does not accept images. Neither publishes its weights, so neither can be trained on a company's own judgments.
Solution
Built on Google Gemma 4 E4B, Reflex-1 is a 4B open model that combines classification with latent reasoning (reasoning inside the model instead of writing the reasoning out as text). Requests use a Jev-compatible shape, and training per judgment (LoRA plus small latent-reasoning modules) lets it learn judgments specific to a company and handle them at the same speed. Each reasoning step can be decoded back into words afterwards.
Result
On classification, Reflex-1 stays in the same accuracy band as frontier LLMs with a median latency of 0.10 s per decision: 12× faster than GPT-5.6 and faster than Jev (0.30 s). On a trained reasoning decision (a 10-hop logic puzzle), it reaches 100% in 0.8 s, where frontier LLMs take 2.7 s. The weights are 9.8 GB at 4-bit and fit one GPU, and the model reads images. It was built on a personal setup (one GX10) in 4 days, with about USD 2.5 in external API costs for the comparisons.
![]()
Reflex-1 is a derived result of the research project NerveReflex (an AI architecture in which models cooperate through internal representations rather than language), turning its ideas into an open model that can be used in practice.
On classification it returns answers in the same accuracy band as frontier LLMs such as GPT-5.6 and Gemini 3.8, faster than the dedicated judgment API Jev, and once trained on your own judgments it reaches frontier-level accuracy at Jev-level speed even on decisions that require reasoning. It runs on a local GPU and reads images. Built on a personal setup, in a few days, for a few dollars.
Classification: same accuracy, a different speed class
Business classification tasks (routing inquiries, emotion of posts, moderation) were compared with identical inputs and options. Accuracy lands in the same band for all four; only the time to answer differs, by an order of magnitude.
Reasoning too: train it, and get frontier-LLM accuracy at Jev speed
Jev is fast but breaks down on decisions that require reasoning. Frontier LLMs are accurate with reasoning on, but take seconds per question. Reflex-1 gets both through latent reasoning: it reasons inside the model instead of writing its reasoning out as text. Train it on a judgment and it returns frontier-LLM accuracy even on decisions that require reasoning, in about a second, like Jev. Each step can be decoded back into words afterwards. The chart shows the result of training on five logic tasks that require reasoning.
Judgments specific to your company can be learned, too
This reasoning was acquired by per-task training; it does not transfer to arbitrary reasoning as-is. The flip side: judgments specific to your company — matching invoices against orders, deciding approval routes, checking policy compliance — can be trained the same way, and Reflex-1 then reasons through them at the same speed. Closed models like Jev and frontier LLMs cannot be trained on your own judgments like this.
How Reflex-1 differs from frontier LLMs and Jev
| frontier LLM GPT-5.6 / Gemini 3.8 | Jev typesafe/jev-1.13 | Reflex-1 | |
|---|---|---|---|
| Instant decisions | ▲accurate, but 1–3 s | ●same band, 0.3 s | ●same band, 0.1–0.2 s |
| Decisions that require reasoning | ▲accurate with thinking, seconds | ×breaks down | ●~1 s (trained judgments) |
| Train on your own judgments | ×not possible | ×not possible | ●trainable |
| Inspect the reasoning | ●read thinking tokens | ×— | ●decode latent steps |
| Images | ●supported | ×not supported | ●~0.2 s per image |
| Runs locally, footprint | ×cloud only | ×cloud only | ●one GPU, 9.8 GB at 4-bit |
| Open weights | ×closed | ×closed | ●open |
| Per-call cost | ×pay per token | ×pay per call | ●none |
| Request format | ▲chat | ●Decisions API | ●same as Jev |
Evaluations: frontier LLMs vs Jev vs Reflex-1
Same question sets and options. Jev and Reflex-1: full sets (200 per classification task, 100 per reasoning task); frontier LLMs: reference values on subsets drawn from them (100 and 50). Numbers are accuracy %, latency in parentheses.
Instant decisions (text)
Give a text and options; pick one. Which desk should handle an inquiry / intent of a Japanese voice-assistant utterance / emotion (4 classes) and positive / negative / neutral (3 classes) of a social post / whether a comment insults someone (moderation).
| Task | frontier LLM | Jev | Reflex-1 | |
|---|---|---|---|---|
| GPT-5.6 Sol | Gemini 3.8 Flash | |||
| Intent (CLINC150) | 97 | 96 | 96.5 | 93.5 |
| Japanese utterance intent (MASSIVE) | 96 | 96 | 92.5 | 90.0 |
| Emotion (tweet_eval) | 84 | 79 | 81.0 | 82.0 |
| Sentiment (tweet_eval) | 78 | 77 | 74.5 | 74.5 |
| Moderation: does it insult someone? (civil_comments) | 73 | 76 | 73.5 | 80.0 |
| median latency (s) | 1.1–1.3 | 2.1–2.6 | 0.30 | 0.09–0.21 |
Instant decisions (images)
Give one photo; pick its category. Food photos (10 dishes) / e-commerce product photos (10 categories) / scanned document types (16 classes).
| Task | frontier LLM | Jev | Reflex-1 | |
|---|---|---|---|---|
| GPT-5.6 Sol | Gemini 3.8 Flash | |||
| Food photos (Food-101, 10 classes) | 98 | 98 | — | 98.5 |
| E-commerce products (10 classes) | 90 | 96 | — | 83.0 |
| Document type (RVL-CDIP, 16 classes) | — | — | — | 58.0 |
| median latency (s) | 1.8–2.0 | 2.9–3.1 | — | 0.21–0.26 |
Reasoning (trained logic tasks)
Logic puzzles that require reasoning (Reflex-1 trained per task). Web of Lies: 10 or 20 people each say whether the previous one lies; is the last one truthful? / Object tracking: after 7 people swap balls 14 times, which ball does a given person hold? / Boolean expressions: evaluate a not/and/or expression / Navigation: after following the steps, are you back at the start?
| Task | frontier LLM | Jev | Reflex-1 | |
|---|---|---|---|---|
| GPT-5.6 Sol reasoning off | Gemini 3.8 Flash reasoning on | |||
| Web of Lies, 10 hops | 46 | 100 (2.7 s) | 59.0 | 100 (0.8 s) |
| Web of Lies, 20 hops | 50 | 100 (3.1 s) | 58.0 | 72 (1.5 s) |
| Object tracking, 7 people / 14 swaps | 10 | 100 (3.5 s) | 26.2 | 70 (1.2 s) |
| Boolean expressions (BBH) | 92 | 100 (2.4 s) | 98.9 | 100 (0.4 s) |
| Navigation (BBH) | 78 | 100 (2.5 s) | 98.2 | 80.8 (0.5 s) |
| median latency (s) | 1.2–1.3 | 2.4–3.5 | 0.30 | 0.4–1.5 |
Try it
Requests use a Jev-compatible shape (state, questions, instructions, criteria; choice type). mode, skill, trace, image and the text type are Reflex-1 extensions.
POST /v1/decisions
{"state": "I want to know my credit rating",
"questions": {"desk": {"type": "choice",
"instructions": "Which desk should handle this inquiry?",
"criteria": {"credit_score": "", "transfer": "", "none": "none of these"}}}}
→ {"answers": {"desk": {"choice": "credit_score", "confidence": 0.97,
"probs": {"credit_score": 0.97, "transfer": 0.01, "none": 0.02}}}, "latency_ms": 90}
{"state": "Question: Ryan tells the truth. Michael says Ryan lies. ...",
"mode": "deep", "skill": "web_of_lies", "trace": true, "questions": {...}}
→ {"answers": {"a": {"choice": "No", "confidence": 1.0,
"trace": ["start -> Ryan tells the truth", "Michael says Ryan lies -> Michael lies", "..."]}},
"latency_ms": 820}
- GitHub — matu79go/reflex-1
- Hugging Face — matu79go/Reflex-1-4B (4-bit: Reflex-1-4B-bnb-4bit, Collection)
- Open in Colab
Cite: Suzuki, G. (2026). Reflex-1: frontier-LLM accuracy at Jev speed with a 4B open model. Zenodo. doi:10.5281/zenodo.23005201
How it was built
No large compute cluster was used; Reflex-1 was developed on a single local GPU machine (ASUS Ascent GX10). Training, evaluation and the videos were all done on this machine.
| Item | Details |
|---|---|
| Machine | 1 × ASUS Ascent GX10 (NVIDIA GB10 Grace Blackwell, 20-core CPU, 128 GB unified CPU/GPU memory) |
| OS / software | Ubuntu 24.04 (aarch64) · PyTorch 2.14 (CUDA 13) · Transformers 5.17 · PEFT 0.21 (LoRA) · bitsandbytes 0.50 · vLLM 0.19 |
| Base model | Google Gemma 4 E4B (Apache 2.0) |
| Training | LoRA plus small latent-reasoning modules. 1.4–2.2 hours per task on one GPU |
| Development time | 4 days, including the comparison measurements |
| External cost | About $2.5 in API calls for the comparisons (Jev, GPT-5.6, Gemini 3.8). No cloud compute for training |
Notes
- Decisions that require reasoning need per-judgment training. Reflex-1 is strong on the kinds and depths of judgment it was trained on, and does not transfer as-is to untrained kinds of reasoning. Classification works without any training.
- Reasoning mode currently runs on NVIDIA GPUs (CUDA). Support for common runtimes such as llama.cpp is coming shortly.
The ideas behind Reflex-1 and the earlier experiments are summarized on the NerveReflex project page:
- NerveReflex — Research on an AI Architecture That Cooperates Without Language
- Part 1. What Makes the New AI Model “Jev” Special? Reproduced with NerveReflex and an Open Model
Measured in September 2026 on one GX10 (NVIDIA GB10). Jev (typesafe/jev-1.13) and frontier models via OpenRouter; total API cost about $2.5. Reflex-1 is fine-tuned from Google Gemma 4 E4B (Apache 2.0); not affiliated with or endorsed by Google. Jev is a product of TypeSafe AI; results were measured by the author via OpenRouter and are not endorsed by TypeSafe AI. Data: CLINC150 (Larson et al. 2019, CC BY 3.0), MASSIVE (FitzGerald et al. 2022, CC BY 4.0), Civil Comments (Jigsaw, CC0), BIG-Bench Hard (Suzgun et al. 2022, MIT), tweet_eval (Barbieri et al. 2020), Food-101 (Bossard et al. 2014), Fashion Product Images, RVL-CDIP (Harley et al. 2015). Photos in the videos: CC0 / public domain from Wikimedia Commons. Concept: NerveReflex.
Join the conversation on LinkedIn — share your thoughts and comments.
Discuss on LinkedIn