Reflex-2: Video Judgment on a Phone, Improved from EmbeddingGemma 2 — Jev Speed, Frontier-LLM Accuracy or Better
Challenge
To judge events such as a fall on an elderly-care camera, an anomaly on a security camera, or the sound of breaking glass, frontier LLMs such as Gemini can read video, but each judgment takes several to more than ten seconds and the footage has to be sent to the cloud. The dedicated judgment API Jev is fast but cannot read video. Neither publishes its weights, so neither can be trained on footage or sounds from the actual site.
Solution
Starting from EmbeddingGemma 2 (740M parameters), the open retrieval model released by Google, Reflex-2 is trained on examples of what should be detected, turning it into a video and audio judgment model. Training takes a few minutes per judgment, and the model can also watch a camera feed continuously. Requests use the same shape as Jev.
Result
Fall judgment takes 0.28 s per clip, 25 times faster than Gemini 3.8 Flash (6.95 s), at 97.5% accuracy (Gemini: 85%). Audio judgment takes 0.05 s at 100%. Watching a camera feed, it noticed falls 1–2 seconds after they happened, while Gemini noticed after 8 seconds or missed them. The weights are 1.5 GB, small enough to run on a phone. Development took 2 days on a single local GPU machine, and the only external cost was about 1.3 dollars of API fees for the comparison.
![]()
EmbeddingGemma 2, the open retrieval model Google released on October 6, 2026, has been improved so that it can judge video. It judges on the spot, in an instant, without sending the footage to the cloud. It can also judge sounds.
Reflex-2 is a derived result of the research project NerveReflex, turning its ideas into an open model that can be used in practice, and the sister model of Reflex-1, which handles text and images. The Reflex-2 weights are available on Hugging Face and the code on GitHub.
The two videos below give Reflex-2 and Gemini 3.8 Flash the same footage and the same sounds, and compare how quickly and how correctly they answer.
What is Reflex-2?
Reflex-2 is an open model that improves Google’s EmbeddingGemma 2 (740M parameters, Apache 2.0) so that it can judge video and audio. It distinguishes what happens on site in an instant: a fall on an elderly-care camera, an anomaly on a security camera, a crying baby, breaking glass.
The base model, EmbeddingGemma 2, is a retrieval model that turns text, images, video, and audio into the same kind of “list of numbers” (a vector) (Google’s announcement, developer guide). It is designed to run on phones and laptops. However, it is originally a model for “finding similar things,” and as it is, it can hardly be used for judgment. Reflex-2 is what makes it usable for judgment.
What makes it stand out
The dedicated judgment API Jev is fast but cannot read video. Frontier LLMs such as Gemini can read video, but each judgment takes several to more than ten seconds and the footage has to be sent to the cloud. Reflex-2 judges video at Jev-like instant speed, with accuracy equal to or better than frontier LLMs, at a size that runs on a phone.
Build a judgment for your site from examples, in seconds
Show a few dozen examples of what you want to detect, such as falls, intrusions, or the sound of glass, and a classifier is ready on the spot. Jev and frontier LLMs do not publish their weights, so they cannot be trained on footage or sounds from your own site.
How it became a video judgment model
The original EmbeddingGemma 2 is a retrieval model for finding similar things, so on its own it cannot make a judgment such as “is this a fall?” Reflex-2 therefore trains it on examples of what should be detected.
Without training, accuracy does not come
Even with the same EmbeddingGemma 2, accuracy is completely different before and after training on examples. Training finishes in a few minutes, and the trained judgment can be used as is from then on.
The table below lists, for four tasks, the number of examples used for training, the time training took, and the accuracy before training (the original EmbeddingGemma 2) and after training (Reflex-2).
| Task | Training examples | Training time | Original EmbeddingGemma 2 (no training) | Reflex-2 (trained) |
|---|---|---|---|---|
| Falls | about 120 | about 1 min | accuracy 53.8% | accuracy 95.7% |
| Security-camera anomalies | 120 | about 3 min | accuracy 63.5% | accuracy 92.5% |
| 10 kinds of sounds | 320 | about 30 s | accuracy 35% | accuracy 96.3% |
| 8 vocal emotions | 1,200 | about 1 min | accuracy 14% | accuracy 52% |
How it can be applied
The following three examples, each modeled on a real-world setting, compare Reflex-2 and Gemini 3.8 Flash under the same conditions. All are measurements scored on footage and sounds not used for training.
Example 1 — Elderly care: alerting to a fall right away
Using indoor camera footage (GMDCSA24: 3 homes, 4 people, 79 falls and 81 normal activities), the models judged whether each clip was a fall. Across all 160 clips (scored on footage of people not seen in training), Reflex-2 is 95.7% accurate.
| Task (same 40 clips) | Gemini 3.8 Flash | Reflex-2 |
|---|---|---|
| Fall or normal activity | 85.0% | 97.5% |
| Time per judgment (median) | 6.95 s | 0.28 s |
The models also watched the footage continuously, as a camera would, to compare how many seconds after the fall each one noticed it.
| Watching a camera feed (same footage) | Actual fall | Gemini 3.8 Flash | Reflex-2 |
|---|---|---|---|
| Falls from a chair | 3.4 s | 8.3 s after the fall | 1.9 s after the fall |
| Falls while walking | 2.7 s | missed (“normal” all 4 times) | 2.1 s after the fall |
| Collapses forward | 2.6 s | missed (“normal” all 4 times) | 1.2 s after the fall |
| Lies down on a bed (normal) | — | no false alarm | no false alarm |

Example 2 — Security cameras: noticing intrusions and anomalies
The models judged whether security-camera footage (UCF-Crime) contains an anomaly such as an intrusion or an assault, or is ordinary everyday footage. Across all 110 evaluation clips, Reflex-2 is 92.5% accurate.
| Task (same 40 clips) | Gemini 3.8 Flash | Reflex-2 |
|---|---|---|
| Is the security-camera footage anomalous or normal? | 95% | 95% |
| Time per judgment (median) | 16.1 s | 1.4 s |


Example 3 — Sound: telling apart crying, glass, and alarms
The models judged 10 kinds of sounds that matter for care and security (baby crying, glass breaking, siren, dog, door knock, footsteps, coughing, snoring, alarm clock, crackling fire). Across all 400 clips (scored on sounds not used for training), Reflex-2 is 96.3% accurate.
| Task (same 40 clips) | Gemini 3.8 Flash | Reflex-2 |
|---|---|---|
| Telling apart 10 kinds of sounds | 85% | 100% |
| Time per judgment (median) | 3.66 s | 0.05 s |

Where it can go from here
Video and sound can be judged in an instant, on a phone or a device next to the camera, against criteria learned from examples taken on site. Items marked Measured were verified this time; the others can be built the same way.
- FallsMeasured
- Night-time wandering, long periods without movement
- Coughing and snoringMeasured, choking and groaning
- Getting up from bed
- A baby cryingMeasured
- Glass breaking while awayMeasured, intrusions
- Crackling fireMeasured, alarm clocksMeasured
- Pets barkingMeasured, unusual behavior
- Anomalies such as intrusions and assaultsMeasured
- Loitering after closing
- Movements that suggest shoplifting
- Responding to screams or breaking glass
- Workers falling or dropping from heights
- Entry into restricted areas
- Abnormal machine sounds and alarms
- Forklifts approaching people
- Accidents and fallen objects
- Falls from station platforms
- SirensMeasured, alarms
- Fights and abnormal crowd movement
- Abandoned luggage
- People who collapse from illness
- Abnormal livestock behavior, births
- Intrusion by pest animals
- Changes in health from vocalizations
- Automatic sorting of large volumes of video and audio
- Automatic detection of dangerous scenes
- Searching for events within video
The footage does not have to leave the premises
Footage from homes, hospitals, shops, and factories often should not leave the premises. Reflex-2 judges next to the camera or inside the phone, so there is no need to send footage to the cloud, and there are no communication costs or per-call fees.
How it differs from frontier LLMs, Jev, and Reflex-1
| frontier LLM Gemini 3.8 / GPT-5.6 | Jev typesafe/jev-1.13 | Reflex-1 text and images | Reflex-2 | |
|---|---|---|---|---|
| Video judgment | ▲accurate, but 7–16 s per clip | ×not supported | ×images only | ●0.28 s per clip |
| Audio judgment | ▲2–13 s per clip | ×not supported | ×not supported | ●0.05 s per clip |
| Time from a fall to noticing it | ▲8 s, or missed | ×— | ×— | ●1–2 s |
| How a judgment is defined | ●write a question | ●write a question | ●write a question, or train | ▲train on a few dozen examples |
| Training on your own site’s judgments | ×not possible | ×not possible | ●can be fine-tuned | ●seconds from examples |
| Runs locally, footprint | ×cloud only | ×cloud only | ●one GPU, 9.8 GB | ●runs on a phone, 1.5 GB |
| Open weights | ×closed | ×closed | ●open | ●open |
| Per-call fees | ×pay per use | ×pay per use | ●none | ●none |
| Request format | ▲chat | ●Decisions API | ●same shape as Jev | ●same shape as Jev |
Evaluation: comparison with Gemini 3.8 Flash
Both models receive the same input (the same frames, the same 16 kHz audio). Because Gemini’s thinking cannot be turned off, its thinking was set to the minimum and it was asked to answer with a single number or a single option. Reflex-2 is a classifier trained on examples; Gemini has no training. Reflex-2 ran on one GPU of a GX10 (NVIDIA GB10), Gemini via OpenRouter (including network time), and both were measured one clip at a time.
| Task (same 40 clips) | Gemini 3.8 Flash | Reflex-2 |
|---|---|---|
| Falls: accuracy / missed | 85.0% / 6 | 97.5% / 0 |
| Falls: time per judgment (median) | 6.95 s | 0.28 s |
| Security cameras: accuracy | 95% | 95% |
| Security cameras: time per judgment (median) | 16.1 s | 1.4 s |
| 10 kinds of sounds: accuracy | 85% | 100% |
| Sounds: time per judgment (median) | 3.66 s | 0.049 s |
| Cost per clip | $0.0003–0.023 | 0 |
Usage
There are two ways to try it. The easiest is Google Colab, which runs in the browser with nothing to install.
Option 1: try it on Google Colab (no installation)
- Open the Colab notebook.
- From the menu, choose “Runtime” → “Change runtime type” and select a GPU (T4 or better).
- Run the cells from the top. On the first run, the base model EmbeddingGemma 2 (1.5 GB) is downloaded automatically.
The notebook walks through four things in order:
- Judging one fall video and one normal-activity video from public data (GMDCSA24)
- Watching the fall video like a camera and plotting how the fall score changes
- Building your own classifier in seconds from 12 videos of three other people (6 falls, 6 normal)
- Trying a simple question without training, such as “is this a bedroom or a kitchen?”
Option 2: try it on your own computer
Python 3.10 or later is required. It runs faster with a GPU, but also works without one.
git clone https://github.com/matu79go/reflex-2
cd reflex-2
pip install -r requirements.txt
1. Judge a single video. The repository includes a trained fall classifier (heads/fall.json). Load it, pass a video file, and it answers whether the clip shows a fall.
from reflex2 import Reflex2
rf = Reflex2().load_heads("heads") # load the trained classifiers (the base model is downloaded on first run)
rf.decide("clip.mp4", "fall") # judge whether clip.mp4 shows a fall
# {'choice': 'fall', 'confidence': 0.97, 'probs': {'adl': 0.03, 'fall': 0.97}, 'latency_ms': 290}
In the result, choice is the judgment (fall = a fall, adl = normal activity), probs is the likelihood of each option, and latency_ms is the time the judgment took (in milliseconds). In this example, it judged a 97% likelihood of a fall in 0.29 seconds.
2. Watch like a camera. Given a longer video, it judges the last 4 seconds every 0.5 seconds and returns the fall score at each point in time. The line in the home-camera video above is drawn from this output.
rf.watch("camera.mp4", "fall")
# [(1.0, {'adl': 0.96, 'fall': 0.04}), ..., (4.5, {'adl': 0.03, 'fall': 0.97}), ...]
3. Build your own classifier. Pass examples of what you want to detect, as a list of files for each option. Video, images, and audio all work. The example below builds and saves a classifier that tells the sound of breaking glass (glass) apart from other sounds (other).
head = rf.train_head({"glass": ["g1.wav", "g2.wav", ...], "other": ["o1.wav", ...]}, name="glass")
head.save("heads/glass.json") # once saved, it can be loaded with load_heads next time
rf.decide("new_sound.wav", "glass") # judge a new sound
A few dozen examples are enough, and training finishes in seconds to a few minutes. For real deployments, using footage and recordings from the cameras and microphones at that site as examples is recommended.
Links
- GitHub — matu79go/reflex-2
- Hugging Face — matu79go/Reflex-2
- Open in Colab
- EmbeddingGemma 2 (Hugging Face)
- EmbeddingGemma 2 announcement (Google)
Development environment
No large compute cluster was used; development took place on a single local GPU machine (ASUS Ascent GX10). Training, evaluation, and video production were all done on this machine.
| Item | Details |
|---|---|
| Machine | 1 × ASUS Ascent GX10 (NVIDIA GB10 Grace Blackwell, 20 CPU cores, 128 GB unified memory shared by CPU and GPU) |
| OS / software | Ubuntu 24.04 (aarch64) · PyTorch 2.14 (CUDA 13) · Transformers 5.19 · sentence-transformers 6.1 · PyAV |
| Base model | Google EmbeddingGemma 2 (Apache 2.0) |
| Improvements | Training on examples (a few minutes per judgment), continuous camera watching, a request interface with the same shape as Jev |
| Development time | 2 days, including comparison measurements and video production |
| External cost | About $1.3 in total for comparison API fees (Gemini 3.8 Flash). No cloud was used for training |
Notes
- Each judgment is used after training on examples of what should be detected. The original EmbeddingGemma 2 on its own does not reach usable judgment accuracy.
- What it tells you is “when and what happened.” To identify which person on screen is involved, combine it with a person-detection model.
- Evaluation used public data (falls performed by actors, security-camera footage, sound effects). In real deployments, using footage and sounds from that site as examples is recommended.
The ideas behind Reflex-2 and its sister model Reflex-1 are covered on the following pages:
- NerveReflex — Research on an AI Architecture That Cooperates Without Language
- Reflex-1: Frontier-LLM Accuracy at Jev Speed with a 4B Open Model
- Part 1. What Makes the New AI Model “Jev” Special? Reproduced with NerveReflex and an Open Model
Measured in October 2026 on a single GX10 (NVIDIA GB10). Gemini 3.8 Flash via OpenRouter; total API cost about $1.3. Reflex-2 is an improvement of Google EmbeddingGemma 2 (Apache 2.0) and is not endorsed by or affiliated with Google. The Gemini results were measured by the author. Data: GMDCSA24 (Data in Brief 2024, MIT License, the 4 participants consented to public release), UCF-Crime (Sultani et al. 2018, research use), ESC-50 (Piczak 2015, CC BY-NC 3.0), RAVDESS (Livingstone and Russo 2018, CC BY-NC-SA 4.0). Concept: NerveReflex.
Join the conversation on LinkedIn — share your thoughts and comments.
Discuss on LinkedIn