Reflex-2: Video Judgment on a Phone, Improved from EmbeddingGemma 2 — Jev Speed, Frontier-LLM Accuracy or Better

View on GitHub

Challenge

To judge events such as a fall on an elderly-care camera, an anomaly on a security camera, or the sound of breaking glass, frontier LLMs such as Gemini can read video, but each judgment takes several to more than ten seconds and the footage has to be sent to the cloud. The dedicated judgment API Jev is fast but cannot read video. Neither publishes its weights, so neither can be trained on footage or sounds from the actual site.

Solution

Starting from EmbeddingGemma 2 (740M parameters), the open retrieval model released by Google, Reflex-2 is trained on examples of what should be detected, turning it into a video and audio judgment model. Training takes a few minutes per judgment, and the model can also watch a camera feed continuously. Requests use the same shape as Jev.

Result

Fall judgment takes 0.28 s per clip, 25 times faster than Gemini 3.8 Flash (6.95 s), at 97.5% accuracy (Gemini: 85%). Audio judgment takes 0.05 s at 100%. Watching a camera feed, it noticed falls 1–2 seconds after they happened, while Gemini noticed after 8 seconds or missed them. The weights are 1.5 GB, small enough to run on a phone. Development took 2 days on a single local GPU machine, and the only external cost was about 1.3 dollars of API fees for the comparison.

Reflex-2: video judgment that runs on a phone. A fall detected on a home camera with a fall score of 98%

EmbeddingGemma 2, the open retrieval model Google released on October 6, 2026, has been improved so that it can judge video. It judges on the spot, in an instant, without sending the footage to the cloud. It can also judge sounds.

Reflex-2 is a derived result of the research project NerveReflex, turning its ideas into an open model that can be used in practice, and the sister model of Reflex-1, which handles text and images. The Reflex-2 weights are available on Hugging Face and the code on GitHub.

The two videos below give Reflex-2 and Gemini 3.8 Flash the same footage and the same sounds, and compare how quickly and how correctly they answer.

When someone falls at home, how many seconds until it is noticed? — The clock starts at the moment of the fall. Reflex-2 catches it 1–2 seconds after the fall. Gemini takes 4–8 seconds per judgment, and either caught the fall after 8 seconds or missed it in 2 of 3 scenes.
Can it tell a baby’s cry or breaking glass right after hearing it? (with sound) — Reflex-2 answered all 4 correctly in 0.05–0.09 s. Gemini took 3–10 s and answered “footsteps” for the glass and “siren” for the alarm clock.

What is Reflex-2?

Reflex-2 is an open model that improves Google’s EmbeddingGemma 2 (740M parameters, Apache 2.0) so that it can judge video and audio. It distinguishes what happens on site in an instant: a fall on an elderly-care camera, an anomaly on a security camera, a crying baby, breaking glass.

The base model, EmbeddingGemma 2, is a retrieval model that turns text, images, video, and audio into the same kind of “list of numbers” (a vector) (Google’s announcement, developer guide). It is designed to run on phones and laptops. However, it is originally a model for “finding similar things,” and as it is, it can hardly be used for judgment. Reflex-2 is what makes it usable for judgment.

What makes it stand out

The dedicated judgment API Jev is fast but cannot read video. Frontier LLMs such as Gemini can read video, but each judgment takes several to more than ten seconds and the footage has to be sent to the cloud. Reflex-2 judges video at Jev-like instant speed, with accuracy equal to or better than frontier LLMs, at a size that runs on a phone.

~2 sfrom a fall to the alert. Gemini 3.8 Flash: 8 s, or missed
25×faster video judgment than Gemini (0.28 s vs. 6.95 s per clip). Audio: 75× (0.05 s vs. 3.66 s)
97.5% · 100%accuracy on falls and sounds. Gemini: 85% and 85% on the same questions
0.7B · 1.5 GBopen weights small enough for a phone. Footage never leaves for the cloud. No per-call fee

Build a judgment for your site from examples, in seconds

Show a few dozen examples of what you want to detect, such as falls, intrusions, or the sound of glass, and a classifier is ready on the spot. Jev and frontier LLMs do not publish their weights, so they cannot be trained on footage or sounds from your own site.

How it became a video judgment model

The original EmbeddingGemma 2 is a retrieval model for finding similar things, so on its own it cannot make a judgment such as “is this a fall?” Reflex-2 therefore trains it on examples of what should be detected.

Without training, accuracy does not come

Even with the same EmbeddingGemma 2, accuracy is completely different before and after training on examples. Training finishes in a few minutes, and the trained judgment can be used as is from then on.

The table below lists, for four tasks, the number of examples used for training, the time training took, and the accuracy before training (the original EmbeddingGemma 2) and after training (Reflex-2).

TaskTraining examplesTraining timeOriginal EmbeddingGemma 2
(no training)
Reflex-2
(trained)
Fallsabout 120about 1 minaccuracy 53.8%accuracy 95.7%
Security-camera anomalies120about 3 minaccuracy 63.5%accuracy 92.5%
10 kinds of sounds320about 30 saccuracy 35%accuracy 96.3%
8 vocal emotions1,200about 1 minaccuracy 14%accuracy 52%
Training times are approximate, on one GPU (GX10). Scoring uses only videos and sounds that were not used for training (for falls, footage of people who did not appear in training). Chance accuracy is 50% for falls and security cameras, 10% for sounds, and 12.5% for vocal emotions.

How it can be applied

The following three examples, each modeled on a real-world setting, compare Reflex-2 and Gemini 3.8 Flash under the same conditions. All are measurements scored on footage and sounds not used for training.

Example 1 — Elderly care: alerting to a fall right away

Using indoor camera footage (GMDCSA24: 3 homes, 4 people, 79 falls and 81 normal activities), the models judged whether each clip was a fall. Across all 160 clips (scored on footage of people not seen in training), Reflex-2 is 95.7% accurate.

Task (same 40 clips)Gemini 3.8 FlashReflex-2
Fall or normal activity85.0%97.5%
Time per judgment (median)6.95 s0.28 s

The models also watched the footage continuously, as a camera would, to compare how many seconds after the fall each one noticed it.

Watching a camera feed (same footage)Actual fallGemini 3.8 FlashReflex-2
Falls from a chair3.4 s8.3 s after the fall1.9 s after the fall
Falls while walking2.7 smissed (“normal” all 4 times)2.1 s after the fall
Collapses forward2.6 smissed (“normal” all 4 times)1.2 s after the fall
Lies down on a bed (normal)—no false alarmno false alarm
Falls: accuracy × speed (per judgment)
Falls: accuracy × speed (per judgment) — the same 40 clips (20 falls / 20 normal). Gemini missed 6 falls.
Watching a home camera feed — The line at the bottom is the fall score, and the light red band is the actual fall. Lying down on a bed does not trigger a false alarm (GMDCSA24, MIT License; the participants consented to public release).

Example 2 — Security cameras: noticing intrusions and anomalies

The models judged whether security-camera footage (UCF-Crime) contains an anomaly such as an intrusion or an assault, or is ordinary everyday footage. Across all 110 evaluation clips, Reflex-2 is 92.5% accurate.

Task (same 40 clips)Gemini 3.8 FlashReflex-2
Is the security-camera footage anomalous or normal?95%95%
Time per judgment (median)16.1 s1.4 s
Security cameras: accuracy × speed
Security cameras: accuracy × speed — the same 40 clips (20 anomalous / 20 normal). Each clip is 32 frames from footage with a median length of 75 seconds.
Judgment examples: an ordinary building entrance and a night-time intrusion at a front door
Judgment examples — left: normal (a building entrance) / right: a night-time intrusion at a front door (faces pixelated).

Example 3 — Sound: telling apart crying, glass, and alarms

The models judged 10 kinds of sounds that matter for care and security (baby crying, glass breaking, siren, dog, door knock, footsteps, coughing, snoring, alarm clock, crackling fire). Across all 400 clips (scored on sounds not used for training), Reflex-2 is 96.3% accurate.

Task (same 40 clips)Gemini 3.8 FlashReflex-2
Telling apart 10 kinds of sounds85%100%
Time per judgment (median)3.66 s0.05 s
10 kinds of sounds: accuracy × speed
10 kinds of sounds: accuracy × speed — the same 40 clips (10 kinds × 4).
Telling sounds apart (with sound) — a baby crying, glass breaking, an alarm clock, and coughing (ESC-50). Reflex-2 got all 4 right; Gemini got the glass and the alarm clock wrong.

Where it can go from here

Video and sound can be judged in an instant, on a phone or a device next to the camera, against criteria learned from examples taken on site. Items marked Measured were verified this time; the others can be built the same way.

Care and medicine
  • FallsMeasured
  • Night-time wandering, long periods without movement
  • Coughing and snoringMeasured, choking and groaning
  • Getting up from bed
Home
  • A baby cryingMeasured
  • Glass breaking while awayMeasured, intrusions
  • Crackling fireMeasured, alarm clocksMeasured
  • Pets barkingMeasured, unusual behavior
Security and retail
  • Anomalies such as intrusions and assaultsMeasured
  • Loitering after closing
  • Movements that suggest shoplifting
  • Responding to screams or breaking glass
Factories, construction, logistics
  • Workers falling or dropping from heights
  • Entry into restricted areas
  • Abnormal machine sounds and alarms
  • Forklifts approaching people
Transport and infrastructure
  • Accidents and fallen objects
  • Falls from station platforms
  • SirensMeasured, alarms
Public spaces and events
  • Fights and abnormal crowd movement
  • Abandoned luggage
  • People who collapse from illness
Agriculture and animals
  • Abnormal livestock behavior, births
  • Intrusion by pest animals
  • Changes in health from vocalizations
Broadcasting and content
  • Automatic sorting of large volumes of video and audio
  • Automatic detection of dangerous scenes
  • Searching for events within video

The footage does not have to leave the premises

Footage from homes, hospitals, shops, and factories often should not leave the premises. Reflex-2 judges next to the camera or inside the phone, so there is no need to send footage to the cloud, and there are no communication costs or per-call fees.

How it differs from frontier LLMs, Jev, and Reflex-1

frontier LLM
Gemini 3.8 / GPT-5.6
Jev
typesafe/jev-1.13
Reflex-1
text and images
Reflex-2
Video judgment▲accurate, but 7–16 s per clip×not supported×images only●0.28 s per clip
Audio judgment▲2–13 s per clip×not supported×not supported●0.05 s per clip
Time from a fall to noticing it▲8 s, or missed×—×—●1–2 s
How a judgment is defined●write a question●write a question●write a question, or train▲train on a few dozen examples
Training on your own site’s judgments×not possible×not possible●can be fine-tuned●seconds from examples
Runs locally, footprint×cloud only×cloud only●one GPU, 9.8 GB●runs on a phone, 1.5 GB
Open weights×closed×closed●open●open
Per-call fees×pay per use×pay per use●none●none
Request format▲chat●Decisions API●same shape as Jev●same shape as Jev
● capable / strong ▲ with conditions × not possible. Use Reflex-1 when you want to write the question as text, and Reflex-2 for judging video and audio.

Evaluation: comparison with Gemini 3.8 Flash

Both models receive the same input (the same frames, the same 16 kHz audio). Because Gemini’s thinking cannot be turned off, its thinking was set to the minimum and it was asked to answer with a single number or a single option. Reflex-2 is a classifier trained on examples; Gemini has no training. Reflex-2 ran on one GPU of a GX10 (NVIDIA GB10), Gemini via OpenRouter (including network time), and both were measured one clip at a time.

Task (same 40 clips)Gemini 3.8 FlashReflex-2
Falls: accuracy / missed85.0% / 697.5% / 0
Falls: time per judgment (median)6.95 s0.28 s
Security cameras: accuracy95%95%
Security cameras: time per judgment (median)16.1 s1.4 s
10 kinds of sounds: accuracy85%100%
Sounds: time per judgment (median)3.66 s0.049 s
Cost per clip$0.0003–0.0230

Usage

There are two ways to try it. The easiest is Google Colab, which runs in the browser with nothing to install.

Option 1: try it on Google Colab (no installation)

  1. Open the Colab notebook.
  2. From the menu, choose “Runtime” → “Change runtime type” and select a GPU (T4 or better).
  3. Run the cells from the top. On the first run, the base model EmbeddingGemma 2 (1.5 GB) is downloaded automatically.

The notebook walks through four things in order:

  • Judging one fall video and one normal-activity video from public data (GMDCSA24)
  • Watching the fall video like a camera and plotting how the fall score changes
  • Building your own classifier in seconds from 12 videos of three other people (6 falls, 6 normal)
  • Trying a simple question without training, such as “is this a bedroom or a kitchen?”

Option 2: try it on your own computer

Python 3.10 or later is required. It runs faster with a GPU, but also works without one.

git clone https://github.com/matu79go/reflex-2
cd reflex-2
pip install -r requirements.txt

1. Judge a single video. The repository includes a trained fall classifier (heads/fall.json). Load it, pass a video file, and it answers whether the clip shows a fall.

from reflex2 import Reflex2

rf = Reflex2().load_heads("heads")      # load the trained classifiers (the base model is downloaded on first run)
rf.decide("clip.mp4", "fall")           # judge whether clip.mp4 shows a fall
# {'choice': 'fall', 'confidence': 0.97, 'probs': {'adl': 0.03, 'fall': 0.97}, 'latency_ms': 290}

In the result, choice is the judgment (fall = a fall, adl = normal activity), probs is the likelihood of each option, and latency_ms is the time the judgment took (in milliseconds). In this example, it judged a 97% likelihood of a fall in 0.29 seconds.

2. Watch like a camera. Given a longer video, it judges the last 4 seconds every 0.5 seconds and returns the fall score at each point in time. The line in the home-camera video above is drawn from this output.

rf.watch("camera.mp4", "fall")
# [(1.0, {'adl': 0.96, 'fall': 0.04}), ..., (4.5, {'adl': 0.03, 'fall': 0.97}), ...]

3. Build your own classifier. Pass examples of what you want to detect, as a list of files for each option. Video, images, and audio all work. The example below builds and saves a classifier that tells the sound of breaking glass (glass) apart from other sounds (other).

head = rf.train_head({"glass": ["g1.wav", "g2.wav", ...], "other": ["o1.wav", ...]}, name="glass")
head.save("heads/glass.json")           # once saved, it can be loaded with load_heads next time
rf.decide("new_sound.wav", "glass")     # judge a new sound

A few dozen examples are enough, and training finishes in seconds to a few minutes. For real deployments, using footage and recordings from the cameras and microphones at that site as examples is recommended.

Links

Development environment

No large compute cluster was used; development took place on a single local GPU machine (ASUS Ascent GX10). Training, evaluation, and video production were all done on this machine.

ItemDetails
Machine1 × ASUS Ascent GX10 (NVIDIA GB10 Grace Blackwell, 20 CPU cores, 128 GB unified memory shared by CPU and GPU)
OS / softwareUbuntu 24.04 (aarch64) · PyTorch 2.14 (CUDA 13) · Transformers 5.19 · sentence-transformers 6.1 · PyAV
Base modelGoogle EmbeddingGemma 2 (Apache 2.0)
ImprovementsTraining on examples (a few minutes per judgment), continuous camera watching, a request interface with the same shape as Jev
Development time2 days, including comparison measurements and video production
External costAbout $1.3 in total for comparison API fees (Gemini 3.8 Flash). No cloud was used for training

Notes

  • Each judgment is used after training on examples of what should be detected. The original EmbeddingGemma 2 on its own does not reach usable judgment accuracy.
  • What it tells you is “when and what happened.” To identify which person on screen is involved, combine it with a person-detection model.
  • Evaluation used public data (falls performed by actors, security-camera footage, sound effects). In real deployments, using footage and sounds from that site as examples is recommended.

The ideas behind Reflex-2 and its sister model Reflex-1 are covered on the following pages:

Measured in October 2026 on a single GX10 (NVIDIA GB10). Gemini 3.8 Flash via OpenRouter; total API cost about $1.3. Reflex-2 is an improvement of Google EmbeddingGemma 2 (Apache 2.0) and is not endorsed by or affiliated with Google. The Gemini results were measured by the author. Data: GMDCSA24 (Data in Brief 2024, MIT License, the 4 participants consented to public release), UCF-Crime (Sultani et al. 2018, research use), ESC-50 (Piczak 2015, CC BY-NC 3.0), RAVDESS (Livingstone and Russo 2018, CC BY-NC-SA 4.0). Concept: NerveReflex.

Share this article

Join the conversation on LinkedIn — share your thoughts and comments.

Discuss on LinkedIn