Fall detection benchmark · senior safety vision

Mercury caught 93.8% of falls, compared with just 41.6% for Claude Opus 5.

Across 25,814 labeled images from 23 camera views, Mercury achieved 93.83% fall recall versus Claude’s 41.56%, with 27× faster inference times and approximately 98% lower estimated cost per inference.

25,814held out test images
23camera views
+52.27 ptsFall recall advantage
≈98% lowermodeled inference cost
The test set

Chosen to test new rooms, camera views, and hard posture boundaries.

Every image from each selected camera view was reserved for evaluation. This camera level separation is more demanding than randomly splitting neighboring images from the same view. The set covers 22 visual environments across varied hardware, placement, image quality, and lighting.

23

Camera views

Complete held out views reduce leakage from visually similar neighboring images.

22

Visual environments

Bedrooms, living areas, care rooms, transitions, distant views, and additional settings.

5

Human labeled states

Away, Fall, In bed, Sit, and Stand, evaluated on the same images for both systems.

Bedrooms and care roomsBeside bed falls, bedding occlusion, mobility equipment, and assisted movement.
Living and open roomsRecliners, floor postures, multiple people, and distant subjects.
Transitions and sparse viewsHallways, partial camera views, small subjects, and ambiguous silhouettes.
Additional camera domainsKitchen, dining, low light, and exterior night scenes.

The public test manifest records test composition, class totals, scenario coverage, and protocol without disclosing Mercury’s private training corpus or model design.

Benchmarking results

Mercury led on macro F1 and every class F1.

Precision measures how often a predicted class is correct. Recall measures how much of that true class is found. F1 summarizes both. Macro F1 gives each of the five classes equal weight.

ClassMercury PMercury RMercury F1Claude PClaude RClaude F1
Away83.55%77.83%80.59%65.19%93.65%76.87%
Fall93.24%93.83%93.53%90.33%41.56%56.93%
In bed82.34%63.36%71.62%74.34%31.07%43.83%
Sit72.68%79.75%76.05%62.65%72.73%67.32%
Stand64.83%85.06%73.58%66.76%79.20%72.45%
Macro average79.33%79.97%79.07%71.85%63.64%63.48%

Falls and safety scenarios Mercury caught and Claude missed

These verified benchmark disagreements show people reclined, prone, or down at floor level. Mercury matched the human label in each case; Claude Opus 5 did not.

Lightly blurred benchmark scene showing an older adult reclined and partially hidden on a sofa
Partially hidden person

Older adult reclined on couch

MercuryIn bed · matched
Claude Opus 5Away · missed
Lightly blurred benchmark Fall scene showing an older adult at floor level beside a sofa
Low posture ambiguity

Older adult down at floor level

MercuryFall · matched
Claude Opus 5Sit · missed
Lightly blurred benchmark scene showing an older adult reclined at floor level among bedding
Floor bedding ambiguity

Older adult on floor under bedding

MercuryFall · matched
Claude Opus 5In bed · missed
Lightly blurred benchmark scene showing an older adult in bed with a mobility aid nearby
Low light bedroom

Older adult hidden under bedding

MercuryIn bed · matched
Claude Opus 5Away · missed
Lightly blurred benchmark Fall scene showing one person down beside a chair while another person stands nearby
Multiple people

One person down while another stands

MercuryFall · matched
Claude Opus 5Stand · missed
Lightly blurred benchmark Fall scene showing a person down beside a bed
Distant bedside view

Person down beside a bed

MercuryFall · matched
Claude Opus 5Stand · missed

Images are lightly blurred for privacy. These are verified image level disagreements selected from the saved benchmark records and are not independent incidents.

Where Claude missed Falls

Claude recognized 41.56% of the fall labeled examples. Its dominant error was Sit, accounting for 83.62% of missed falls. Low or reclined postures were often interpreted as ordinary sitting.

Where Mercury still needs work

Mercury recognized 93.83% of the fall labeled examples. Remaining errors clustered in sparse, distant, and cross camera views. The benchmark supports an aggregate result, not superiority in every environment.

Selected fall scenarios

Strong aggregate recall with visible variation across real world scenes.

These held out examples show how posture, distance, occlusion, and scene complexity affect performance. Smaller scenario samples are descriptive and carry more uncertainty.

Assistance and multiple people

99.05%

Mercury Fall recall across 524 fall labeled images from a representative held out view.

Partial furniture occlusion

91.19%

Mercury Fall recall across 511 fall labeled images with partial bed or furniture obstruction.

Distant subject

89.08%

Mercury Fall recall across 513 fall labeled images where the person occupies less of the scene.

Distant sparse view

42.86%

A clear weakness measured on only 14 fall labeled images.

Bedside low posture

76.47%

A moderate result measured on 17 fall labeled images.

Floor sitting boundary

88.89%

A promising result measured on only 9 fall labeled images.

Claude’s saved benchmark output contains an overall confusion matrix but no camera keyed matrix. A reliable scenario by scenario Claude comparison cannot be reconstructed from the aggregate result and is therefore not shown.

Purpose built efficiency

A continuous monitoring system, not a vision language model request for every image.

Mercury Instant

$0.036per 1,000 observed production inferences

Measured endpoint spend divided by observed traffic. The deployed controller also uses motion and state logic to avoid unnecessary repeated inference.

Claude Opus 5 wrapper

$2.30per 1,000 modeled requests

Estimated from benchmark image and prompt token counts, a minimal response, and standard API pricing.

Measurement boundary: cost is a measured Mercury allocation compared with a modeled Claude request, not two invoices. Median latency was 136 ms for 100 warm Mercury model server observations and 3,681 ms for 10 successful Claude requests on the same 480 by 480 image. The Claude timing includes image retrieval and its API round trip, so the 27× comparison is operational rather than a laboratory measure of model execution alone.

Technical boundary

Useful evidence, stated at the right level.

On this image benchmark, Mercury delivered substantially higher Fall recall and macro F1 than the evaluated Claude Opus 5 configuration. The finding applies to these systems and this camera sample; it does not establish superiority over every Claude configuration or in every deployment.

Study limitations: Images from the same capture sequence are correlated, and larger camera samples contribute more to pooled results. These are image classification results, not measurements of independent fall incidents or deployment alert precision. The paper includes the full limitations.

Download the results data and camera inventory, or recompute the aggregate metrics.

Read the methods, class results, scenario findings, and cost assumptions.

Read the benchmarking paper

Real-time person activity and emergency monitoring with a Ring camera, powered by Mercury Alert AI’s custom model.