Mercury caught 93.8% of falls, compared with just 41.6% for Claude Opus 5.
Across 25,814 labeled images from 23 camera views, Mercury achieved 93.83% fall recall versus Claude’s 41.56%, with 27× faster inference times and approximately 98% lower estimated cost per inference.
Chosen to test new rooms, camera views, and hard posture boundaries.
Every image from each selected camera view was reserved for evaluation. This camera level separation is more demanding than randomly splitting neighboring images from the same view. The set covers 22 visual environments across varied hardware, placement, image quality, and lighting.
Camera views
Complete held out views reduce leakage from visually similar neighboring images.
Visual environments
Bedrooms, living areas, care rooms, transitions, distant views, and additional settings.
Human labeled states
Away, Fall, In bed, Sit, and Stand, evaluated on the same images for both systems.
The public test manifest records test composition, class totals, scenario coverage, and protocol without disclosing Mercury’s private training corpus or model design.
Mercury led on macro F1 and every class F1.
Precision measures how often a predicted class is correct. Recall measures how much of that true class is found. F1 summarizes both. Macro F1 gives each of the five classes equal weight.
| Class | Mercury P | Mercury R | Mercury F1 | Claude P | Claude R | Claude F1 |
|---|---|---|---|---|---|---|
| Away | 83.55% | 77.83% | 80.59% | 65.19% | 93.65% | 76.87% |
| Fall | 93.24% | 93.83% | 93.53% | 90.33% | 41.56% | 56.93% |
| In bed | 82.34% | 63.36% | 71.62% | 74.34% | 31.07% | 43.83% |
| Sit | 72.68% | 79.75% | 76.05% | 62.65% | 72.73% | 67.32% |
| Stand | 64.83% | 85.06% | 73.58% | 66.76% | 79.20% | 72.45% |
| Macro average | 79.33% | 79.97% | 79.07% | 71.85% | 63.64% | 63.48% |
Falls and safety scenarios Mercury caught and Claude missed
These verified benchmark disagreements show people reclined, prone, or down at floor level. Mercury matched the human label in each case; Claude Opus 5 did not.
Older adult down at floor level
Older adult on floor under bedding
Older adult hidden under bedding
One person down while another stands
Person down beside a bed
Images are lightly blurred for privacy. These are verified image level disagreements selected from the saved benchmark records and are not independent incidents.
Where Claude missed Falls
Claude recognized 41.56% of the fall labeled examples. Its dominant error was Sit, accounting for 83.62% of missed falls. Low or reclined postures were often interpreted as ordinary sitting.
Where Mercury still needs work
Mercury recognized 93.83% of the fall labeled examples. Remaining errors clustered in sparse, distant, and cross camera views. The benchmark supports an aggregate result, not superiority in every environment.
Strong aggregate recall with visible variation across real world scenes.
These held out examples show how posture, distance, occlusion, and scene complexity affect performance. Smaller scenario samples are descriptive and carry more uncertainty.
Assistance and multiple people
99.05%Mercury Fall recall across 524 fall labeled images from a representative held out view.
Partial furniture occlusion
91.19%Mercury Fall recall across 511 fall labeled images with partial bed or furniture obstruction.
Distant subject
89.08%Mercury Fall recall across 513 fall labeled images where the person occupies less of the scene.
Distant sparse view
42.86%A clear weakness measured on only 14 fall labeled images.
Bedside low posture
76.47%A moderate result measured on 17 fall labeled images.
Floor sitting boundary
88.89%A promising result measured on only 9 fall labeled images.
Claude’s saved benchmark output contains an overall confusion matrix but no camera keyed matrix. A reliable scenario by scenario Claude comparison cannot be reconstructed from the aggregate result and is therefore not shown.
A continuous monitoring system, not a vision language model request for every image.
Mercury Instant
Measured endpoint spend divided by observed traffic. The deployed controller also uses motion and state logic to avoid unnecessary repeated inference.
Claude Opus 5 wrapper
Estimated from benchmark image and prompt token counts, a minimal response, and standard API pricing.
Measurement boundary: cost is a measured Mercury allocation compared with a modeled Claude request, not two invoices. Median latency was 136 ms for 100 warm Mercury model server observations and 3,681 ms for 10 successful Claude requests on the same 480 by 480 image. The Claude timing includes image retrieval and its API round trip, so the 27× comparison is operational rather than a laboratory measure of model execution alone.
Useful evidence, stated at the right level.
On this image benchmark, Mercury delivered substantially higher Fall recall and macro F1 than the evaluated Claude Opus 5 configuration. The finding applies to these systems and this camera sample; it does not establish superiority over every Claude configuration or in every deployment.
Study limitations: Images from the same capture sequence are correlated, and larger camera samples contribute more to pooled results. These are image classification results, not measurements of independent fall incidents or deployment alert precision. The paper includes the full limitations.
Download the results data and camera inventory, or recompute the aggregate metrics.