Mercury Instant versus Claude Opus 5: A Fall Detection Benchmark

A camera held out, five class image benchmark for continuous senior safety monitoring

Brandon Peck
3 September 2026 · 25,814 test images · 23 camera views
Abstract. This benchmark compares Mercury Instant, a proprietary task specific vision classifier, with Claude Opus 5 used for zero shot five class image inference. Both systems scored the same 25,814 human labeled images from 23 held out camera views. Mercury achieved 79.07% macro F1 versus 63.48% for Claude. Fall recall was 93.83% versus 41.56%, a 52.27 percentage point difference, while Fall precision was 93.24% versus 90.33%. Mercury also delivered 27 times lower median operational latency and approximately 98.4% lower modeled inference cost.

1. Research question and metric priority

The question is narrow: on the same camera images, does a specialist monitoring model provide a better precision and recall balance than zero shot inference with a vision language model? The five labels are Away, Fall, In bed, Sit, and Stand. Fall recall is the primary safety measure because it is the fraction of fall labeled images the system catches. Fall precision is reported alongside it to reveal whether higher sensitivity simply produces more false alarms. Macro F1 gives every class equal influence regardless of image count.

2. Test set design

The test set reserves every image from 23 camera views representing 22 visual environment groups. Keeping complete views out of training reduces leakage from neighboring images and tests transfer to new scenes. Two cameras showing the same room were treated as one environment group.

ClassImagesShareSelection coverage
Away7,13727.65%Empty rooms, partial views, transitions
Fall3,43813.32%Bedside, floor level, distant, assisted, occluded
In bed5,10119.76%Bedding, recliners, low light, partial visibility
Sit7,08527.45%Chairs, recliners, floor sitting boundaries
Stand3,05311.83%Near camera, distant, hallway, multi person
Total25,814100%Bedrooms, care rooms, living areas, sparse rooms, kitchen, exterior

The camera views were chosen to represent difficult posture boundaries and visual domains, not to balance the number of images per view. Selection criteria included furniture occlusion, mobility equipment, multiple people, distant subjects, hallway crops, floor sitting, recliners, low light, and varied camera hardware. The machine readable test manifest records composition, counts, and protocol. Mercury’s training corpus, architecture, and initialization source are intentionally not disclosed.

3. Model development and evaluation protocol

Internal model selection examined capacity, initialization families, learning rate, regularization, sampling, augmentation, and training duration across a larger experiment program. The selected configuration showed stable optimization in the benchmark run: training loss declined from 0.991 to 0.198 and training accuracy rose from 60.69% to 93.13% over 14 epochs. Test macro F1 reached 79.07% at epoch 8, then remained below that peak through epoch 14.

Mercury and Claude received the same images and the same five target labels. Claude used model claude-opus-5 for zero shot inference with a frozen five state prompt and JSON parser. It received no labeled examples, task specific fine tuning, or benchmark feedback. All 25,814 Claude outputs were recorded; 25,628 were already cached and 186 were completed for this run.

3.1 Claude comparator prompt

The comparator used no system prompt. The current prompt is shown below. The evaluated run used a legacy product specific phrase in the opening sentence; the label definitions and response schema are unchanged. The image follows the text, with max_tokens=1024:

Classify the accompanying image into exactly one person position.
Allowed labels:
- away: no person is visible
- fall: a person is down on the floor or ground
- in bed: a person is lying in a bed
- sit: a person is seated
- stand: a person is upright or standing
The image may be grayscale, low light, or partially occluded.
Return only valid JSON with this shape: {"label":"away|fall|in bed|sit|stand"}.

4. Overall results

Primary result: Mercury correctly classified 3,226 of 3,438 fall labeled images. Claude correctly classified 1,429. Mercury therefore identified 1,797 more fall labeled images, with 2.91 percentage points higher Fall precision.
ClassMercury PMercury RMercury F1Claude PClaude RClaude F1F1 Δ
Away83.55%77.83%80.59%65.19%93.65%76.87%+3.72
Fall93.24%93.83%93.53%90.33%41.56%56.93%+36.60
In bed82.34%63.36%71.62%74.34%31.07%43.83%+27.79
Sit72.68%79.75%76.05%62.65%72.73%67.32%+8.73
Stand64.83%85.06%73.58%66.76%79.20%72.45%+1.13
Macro average79.33%79.97%79.07%71.85%63.64%63.48%+15.59

5. Failure patterns and camera scenarios

Claude Fall errors

Claude missed 2,009 fall labeled images. It labeled 1,680 of those images Sit, 195 Stand, 126 Away, and 8 In bed. Sit therefore represents 83.62% of Claude’s Fall misses. This is consistent with the central visual ambiguity in senior monitoring: a person low beside furniture can resemble an ordinary seated posture.

Mercury Fall errors

Mercury missed 212 fall labeled images: 88 were labeled Away, 58 Stand, 50 Sit, and 16 In bed. The smaller miss count does not eliminate variation across scenes. In particular, sparse and distant camera views remained difficult.

Lightly blurred benchmark Fall scene showing an older adult at floor level beside a sofa
Low posture ambiguityReference: FallMercury: Fall · matchedClaude Opus 5: Sit · missed
Lightly blurred benchmark Fall scene showing one person down beside a chair while another person stands nearby
Multiple peopleReference: FallMercury: Fall · matchedClaude Opus 5: Stand · missed
Lightly blurred benchmark scene showing an older adult on the floor under bedding
Older adult on floor under beddingReference: FallMercury: Fall · matchedClaude Opus 5: In bed · missed

Figure 1. Verified image level disagreements selected from the saved benchmark records. Images are lightly blurred for privacy. They are not independent incidents and are not used as a qualitative score.

Test scenarioFall labeled imagesMercury Fall recallInterpretation
Assistance and multiple people52499.05%Strong
Partial bed or furniture occlusion51191.19%Strong
Distant subject51389.08%Strong Fall result; weaker overall state accuracy
Distant sparse view1442.86%Clear weakness; small Fall sample
Clear bedside low posture1776.47%Moderate; small sample
Floor sitting boundary988.89%Promising; very small sample

Claude’s stored result contains an overall confusion matrix but no camera keyed matrix. The evidence cannot support a camera by camera Claude comparison without rescoring or recovering the individual prediction records. Scenario results above therefore describe Mercury only.

6. Compute and inference economics

Mercury’s observed production traffic was served by two CPU endpoints handling approximately 223,510 model invocations per day. Allocating full endpoint spend across that traffic gives an average of about $0.0000362 per invocation, or $0.0362 per 1,000. Mercury’s batched benchmark evaluation reached approximately 263 images per second on one NVIDIA L4. This throughput value is not single image production latency.

A warm operational timing study found median latency of 136 ms for Mercury and 3,681 ms for Claude Opus 5, making Mercury 27 times faster by the median. Mercury used 100 live model server observations with a 146 ms 95th percentile. Claude used 10 successful requests on one 480 by 480 benchmark image with a 4,140 ms 95th percentile.

The Claude request was modeled using 324 visual tokens, 99 prompt tokens, a minimal output, and Anthropic’s standard Opus 5 pricing of $5 per million input tokens and $25 per million output tokens. The resulting estimate is approximately $0.0023 per image, or $2.30 per 1,000. Across the full test set, that is about $59.37 for Claude versus $0.93 at Mercury’s observed allocation, a ratio of approximately 63.5 to 1. Pricing source: Anthropic, Claude Opus.

The efficiency comparisons are directional. Mercury cost depends on endpoint utilization. Claude cost depends on tokenization, output length, caching, discounts, region, and future pricing. Mercury latency is model server time, while Claude latency includes S3 image retrieval and the Anthropic API round trip. Sample sizes also differ, so 27 times is an operational comparison rather than isolated model execution.

7. Limitations and claim boundary

8. Conclusion and availability

On this image benchmark, Mercury produced a materially better precision and recall balance than the evaluated Claude Opus 5 configuration. Fall recall was 93.83% for Mercury and 41.56% for Claude, with Mercury also maintaining higher Fall precision. This finding applies to the evaluated systems, labels, and camera sample; it does not establish superiority over every Claude configuration or in every deployment. Mercury was also 27 times faster by measured median operational latency and approximately 98% less expensive under the stated cost model.

The public camera inventory documents anonymized camera membership, class totals, and environment groups. The results data contains both overall confusion matrices, Mercury’s per-camera metrics, and the epoch history. The scoring script recomputes the aggregate metrics from those matrices. Full image membership and individual prediction records are not included. The underlying images require a privacy and licensing review before public distribution. Mercury’s training set, model architecture, initialization source, and detailed internal model selection records remain confidential.