A camera held out, five class image benchmark for continuous senior safety monitoring
The question is narrow: on the same camera images, does a specialist monitoring model provide a better precision and recall balance than zero shot inference with a vision language model? The five labels are Away, Fall, In bed, Sit, and Stand. Fall recall is the primary safety measure because it is the fraction of fall labeled images the system catches. Fall precision is reported alongside it to reveal whether higher sensitivity simply produces more false alarms. Macro F1 gives every class equal influence regardless of image count.
The test set reserves every image from 23 camera views representing 22 visual environment groups. Keeping complete views out of training reduces leakage from neighboring images and tests transfer to new scenes. Two cameras showing the same room were treated as one environment group.
| Class | Images | Share | Selection coverage |
|---|---|---|---|
| Away | 7,137 | 27.65% | Empty rooms, partial views, transitions |
| Fall | 3,438 | 13.32% | Bedside, floor level, distant, assisted, occluded |
| In bed | 5,101 | 19.76% | Bedding, recliners, low light, partial visibility |
| Sit | 7,085 | 27.45% | Chairs, recliners, floor sitting boundaries |
| Stand | 3,053 | 11.83% | Near camera, distant, hallway, multi person |
| Total | 25,814 | 100% | Bedrooms, care rooms, living areas, sparse rooms, kitchen, exterior |
The camera views were chosen to represent difficult posture boundaries and visual domains, not to balance the number of images per view. Selection criteria included furniture occlusion, mobility equipment, multiple people, distant subjects, hallway crops, floor sitting, recliners, low light, and varied camera hardware. The machine readable test manifest records composition, counts, and protocol. Mercury’s training corpus, architecture, and initialization source are intentionally not disclosed.
Internal model selection examined capacity, initialization families, learning rate, regularization, sampling, augmentation, and training duration across a larger experiment program. The selected configuration showed stable optimization in the benchmark run: training loss declined from 0.991 to 0.198 and training accuracy rose from 60.69% to 93.13% over 14 epochs. Test macro F1 reached 79.07% at epoch 8, then remained below that peak through epoch 14.
Mercury and Claude received the same images and the same five target labels. Claude used model claude-opus-5 for zero shot inference with a frozen five state prompt and JSON parser. It received no labeled examples, task specific fine tuning, or benchmark feedback. All 25,814 Claude outputs were recorded; 25,628 were already cached and 186 were completed for this run.
The comparator used no system prompt. The current prompt is shown below. The evaluated run used a legacy product specific phrase in the opening sentence; the label definitions and response schema are unchanged. The image follows the text, with max_tokens=1024:
Classify the accompanying image into exactly one person position.
Allowed labels:
- away: no person is visible
- fall: a person is down on the floor or ground
- in bed: a person is lying in a bed
- sit: a person is seated
- stand: a person is upright or standing
The image may be grayscale, low light, or partially occluded.
Return only valid JSON with this shape: {"label":"away|fall|in bed|sit|stand"}.
| Class | Mercury P | Mercury R | Mercury F1 | Claude P | Claude R | Claude F1 | F1 Δ |
|---|---|---|---|---|---|---|---|
| Away | 83.55% | 77.83% | 80.59% | 65.19% | 93.65% | 76.87% | +3.72 |
| Fall | 93.24% | 93.83% | 93.53% | 90.33% | 41.56% | 56.93% | +36.60 |
| In bed | 82.34% | 63.36% | 71.62% | 74.34% | 31.07% | 43.83% | +27.79 |
| Sit | 72.68% | 79.75% | 76.05% | 62.65% | 72.73% | 67.32% | +8.73 |
| Stand | 64.83% | 85.06% | 73.58% | 66.76% | 79.20% | 72.45% | +1.13 |
| Macro average | 79.33% | 79.97% | 79.07% | 71.85% | 63.64% | 63.48% | +15.59 |
Claude missed 2,009 fall labeled images. It labeled 1,680 of those images Sit, 195 Stand, 126 Away, and 8 In bed. Sit therefore represents 83.62% of Claude’s Fall misses. This is consistent with the central visual ambiguity in senior monitoring: a person low beside furniture can resemble an ordinary seated posture.
Mercury missed 212 fall labeled images: 88 were labeled Away, 58 Stand, 50 Sit, and 16 In bed. The smaller miss count does not eliminate variation across scenes. In particular, sparse and distant camera views remained difficult.



Figure 1. Verified image level disagreements selected from the saved benchmark records. Images are lightly blurred for privacy. They are not independent incidents and are not used as a qualitative score.
| Test scenario | Fall labeled images | Mercury Fall recall | Interpretation |
|---|---|---|---|
| Assistance and multiple people | 524 | 99.05% | Strong |
| Partial bed or furniture occlusion | 511 | 91.19% | Strong |
| Distant subject | 513 | 89.08% | Strong Fall result; weaker overall state accuracy |
| Distant sparse view | 14 | 42.86% | Clear weakness; small Fall sample |
| Clear bedside low posture | 17 | 76.47% | Moderate; small sample |
| Floor sitting boundary | 9 | 88.89% | Promising; very small sample |
Claude’s stored result contains an overall confusion matrix but no camera keyed matrix. The evidence cannot support a camera by camera Claude comparison without rescoring or recovering the individual prediction records. Scenario results above therefore describe Mercury only.
Mercury’s observed production traffic was served by two CPU endpoints handling approximately 223,510 model invocations per day. Allocating full endpoint spend across that traffic gives an average of about $0.0000362 per invocation, or $0.0362 per 1,000. Mercury’s batched benchmark evaluation reached approximately 263 images per second on one NVIDIA L4. This throughput value is not single image production latency.
A warm operational timing study found median latency of 136 ms for Mercury and 3,681 ms for Claude Opus 5, making Mercury 27 times faster by the median. Mercury used 100 live model server observations with a 146 ms 95th percentile. Claude used 10 successful requests on one 480 by 480 benchmark image with a 4,140 ms 95th percentile.
The Claude request was modeled using 324 visual tokens, 99 prompt tokens, a minimal output, and Anthropic’s standard Opus 5 pricing of $5 per million input tokens and $25 per million output tokens. The resulting estimate is approximately $0.0023 per image, or $2.30 per 1,000. Across the full test set, that is about $59.37 for Claude versus $0.93 at Mercury’s observed allocation, a ratio of approximately 63.5 to 1. Pricing source: Anthropic, Claude Opus.
The efficiency comparisons are directional. Mercury cost depends on endpoint utilization. Claude cost depends on tokenization, output length, caching, discounts, region, and future pricing. Mercury latency is model server time, while Claude latency includes S3 image retrieval and the Anthropic API round trip. Sample sizes also differ, so 27 times is an operational comparison rather than isolated model execution.
On this image benchmark, Mercury produced a materially better precision and recall balance than the evaluated Claude Opus 5 configuration. Fall recall was 93.83% for Mercury and 41.56% for Claude, with Mercury also maintaining higher Fall precision. This finding applies to the evaluated systems, labels, and camera sample; it does not establish superiority over every Claude configuration or in every deployment. Mercury was also 27 times faster by measured median operational latency and approximately 98% less expensive under the stated cost model.
The public camera inventory documents anonymized camera membership, class totals, and environment groups. The results data contains both overall confusion matrices, Mercury’s per-camera metrics, and the epoch history. The scoring script recomputes the aggregate metrics from those matrices. Full image membership and individual prediction records are not included. The underlying images require a privacy and licensing review before public distribution. Mercury’s training set, model architecture, initialization source, and detailed internal model selection records remain confidential.