step3a-2 emits a box whenever exactly one camera clears the confidence bar
(edge ≤ 0.15). These are those frames. Every camera's own pose
hypothesis is projected into all three views, so you can see whether the
cameras that were filtered out would have agreed.
The yellow contour is the SAM3 v2 mask — the edge metric is measured against it. If the mask is missing or sitting on the wrong object, that camera abstains for a good reason; if the mask is clean but the box is off it, the pose is bad.
Edge values of the cameras that were filtered out on these frames. If they clustered just above 0.15 a looser bar would recover real consensus — they don't.
For every filtered camera on a single-cam frame: how far is its own pose from the emitted one, and would it have agreed under the pipeline's own rule (≤ 5 cm and ≤ 45°) had it not been gated on edge first?