← All episodes

Single-camera consensus — is the lone box trustworthy?

step3a-2 emits a box whenever exactly one camera clears the confidence bar (edge ≤ 0.15). These are those frames. Every camera's own pose hypothesis is projected into all three views, so you can see whether the cameras that were filtered out would have agreed.

The yellow contour is the SAM3 v2 mask — the edge metric is measured against it. If the mask is missing or sitting on the wrong object, that camera abstains for a good reason; if the mask is clean but the box is off it, the pose is bad.

confident camera — this pose was emitted had a pose, filtered on edge SAM3 v2 mask contour

Would a looser threshold help?

Edge values of the cameras that were filtered out on these frames. If they clustered just above 0.15 a looser bar would recover real consensus — they don't.

Do the filtered cameras actually corroborate the emitted pose?

For every filtered camera on a single-cam frame: how far is its own pose from the emitted one, and would it have agreed under the pipeline's own rule (≤ 5 cm and ≤ 45°) had it not been gated on edge first?