A mismatch on its own line needs a third eye
Worth reading first: A wrong match is not a small error · A point is a line over there.
A wrong match is not a small error moved one correspondence of forty-four by thirty pixels and watched the damage spread: the refitted epipolar geometry was wrong for every other point, the typical one by half a pixel, from a fit that had been exact to a part in ten trillion. A mismatch, that essay concluded, is not a small error in one place but a contamination of the whole estimate.
Near its end it named the exception, in one paragraph and without a figure. A wrong match that happens to lie on the correct epipolar line agrees with the geometry perfectly. It contaminates nothing, and no amount of resampling finds it, because there is nothing to find. Two rays that do not meet had said the same thing from the triangulation side: slide a mark along its epipolar line and the gap between the two rays stays at zero.
Both sentences are true, and neither measured the consequence. This essay does, and then asks the question they leave open: if two pictures cannot see such a mismatch, what does a third see, and from where?
Two kinds of wrong match
A mismatch is a pair of marks that are not images of one world point. A wrong match is not a small error moved its bad mark diagonally, partly along its epipolar line and mostly across it, and it is the across component that did the damage: a mark off its line is inconsistent with every epipolar geometry that fits the other forty-three, so a least-squares fit has to bend the geometry to accommodate it.
Move the mark purely along its line and none of that happens. The figure below takes the same correspondence, number 9 of the standing courtyard, and slides its second mark thirty pixels along the epipolar line its first mark fixes, in the direction that increases the separation between the two marks.
Then it runs every test a pair of photographs permits. The moved mark’s distance from its epipolar line is 2.2 × 10⁻¹⁴ px. The epipolar geometry refitted from all forty-four correspondences, the bad one included, is exact for all of them to 3.8 × 10⁻¹³ px. The two rays through the correspondence meet, to 1.5 × 10⁻¹⁵ m. And the point they meet at reprojects into both pictures to 1.1 × 10⁻¹³ px. Four tests, four results at the floor of double-precision arithmetic, and a bar for one pixel beside them that is sixteen decades longer.
The point is 0.50 m from where it belongs: 6.20 m from the first eye in the world, 5.70 m in the reconstruction.
Why nothing in two pictures can object
This is not an estimator failing to notice something. It is two photographs containing an exact explanation for a wrong answer, and it is worth seeing why that explanation must exist.
The first mark fixes a ray from the first eye: every world point that could have produced it lies on that ray. A point is a line over there is the statement that those points project into the second picture along a single line — the epipolar line — because the ray and both camera centres lie in one plane, and that plane cuts the second picture in a line.
So every position on the epipolar line is the image, in the second picture, of some point on the first ray. The slid mark is at one such position. It is exactly the mark the second eye would have made if the world point had been 5.70 m from the first eye instead of 6.20 m. The pair of marks is not an inconsistent measurement of the true point; it is a perfectly consistent measurement of a different point on the same line of sight.
Two photographs cannot tell those situations apart, because they are the same two photographs. A scene with a point at 5.70 m, correctly matched, and a scene with a point at 6.20 m, mismatched along its line, produce identical marks. Any test computed from the marks gives identical answers for both, and for the first scene every one of those answers is correct.
Silent, and local
The two kinds of mismatch are not merely different in size. They behave in opposite ways, and each is the other’s mirror image.
A mark moved across its line is loud and global. It cannot be explained by any point, so it forces the fit to compromise, and the compromise shows up at every correspondence as a raised residual. The loud case’s figure is that shape: nothing is exactly right any more.
A mark moved along its line is silent and local. It is explained exactly by a point, so the fit has nothing to compromise. The refitted geometry stays exact for the other forty-three — its worst residual is still 3.8 × 10⁻¹³ px — and so does every other reconstructed point. The only wrong thing in the reconstruction is the one point, and it is wrong by half a metre.
That second property is what makes it dangerous rather than merely hard to see. A robust estimator, of the kind that essay described and that is not the subject here, works by finding the geometry most correspondences agree with and rejecting those that disagree. The slid match agrees. It is kept, its point is reconstructed confidently, and the diagnostics that would have flagged a loud mismatch report a reconstruction that is exact everywhere.
How much depth a small slide is worth
Thirty pixels was chosen to be visible. The mismatches that repeated structure produces are usually smaller — the next post of a railing, the next window of a row — so it is worth running the same tests on a slide one sixth as large.
The tests are exactly as blind at five pixels as at thirty: 2.3 × 10⁻¹⁴ px from the line, rays meeting to 5.0 × 10⁻¹⁶ m, reprojection at 4.0 × 10⁻¹⁴ px. The point is 9 cm too near.
The depth error is the reciprocal law of depth is a reciprocal with the reading error replaced by the slide. A slide along the line is a change in disparity, and the courtyard pair’s disparity at six metres makes five pixels worth about nine centimetres and thirty pixels worth half a metre. The relation is nearly linear over that range and would steepen further out, as every depth measurement from a pair does. The difference from the reading-error case is that nothing about the slide is random: it does not average away over repeated photographs, and it is not reported by any interval, because no test has registered that anything is uncertain.
A third picture
A third camera looks at the courtyard from somewhere else. The slid correspondence has produced a point, and that point projects into the third picture somewhere. The true point projects somewhere too. If those two places differ, the third picture contradicts the pair.
The figure places a third eye at the first eye’s height and distance from the true point, swung 20° round it, and draws the courtyard as that eye sees it. The true point is where the filled mark is. The pair’s reconstructed point — half a metre nearer the first eye, along the first eye’s ray — lands 21.4 px away. That is not a floor-level residual; it is a displacement of about three percent of the picture’s width, and any comparison of the third picture with the pair’s reconstruction finds it.
The figure finds it a second way as well, and the two are worth distinguishing. Transfer through the point reconstructs the pair’s point and projects it into the third picture. Transfer by lines never builds a point: the first mark draws an epipolar line in the third picture and so does the second mark, and where those two lines cross is where the pair says the third picture should see its match. At 20° the lines cross at 14.3°, and the crossing lands on the pair’s projected point to within the arithmetic floor. Two routes, one of which never touches three dimensions, agree about what the third picture should contain, and both disagree with what it does contain.
Where the third eye has to stand
A third view does not detect the mismatch from everywhere. The figure below swings the third eye round the true point from the first eye’s own direction outward, and plots the exposure against how far the third eye is from the first eye’s line of sight.
The first point is the control, and it is the essay’s sharpest single statement. A third eye placed on the first eye’s line of sight, a metre and a half behind the first eye, sees the true point and the pair’s point one exactly behind the other, and reports an exposure of 8 × 10⁻¹⁴ px. The reason is the one above, read backwards: both points lie on the first eye’s ray, and any eye on that ray sees them in the same direction. A third camera there is not a third view of this question at all.
Off that line the exposure grows with the distance: 10.9 px with the third eye 1.04 m from the ray, 21.4 px at 2.05 m, 31.1 px at 3.00 m. Near the ray it grows at about ten and a half pixels per metre of offset, and further out it levels, reaching 59.2 px when the third eye has swung 80° and stands six metres from the ray — the exposure is an angle, and swinging further buys less of it.
So the geometry answers the question of where to put a third camera precisely. What matters is not how far the third eye is from the first two, or from the scene, but how far it is from the one line along which the mismatch moved its point. A third camera mounted right behind the first, or looking over its shoulder, contributes almost nothing to catching this kind of error, however good its picture is.
The third eye is a baseline of its own
The shape of that curve has a reason, and the reason turns the third view into a measurement of the same kind as every other depth a pair of eyes reports.
The true point and the pair’s wrong point both lie on the first eye’s ray, 6.20 m and 5.70 m from the first eye. A third eye standing a distance d off that ray sees the two points in slightly different directions, and the angle between those directions is a parallax: the half-metre displacement Δ between the points, seen across a baseline d. For a third eye near the ray, and much further from the points than d, that angle is about Δd divided by the product of the two points’ distances, and a picture with focal length f turns it into
With the third camera’s 740 px of focal length, Δ of half a metre and distances of 6.20 m and 5.70 m, that is 10.5 px for every metre the third eye stands off the ray — the slope the sweep measures near its start. For a five-pixel slide, whose point is nine centimetres out, the same expression predicts 1.7 px per metre, and the sweep for that slide measures 1.8 px at its first step. Further out the approximation stops holding, because d is no longer small beside the distances and the angle stops growing in proportion to it; that is the levelling toward 59 px.
Read the other way round, the expression is a detection limit, and it has exactly the form the range a pair cannot see past gave for depth. A third picture read to one pixel sees the slide only if e exceeds that pixel, which needs a depth error larger than about Z²/(f d). For a third eye two metres off the ray at six metres, that is about 2.6 cm. Halving the offset doubles the smallest error it can catch, and a third eye on the ray, where d is zero, catches none.
So the third view has not escaped the reciprocal law that governs every other depth read from two eyes; it has moved it. The error a pair measures is bounded by the baseline between its two eyes. The error a third eye can expose is bounded by that eye’s own baseline to the first line of sight — a different length, set by a different decision about where to put a camera, and invisible in any description of the rig that lists only how far apart its cameras are.
Where one of the two routes fails
The open marks in the sweep are three positions — the third eye swung 40°, 50° and 60° — where the exposure is large and yet the transfer by lines is refused.
At those positions the two epipolar lines the pair’s marks draw in the third picture are within two degrees of each other — 0.96° at 50° — and two nearly coincident lines cross at a place that the slightest error moves a long way. The construction declines to report that crossing rather than return a wandering number. The route through the point is unaffected and reports 47.0 px.
The geometry of the failure is specific. The two epipolar lines coincide when the third eye lies in the plane that contains both of the first two eyes and the world point, because all three centres and the point then share one epipolar plane and every constraint between the pictures collapses onto one line. As the third eye swings round the point at the first eye’s height, it passes close to that plane, and the line route degrades there while the point route does not.
That is worth knowing because the two routes are not interchangeable in practice. A check that works purely in the pictures — transferring matches between views by their epipolar lines, without reconstructing anything — has blind spots of its own, at third-eye positions near the plane of the first two eyes and the point, which a check through the reconstructed point does not have. A system that relies on the first to catch mismatches along epipolar lines can fail to catch them from a third camera that is perfectly placed for the second.
A small slide from the same positions
The sweep’s shape does not depend on the size of the slide; its scale does.
With a five-pixel slide the control is again at the floor, the exposure rises from 1.8 px at a metre off the ray to 10.4 px at six metres, and the same three positions refuse the line transfer. The curve is the thirty-pixel one scaled down by a little under six, which is roughly the ratio of the two depth errors — nine centimetres against fifty.
That scaling sets a practical threshold. A mismatch that moves its point by a few centimetres needs a third eye well off the first line of sight before its exposure clears a pixel, and a third picture read to a pixel will not see it from closer. The detection is geometric and exact in principle, and it is limited in practice by the same reading precision that limits every other depth taken from pictures.
What a third view does not settle
Three limits, because the result is easy to overstate.
It detects and does not identify. The third picture says the three marks are not consistent with one world point. It does not say which of the pair’s two marks was wrong, or whether the third picture’s own mark is. Deciding that needs a fourth view, a prior, or the kind of consensus over many points that the track and the scene together builds, and each of those can be fooled in its own way.
It needs a correct match in the third picture. Everything above assumes the third camera’s mark of the true point is right. If repeated structure fooled the pair, it may fool the third picture too, and a second wrong match that happens to land where the pair predicts would confirm the error rather than expose it.
And it does not repair the reciprocal law. A third view removes this particular silent error. It does not change what the range a pair cannot see past or whole pixels cut space into shells say about how well a correct match is placed, nor what another picture of the same sweep found about views taken from nearly the same direction, which is the same lesson as the sweep above arrived at from the other side: a view is worth what it adds in angle, not in count.
Still open: when a third eye lands on the next post
The second of those limits is the open question, and it can be measured without any matching at all. Repeated structure produces mismatches spaced by the pattern’s period — the next post of a railing, the next window along a facade — and a third picture sees the pattern repeated too. The pair’s wrong point projects into the third picture a stated number of pixels from the true one; if that displacement is close to a whole number of the pattern’s periods as the third picture draws it, the wrong point lands on another post and the third view confirms the mismatch instead of exposing it. The measurement that settles it builds a railing in the courtyard, slides a match by one period along its epipolar line, and sweeps the third eye round the point as this essay did — measuring the positions at which a perfectly good third picture is blind because the exposure it produces is a multiple of what the railing repeats by, and how narrow those blind positions become as the railing’s period shrinks.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A shadow edge read as a profile — both name baseline, residual, triangulation
- Eight points and the basis they are read in — both name baseline, correspondence, fundamental matrix
- Four cameras fit, and one of them can see — both name baseline, correspondence, fundamental matrix
- One shutter, two views — both name baseline, correspondence, fundamental matrix
- The image of the other eye — both name baseline, correspondence, fundamental matrix
- The lamp is the second eye — both name baseline, correspondence, triangulation
Named objects
A flat tag is an object no other essay names yet.
BaselineCorrespondenceFundamental matrixleast squaresOutlierResidualrobust estimationTriangulation