What a pair is for

A third eye that lands on the next post

Match one post of a railing to its neighbour and the pair reports it at 19.8 metres instead of 9.0, with every test two photographs can run at the arithmetic floor. A third picture usually exposes that by hundreds of pixels — but at five azimuths in seventy-eight degrees the wrong point lands within three pixels of another post, and the third view confirms the mistake. Narrow the railing to twenty centimetres and those places cover 28 per cent of the arc.

Worth reading first: A wrong match is not a small error · Two rays that do not meet.

A mismatch on its own line needs a third eye slides one mark thirty pixels along the epipolar line its partner fixes, and finds that every test two photographs can run stays at the arithmetic floor while the point is reported half a metre too near. Its conclusion is that a third picture is the instrument: 21.4 pixels of exposure from an eye two metres off the first line of sight, and exactly nothing from an eye on that line.

That conclusion has an exception, and the exception is the case where the mismatch happens in the first place.

A mismatch that happens by itself

Nothing has to be contrived for a mark to slide along its own epipolar line. A railing does it.

A row of posts at one depth and one height, photographed by a pair with a horizontal baseline, puts every post’s image on one row of both pictures — which is the epipolar row. So matching a post in the left picture to any post in the right picture is epipolar-consistent, and a matcher comparing appearances finds every candidate equally convincing, because the posts are identical.

Take posts sixty centimetres apart at nine metres, a baseline of 1.1 metres and nine hundred pixels of focal length. The posts are sixty pixels apart in the picture and the true disparity is a hundred and ten. Match one post to the next along and the disparity becomes fifty, which puts the post at 19.8 metres instead of 9.0.

Every test two pictures can run, on a match slid 30 px along its own lineCorrespondence 9's second mark moved 30 px along the epipolar line its first mark fixes. Bars are drawn on a logarithmic scale from 10⁻¹⁶ to one pixel. Its epipolar distance is 2.2e-14 px, the refitted geometry's worst is 3.8e-13 px, its two rays meet to 1.5e-15 m, and its triangulated point reprojects to 1.1e-13 px. The point is 0.50 m from where it belongs — 6.20 m from the first eye, reported at 5.70 m.its distance from its epipolar line2.2e-14 pxworst epipolar distance, fit including it3.8e-13 pxgap between its two rays1.5e-15 mreprojection error of its point1.1e-13 pxone pixel, for scale1 pxdepth wrong by 0.50 mlog scale, 10⁻¹⁶ to 1 px
Fig. 1 What the pair can say about such a match, from the case where the slide was imposed rather than natural: every two-view test at the arithmetic floor, and the point in the wrong place.

What a third picture is supposed to do

A third camera sees the scene from somewhere else, so it sees the true post and the pair’s wrong point in different places. The distance between those two places in the third picture is the exposure, and it is what a three-view consistency check measures.

Swept round the post at nine metres, the exposure rises from 82 pixels at six degrees to 1,027 at eighty-four. Any of those is enormous by the standards of a matcher’s tolerance, and on that evidence a third picture settles the question everywhere.

Where it does not

The third picture sees the railing too.

If the wrong point lands, in the third picture, within a matcher’s tolerance of some post of the same railing, then the third view does not contradict the pair. It agrees: there is a post there, the three pictures are consistent, and the mismatch is confirmed rather than exposed.

That happens whenever the exposure is close to a whole number of the railing’s periods as the third picture draws them — and the third picture draws them at its own scale, which changes as the camera moves round. Swept through seventy-eight degrees, the wrong point falls within three pixels of a post at five azimuths, covering 2.2 per cent of the arc. At those five places the exposure is between about a hundred and eight hundred pixels, so nothing about the size of the disagreement warns anybody.

5 places in 78° where a good third picture confirms the mistakeA third camera swept round the post, at nine metres. The upper series is how far the pair's wrong point lands from the true one in the third picture — the displacement a third view offers, rising from 82 to 1027 pixels. The lower series is how far that landing falls from the image of *some* post, and it is the one that matters: where it dips to nothing, the wrong point has landed on another post of the same railing, and a matcher looking for agreement between the three pictures finds it. That happens at 5 azimuths here, covering 3.8 per cent of the sweep, and the exposure at those azimuths is hundreds of pixels.0.1110100100020406080where the third eye stands, in degrees round the pointpixelshow far the wrong point landshow far from the nearest postthe vertical rules are the blind azimuths3.8% of the sweep
Fig. 2 The third eye swept round the post. The upper series is how far the wrong point lands from the true one; the lower is how far it lands from the nearest post, and the dips are where the third picture agrees with the mistake.

The finer the structure, the worse it gets

The blind places are not a fixed hazard. They depend on how densely the structure repeats, and the dependence runs the wrong way.

A railing with posts a metre apart is blind at two azimuths covering 1.9 per cent of the arc. At sixty centimetres it is five places and 2.2 per cent. At thirty, twelve places and 14 per cent. At twenty, nineteen places and 28 per cent.

The reason is direct. A fine railing puts its posts close together in every picture, so a wrong point has more posts to land near and a fixed three-pixel tolerance covers a larger share of each period. Both effects push the same way.

So the structure a matcher finds hardest — a fine, regular, repeating pattern — is also the structure a third picture is least likely to rescue. That is the wrong way round for an instrument to behave, and it is worth stating plainly because the usual advice is “add a view”.

A railing at 20 cm is blind over 28% of the arc; at a metre, 1.9%The share of third-eye positions at which the pair's wrong point lands within three pixels of some post of the railing, against how far apart the posts stand. A coarse railing gives the wrong point few chances to land on one, and those chances are narrow: 4 places in seventy-eight degrees, covering 1.9 per cent. A fine one gives many and they are wide: 19 places covering 28 per cent. So the structure hardest to match is also the structure a third picture is least likely to rescue, which is the wrong way round for an instrument to behave.01020300.2000.4000.6000.8001how far apart the posts stand, in metresshare of third-eye positions that confirm the mistake, in per cent5 places12 placesa three-pixel matching tolerance19 places at 20 cm
Fig. 3 The share of third-eye positions that confirm the mistake, against how far apart the posts stand. It rises as the railing gets finer, which is the opposite of what more evidence is supposed to do.

The condition, written out

The blind azimuths are not found by search in principle, only in practice, and writing the condition down says what each ingredient does.

Let Δ\Delta be the exposure — how far the wrong point lands from the true one in the third picture — and pp the railing’s period as that picture draws it. The third view is blind when

Δkp  <  τfor some whole k,\left|\Delta - k\,p\right| \;<\; \tau \quad\text{for some whole } k,

with τ\tau the matcher’s tolerance. Both Δ\Delta and pp change with the third camera’s azimuth, and they change in opposite directions: as the camera swings round, the exposure grows because the two points separate, and the period shrinks because the railing is seen more obliquely. Measured here, the exposure runs from 82 to 1,027 pixels while the period falls from 60 to 6.

So Δ/p\Delta/p is a quantity that rises steeply, from about 1.4 at six degrees to about 160 at eighty-four, and the blind azimuths are where it passes a whole number. That is why they are isolated, why there are more of them at large azimuth, and why they get closer together as the camera swings out — the ratio’s own steepness sets the spacing.

It also says what the tolerance buys. The blind share is roughly 2τ2\tau divided by the local spacing of the whole numbers of Δ/p\Delta/p, which is why halving a matcher’s tolerance halves the blind share and why a fine railing, whose pp is small and whose ratio therefore climbs faster, is blind more often.

Why the depth error is so large

The mismatch moves the point ten metres and it is worth a sentence to say why a sixty-pixel error in a mark does that, because the size is what makes the failure worth finding.

Disparity and depth are reciprocal. Depth is a reciprocal is the standing account: a fixed error in what is read maps onto an interval in what is reported that is not centred on the answer, and at forty metres runs sixteen metres nearer and eighty-six further. Here the reading changes from 110 pixels to 50, which is a factor of 2.2 in disparity and therefore a factor of 2.2 in depth — 9.0 metres becoming 19.8.

The direction matters as much as the size. Matching to the post on the far side makes the disparity smaller and the point further away; matching to the near side makes it larger and the point nearer. So a railing generates errors in both directions, and a reconstruction contaminated by them is spread along the line of sight rather than scattered — which is exactly the shape whole pixels cut space into shells describes for the quantisation, arriving here from a different cause.

And this is a case where the estimator cannot help. A third ray is worth what its picture is worth separates a ray that is correct and imprecise from one that is not a ray to the point at all, and weights the first while requiring the second to be rejected. This is the second kind, and at a blind azimuth it produces a small residual in every picture, so no weighting distinguishes it.

What a matcher actually compares

The posts are identical in the model and nearly identical in a photograph, and the difference between those is where the practical margin lives.

A real matcher compares a patch of image around each candidate. Two posts of the same railing at the same depth present patches that differ only in what is behind them and how the light falls, and on a plain sky or a plain wall the difference can be a fraction of a grey level. That is the case the model here represents exactly.

Where it breaks down is at the ends of the railing, where a post has a neighbour on one side only, and where the background changes. Those posts are matchable, and a matcher that propagates from them inward — an ordering or smoothness constraint along the row — gets the rest right. So the repair for a railing is a constraint along the structure rather than a better test on each post — an assumption imported from outside the pictures, which is the currency every recovery in this subject is bought with — and two rays that do not meet is the reminder of why the per-point residual will not carry it: the residual of a wrong match here is not large.

The blind places are a set of azimuths, not a region

The shape of the failure matters for what can be done about it, and it is a good deal friendlier than the share alone suggests.

The blind azimuths are isolated and roughly evenly spaced — about seven degrees apart for the sixty-centimetre railing — rather than forming a band. So a third camera moved a few degrees is almost certainly out of one, and the failure is not a direction to be avoided but a coincidence to be broken.

That is a different shape from the failure a mismatch on its own line needs a third eye records, where a third eye on the first line of sight sees nothing whatever and the blind set is a whole line. Here the geometry is fine everywhere; what fails is the appearance test, at a discrete set of places where the arithmetic happens to line up.

It is also a different shape from an epipole in the picture leaves a blind disc, where a region of the picture is unusable for a reason that has nothing to do with what is in it. Here the whole picture is usable and a few positions of the camera are not.

Two views of the same failure

Reading the three numbers together says what kind of failure this is, and it is not an outlier in the usual sense.

A wrong match is not a small error measures what a bad correspondence does to a fit — move one of forty-four by thirty pixels, and the typical remaining point is out by half a pixel from a fit that had been exact to a part in ten trillion — and the repair there is detection followed by rejection, because least squares has nowhere to put a bad row except across all of them.

Detection is what fails here. The mismatch produces no large residual in two views, because it is epipolar-consistent; and at the blind azimuths it produces no large residual in three, because the third picture has a post where the wrong point lands. A robust estimator looks for the row that disagrees, and no row disagrees.

What does still work is the depth. The wrong point sits at 19.8 metres against the railing’s 9.0, and a scene model that knows a railing is a plane, or a straight line, or roughly one depth, has an easy rejection. So the repair is not a better outlier test but a constraint about the structure — which is what the structure being repeated gave away in the first place.

The other eye, found in the picture — 2.6 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 1.1e-9 px on the left and 3.7e-10 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 1.1e-9 px2.60 m between the eyes
Fig. 4 The point every joining line runs through, which is what makes a railing’s posts share one epipolar row at all — and the reason a mismatch along that row costs nothing in two views.
One correspondence moved 30 px, and where the damage wentThe clean fit is exact to 2.3e-13 px. Moving correspondence 9 by 30 px leaves every other point wrong too — the typical one by 0.54 px and the worst by 1.7 px — because a least-squares fit has nowhere to put a bad row except across all of them. Here the largest residual does fall on the culprit; it is not obliged to.00.50011.50010203040correspondenceepipolar error at every OTHER point, after one match is movedmedian 0.54 pxthe moved oneevery point wrong: median 0.54 px, worst 1.7 pxclean fit 2.3e-13 px
Fig. 5 The ordinary case for comparison: one contaminated correspondence, and the way least squares distributes it across every other point rather than isolating it.

What a fourth view buys

The natural next move is a fourth camera, and the arithmetic says it should be decisive.

The blind azimuths are set by the third camera’s own position, through the period the railing subtends in its picture. A fourth camera at a different azimuth has a different period and therefore a different set of blind places, and two sets of isolated azimuths a few degrees wide coincide only by a much rarer accident. With a blind share of 2.2 per cent each and no reason for correlation, two cameras are blind together about five parts in ten thousand of the time.

That estimate is honest only if the two sets are independent, and they are not exactly: both sets are generated by the same railing and the same wrong point, and both are roughly periodic in azimuth. Whether their periods are commensurate is a question about the arrangement rather than about the structure, and it is not answered here.

The practical form of the advice is therefore about where rather than how many: two extra views separated by an azimuth that is not a multiple of the blind spacing. Seven degrees apart is the worst choice available and it is also a very natural one.

What this does not settle

The matcher is taken as exact. Every mark here is placed by projection and every tolerance is a stated three pixels. A real matcher’s appearance test would often prefer the true post slightly, since the posts are not quite identical and the lighting is not quite the same, and nothing here measures how much of a preference is enough.

The railing is perfectly regular and perfectly planar. Real repeated structure is not, and an irregular row is measurably better: the same sweep on posts whose spacings differ by up to seventeen centimetres finds one blind place instead of five. How irregular a structure has to be before the effect goes away is a measurement this does not make.

Nothing here is about the pair’s own geometry being uncertain. Both cameras are exactly known, so the wrong point is computed exactly and the third picture’s prediction of it is exact. In a reconstruction where the poses are themselves recovered, a railing contributes many correspondences and a set of them shifted by one period would drag the recovered poses as well — which is the shape a flat scene fixes no second eye records for a different degeneracy, where the failure is silent because the residual stays at the floor throughout.

And only one mismatch is considered at a time. A railing produces many, one per post, and they are correlated — a matcher that shifts one post by a period usually shifts its neighbours by the same period, which produces a wholly consistent wrong reconstruction of the entire railing at a wrong depth. That is a worse failure than the single bad point and it is not measured here.

Still open: how far a wrong railing can be shifted whole

The failure measured above moves one post. A matcher that shifts every post of a railing by the same period produces a reconstruction in which the whole railing stands at 19.8 metres instead of 9.0 — internally consistent, straight, evenly spaced, and wrong.

That reconstruction is not detectable by any test on the railing alone, since a railing at nineteen metres is a perfectly good railing. What would detect it is the rest of the scene: the ground the posts stand on, the wall behind them, the point where the railing meets something that is not a railing. So the open measurement places a railing in a scene with a ground plane and a few unique features, shifts the whole railing by one period, and asks how far from the railing a unique feature has to be before the inconsistency it creates exceeds the matching tolerance — and whether the answer depends on the railing’s length, which decides how far the wrong reconstruction has to be dragged to meet the rest of the scene.

The short version

A row of identical posts at one depth puts every post on one epipolar row, so matching one to its neighbour is a mismatch no two-view test can see. Sixty-centimetre posts at nine metres come back at 19.8 metres.

A third picture usually exposes that by between eighty and a thousand pixels. At five azimuths in seventy-eight degrees it does not, because the wrong point lands within three pixels of another post of the same railing and the third view agrees with the mistake. Those places cover 2.2 per cent of the arc at sixty centimetres and 28 per cent at twenty — so the finer the repeating structure, the less a third view is worth, which is the opposite of how more evidence usually behaves.

A railing 30 px between posts, matched across a disparity of 110 pxThe two pictures of a row of posts 30 centimetres apart at nine metres, drawn as strips along the one image row they all lie on. In the left picture the marked post is the one being matched; in the right, the filled mark is its true partner and the open one is the post next along, which is what a matcher confuses it with. Both candidates lie on the same epipolar row, so the pair has no way to prefer one: the wrong match satisfies every test two photographs can run, and puts the post at 12.4 metres instead of 9.0.left pictureright picturethe post it is matched toone image row, drawn as two strips12.4 m instead of 9.0
Fig. 6 And the same two pictures with the posts half as far apart: a shorter period, a nearer wrong point, and nineteen places in the arc where a third camera would agree with the mistake instead of exposing it.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineCorrespondenceDegenerate configurationDemonstrationEpipolar lineinstrument limitOutlierResidualrobust estimationTriangulation