What a pair is for

A third ray is worth what its picture is worth

Three eyes on one point, two at seven metres and one walked back to seventy. The point nearest all three rays in metres is 132 millimetres from the truth and the point of least reprojection error is 34 — the same 34 the near pair gives alone — and the first is pulled 12 millimetres along the line to the distant eye. And arrangement beats count outright — two rays spread over fifty-five degrees beat eight rays inside four, by a factor of 4.4.

Worth reading first: Two rays that do not meet · Another picture of the same sweep.

The midpoint is a choice of ruler finds that the point halfway between two skew rays is not an answer the photographs give: stretch the world threefold and it moves 0.203 millimetres, put it in a projective frame and it moves 1.503, while the point that minimises reprojection error stays put to a part in a thousand million million.

Two rays is the case where the difference is small enough to be a curiosity. With more than two it stops being one, and the reason is a change of character: with two rays the two estimators differ in where they put a point, and with three or more they differ in what they think the third picture is worth.

Two rays, 2.77 mm apart, in the plane that contains bothThe ray from the left eye through its mark and the ray from the right eye through its. With the marks placed exactly they meet, to 2.2e-16 m. With the same marks read to 1 px they miss by 2.77 mm at a range of 7.22 m. Triangulation is not an intersection; the reported point is a choice about what to minimise, and the gap is the part a residual alone will not tell you.midpoint — 2.77 mm gapfrom the left eyefrom the right eyegap 2.77 mm at 7.22 mexact marks: 2.2e-16 m
Fig. 1 The case this starts from: two rays read to a whole pixel, and the millimetres they miss each other by at seven metres.

Two ways of counting a miss

The point nearest a bundle of rays in metres is a natural object and an easy one. Each ray contributes the squared distance from the candidate point to that ray, the total is a quadratic, and the minimum is three linear equations. It is the direct generalisation of the midpoint, and it is what a reader writes down when there are more than two rays and the midpoint has no meaning.

What it does is count every ray equally per metre of miss. A ray from a camera whose pixel spans a metre at the point is asked to be as close to the answer as a ray from a camera whose pixel spans a centimetre, and it will not be — so the answer moves to meet it.

The other estimate is the one the pictures themselves state. Project the candidate point into every camera, take the distance in pixels between where it lands and where the mark was read, and minimise the total. Every ray is then weighted by what its own picture says about it, which is roughly one over its distance, and a picture that says little is allowed to say little.

Both are least squares. They differ only in the units the residual is measured in, and the units are the whole of it.

One eye, walked away

The measurement is a rig of three eyes on one point, with two held at seven metres and the third walked back along the line of sight. Every mark is read half a pixel out, and a hundred and twenty noise draws are averaged at each station so that what is compared is a distribution rather than an accident.

With the three eyes together the two estimates agree: 34.6 millimetres of error against 34.6. That is the control, and without it every number below would be a statement about the estimators rather than about the arrangement.

Walk the third eye to seventy metres — ten times the pair’s own distance — and the two part company. The point nearest all three rays in metres sits 132.3 millimetres from the truth. The point of least reprojection error sits 34.4 millimetres from it, which is within two tenths of a millimetre of what the near pair gives with no third eye at all.

A third eye at 70 m makes one answer 3.9× worse and the other not at allThree eyes on one point, two of them held at seven metres and the third walked back along the horizontal axis, with every reading half a pixel out and 120 noise draws averaged at each station. The upper series is the point nearest all three rays in metres; the lower is the point that minimises reprojection error. They agree while the three eyes are together — 34.6 against 34.6 millimetres — and part company as the third recedes, because the metric answer counts every ray equally per metre of miss and the reprojection answer counts each by what its own picture says. The flat lower line is also the answer from the near pair alone, 34.2 millimetres, which is what says the distant eye contributes nothing rather than being mishandled.050100204060how far the third eye stands from the point, in metreshow far the answer is from the truth, in millimetresnearest in metresleast reprojection error120 noise draws averaged at each station132 vs 34 mm
Fig. 2 The two estimators against how far the third eye stands, with the near pair’s own answer drawn flat underneath them.
The two answers, 10.2 mm apart at the same pointA few centimetres of the scene around the point, in the plane of the near pair's rays, with the three rays drawn through it. None of the three passes through the truth, because every reading is half a pixel out. The point nearest all three in metres and the point of least reprojection error are marked, and the arrow shows the direction of the third eye — 70 metres away against the pair's 7. The first answer sits along that direction from the truth and the second does not: a ray whose own picture is worth less is being given the same weight per metre as the two that are worth more.the truthnearest in metresa few centimetres of scene, in the pair's own planethe arrow points at the third eye
Fig. 3 A few centimetres of scene around the point, with the three rays through it and both answers marked. The arrow points at the distant eye, and one of the two answers has moved along it.

The direction matters as much as the size

A worse answer is not evidence of a mechanism. What makes this a mechanism rather than a coincidence is where the error goes.

Taking the component of each error along the direction from the point to the distant eye: the metric answer is displaced 12.1 millimetres along that line, and the reprojection answer 0.26 millimetres. So the metric estimate is not merely noisier; it is pulled toward the camera whose picture is worth least, which is exactly what equal weighting per metre predicts.

The size of the pull tracks the arrangement too. At a third eye of 10.5 metres it is 1.5 millimetres; at 21 metres, 3.5; at 49 metres, 8.5; at 70, 12.1. Roughly proportional to how far the distant eye is, which is what happens when a ray’s angular reading error spans a proportionally larger distance at the point.

The distant picture is not merely discounted, it is worth nothing

There is a stronger statement in the numbers than “one estimator handles it better”, and it is worth separating out.

The reprojection answer with three eyes — 34.35 millimetres — is not better than the near pair’s answer alone, which is 34.17. It is very slightly worse. So the third picture, taken from ten times the distance, adds nothing whatever to the reconstruction of that point.

That is not a failure of the estimator. It is a fact about the picture: a camera ten times further away reads the same angular error over ten times the distance, so its ray constrains the point ten times less tightly, and its contribution is swamped. The right weighting does not extract value that is not there; it declines to be damaged by a picture that has none.

Another picture of the same sweep reaches the same conclusion from the other side, on whole reconstructions rather than single points: going from three views to seven across one sixty-degree arc leaves the answer where it started and at one point makes it worse. The quantity that was short there was angular spread, and it is short here too.

Depth from disparity, with the 1 px the reading is worthZ = fB/d on a 65 mm baseline at 900 px. The line is exact — it returns the camera's own depth to 1e-12 m. The band is what 1 px of disparity error costs, and it stops being a ±. At 6.5 m it runs 5.86–7.33 m, lopsided by 1.25, and the textbook ±Z²δ/fB is 1.2% out. At 40 m it runs 23.8–126.5 m — 86.5 m beyond the estimate against 16.2 m before it, a lopsidedness of 5.32 — and the same formula is 47% out. Past 58.5 m the far edge is infinity.025507510010203040true depth (m)depth reported from the disparity, with a 1 px reading error5.86–7.33 m13.96–26.69 m23.76–126.49 mat 40 m: +86.5 m against −16.2 munbounded past 58.5 m
Fig. 4 Why a distant picture says so little: depth is a reciprocal of what is read, so a fixed error in the reading maps onto an interval in the answer that grows with the square of the distance.

Spread against count, at equal evidence

The last claim deserves its own measurement rather than an inference, because “spread rather than photographs” is easy to say and easy to say loosely.

Take one point, one reading error, one estimator, and vary only the arrangement: n rays crowded inside four degrees of arc, against n rays spread evenly over fifty-five. At every count the spread arrangement wins, and the margin is not small. Eight tight rays leave 54.2 millimetres of error; two spread rays leave 12.3.

So two pictures from different places beat eight from nearly the same place by a factor of four and a half, with four times as much evidence on the losing side. Adding tight rays does help a little — 62.9 millimetres at two, 54.2 at eight — but the improvement over a fourfold increase in evidence is fourteen per cent, and the whole of that is the square-root improvement that averaging any independent readings gives.

Two rays spread over 55° beat eight rays within 4°, by 4.4×The same point, the same reading error, and the same estimator, with the rays arranged two ways: all of them inside four degrees of arc, or spread evenly over fifty-five. The upper series is the tight arrangement and the lower is the spread one. Eight tight rays leave 54.2 millimetres of error where two spread ones leave 12.3, so the count is worth far less than the arrangement. The reason is the shape of the region two rays leave: it is long along the line of sight and short across it, and a ray from nearly the same place shortens only the direction that was already short.020406023468how many rayshow far the answer is from the truth, in millimetresall within four degreesspread over fifty-five120 noise draws averaged at each count54 vs 12 mm
Fig. 5 The same point and the same reading error, with the rays crowded into four degrees or spread over fifty-five. The count moves each series a little; moving between the series is worth four and a half times as much.

The shape of the region two rays leave is the reason. It is a needle — long along the line of sight, short across it — and a ray arriving from nearly the same direction is nearly parallel to the needle’s long axis, so it shortens the direction that was already short and does nothing to the one that was not. A ray from a different direction cuts the needle across, and that is the only thing that shortens it.

Why the weight is one over the distance

The factor the reprojection residual applies is worth deriving rather than stating, because it is the whole of the difference and it is one line.

A camera reads a mark to some number of pixels. That reading error is an angle — the pixel subtends an angle of one over the focal length, whatever the scene is — and an angle at distance ZZ spans a length Z/fZ/f at the point. So a ray’s positional uncertainty across the line of sight grows in proportion to the distance, and its inverse, which is the weight a least-squares fit should give it, falls as one over the distance.

Measuring the residual in pixels applies exactly that weight, because a pixel is the angle: a candidate point a given number of metres off the ray produces a reprojection residual of that many metres times f/Zf/Z, so the far camera’s metres are automatically divided by its own distance. Measuring the residual in metres applies no weight at all.

That also says how far the two answers can diverge, which is a useful bound to have. A bundle of rays whose distances span a factor of kk is weighted by a factor of kk differently under the two estimators, so the disagreement grows in proportion to kk — which is the near-linear growth of the drag measured above, ten times the pair’s depth giving roughly ten times the pull. It is not an effect that saturates.

Depth is a reciprocal is the same arithmetic run forwards rather than backwards: a fixed error in what is read maps onto an interval in depth that is not centred on the answer and grows with the square of the distance. One over the distance for the position across the line of sight, one over the distance squared for the depth along it — the two exponents are the needle’s two axes.

The same accounting, on a rig nobody built

The result is about weighting rather than about cameras, and it applies to any bundle of rays with different distances in it, however the rays arose.

A photograph containing two mirrors supplies exactly such a bundle. Two mirrors are three cameras counts three views in one exposure, at separations of 2.90, 3.10 and 2.26 metres — and the three virtual eyes are at three different distances from the scene, because each is the lens reflected in a mirror at its own distance. The three rays to a mark are therefore worth different amounts, and combining them by the nearest point in metres would over-weight whichever mirror is furthest away.

The same holds for a pair of eyes reading through water, for a scene photographed once from a hilltop and once from the doorway, and for any reconstruction assembled from pictures taken for other purposes. Two views give shape and no size established that the answer is a shape rather than a survey; what this adds is that the shape’s accuracy depends on a choice of units in the residual that has nothing to do with the shape at all.

Which estimator to use, stated plainly

The practical conclusion is short and it is not “always minimise reprojection error”, because the metric answer has a use.

Minimise reprojection error when the cameras differ in what they can see. Different distances, different focal lengths, different marking precision: any of those makes a metre of miss mean different things in different pictures, and the reprojection residual is the only one of the two that knows it.

The metric answer is right when every ray is worth the same, which is the case a two-camera rig of matched cameras at one distance actually is — and it is the case the midpoint formula is usually met in. It is also cheaper, closed-form, and a perfectly good starting point for the iteration that finds the other, which is what it is used for here.

What is not defensible is using the metric answer on a bundle whose rays are not comparable, and quoting the residual it produces as a quality figure. The residual there is in metres, it is dominated by the worst ray, and it is both the wrong number and a plausible-looking one.

Where a stereo pair stops being able to say "no further than"The disparity of a point at range Z is fB/Z, so once it falls below the 1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 900 px it is 58.5 m. Doubling the baseline doubles it; nothing about the estimator does.123-2-10baseline (m, log scale)range past which 1 px of disparity error is unbounded (m, log scale)65 mm → 58.5 mrange = fB/δ — 58.5 m at a 65 mm baseline900 px focal length, 1 px reading
Fig. 6 The other end of the same accounting: the distance past which a pair of eyes cannot say “no further than”, which is set before anything is built.

What a bad ray does, as against a distant one

A distant ray and a wrongly-matched ray are two different failures and they need different answers, which is worth saying because both look like “one ray disagreeing”.

A distant ray is correct and imprecise. Its error is drawn from the same distribution as every other ray’s, scaled up, so it belongs in the estimate at a reduced weight — which is what the reprojection residual gives it automatically, with no outlier test at all.

A wrongly-matched ray is not a ray to that point. Its error is drawn from no distribution the model contains, and reducing its weight does not help, because the weight it ought to have is zero. A wrong match is not a small error measures what one of those does to a fit — move one correspondence of forty-four by thirty pixels, and the typical remaining point is out by half a pixel from a fit that had been exact to a part in ten trillion — and the repair there is rejection rather than weighting.

The distinction is testable on the numbers above. The distant ray’s reprojection residual is small, because a distant camera’s picture is nearly satisfied by any nearby point; a mismatched ray’s is large. So a reader looking at residuals can tell the two apart, and a reader looking at the metric misses in metres cannot, since both produce a large miss.

Where the third eye is worth having

None of this says a third picture is useless. It says a third picture from further along the same line of sight is, and there is a case in the collection where a third eye is decisive.

A mismatch on its own line needs a third eye slides one mark thirty pixels along the epipolar line its partner fixes, and finds that every test two photographs can run stays at the arithmetic floor while the point is reported half a metre too near. A third picture exposes it — by 21.4 pixels from an eye two metres off the first line of sight, and by exactly nothing from an eye on that line.

Read beside the measurement here, the two say the same thing about the same quantity. A third eye buys a new direction or it buys nothing, and whether the job is precision or detection does not change the answer.

What this does not settle

Every reading here is independently and identically wrong, and continuously so. Half a pixel, Gaussian in effect, uncorrelated between cameras — and a real matcher reading to whole pixels does not produce a continuous error at all. Whole pixels cut space into shells measures what that does to a pair, which is to replace a distribution by a set of allowed depths; a bundle of rays at different distances has shells of different spacings, and how those intersect is not the same question as how two distributions combine. Beyond that, real marking error is none of those: it is correlated with the image content, it is worse where an edge is oblique, and a feature detector’s errors on two views of the same corner are not independent. Weighting by what a picture says assumes the picture says it honestly.

The cameras are taken as exactly known. In a real reconstruction they are not, and a ray’s uncertainty then has the camera’s own uncertainty in it — which is larger for the camera that is hardest to place, and that is often the distant one. So the effect measured here is a floor rather than the whole of it.

And only one point is triangulated. A bundle adjustment solves for every point and every camera at once, and the weighting question there is the same question at a different scale, with the additional feature that a badly-weighted point drags the cameras which then drag every other point.

Still open: what weight a ray deserves when the camera is uncertain too

The measurement above weights each ray by what its own picture says about the point, holding the camera fixed. The case that matters weights it by what the picture says and by what is known about where the camera was.

Those two pull in the same direction when the distant camera is also the badly-placed one, which is the common case: a camera far from the scene is hard to place because it sees little parallax, so its ray is both imprecise and mis-aimed. The question with a measurement in it asks how the two contributions compose. Each ray carries a direction uncertainty from its own reading and a position-and-orientation uncertainty from its camera, and those combine into an uncertainty about the point that is not simply the larger of the two. Building a rig in which the camera uncertainty is stated rather than assumed — a third eye whose pose is itself recovered from the same marks — and comparing the point’s error against a weighting that uses only the reading, would say whether ignoring the camera’s own uncertainty is a small correction or the dominant term once the third eye is far enough away to be worth adding at all.

The short version

More rays are more evidence only if they come from somewhere else, and the two ways of combining them disagree about a bad ray rather than about a good one.

Three eyes on a point at seven metres, with the third walked to seventy: the point nearest all three rays in metres is 132.3 millimetres from the truth and is pulled 12.1 millimetres along the line to the distant eye; the point of least reprojection error is 34.4 millimetres from it, is pulled 0.26, and is within two tenths of a millimetre of what the near pair gives alone. With the three eyes together both give 34.6, which is what makes the rest a measurement.

And the arrangement outweighs the count. Two rays spread over fifty-five degrees beat eight rays crowded into four, by a factor of 4.4, because the region two rays leave is a needle and only a ray from a different direction cuts across it.

The other eye, found in the picture — 2.6 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 1.1e-9 px on the left and 3.7e-10 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 1.1e-9 px2.60 m between the eyes
Fig. 7 The mark a third eye has to be placed away from: the image of one camera in another’s picture, and the line along which a third eye buys nothing.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineConditioningCovarianceDemonstrationDepth uncertaintyleast squaresReprojection errorResidualskew raysTriangulation