A third ray is worth what its picture is worth
Worth reading first: Two rays that do not meet · Another picture of the same sweep.
The midpoint is a choice of ruler finds that the point halfway between two skew rays is not an answer the photographs give: stretch the world threefold and it moves 0.203 millimetres, put it in a projective frame and it moves 1.503, while the point that minimises reprojection error stays put to a part in a thousand million million.
Two rays is the case where the difference is small enough to be a curiosity. With more than two it stops being one, and the reason is a change of character: with two rays the two estimators differ in where they put a point, and with three or more they differ in what they think the third picture is worth.
Two ways of counting a miss
The point nearest a bundle of rays in metres is a natural object and an easy one. Each ray contributes the squared distance from the candidate point to that ray, the total is a quadratic, and the minimum is three linear equations. It is the direct generalisation of the midpoint, and it is what a reader writes down when there are more than two rays and the midpoint has no meaning.
What it does is count every ray equally per metre of miss. A ray from a camera whose pixel spans a metre at the point is asked to be as close to the answer as a ray from a camera whose pixel spans a centimetre, and it will not be — so the answer moves to meet it.
The other estimate is the one the pictures themselves state. Project the candidate point into every camera, take the distance in pixels between where it lands and where the mark was read, and minimise the total. Every ray is then weighted by what its own picture says about it, which is roughly one over its distance, and a picture that says little is allowed to say little.
Both are least squares. They differ only in the units the residual is measured in, and the units are the whole of it.
One eye, walked away
The measurement is a rig of three eyes on one point, with two held at seven metres and the third walked back along the line of sight. Every mark is read half a pixel out, and a hundred and twenty noise draws are averaged at each station so that what is compared is a distribution rather than an accident.
With the three eyes together the two estimates agree: 34.6 millimetres of error against 34.6. That is the control, and without it every number below would be a statement about the estimators rather than about the arrangement.
Walk the third eye to seventy metres — ten times the pair’s own distance — and the two part company. The point nearest all three rays in metres sits 132.3 millimetres from the truth. The point of least reprojection error sits 34.4 millimetres from it, which is within two tenths of a millimetre of what the near pair gives with no third eye at all.
The direction matters as much as the size
A worse answer is not evidence of a mechanism. What makes this a mechanism rather than a coincidence is where the error goes.
Taking the component of each error along the direction from the point to the distant eye: the metric answer is displaced 12.1 millimetres along that line, and the reprojection answer 0.26 millimetres. So the metric estimate is not merely noisier; it is pulled toward the camera whose picture is worth least, which is exactly what equal weighting per metre predicts.
The size of the pull tracks the arrangement too. At a third eye of 10.5 metres it is 1.5 millimetres; at 21 metres, 3.5; at 49 metres, 8.5; at 70, 12.1. Roughly proportional to how far the distant eye is, which is what happens when a ray’s angular reading error spans a proportionally larger distance at the point.
The distant picture is not merely discounted, it is worth nothing
There is a stronger statement in the numbers than “one estimator handles it better”, and it is worth separating out.
The reprojection answer with three eyes — 34.35 millimetres — is not better than the near pair’s answer alone, which is 34.17. It is very slightly worse. So the third picture, taken from ten times the distance, adds nothing whatever to the reconstruction of that point.
That is not a failure of the estimator. It is a fact about the picture: a camera ten times further away reads the same angular error over ten times the distance, so its ray constrains the point ten times less tightly, and its contribution is swamped. The right weighting does not extract value that is not there; it declines to be damaged by a picture that has none.
Another picture of the same sweep reaches the same conclusion from the other side, on whole reconstructions rather than single points: going from three views to seven across one sixty-degree arc leaves the answer where it started and at one point makes it worse. The quantity that was short there was angular spread, and it is short here too.
Spread against count, at equal evidence
The last claim deserves its own measurement rather than an inference, because “spread rather than photographs” is easy to say and easy to say loosely.
Take one point, one reading error, one estimator, and vary only the arrangement: n rays crowded inside four degrees of arc, against n rays spread evenly over fifty-five. At every count the spread arrangement wins, and the margin is not small. Eight tight rays leave 54.2 millimetres of error; two spread rays leave 12.3.
So two pictures from different places beat eight from nearly the same place by a factor of four and a half, with four times as much evidence on the losing side. Adding tight rays does help a little — 62.9 millimetres at two, 54.2 at eight — but the improvement over a fourfold increase in evidence is fourteen per cent, and the whole of that is the square-root improvement that averaging any independent readings gives.
The shape of the region two rays leave is the reason. It is a needle — long along the line of sight, short across it — and a ray arriving from nearly the same direction is nearly parallel to the needle’s long axis, so it shortens the direction that was already short and does nothing to the one that was not. A ray from a different direction cuts the needle across, and that is the only thing that shortens it.
Why the weight is one over the distance
The factor the reprojection residual applies is worth deriving rather than stating, because it is the whole of the difference and it is one line.
A camera reads a mark to some number of pixels. That reading error is an angle — the pixel subtends an angle of one over the focal length, whatever the scene is — and an angle at distance spans a length at the point. So a ray’s positional uncertainty across the line of sight grows in proportion to the distance, and its inverse, which is the weight a least-squares fit should give it, falls as one over the distance.
Measuring the residual in pixels applies exactly that weight, because a pixel is the angle: a candidate point a given number of metres off the ray produces a reprojection residual of that many metres times , so the far camera’s metres are automatically divided by its own distance. Measuring the residual in metres applies no weight at all.
That also says how far the two answers can diverge, which is a useful bound to have. A bundle of rays whose distances span a factor of is weighted by a factor of differently under the two estimators, so the disagreement grows in proportion to — which is the near-linear growth of the drag measured above, ten times the pair’s depth giving roughly ten times the pull. It is not an effect that saturates.
Depth is a reciprocal is the same arithmetic run forwards rather than backwards: a fixed error in what is read maps onto an interval in depth that is not centred on the answer and grows with the square of the distance. One over the distance for the position across the line of sight, one over the distance squared for the depth along it — the two exponents are the needle’s two axes.
The same accounting, on a rig nobody built
The result is about weighting rather than about cameras, and it applies to any bundle of rays with different distances in it, however the rays arose.
A photograph containing two mirrors supplies exactly such a bundle. Two mirrors are three cameras counts three views in one exposure, at separations of 2.90, 3.10 and 2.26 metres — and the three virtual eyes are at three different distances from the scene, because each is the lens reflected in a mirror at its own distance. The three rays to a mark are therefore worth different amounts, and combining them by the nearest point in metres would over-weight whichever mirror is furthest away.
The same holds for a pair of eyes reading through water, for a scene photographed once from a hilltop and once from the doorway, and for any reconstruction assembled from pictures taken for other purposes. Two views give shape and no size established that the answer is a shape rather than a survey; what this adds is that the shape’s accuracy depends on a choice of units in the residual that has nothing to do with the shape at all.
Which estimator to use, stated plainly
The practical conclusion is short and it is not “always minimise reprojection error”, because the metric answer has a use.
Minimise reprojection error when the cameras differ in what they can see. Different distances, different focal lengths, different marking precision: any of those makes a metre of miss mean different things in different pictures, and the reprojection residual is the only one of the two that knows it.
The metric answer is right when every ray is worth the same, which is the case a two-camera rig of matched cameras at one distance actually is — and it is the case the midpoint formula is usually met in. It is also cheaper, closed-form, and a perfectly good starting point for the iteration that finds the other, which is what it is used for here.
What is not defensible is using the metric answer on a bundle whose rays are not comparable, and quoting the residual it produces as a quality figure. The residual there is in metres, it is dominated by the worst ray, and it is both the wrong number and a plausible-looking one.
What a bad ray does, as against a distant one
A distant ray and a wrongly-matched ray are two different failures and they need different answers, which is worth saying because both look like “one ray disagreeing”.
A distant ray is correct and imprecise. Its error is drawn from the same distribution as every other ray’s, scaled up, so it belongs in the estimate at a reduced weight — which is what the reprojection residual gives it automatically, with no outlier test at all.
A wrongly-matched ray is not a ray to that point. Its error is drawn from no distribution the model contains, and reducing its weight does not help, because the weight it ought to have is zero. A wrong match is not a small error measures what one of those does to a fit — move one correspondence of forty-four by thirty pixels, and the typical remaining point is out by half a pixel from a fit that had been exact to a part in ten trillion — and the repair there is rejection rather than weighting.
The distinction is testable on the numbers above. The distant ray’s reprojection residual is small, because a distant camera’s picture is nearly satisfied by any nearby point; a mismatched ray’s is large. So a reader looking at residuals can tell the two apart, and a reader looking at the metric misses in metres cannot, since both produce a large miss.
Where the third eye is worth having
None of this says a third picture is useless. It says a third picture from further along the same line of sight is, and there is a case in the collection where a third eye is decisive.
A mismatch on its own line needs a third eye slides one mark thirty pixels along the epipolar line its partner fixes, and finds that every test two photographs can run stays at the arithmetic floor while the point is reported half a metre too near. A third picture exposes it — by 21.4 pixels from an eye two metres off the first line of sight, and by exactly nothing from an eye on that line.
Read beside the measurement here, the two say the same thing about the same quantity. A third eye buys a new direction or it buys nothing, and whether the job is precision or detection does not change the answer.
What this does not settle
Every reading here is independently and identically wrong, and continuously so. Half a pixel, Gaussian in effect, uncorrelated between cameras — and a real matcher reading to whole pixels does not produce a continuous error at all. Whole pixels cut space into shells measures what that does to a pair, which is to replace a distribution by a set of allowed depths; a bundle of rays at different distances has shells of different spacings, and how those intersect is not the same question as how two distributions combine. Beyond that, real marking error is none of those: it is correlated with the image content, it is worse where an edge is oblique, and a feature detector’s errors on two views of the same corner are not independent. Weighting by what a picture says assumes the picture says it honestly.
The cameras are taken as exactly known. In a real reconstruction they are not, and a ray’s uncertainty then has the camera’s own uncertainty in it — which is larger for the camera that is hardest to place, and that is often the distant one. So the effect measured here is a floor rather than the whole of it.
And only one point is triangulated. A bundle adjustment solves for every point and every camera at once, and the weighting question there is the same question at a different scale, with the additional feature that a badly-weighted point drags the cameras which then drag every other point.
Still open: what weight a ray deserves when the camera is uncertain too
The measurement above weights each ray by what its own picture says about the point, holding the camera fixed. The case that matters weights it by what the picture says and by what is known about where the camera was.
Those two pull in the same direction when the distant camera is also the badly-placed one, which is the common case: a camera far from the scene is hard to place because it sees little parallax, so its ray is both imprecise and mis-aimed. The question with a measurement in it asks how the two contributions compose. Each ray carries a direction uncertainty from its own reading and a position-and-orientation uncertainty from its camera, and those combine into an uncertainty about the point that is not simply the larger of the two. Building a rig in which the camera uncertainty is stated rather than assumed — a third eye whose pose is itself recovered from the same marks — and comparing the point’s error against a weighting that uses only the reading, would say whether ignoring the camera’s own uncertainty is a small correction or the dominant term once the third eye is far enough away to be worth adding at all.
The short version
More rays are more evidence only if they come from somewhere else, and the two ways of combining them disagree about a bad ray rather than about a good one.
Three eyes on a point at seven metres, with the third walked to seventy: the point nearest all three rays in metres is 132.3 millimetres from the truth and is pulled 12.1 millimetres along the line to the distant eye; the point of least reprojection error is 34.4 millimetres from it, is pulled 0.26, and is within two tenths of a millimetre of what the near pair gives alone. With the three eyes together both give 34.6, which is what makes the rest a measurement.
And the arrangement outweighs the count. Two rays spread over fifty-five degrees beat eight rays crowded into four, by a factor of 4.4, because the region two rays leave is a needle and only a ray from a different direction cuts across it.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A fitted radius is wrong before it is uncertain — both name conditioning, demonstration, least squares, residual
- A shadow edge read as a profile — both name baseline, conditioning, residual, triangulation
- A turn of the head is not a step sideways — both name baseline, conditioning, reprojection error, residual
- An error with two terms — both name conditioning, demonstration, least squares, residual
- Square to the camera is the worst mirror — both name baseline, conditioning, depth uncertainty, triangulation
- The curvature a shadow reports — both name conditioning, demonstration, least squares, residual
Named objects
A flat tag is an object no other essay names yet.
BaselineConditioningCovarianceDemonstrationDepth uncertaintyleast squaresReprojection errorResidualskew raysTriangulation