Many pictures at once

A narrow view keeps a second answer, inside out

Six pictures of a courtyard through a 3° field, every mark exact. Started from the scene turned inside out, the adjustment settles there — near points far, far points near — and misfits the marks by under a pixel. The misfit grows in proportion to the field and to the sweep of the cameras, and vanishes only where perspective does.

Worth reading first: The track and the scene together · Seven numbers no picture can name.

A bundle adjustment is a descent. It starts somewhere, moves every camera and every point downhill in reprojection error, and stops where nothing downhill is left. Where the adjustment stops measured the stopping place for a good start: five or six iterations from a chained initialisation to a residual set by how finely the marks were read. The same essay noted in passing that from a random start the descent does not converge at all, because the error surface over a hundred and sixty-eight parameters has a great many valleys and almost none of them are reconstructions.

Almost none is not none. One wrong start is special, because it has a valley of its own that is a reconstruction, of a scene that does not exist, and a narrow field of view makes that valley almost as deep as the right one.

The wrong start is the scene turned inside out.

The courtyard and its inside-out twin, seen from above the first camera at a 3° fieldBoth answers the adjustment settles in, drawn in the first camera's own frame: across the picture horizontally, and depth toward the camera upward, with the second answer scaled to the same mean distance. The true courtyard's near points are the inside-out answer's far points, by up to 8.66 m, and the depths of the two correlate at -1.000. Both reproject exact marks: the true answer to 2.4e-13 px, the inside-out one to 0.946 px.each point joined to its place in the inside-out answer — nearer the camera is uptoward the cameratrue 2.4e-13 px · inside out 0.946 px3° field, 60° of arc
Fig. 1 Both answers the adjustment settles in at a 3° field, drawn in the first camera’s own frame with depth toward the camera upward. Each true point is joined to its place in the second answer: near points have become far and far points near, by up to 8.66 m, and the two sets of depths correlate at −1.000. The true answer reprojects the exact marks to 2.4e-13 px; the inside-out one to 0.946 px.

An ambiguity older than photography

A drawing of a wire cube with no hidden lines removed can be seen two ways, and the eye flips between them: one face is in front, then the other. The flip is not a failure of the eye. The drawing genuinely does not contain the information to decide, because a parallel projection of a shape and a parallel projection of its mirror image in depth are the same drawing.

That statement holds for any number of views, not just one, and the reason fits in a line. Let M=diag(1,1,1)M = \mathrm{diag}(1, 1, -1) reflect depth. Reflect every point of the scene, X=MXX' = MX, and replace every camera’s rotation by R=MRMR' = MRM, which is still a rotation. Then RX=MRMMX=M(RX)R'X' = MRMMX = M(RX). The reflected scene seen by the reflected cameras lands at MM times where the true scene seen by the true cameras landed, and MM changes only the depth coordinate. An orthographic camera keeps the first two coordinates and throws the depth away, so every picture is identical. Nothing a set of parallel projections contains can tell the scene from its twin.

A pinhole does not throw depth away. It divides by it, so a point’s image is scaled by one over its distance, and a point that has moved from near to far is drawn smaller than it was. The twin now produces different pictures from the scene. How different depends on how much the division by depth varies across the scene, which is the ratio of the scene’s depth to its distance — and for a camera that keeps the scene filling its frame, that ratio is set by the field of view.

So a pinhole breaks the ambiguity, and by an amount the field of view controls. The question is whether it breaks it enough to matter, and the only way to find out is to put the twin into the solver and see what happens.

Putting the twin into the solver

The six cameras here are the courtyard’s standing sequence, moved back as the field narrows so the courtyard fills the same share of each frame: at a 50° field they are exactly the standing arc 7.2 m out, and at 3° they stand 120 m away. The marks are exact. The start is the reflected scene and the reflected cameras, each camera kept at its own distance from the middle of the courtyard.

From that start the descent runs for a handful of iterations and stops. It does not climb back out toward the true scene. It settles on a reconstruction whose depth order is the true depth order reversed, point for point — a correlation of −1.000 between the true depths and the twin’s, measured in the first camera’s frame.

The first picture, and where the inside-out answer puts every markThe first camera's picture at a 3° field, with each mark and the place the inside-out reconstruction reprojects it to. The two differ by at most 3.65 px in this picture, and the arrows are drawn 20 times their length so they can be seen at all. Across all six pictures the inside-out answer misfits by 0.946 px root-mean-square, against 2.4e-13 px for the true one.marks, and the inside-out answer's reprojections drawn 20× further awayworst in this picture 3.65 px · all six 0.946 px3° field
Fig. 2 The first camera’s picture at a 3° field, with each exact mark and the place the inside-out reconstruction reprojects it to, the difference drawn twenty times its length. The worst mark in this picture is off by 3.65 px; across all six pictures the twin misfits by 0.946 px root-mean-square.

The misfit is the one thing that separates the two answers. Across all six pictures the twin reprojects the marks to 0.946 px root-mean-square, with the worst mark in the first picture off by 3.65 px. The true answer reprojects them to 2.4×10132.4 \times 10^{-13} px.

A picture drawn with every reprojection moved by the misfit shows how little that is. The arrows had to be drawn twenty times their length to be visible at all, and they are arranged the way a perspective difference would arrange them — small near the middle of the scene, larger at the points whose depth changed most. Laid over the photograph at their true length they would be invisible.

In proportion to the field

Repeating the experiment across fields of view gives the relation the argument above predicts.

The inside-out answer, and how badly it fits as the field widensSix pictures of the courtyard, every mark exact, the cameras moved back as the field narrows so the scene fills the same share of each frame. Started from the scene turned inside out, the adjustment settles in a second minimum wherever the field is narrow enough, with the depth order still reversed. Its misfit is 0.474 px at 1.5°, 0.946 px at 3°, 1.884 px at 6°, 3.123 px at 10°, 4.643 px at 15°, 7.528 px at 25° along a 60° arc — a slope of 0.98 on logarithmic axes — and 0.136 px at 1.5°, 0.270 px at 3°, 0.523 px at 6° along 20°. The solver does not settle with the depth order reversed at 10° at 20° of arc, 15° at 20° of arc, 25° at 20° of arc within eighty iterations. The true answer fits every one of them to 5e-13 px.1.5°10°15°25°0.10.31310field of view (degrees, log scale)misfit of the inside-out answer, exact marks (px, log scale)60° of arc20° of arcslope 0.98 along a 60° arcexact marks
Fig. 3 The misfit of the inside-out answer against the field of view, exact marks, for the standing 60° arc of cameras and for a 20° arc. Along 60° it runs from 0.474 px at 1.5° to 7.528 px at 25°, a slope of 0.98 on logarithmic axes; along 20° it is 0.136 px at 1.5° and 0.523 px at 6°. The true answer fits every configuration to 5e-13 px.

Along the standing arc the twin misfits the marks by 0.474 px at a 1.5° field, 0.946 at 3°, 1.884 at 6°, 3.123 at 10°, 4.643 at 15° and 7.528 at 25°. On logarithmic axes those six points lie on a line of slope 0.98. Halving the field halves the misfit, over a range of seventeen to one.

That is the first-order effect of perspective showing up where it should. With the scene kept filling the frame, the camera’s distance is inversely proportional to the tangent of half the field, so the ratio of the courtyard’s depth to its distance is proportional to that tangent, and over these fields the tangent is proportional to the angle. The twin’s error is the part of perspective an orthographic camera would not have, and there is exactly that much of it.

It also means the misfit does not reach zero at any field a real lens has. It reaches zero in the limit, where the camera is infinitely far away and the pinhole has become a parallel projection. Every finite field keeps the two answers distinct, by an amount that can be computed, and every narrow one keeps them close.

And in proportion to the sweep

The field is one of two things that separate the scene from its twin. The other is how differently the six cameras see it.

How much arc it takes to tell the scene from its twin, at a 3° fieldThe field held at 3° and the six cameras spread over wider arcs. The inside-out answer misfits exact marks by 0.128 px at 10°, 0.270 px at 20°, 0.579 px at 40°, 0.946 px at 60°, 1.607 px at 90°. A wider sweep separates the two answers, but a narrow field keeps them close whatever the sweep.00.50011.501020406090angle swept by the six cameras (degrees)misfit of the inside-out answer, exact marks (px)3° field · exact marks0.128 → 1.607 px
Fig. 4 The field held at 3° and the six cameras spread over wider arcs. The inside-out answer misfits exact marks by 0.128 px when the cameras sweep 10°, 0.270 px at 20°, 0.579 px at 40°, 0.946 px at 60° and 1.607 px at 90°.

Holding the field at 3° and widening the arc the six cameras stand on, the twin misfits the marks by 0.128 px across a 10° sweep, 0.270 px across 20°, 0.579 across 40°, 0.946 across 60° and 1.607 across 90°. Nearly proportional again, rising a little faster than the arc.

Read together, the two figures say the separation is close to the product of the two angles. At 3° of field and 60° of sweep it is 0.946 px, 0.0053 px per degree of each; at 1.5° and 20° it is 0.136 px, 0.0045; at 25° and 60° it is 7.528 px, 0.0050; at 3° and 90° it is 1.607, 0.0060. The constant drifts by a third over configurations whose product varies fifty-fold, which is close enough to call the relation a product and far enough from exact to leave the constant as a property of this courtyard.

The product has a direct reading. The twin is hidden by two things at once: each picture carries little perspective, and the pictures differ little from each other. Either alone would betray it. Wide fields with a narrow sweep, or narrow fields with a wide one, each keep some leverage. Narrow fields with a narrow sweep keep almost none — which is the arrangement of a long lens following a distant subject, or a camera looking down from high above a small patch of ground.

The 20° arc has one feature this measurement does not explain. At fields of 10° and wider, a start from the twin does not settle within eighty iterations along that arc: the depth order stays mostly reversed but keeps changing, and the solver is still improving slowly when its budget runs out. Whether that is a very flat valley, a slow slide back toward the true scene, or a continuous family of nearly equivalent answers rather than a single twin is not settled by the runs made here. What is settled is that along the standing arc the twin is a genuine minimum at every field from 1.5° to 25°.

Against marks read to a pixel

Exact marks make the twin look easy to reject: 0.474 px of misfit against 101310^{-13} is a difference of twelve orders of magnitude. Real marks carry their own error, and the twin only has to hide inside it.

Both answers on marks read to a pixelThe same two minima with every mark rounded to 1 px, along a 60° arc. The true answer settles where the rounding lets it, 0.325, 0.353, 0.333, 0.350 px; the inside-out answer at 0.541, 1.012, 1.901, 3.157 px, at fields of 1.5°, 3°, 6°, 10°. The margin between them is 0.216, 0.659, 1.568, 2.807 px. A solver that lands in the wrong one reports a residual larger than the right one's by that margin and no more.1.5°10°123field of view (degrees, log scale)reprojection error, marks read to 1 px (px)the true answerinside outmargin 0.216 px at 1.5° → 2.807 px at 10°marks read to 1 px
Fig. 5 Both answers with every mark rounded to a whole pixel, along the 60° arc. The true answer settles at 0.325, 0.353, 0.333 and 0.350 px; the twin at 0.541, 1.012, 1.901 and 3.157 px, at fields of 1.5°, 3°, 6° and 10°. The margin between them is 0.216, 0.659, 1.568 and 2.807 px.

Rounded to a whole pixel, the marks leave the true answer at 0.325 to 0.353 px, which is the floor where the adjustment stops predicted from a count of residuals and parameters. The twin settles at 0.541 px at 1.5°, 1.012 at 3°, 1.901 at 6° and 3.157 at 10°. The margins between them are 0.216, 0.659, 1.568 and 2.807 px.

Two things about those numbers matter. The margins are slightly smaller than the exact-mark misfits would suggest, because the rounding is partly absorbed by the twin as well as by the truth. And the margin at 1.5° is two-thirds of the floor itself. A solver that landed in the twin at that field would report a residual of 0.54 px, which is a residual somebody might easily accept from marks read to a pixel. Nothing but a comparison with the other minimum says it should have been 0.33.

Extrapolating the slope, the margin would fall to a tenth of the floor somewhere below a degree of field on this courtyard. That is an extrapolation and is quoted as one; nothing narrower than 1.5° was run.

Why seeing the scene does not rule it out

A two-view reconstruction has a discrete ambiguity that looks similar and is removed by a simple test. Four cameras fit, and one of them can see found that an essential matrix yields four camera pairs, all of which reproject the marks exactly, and that the right one is picked by requiring the scene to be in front of both cameras. Three of the four put the scene behind somebody.

The twin passes that test. Its points are reflected about the middle of the scene, not about the cameras, so every point that was in front of every camera is still in front of every camera — nearer or further than it should be, but ahead. At 120 m away the courtyard’s few metres of depth reflect to a few metres of depth the other way, and no point comes anywhere near the cameras. The cheirality test asks a question the twin answers correctly.

It is also not one of the ambiguities that come from the scene’s arrangement. The surface two pictures cannot separate measured a quadric on which two different motions give identical pictures, and far enough away, a pair is one eye measured the parallax a pair loses as the scene recedes. Both are about where the scene is relative to the cameras. The twin exists for a courtyard in general position, with cameras on a wide arc, and needs only that the cameras see it through a narrow field.

And it is not the gauge. Seven numbers no picture can name walked the reconstruction along seven directions that change no picture at all, and a reflection is not among them — it reverses handedness, which no rotation, shift or scale does. The twin is a separate answer, not another placement of the same one, and an uncertainty is quoted from something cannot report it: a covariance describes the valley the solver is in, and says nothing about the valley next door.

The twin at a wider field

The danger is specific to narrow fields, and a wider one shows why.

The courtyard and its inside-out twin, seen from above the first camera at a 25° fieldBoth answers the adjustment settles in, drawn in the first camera's own frame: across the picture horizontally, and depth toward the camera upward, with the second answer scaled to the same mean distance. The true courtyard's near points are the inside-out answer's far points, by up to 13.40 m, and the depths of the two correlate at -0.950. Both reproject exact marks: the true answer to 8.8e-14 px, the inside-out one to 7.528 px.each point joined to its place in the inside-out answer — nearer the camera is uptoward the cameratrue 8.8e-14 px · inside out 7.528 px25° field, 60° of arc
Fig. 6 The two answers at a 25° field along the standing arc, drawn as before in the first camera’s frame. The twin is still a minimum with its depth order reversed — the depths correlate at −0.950 — but it now misfits the exact marks by 7.528 px, a difference no residual would miss.

At a 25° field the twin still exists as a settled minimum, its depth order still reversed. But it misfits the marks by 7.528 px, and it is no longer an exact mirror.

The correlation between its depths and the true ones has weakened from −1.000 at 3° to −0.950, and points move by up to 13.40 m between the two answers where at 3° they moved by 8.66. Both changes say the same thing. At a narrow field the solver can reverse the scene almost rigidly, because a rigid reversal costs almost nothing; at a wide field a rigid reversal costs a great deal, and the minimum the solver finds is a compromise that is reversed in its overall depth order and bent away from a mirror image to recover some of the misfit. The twin a wide field keeps is a distorted twin, and the distortion is what the extra perspective bought. A reconstruction reporting seven and a half pixels of reprojection error on marks read to one would not be accepted by anyone looking at the number, and the solver that produced it would be restarted.

So the arrangement to be careful with is not every reconstruction. It is the one where both angles are small: a long lens, cameras that barely move relative to the subject, a subject far away. That is a common arrangement — telephoto sequences, aerial photographs of small areas from height, microscopes and anything else that looks at a small thing from far off — and it is exactly where the twin hides inside a residual that looks fine.

A test that costs one solve

The practical defence follows from how the twin was found. It was found by constructing it: reflect the reconstruction about its middle, reflect the cameras’ rotations, and adjust from there. That is one extra solve, and it produces the other minimum directly.

If the reflected start climbs back to the same answer, there was no twin in reach. If it settles somewhere else, both answers are now in hand, and the margin between their residuals is the evidence — measured against what the marks’ precision says the true residual should be, as the accounting in where the adjustment stops sets out. A margin much larger than that floor decides the question. A margin comparable to it says the pictures do not decide it, and the honest report is both reconstructions with a statement that the data cannot choose.

What the defence cannot do is make the choice when the pictures cannot. At a 1.5° field and whole-pixel marks the margin is 0.216 px; a solver can report it, but only information from outside the pictures — which surfaces face the camera, which way the light falls, what the scene is — can say which answer is the scene.

What this does not settle

One wrong start. The twin is the one wrong start with a known reason to be a valley. Other starts were not tried, and a general account of which starts fall into which valley is not attempted here.

One scene and one sequence. The courtyard’s depth is a few metres against its width of a few metres, and a flatter scene has less depth to reverse and a smaller misfit at every field. The constant of about 0.005 px per degree squared belongs to this courtyard.

The 20° arc. The failure to settle at wider fields along that arc is recorded above and not explained.

Rounding. As throughout, the marks’ error is rounding to a grid, which a chain and an adjustment and its neighbours describe as tidier than a real matcher’s; a less tidy error model would narrow every margin in the last figure.

Still open: which starts fall into the twin

The twin was reached by starting exactly on it. A real solver starts from an initialisation — a chain of pairs, as in the track and the scene together — that is somewhere between the true scene and its twin, closer to one or the other depending on how well the initialisation recovered depth.

The question that leaves is where the boundary between the two valleys runs. A start can be built as a blend: the true reconstruction moved a fraction of the way toward its reflection, depth by depth. Sweeping that fraction from nothing to all the way at a narrow field and a wide one, and recording which answer each start settles in, would find the fraction at which a start stops returning to the truth. Whether that boundary sits at halfway, or moves toward the truth as the field narrows — so that a slightly bad initialisation at a long focal length is enough to fall into the twin — is the measurement that decides how much a narrow-field pipeline should worry about its initialisation.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

bundle adjustmentDepth reversalfield of viewOrthographicReprojection errorResidual