Many pictures at once

Another picture of the same sweep

Going from three views to seven across the same sixty degrees leaves the reconstruction exactly where it started, and at one point makes it worse. What a reconstruction is short of is angular spread, not photographs.

Worth reading first: The track and the scene together · Seven numbers no picture can name.

The straightforward expectation about a reconstruction is that more photographs make it better. Each one adds observations, observations constrain parameters, and a well-posed least-squares problem improves as the square root of the count.

The figure was built to show that, with the flat count staying at seven alongside it so that “more data buys precision and never buys the gauge” could be said with two curves rather than one.

It does not show it.

Another picture of the same sweep buys nothingTwo ways of adding views to the courtyard, every mark read to 1 px. Filling in a fixed 60° sweep leaves the worst camera-centre error at 2.0e-3 where 3 views gave 1.1e-3. Widening the sweep by 12° per view improves it from 3.2e-3 to 1.3e-3 and then flattens as well. What the reconstruction is short of is angular spread, not pictures. At every point on both curves the Jacobian has exactly 7 flat directions.-2.80-2.6034567number of viewsworst camera-centre error (fraction of the track's mean radius, log scale)7 flat7 flat7 flat7 flat7 flat60° sweep, filled in12° per view, wideningfilled in: 1.1e-3 → 2.0e-3widened: 3.2e-3 → 1.8e-3
Fig. 1 Two ways of adding views to the courtyard, every mark read to 1 px. Filling in a fixed 60° sweep leaves the worst camera-centre error at 2.0e-3 where 3 views gave 1.1e-3. Widening the sweep by 12° per view improves it from 3.2e-3 to 1.3e-3 and then flattens as well. What the reconstruction is short of is angular spread, not pictures. At every point on both curves the Jacobian has exactly 7 flat directions.

The two ways of adding a view

The distinction the figure turns on is between adding a photograph and adding a viewpoint, and they are not the same thing.

Filling in. Keep the sweep at sixty degrees and put more cameras along it. Three views become seven, spaced more closely, over the same arc.

Widening. Keep the angular step at twelve degrees and let the sweep grow. Three views cover twenty-four degrees; seven cover seventy-two.

Both add exactly the same number of photographs and the same number of observations. They differ only in where the new cameras stand.

Filling in leaves the worst camera-centre error where it started — 1.1 × 10⁻³ at three views, 2.0 × 10⁻³ at seven. Widening improves it, from 3.2 × 10⁻³ at three views to 1.3 × 10⁻³ at four, and then flattens too.

How far apart the eyes were, read out of the two picturesThe eight-point system's second-smallest singular value against the distance between the cameras, over a range of 100×. The fitted slope is 0.982 — proportional — with every point within 0.016 of a decade of the line. Nothing in the computation knows where either camera was: this is a distance in the world, recovered from ink.-4-3.50-3-2.50-2-2-1.50-1-0.500baseline between the eyes (m, log scale)σ₈ / σ₁ of the design matrix — computed from the two pictures aloneslope 0.98log–log slope 0.982, worst residual 0.016 decades8 mm to 0.8 m
Fig. 2 The two-view quantity this essay’s finding reduces to, over a hundredfold range of separation. Filling in a sweep shortens every consecutive baseline, which moves every pair down this line.

The same statement, in the two-view field’s language

The finding has a translation into the previous field’s vocabulary that makes it sound less surprising, and the translation is worth making because it shows the two fields measuring one thing.

The two-view field’s central conditioning result is that the second-smallest singular value of the eight-point design matrix is proportional to the baseline. A short baseline is a nearly degenerate configuration, and its reconstruction is dominated by the reading error rather than by the geometry.

Filling in a sweep shortens every consecutive baseline. So each pair in a filled-in sequence is worse conditioned than each pair in a sparse one covering the same arc, and the fact that there are more of them does not compensate — because what the extra pairs measure is what the existing ones already measured.

Widening the sweep lengthens the baselines. Each pair is better conditioned, and the additional pairs measure directions that were not previously represented.

So the many-view result is the two-view result summed over a sequence, and the quantity it turns on is the same: the angle between viewpoints, compared against the precision of the marks. That is the parallax angle again, which is the quantity every triangulation from any number of views is really reading.

Why filling in buys nothing

A camera placed between two existing ones sees very nearly what they see.

Its observations are not independent measurements in any useful sense: the geometry it constrains was already constrained, from directions barely different from its own. It adds six parameters — its own pose — and adds observations that are close to redundant with observations already present, so the ratio of information to unknowns barely moves.

Worse, filling in a fixed sweep shortens every consecutive baseline. Three cameras across sixty degrees are much further apart than seven are. So while the count of pictures rises, the quantity that actually determines depth for each neighbouring pair falls, and the two effects work against each other.

That is the mechanism, and it explains the mild worsening as well as the flatness. The worst-case camera error is a worst case over more cameras as the count rises, so a metric that takes a maximum over a growing set has a mild upward bias built in — and there is nothing pushing back on it.

3 pictures, one track and one scene, recovered togetherEvery point read to 1 px. Bundle adjustment moves all 150 parameters at once and lands at 0.295 px of reprojection error in 6 iterations. The recovered camera centres, drawn over the true ones after the similarity the pictures cannot fix is removed, agree to 1.1e-3 of the track's own mean radius. Nothing in the solve was ever told where a camera or a point was.view 1view 3the scene, in plan — recovered points and cameras over the true onesreprojection 0.295 px · track 1.1e-3132 observations, 150 parameters
Fig. 3 Three views of the courtyard, which is where both curves start. The track is recovered; what the next four pictures do to it depends entirely on where they are taken from.

Why widening helps, and then stops

Widening improves the reconstruction once, substantially, and then flattens.

The improvement from three views to four is from 3.2 × 10⁻³ to 1.3 × 10⁻³, and the reason is that three views over twenty-four degrees is a genuinely narrow sweep. The needle-shaped uncertainty of every triangulated point is long in the view direction, all the view directions are similar, and nothing cuts the needles crossways. Widening to thirty-six degrees introduces a direction that does.

After that it flattens, and the flattening deserves its own explanation because it is not obviously the same effect as before.

Once the sweep is wide enough that every point is seen from directions genuinely far apart, the reconstruction’s error is limited by how well the individual marks were read rather than by the configuration. Widening further keeps improving the geometry and starts losing observations — points that were visible from every camera fall out of the frame at the extremes, or become invisible behind blocks — and the two roughly cancel.

So both curves flatten, for related but distinct reasons, and neither of them is the square-root improvement that counting observations would predict.

Seven moves that change no picture, and three that change every oneThe whole reconstruction shifted by up to 1.3 m, turned by up to 0.55 rad and scaled by 2.7: every picture stays where it was, to 1.6e-11 px. Move one camera by 50 mm or one point by 50 mm and the pictures move by 0.63 px or more. The seven are not small effects that could be measured with better data; they are exactly zero, and the 4e+10× between the two groups is what makes that a claim rather than a tolerance.shift x1.0e-11 pxshift y1.0e-11 pxshift z1.6e-11 pxturn x2.8e-12 pxturn y8.8e-12 pxturn z1.0e-11 pxscale ×2.71.0e-11 pxturn camera 0 by 0.01 rad1.112 pxmove camera 0 by 50 mm2.529 pxmove point 0 by 50 mm0.627 pxchange in reprojection error, log scalebelow: not gauge directionsflat to 1.6e-11 px · stiff from 0.63 px4e+10× apart
Fig. 4 The quantity that does not move as views are added. Five views, the same seven flat moves and the same three stiff ones, separated by the same enormous factor.

The metric, and what it hides

A word about what “worst camera-centre error” measures, because the choice of statistic is doing real work and a different one tells a slightly different story.

It is the largest disagreement between a recovered camera position and the true one, after the similarity has been removed, expressed as a fraction of the track’s own mean radius. A maximum rather than a mean, and a fraction rather than a length.

The maximum was chosen because it is the quantity a reader of a reconstruction cares about — one badly-placed camera is a problem whether or not the others are fine — and because a mean would flatter the wider sweeps, where the extreme cameras are the ones with the fewest observations.

It has a known bias, and it is the one named above: a maximum over a growing set drifts upward even when nothing degrades. Seven cameras give seven chances to be the worst; three give three. That accounts for part of the mild rise in the filling-in curve and none of its flatness — the mean behaves the same way, staying between eight and twelve parts in ten thousand across the whole range, which is to say also flat.

Quoting the maximum and knowing its bias is the honest arrangement. Quoting the mean would have made the curve look flatter and the claim stronger, and it would have hidden the fact that adding cameras adds opportunities for one of them to be poorly placed.

What the count would have predicted

It is worth doing the naive calculation to see how wrong it is.

Going from three views to seven roughly doubles the observation count. If the errors were independent and the parameters fixed, that predicts an improvement of about a factor of 1.4 in the reconstruction.

The parameters are not fixed: each new view adds six of its own. So the honest version is that observations grow by 4 × 44 while unknowns grow by 4 × 6, which is still strongly in favour of the observations, and still predicts a substantial improvement.

Measured, filling in gives none. The prediction fails because it treats observations as independent measurements of the same quantities, and observations from a camera nearly coincident with an existing one are not.

This is the same shape as a mistake this site recorded in the metrology field. The conditioning of the height recovery was computed correctly, in the cross-ratio, and the conclusion drawn from it was false, because the quantity anybody actually perturbs is a pixel position rather than a cross-ratio and the two do not vary together. Here the count of observations is correct and the conclusion drawn from it is false, because near-duplicate observations are not what the count is a proxy for.

5 pictures, one track and one scene, recovered togetherEvery point read to 1 px. Bundle adjustment moves all 162 parameters at once and lands at 0.318 px of reprojection error in 5 iterations. The recovered camera centres, drawn over the true ones after the similarity the pictures cannot fix is removed, agree to 1.4e-3 of the track's own mean radius. Nothing in the solve was ever told where a camera or a point was.view 1view 5the scene, in plan — recovered points and cameras over the true onesreprojection 0.318 px · track 1.4e-3220 observations, 162 parameters
Fig. 5 Five views of the courtyard, filled in across the same arc. The recovered track sits on the true one; what the two extra cameras over three did not do is make it sit closer.

What a further view does buy

Nothing above says a redundant view is worthless, and it is worth listing what it does buy, since the measurement only covers one thing.

Robustness. A point seen by seven cameras can have one of its observations be wrong and still be well determined by the other six. A point seen by two cannot: an outlier there contaminates everything and there is no majority to detect it against. That gain does not show in a measurement made with correct correspondences throughout, which is what these figures use, and it is real.

Coverage. A view added between two others sees parts of the scene that the geometry hid from both — the far side of a block, a face that was edge-on. Those points get reconstructed where before they were not reconstructed at all, which is an improvement that a metric over the cameras cannot see, because it is about which points exist rather than about how well the cameras are placed.

And the ability to detect its own failure. A point observed many times has a residual per observation, and the spread of those residuals is a diagnostic. With two observations there is nothing to compare, and a wrong match is a fact with no counter-evidence.

So the accurate summary is that a filled-in sweep buys robustness and coverage and buys no accuracy, and the figure measures the third because the third is the one that gets claimed. A practitioner adding frames to a pan for the sake of robustness is doing something sensible; one adding them expecting the camera positions to sharpen is not.

The seven, at every count

The other curve in the figure does not move at all.

At three views the Jacobian has seven flat directions; at four, five, six and seven, likewise, with a gap of at least a million between the seventh singular value and the eighth in every case.

That was the figure’s original purpose and it survived intact when the first claim did not. It is the sharpest available statement of the field’s central result: the similarity freedom is not a shortage of data, because here is data being added, in two different ways, with the freedom not moving.

Setting the two readings side by side gives the field’s summary in one line. More pictures may buy precision, depending entirely on where they are taken from, and never buy the gauge.

The measurement that changed the essay

This is the third time in this phase that a figure was built to show a standard result and measured something else, and the pattern across the three is worth recording because it is a pattern about method rather than about geometry.

The normalisation figure was to show that Hartley normalisation is worth an order of magnitude. At the operating point first chosen it was worth 0.53× — the raw fit better — and the honest result is a curve whose gap opens as the data worsens.

The chain figure was to show errors accumulating along a sequence. They wander instead, and one chain ends closer to the truth than its own worst link.

This one was to show accuracy improving with the number of views. It does not, unless the views are somewhere new.

In all three cases the claim being tested is true in the right regime and was being asserted as a property of the method. Normalisation matters when the data is poor. Chained errors accumulate over hundreds of links. More views help when they widen the sweep. Each of those is a real statement with a condition attached, and each had lost its condition somewhere between the derivation and the folklore.

The mechanism that caught all three was the same, and it is worth naming because it is cheap. Every figure on this site asserts what it claims, in code, at the numbers it actually draws. An assertion written to match the expected result fails when the expected result is not what happens, and a failing assertion is not an inconvenience — it is the only part of the process that has no opinion about what ought to be true.

What this means for taking pictures

The practical version is short and slightly counterintuitive.

Photographs of a subject, taken to reconstruct it, are worth their spread rather than their number. Twenty frames of a slow pan buy little more than five. Five frames from genuinely different places buy a great deal more, and the improvement comes almost entirely from the first few that introduce a direction not already present.

Against that, spread costs overlap. Cameras far apart share less of the scene, points visible in one are hidden in another, and matching between widely separated views is harder in ways this site does not measure. The flattening of the widening curve is partly that cost arriving.

So the shape of the advice is: spread first, until every part of the subject is seen from directions that differ substantially, and then add frames only where the coverage is thin. Which is a rule about where to stand rather than about how many times to press the shutter, and is the same rule this site has been making in every field — the position of the eye is the content of the picture, and adding pictures from where the eye already was adds nothing.

A chain of pairwise poses wanders rather than driftingFour chains through the same 7 views, differing only in how finely each mark is read. Every individual link is recovered to better than 2.35°. The composed pose does not grow steadily away from the truth: it rises and falls, and the 4 chains end between 0.53× and 1.17× their own worst link — one of them closer to the truth at the end than at any single step. Consecutive links have no reason to err in the same direction, so composing them is a walk rather than a sum.00.50011.500123456view along the sequencerotation error of the composed pose (degrees)0.5 px1 px2 px4 pxend of chain against worst link: 0.64×, 1.17×, 0.53×, 0.62×7 views, 6 links
Fig. 6 And the last of the three fields’ recurring shapes. Seven views, four chains, and a composed error that goes up and down rather than away — measured because the assertion that said otherwise refused to pass.

Where the three fields end

This is the last essay of the three fields this phase opened, and they close on a smaller claim than they opened with, which is the right direction for a set of measurements to move.

The opening claim was large: two pictures of a scene contain the other camera, the whole three-dimensional shape, and — with a supplied length — a survey. Every part of that was measured and every part held, to between nine and fourteen digits.

What the fields spent most of their length on is the qualifiers. The shape comes without a size. The pose comes in four candidates and needs a physical assumption to choose. The depth comes as an interval that is not symmetric and past a computable range is not bounded. The reconstruction comes up to a similarity, seven numbers, and no quantity of photographs removes any of them. And the arithmetic that produces all of it fails, in at least four distinct ways, without producing an error.

Read together, those are not a list of caveats attached to a method. They are the method’s actual output. A recovery that reports what it determined and stays silent about what it did not is more useful than one that reports everything and is quietly wrong about some of it — and the difference between the two is entirely a matter of having computed the quantities that tell them apart.

Which is where the next phase starts, and it is a different subject: drawing systems that were never trying to be projections through a centre at all.

What links here

Computed from the collection, not written here: the essays that point at this one.

Named objects

A flat tag is an object no other essay names yet.

Baselinebundle adjustmentcamera trackConditioninggauge freedomRedundancy