The eighth held number bends the scene
Worth reading first: Seven numbers no picture can name · The track and the scene together.
A reconstruction from photographs knows its shape and not its place. Seven numbers no picture can name measured the gap exactly: three numbers of position, three of orientation and one of scale are free, and no quantity of pictures supplies them. The standard remedy is to bring them in from somewhere else — to survey a few points of the scene with an instrument that does measure position, and give the reconstruction those positions.
There are two ways to give them, and they are usually treated as two implementations of one idea.
The first adjusts the pictures alone and then fits a similarity from the reconstructed control points to the surveyed ones: seven numbers, taken from the survey by least squares, applied to everything. The second holds the surveyed coordinates fixed inside the adjustment itself, so the cameras and every other point are solved for around them.
They are not the same idea. The first chooses a frame. The second, as soon as it holds more than seven numbers, makes a statement about the scene’s shape — and a survey’s coordinates, like the pictures, carry errors of their own.
Seven held numbers choose a frame
An uncertainty is quoted from something held seven parameters of one reconstruction in five different ways and found that every one of them was a legitimate gauge. The reason each was legitimate is the reason seven is the number to watch here.
The pictures determine a reconstruction up to a similarity, which means the set of all reconstructions consistent with them is a seven-dimensional family: one shape, placed anywhere, turned any way, at any size. Seven held numbers, chosen so that they are not redundant with each other, pick exactly one member of that family. Any values of them do — the whole point of a family indexed by seven numbers is that every setting of the seven is in it.
The control points in the courtyard make that concrete. Holding all three coordinates of one surveyed point fixes where the reconstruction is. Holding all three of a second fixes which way the line between them points and how long it is, which is two orientations and the scale. Holding one coordinate of a third stops the last rotation, about the line through the first two. Three, three and one: seven.
If the survey got those coordinates wrong, the reconstruction simply lands at a slightly different place, turned and scaled slightly differently. Its shape is the shape the pictures made, because nothing asked it to be otherwise.
The seven are a datum
Surveying has a name for the fitted procedure, and the name says what kind of thing the seven numbers are. Carrying coordinates from one reference system to another — a local site grid onto a national one, one national datum onto another — is done with a seven-parameter Helmert transformation: three shifts, three rotations and a scale, fitted by least squares to points whose coordinates are known in both systems. The two systems are each internally consistent, each carries its own errors, and the transformation between them is fitted rather than imposed, precisely so that neither system’s errors are forced into the other’s shape.
A reconstruction from pictures is, in that language, a network in a datum of its own choosing. The track and the scene together solved for every camera and point at once and left the result wherever the starting point had put it, and that position, orientation and size are the reconstruction’s private datum. A survey is a second network in a second datum. Fitting a similarity between them is a datum transformation. Holding the survey’s coordinates inside the adjustment is something else: it declares the two networks to be one network, and declares the survey’s version of the shared points to be exact.
It is worth noticing how rarely the seven numbers arrive from a survey at all. Two views give shape and no size took the scale from a single measured length. Five facts that close the same gap catalogued the facts that can supply it from inside a photograph — a camera height, an object of known size, a focal length with a horizon — and found their precisions differ tenfold. A sequential pipeline takes all seven from its first pair of cameras, as a chain and an adjustment describes, and never revisits them. Every one of those supplies exactly seven numbers, or fewer, and none of them can bend a shape. Control points are the case where more than seven arrive at once, and that is the case where the difference between fitting and holding stops being a matter of bookkeeping.
The eighth makes a claim
Hold one more coordinate and that stops being true.
Seven held coordinates change the shape by mm, which is the arithmetic floor and is nothing. The eighth changes it by 0.591 mm.
The mechanism is worth stating carefully because the obvious account is wrong. The obvious account is that the eighth coordinate is erroneous and the error enters the reconstruction. But the eighth coordinate held here — the third control point’s position across the courtyard — happens to have been surveyed exactly. Its survey error was entirely in its height. What is wrong is not the eighth number but its relationship to the first seven. The first seven, with their errors, already placed, turned and scaled the reconstruction; that placement puts the third point’s across-coordinate somewhere; the survey says it is somewhere else, by a few millimetres; and there is no similarity that satisfies both. The adjustment has to satisfy both, because both are held, so it changes something that is not a similarity. It changes the shape.
The rest of the bars say that the count is not what matters. The ninth held coordinate raises the change to 0.782 mm, the tenth lowers it to 0.718, the eleventh raises it to 0.769 and the twelfth lowers it to 0.595. There is no trend, because each added coordinate brings its own disagreement with the rest, and disagreements in different directions can partly cancel.
So the useful statement is not “more control bends more”. It is that seven is a threshold: below it a survey’s errors are absorbed into where the reconstruction sits, and above it they have nowhere to go except into its shape.
In proportion to the survey
With the count settled at all twelve coordinates of four points, the size of the bend can be swept against the size of the survey’s error.
Two of the three lines lie on zero for the whole sweep. The fitted similarity never touches the shape, whatever the survey says, because a similarity cannot. The seven held coordinates never touch it either, for the reason above.
The third line is straight. Twelve held coordinates bend the courtyard by an amount proportional to the survey’s error: 0.120 mm at 2 mm, 0.595 mm at 10, 2.899 mm at 50 — about 5.8 per cent of what the survey got wrong, spread through a scene several metres across. Linearity is what a small inconsistency in a least-squares problem should produce, and it means the percentage is the useful number: on this courtyard, with these four points and these directions of error, a twelve-coordinate hold passes a little under a sixteenth of the survey’s error into the shape.
A sixteenth sounds small, and whether it is small depends entirely on what the reconstruction is for. It is 2.9 mm of distortion from a survey that was 50 mm out. It is also 2.9 mm of distortion in a reconstruction whose pictures were exact, which is to say distortion that no amount of care with the photographs would have prevented, introduced by the one step that was supposed to make the reconstruction more accurate.
What a fitted similarity says instead
The fitted similarity does not throw the survey’s error away. It reports it, in the one place where a report can be read.
Ten millimetres of error went into each of the four control points, each in its own direction. After the fit, the four residuals are 8.59, 7.01, 3.23 and 9.51 mm, and their root-mean-square is 7.48 mm.
Every one is smaller than the error that caused it, and that is the correct behaviour rather than a flattering one. A survey error at four points is a pattern of twelve numbers. Part of that pattern is itself a similarity — a common shift, a slight turn, a slight change of size — and the fit absorbs that part by moving the reconstruction, exactly as it should, since a survey that is merely shifted is a survey in a different frame. What is left over cannot be a similarity, and the fit leaves it as residuals. How much of each point’s error is absorbable depends on where the point sits relative to the other three and on the direction its error happens to take, which is why the four residuals differ by a factor of three although the four errors were the same size; the third point, whose error was a change of height, keeps the least of it.
The practical difference between the two procedures is therefore not that one is accurate and the other is not. It is where each puts a survey’s inconsistency. Holding the control puts it into the shape of the reconstruction, silently. Fitting a similarity puts it into four numbers beside four surveyed points, where a surveyor can look at them, notice that one control point is 9.51 mm from where the pictures say it is, and go and check that point.
When the pictures object
A held survey that bends the reconstruction does leave a trace in one place: the pictures no longer fit as well.
With exact marks the trace is perfect. Every bit of the reprojection error is the survey’s, and it rises from zero to 0.957 px at 50 mm of survey error. A reprojection error of a pixel on exact marks is a loud objection.
Real marks are not exact, and where the adjustment stops measured what reading to a whole pixel leaves behind: 0.332 px on this courtyard with no control at all, a floor set by the rounding and predictable from a count of residuals and parameters. The held control has to raise the reprojection error above that floor before anyone can see it. It raises it by a tenth — to a level a careful reader of the residual might notice — only once the survey is 10 mm out.
Below that the survey’s disagreement with the pictures is buried in the pictures’ own noise, and the residual, as a diagnostic, cannot say anything about it. That is the same shape as the silent failures that essay lists: a residual that sits where it should while something the residual cannot see has gone wrong. What is different here, and the next figure is about it, is that the silence below 10 mm is not all bad news.
When holding the control is right
Everything so far has used exact pictures and treated the survey as the only source of error, which is the arrangement that makes holding control look worst. With marks read to a pixel, the pictures have errors too, and a good survey knows things the pictures do not.
The fitted similarity leaves the courtyard 4.67 mm from its true shape at every survey error. That is the pictures’ own shape error, and a similarity cannot improve on it any more than it can worsen it.
Holding the control does better while the survey is good: 4.42 mm with a perfect survey, 4.30 mm with one 20 mm out. The reason is that an accurate surveyed position is information about the scene that the rounded marks do not contain, and holding it lets the adjustment use that information to pull the reconstruction’s shape toward the truth. The two lines cross near 50 mm, where the held reconstruction’s shape error is 4.68 mm against the fit’s 4.67. Past that the survey’s inconsistency outweighs its information: 6.27 mm at 100 mm of survey error, 10.41 mm at 200.
So neither procedure is right in general, and the figure says what the choice depends on. Holding control is worth it when the survey is more accurate than the pictures, in the specific sense of how accurately each pins the scene’s shape. Fitting a similarity is safer when it is not, or when nobody knows. On this courtyard the pictures pin the shape to about 4.7 mm and the crossover is at a survey error roughly ten times that — a ratio that belongs to this arrangement, with its four control points and its fixed directions of error, and should not be carried anywhere else as a rule.
The alarm figure and this one together say something slightly uncomfortable about the residual as a guide. Between 10 mm and 50 mm of survey error the held adjustment’s residual has risen measurably above its floor, which reads as an objection to the control, while the held control is still making the shape better. The residual reports disagreement between the survey and the pictures; it does not report which of them is closer to the truth, and it cannot.
Why this is easy to miss
Three habits make the difference between the two procedures invisible in practice.
Control is described in points, not in numbers. “Four control points” does not sound like twelve constraints on a seven-dimensional freedom, and the question of whether five of those constraints are consistent with the pictures does not arise in that language. Counting in numbers is what makes the threshold at seven visible at all.
The residual looks fine. Below about 10 mm of survey error on this courtyard, a held adjustment’s residual is indistinguishable from an unheld one’s. A pipeline that checks convergence by the residual reports success in both cases. It is the same trap a wrong match is not a small error found for a single bad correspondence and fitting a lens from straightness alone found for an over-parameterised lens: a model that has absorbed something it should not have, reporting a residual that says nothing about it.
The bend is smooth. Two point nine millimetres spread smoothly across a courtyard does not look like an error in a drawing of the reconstruction. It looks like a reconstruction. Nothing in the shape itself says that part of it came from a survey rather than from the pictures.
What this does not settle
One arrangement. The courtyard, six cameras along one arc, four control points, and survey errors in fixed directions chosen not to be parallel. Different directions give different percentages and a different crossover, and the non-monotonic bars in the count figure are a reminder that the size of the bend depends on how the errors line up, not only on how large they are. What carries over is the threshold at seven and the proportionality in the survey’s error, both of which follow from the structure of the problem rather than from this scene.
Hard constraints only. Every held coordinate here is held exactly. A survey adjustment would usually give each control coordinate a weight reflecting how well it was surveyed, and let the pictures and the survey pull against each other in proportion to their stated accuracies. That procedure sits between the two measured here, and the crossover in the last figure is the strongest argument that it should exist. It is not measured.
The pictures’ error is rounding. As in every figure in this field, the marks are read to a grid, which is bounded and independent; a real matcher’s error is neither, and the pictures’ own shape error would be larger and less tidy.
Still open: how much weight a survey deserves
The two procedures measured here are the two ends of one dial. Holding control gives it infinite weight against the pictures; fitting a similarity after the fact gives it only enough to set seven numbers and none to affect the shape. Every intermediate setting is a weighted adjustment in which each surveyed coordinate carries a stated standard deviation and pulls the reconstruction toward its surveyed value in proportion.
The crossover in the last figure says that the best setting is somewhere in the middle, and that where it is depends on the relative accuracy of survey and pictures. The question that leaves is whether the textbook answer is right: whether weighting each control coordinate by the inverse of its survey variance, and each mark by the inverse of its reading variance, actually lands on the smallest shape error when the two are swept against each other on this courtyard — or whether a survey error that is systematic rather than random, like the fixed directions used throughout here, moves the best weight somewhere the textbook does not predict. Measuring that means a weighted adjustment across a range of stated survey accuracies, with the true shape error read beside each, and the answer is a statement about how much a survey should be trusted by the pictures it is meant to correct.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The midpoint is a choice of ruler — both name gauge freedom, reprojection error, residual, similarity
- A narrow view keeps a second answer, inside out — both name bundle adjustment, reprojection error, residual
- A fold names the height — both name gauge freedom, residual
- A start needs the sign of its depths, not their size — both name bundle adjustment, reprojection error
- A third ray is worth what its picture is worth — both name reprojection error, residual
- A turn of the head is not a step sideways — both name reprojection error, residual
Named objects
A flat tag is an object no other essay names yet.
bundle adjustmentControl pointgauge freedomReprojection errorResidualSimilarity