An uncertainty is quoted from something
Worth reading first: Seven numbers no picture can name · The track and the scene together.
A reconstruction from photographs usually arrives with error bars. The solver that produced it has a Jacobian in hand, the Jacobian gives a normal matrix, the normal matrix inverted gives a covariance, and the square roots of its diagonal are printed beside every camera position and every point: this one to 5 mm, that one to 150.
Those numbers look like properties of the reconstruction. They are properties of the reconstruction and of a choice nobody printed, and the choice changes them by a factor of forty on the same pictures, the same marks and the same solution.
The choice is the gauge. Seven numbers no picture can name found that a reconstruction from many views is free along exactly seven directions — three of position, three of orientation, one of scale — and walked a whole courtyard a metre and a half, half a radian and a factor of 2.7 along them without moving a single mark. A covariance cannot be written down until those seven are removed, and there is no one way to remove them.
Why there is no covariance until something is held
The reason is arithmetic, and it is worth having in full because everything afterwards follows from it.
The covariance of a least-squares estimate is the inverse of the normal matrix scaled by the variance of one residual, . For the courtyard, has 528 rows and 168 columns, and is singular by exactly seven: the seven flat directions are directions in which the residuals do not change, so they are null vectors of , so they are null vectors of . A singular matrix has no inverse. The formula every textbook gives cannot be evaluated.
So every solver that reports a covariance has done something first, whether or not it says so. The commonest thing is to hold seven parameters at their current values and invert the normal matrix of the other 161. Holding the first camera’s six pose parameters removes three positions and three orientations; holding one component of the second camera’s position removes the scale. What is left is non-singular and inverts cleanly.
The covariance that comes out is a genuine covariance, correctly computed. It is the covariance of every parameter relative to the ones that were held. The first camera reads zero because it was told to. Every other camera reads its uncertainty as a displacement from a camera that cannot be displaced, measured in a scale set by a baseline that cannot be wrong, and the further a camera is from the held one the more of the held camera’s own uncertainty it inherits.
That is what the first figure shows. Its six cameras, read along the track, carry 0.0, 22.2, 43.9, 64.7, 85.1 and 104.8 mm. The growth is almost exactly linear in position along the arc, which is the signature of an error in orientation and scale at one end being carried out along a lever.
Nothing in the pictures says the far end of the track is worse determined than the near end. The six photographs are of one courtyard from one arc at one spacing, and the marks were rounded to the same grid in all of them. The slope is the gauge’s.
Hold the other end
The quickest demonstration is to hold the other end.
The picture is close to a mirror image of the first, and it has to be. The arc is symmetric and the spacing is uniform, so holding either end puts the held camera at one tip of a lever and the uncertainty at the other. The numbers do not match one for one — the sequence climbs 0.1 m a view and the courtyard is not symmetric about the middle of the arc, so neither end is the other’s reflection — but a reader handed only one of the two figures would draw the same conclusion from either: one end of the track is good, the other bad.
Handed both, the conclusion falls apart. The same camera is the best-determined in one report and the worst in the other, on identical data. Neither report is wrong. They answer two different questions — how far is each camera from the first, how far is each camera from the last — and neither of them is the question a reader thinks the table answers, which is how well the pictures determine each camera.
Two gauges that both hold the first camera
Saying “the reconstruction is expressed in the first camera’s frame” does not even name a single gauge, and the difference is large.
Six of the seven held numbers are the first camera’s pose. The seventh fixes the scale, and it has to be some length. Taking it from the first baseline — the step from the first camera to the second — uses a length a fifth of the track. Taking it from the step between the first camera and the last uses the whole track. Both are the first camera’s frame.
The points go from 52.1–158.0 mm to 4.9–11.5 mm. Same held camera, same pictures.
The mechanism is the lever again, turned around. A scale held on a short baseline is a scale whose error, whatever it is, is multiplied by the ratio of every other distance to that baseline. The courtyard is several baselines across, so a small relative error in how long the first step is becomes a large absolute error at the far side. Holding the scale on the longest length available divides that multiplier out.
The practical form is a sentence that should accompany any table of positional uncertainties and almost never does: which length was the scale taken from. A reader who knows only that “the first camera was held” cannot tell which of these two tables they are looking at, and the two differ by a factor of ten.
Holding nothing in particular
There is a way to remove the seven without holding any parameter at all.
The normal matrix cannot be inverted, but it has a pseudo-inverse: diagonalise it, invert the 161 eigenvalues that are not zero, and leave the seven that are zero at zero. The covariance that results assigns no uncertainty along the flat directions and treats every camera and every point alike. In survey adjustment it goes by the name of inner constraints, and its defining property is that it favours nothing.
The picture is the one a reader expects a reconstruction to have: uncertainty spread evenly across the track, somewhat larger at the points furthest from every camera, and nothing exact.
It is tempting to call this one the true covariance, and it is worth resisting. It does have a distinction, measured here: of the five gauges drawn in this essay, it gives 42 of the 44 points and every camera not held elsewhere its smallest uncertainty, and the two exceptions are the points another gauge held at zero by decree. But smallest is not truest. Its uncertainties are still measured from something — from the centroid, the mean orientation and the mean size of the whole reconstruction, held still in an average sense rather than at any one camera or point. It favours nothing because it holds everything a little, and a reader who wants to know how far the last camera is from the first is not answered by it at all.
The same cameras four ways
Set side by side, the camera uncertainties under four gauges do not resemble each other.
The third curve is the one survey practice would produce. Three points of the courtyard are held — two of them entirely, and one coordinate of a third — which is what fixing a reconstruction to ground control does. The cameras now read 88.3 to 135.4 mm, all of them large, because every camera is a long way from those three points compared with the distance between the two that set the scale. The held points themselves read zero.
One more comparison sits under the figure without being drawn. Reading the marks to half a pixel instead of a whole one halves every uncertainty: the worst camera under the first gauge goes from 104.8 to 53.2 mm, and the inner gauge’s spread from 4.2–5.7 mm to 2.2–2.9 mm. Better marks help, exactly as where the adjustment stops found for the residual, and they help every gauge in proportion. They do nothing to the ratio between gauges. A factor of twenty between two reports of the same camera is not a symptom of poor data that more care would remove.
What does not move
If every coordinate’s uncertainty depends on the gauge, it is fair to ask whether any uncertainty does not. Some quantities do not, and the figure that reads them is the centre of the essay.
A ratio of two distances between points carries a standard deviation of 0.1465 per cent under all four gauges, and the four agree to two parts in a billion — which is as closely as the numerical derivatives feeding them agree with each other. An angle between three points carries 0.0778° under all four. A single coordinate of a single point carries anything from 1.48 mm to 65.14 mm.
The agreement is not a coincidence of the courtyard and it is not approximate. It is an identity, and the argument for it takes three lines.
Any two gauges’ covariances differ only along the seven flat directions. Write those directions as the columns of a matrix , so that . Every gauge’s covariance can be written as , where is the inner-constraint covariance and is a projector that moves a parameter vector along the columns of until the held parameters are satisfied. A quantity that does not change under a similarity has a gradient with no component along any flat direction, , so , and its variance
is the same whichever projector was used. A coordinate is not invariant — shifting the reconstruction changes it — so its gradient does have a component along , the projector does act on it, and its variance depends on the gauge.
So the figure divides the quantities a reconstruction reports into two kinds, and the division is exact. Ratios of lengths, angles, shapes: an uncertainty that belongs to the pictures. Coordinates, positions, distances in metres, heights, orientations relative to an axis of the reconstruction’s own frame: an uncertainty that belongs to the report.
This is the same division seven numbers no picture can name drew for the values of those quantities, now drawn for their uncertainties. The value of a coordinate is a statement about where the solver started; its uncertainty is a statement about what the solver held. The value of a ratio is a property of the scene; its uncertainty is a property of the pictures.
A table of point uncertainties
The sorted uncertainties of all forty-four points make the scale of the difference visible at once.
Three features are worth reading off it.
The curves do not cross in a consistent order. Some points are better determined in the scene gauge than in the inner one, most are worse, and two are exact. A ranking of points by uncertainty — which parts of a reconstruction are “good” — is gauge-dependent too, not just the magnitudes.
The spread within a gauge differs. Under inner constraints the worst point is 2.6 times the best; under the first-camera gauge it is 3.0 times. A reader looking only at the spread would call the first report more uneven, when the unevenness comes from the lever and not from the points.
And the level differs by a factor of twenty between the first-camera and inner curves, with nothing else changed. A reconstruction advertised as good to 5 mm and one advertised as good to 150 mm can be the same reconstruction.
Why the anchored report is the usual one
Given all that, the first-camera gauge is still what most pipelines print, and the reason is historical rather than careless.
A two-view reconstruction is conventionally written in the first camera’s frame — two views give shape and no size is written that way, with the baseline set to unit length — because the first camera is the natural origin and the baseline the only length available. A sequential pipeline then grows the reconstruction from that first pair, one camera at a time, and the first camera stays where it was put. By the time an adjustment runs, the frame has been inherited from the initialisation, and the covariance is computed in it without anybody deciding that it should be.
For some purposes it is exactly the right report. Anyone who needs to know how far each later camera might be from the first one — a vehicle asking how far it has travelled since it started — is asking the anchored question, and the anchored covariance answers it. The failure is not in computing it but in printing it as though it were the uncertainty of the reconstruction, and then comparing it with a survey’s error ellipses, which are written in a gauge fixed at ground control, or with another reconstruction’s, which was written in its own first camera.
A chain and an adjustment ran into the same trap from the solver’s side: comparing a recovered camera track with the true one gives metres of error on a perfect reconstruction unless the similarity is removed first. Comparing two covariances has the same requirement one level up, and it is much less often met, because a covariance looks like a measurement of quality rather than a quantity with a frame.
What to report instead
Three habits remove the trap, and they are not equally expensive.
State the gauge. The cheapest fix is a sentence: which parameters were held, and which length set the scale. It does not make two reports comparable, but it makes their incomparability visible, and a reader can then decide whether the question they have is the anchored one.
Report invariant quantities. Where the reconstruction is going to be used for shape — the proportions of a building, the angle between two walls, the ratio of two spans — report those quantities and their uncertainties directly. Those uncertainties are gauge-free by the identity above, and they are what five facts that close the same gap calls the part of a measurement the pictures supply before any length is brought in from outside.
Transform between gauges before comparing. A covariance in one gauge can be carried into another exactly by the projector above — the operation geodesy calls an S-transformation — and two reconstructions’ uncertainties are comparable only once they are in the same gauge. This is the expensive fix and the only one that keeps coordinates.
What does not work is choosing the gauge that gives the smallest numbers. The inner-constraint report is the smallest here and it is no more the truth than the anchored one; a gauge chosen for its small uncertainties is a gauge chosen to make a report look good.
What this does not settle
The covariances are linearised. Each is the first-order propagation of the marks’ error through the Jacobian at the solution, and it describes small perturbations about that solution. The identity for invariant quantities holds exactly for the linearised covariances and approximately for the true distribution of errors, to the extent the problem is linear over the size of those errors. At a pixel of reading error on this courtyard the non-linearity is small, and nothing here measures how small.
The error model is rounding. The variance of one residual is estimated from the adjustment’s own residuals with 161 parameters counted as fitted, and the marks’ error is a rounding to a whole pixel — bounded, uniform and independent, which where the adjustment stops explains is tidier than a real matcher’s. A different error model changes every magnitude here and none of the ratios between gauges.
It does not say one gauge is better. Each answers a different question, and the essay’s claim is only that the question should be named when the answer is printed.
Still open: what happens when more than seven numbers are held
Every gauge in this essay holds exactly seven numbers, and that is why they all agree on shape: seven held numbers remove the flat directions and nothing else, so any values of them are consistent with some similarity of the one reconstruction the pictures determine.
A survey does not hold seven numbers. It holds every coordinate of every control point it has — twelve for four points, more for more — and those coordinates come from a different instrument with its own errors. Holding eight numbers is no longer choosing a gauge. It is a claim about the scene’s shape that the pictures may disagree with, and if they do, the adjustment has nowhere to put the disagreement except into the reconstruction itself.
The question that leaves is how much. The eighth held number bends the scene holds the courtyard’s control points at surveyed positions put out by a few millimetres in different directions, and measures what happens to the shape: whether seven held coordinates really change nothing at any survey error, how far the eighth bends the reconstruction, and whether the adjustment’s own residual ever notices that it has been made to fit a survey rather than the pictures.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The spread a point gets — both name bundle adjustment, covariance, error propagation
- Another picture of the same sweep — both name bundle adjustment, gauge freedom
- The midpoint is a choice of ruler — both name gauge freedom, similarity
- What one picture of a plane determines — both name gauge freedom, similarity
Named objects
A flat tag is an object no other essay names yet.
bundle adjustmentControl pointCovarianceerror propagationgauge freedomInvariantSimilarity