A scale chain leans rather than wanders
Worth reading first: A chain and an adjustment · The one thing a single view cannot give.
A chain and an adjustment measured a chain of pairwise camera poses and found that its rotation does not drift away from the truth but wanders about it, because consecutive links have no reason to err in the same direction. It handed every link its true baseline length, which kept scale out of the measurement, and then said in prose what scale would do: each link’s length is a factor times the previous one, so the logarithm of the scale performs a random walk, its spread grows as the square root of the chain’s length, and a few hundred frames can end up substantially shrunk or grown.
That paragraph contains a prediction, and it can be measured. Measured, it is wrong in both its halves. The spread does not grow. Something else does.
Handing a scale from one pair to the next
A camera is driven along a street, one pair of pictures every metre, looking sideways at facades between 5 and 14 m away. Each street is drawn from a seeded generator, so the sixteen streets are sixteen genuinely different scenes, and every number here is reproducible.
Each pair of consecutive pictures gives a relative pose, and two views give shape and no size is exact about what that pose contains: a rotation and the direction of the step between the cameras, never its length. So every pair is reconstructed at a baseline of one unit, and the points it sees are placed at distances measured in that unit.
The next pair sees many of the same points. It too reconstructs them at a baseline of one of its own units. The ratio of a shared point’s distance from the middle camera, as the first pair puts it, to its distance as the second pair puts it, is the ratio of the second pair’s unit to the first pair’s. Every shared point gives one such ratio; the tie between the two pairs is some average of them; and multiplying the ties along the chain carries the first pair’s scale to every later pair.
That is the standard way a single moving camera keeps a consistent scale, and with the marks exact it keeps it exactly. With the marks read to a whole pixel, every reconstructed distance carries an error and so does every ratio.
The walk that does not happen
Each tie’s error is easy to measure directly.
Across sixteen streets and 199 ties on each, a single tie’s log-scale error has a standard deviation of 7.6 per cent. If those errors were independent from one tie to the next, the chain’s log-scale would be their sum, its standard deviation would be 7.6 per cent times the square root of the number of ties, and after 199 links it would be 107.7 per cent — a spread in which a quarter of streets end up with scales more than twice or less than half the truth.
The measured spread after 199 links is 7.3 per cent. A quarter of the way along, at link 50, it is 9.5. It is not growing at all, and it is fifteen times smaller than the random walk’s.
The prediction had the right arithmetic for its assumption. The assumption is what fails.
Why each tie pays back the last
The ties are not independent, and one scatter plot says how.
Plotting each tie’s error against the next one’s gives a cloud leaning from upper left to lower right, a correlation of −0.62. A tie that reads high is followed by one that reads low.
The mechanism is in how a tie is built. The tie between links and divides link ’s distances by link ’s. The tie between links and divides link ’s distances by link ’s. Link ’s reconstruction appears in both — underneath in the first, on top in the second. Whatever error link made in its own overall size, it makes the first tie too large by a factor and the second too small by the same factor, and when the two ties are multiplied along the chain that factor cancels exactly.
Written as a sum of logarithms, the chain’s scale after links is
where is link ’s reconstructed distance to the points it shares with the link behind and its distance to the points it shares with the link ahead. An error common to both — a mistake in link ’s size — adds to and equally and vanishes. What survives is only the part of link ’s error that treats its backward points differently from its forward ones: the inconsistency inside each reconstruction, rather than its scale.
That surviving part is much smaller than a tie’s error, which is why the spread sits at a few per cent. It is also still there, and it has a property the random part does not have.
What does build up
The random part cancels. A systematic part does not, because a bias that makes every tie slightly too large makes every tie slightly too large, and a product of factors each slightly too large grows geometrically.
The heavy line in the first figure is that growth. The sixteen streets’ average scale leans steadily away from the truth, reaching 28.6 per cent after 199 links, with a standard error across the streets of 1.8 per cent — sixteen standard errors from zero. The spread stays put and the whole bundle of streets moves.
Where the bias comes from shows up by changing nothing but the average.
The ratios are identical in all five curves. Only the rule that turns a tie’s thirty-odd ratios into one number differs, and the rule decides the drift: +28.6 per cent by the arithmetic mean of the ratios, +25.1 per cent by the upper median, +5.9 per cent by the median, +6.3 per cent by the geometric mean, and −12.2 per cent by the harmonic mean.
The ordering is not an accident of these streets. A ratio of two noisy positive numbers is skewed: an error that shrinks the denominator raises the ratio by more than the same error in the other direction lowers it. The arithmetic mean of skewed ratios is pulled toward the long tail and reads high. The harmonic mean is the reciprocal of the mean of the reciprocal ratios, which are skewed the other way, and it reads low by the mirror-image argument. The geometric mean averages logarithms, in which a ratio and its reciprocal are symmetric, and it sits between the two. This is the same asymmetry depth is a reciprocal found in a single stereo reading, whose depth interval is not centred on the answer, now compounded two hundred times.
The upper median needs a sentence of its own, because it is what a quick implementation produces. The median of an even number of values is conventionally the mean of the two middle ones. Taking the element at position instead gives the upper of the two, and roughly half the ties in a chain have an even count of shared points. That one indexing choice turns a lean of 5.9 per cent into 25.1, on the same street and the same ratios.
The median’s lean, at sixteen streets
The two estimators that lean least deserve a closer look, because a lean of a few per cent is where the measurement’s own resolution starts to matter.
Tied by the median, the average after 199 links is 5.9 per cent from the true scale and the standard error across the sixteen streets is 2.5 per cent. That is a lean of 2.4 standard errors: suggestive, and not something sixteen streets resolve. The geometric mean’s 6.3 per cent is in the same position. What the figure does resolve is the same flat spread as before, 10.1 per cent after 199 links against 10.7 at link 50, so the median tie shares the cancellation even if its bias is small enough to be uncertain.
The honest summary is not that the median is unbiased. It is that the median and the geometric mean lean by an amount this measurement cannot distinguish from a few per cent either way over two hundred links, while the arithmetic mean and the harmonic mean lean by amounts it measures to within a tenth of themselves, in opposite directions.
Finer marks, smaller leans
A bias that comes from noise should shrink as the noise does, and it does.
Read per link, the median’s lean is 0.001 per cent at half-pixel marks, 0.029 per cent at a pixel and 0.221 per cent at two; the harmonic mean’s is −0.018, −0.065 and −0.209. Every estimator’s lean shrinks as the marks get finer, and none changes sign: at every precision drawn, the arithmetic mean of ratios leans outward and the harmonic mean inward.
Per link, those are small numbers. The point of a chain is that it multiplies them, and two hundred links at a pixel turn a lean of a tenth of a per cent a link into a scale error of a quarter. A thousand links, which is a short drive, would compound it further, and the flat spread would not hide it: every street would be wrong in the same direction.
How the lean scales with the reading error
The rate at which each lean shrinks is itself informative, because a bias that comes from the variance of the noise should shrink with the square of the reading error: four times less for marks read twice as finely.
The arithmetic mean of ratios comes close to that. Its lean per link is 0.030 per cent at half a pixel, 0.127 at one and 0.662 at two, falling by a factor of 4.2 from two pixels to one and 5.2 from one to a half. That is the behaviour of a bias produced by averaging a skewed quantity, whose skew grows with the spread of what is averaged.
The harmonic mean falls more slowly, by 3.2 and then 3.6, and the median does something else again: 0.221, 0.029 and 0.001 per cent, a factor of 7.6 and then of 29. A median does not average the ratios at all; it picks one, and which one it picks depends on how the few ratios near the middle happen to be arranged, so its lean need not follow any power of the noise. The measurement does not say what law it follows, only that it is not the variance law and that it falls fastest of the five.
So the five estimators differ not only in the size and sign of their lean at a given precision but in how much a finer instrument helps them. Improving the marks from one pixel to a half cuts the arithmetic mean’s lean fivefold and the harmonic mean’s threefold. A choice of estimator made at one precision does not stay the best choice at another.
What a lean does to a street
A scale error is abstract until it is carried into the positions it produces, and the arithmetic is short enough to do in full.
Each link’s recovered length is its true metre times the scale carried to it. With a lean that grows the log-scale by a fixed amount a link, the recovered distance to link is the sum , which is . At the arithmetic mean’s lean of 0.127 per cent a link, the 199 m of street to the last camera comes out at about 227 m. The street is 28 m too long, but the error is not spread evenly along it: the first 50 m are recovered to within a metre and a half, and most of the excess accumulates in the last hundred.
That is an idealisation — it treats every link as leaning by the average amount, which the streets individually do not — and it is quoted as arithmetic on the measured average rather than as a measurement of positions. What it does show is the shape a lean gives a reconstruction: good near where the chain started, increasingly stretched along its length, and consistent enough from run to run that repeating the drive would not reveal it. A reconstruction that looks the same every time is precisely the one the spread-based check passes.
Which arithmetic was wrong
The random-walk account in a chain and an adjustment was correct arithmetic applied to independent ties, and it predicted the one thing a chain built this way does not do. It also missed the one thing the chain does, because a random walk has no direction and the chain has one.
Two consequences follow for anyone carrying scale along a sequence.
Spread across runs is the wrong diagnostic. A set of repeated runs whose scales agree with each other to a few per cent looks consistent, and it can be consistently wrong by a quarter. The agreement is the cancellation at work, and it says nothing about the bias. The number that would catch the bias is a comparison with a length known from outside — the one length that the eighth held number bends the scene counts as one of the seven numbers a survey supplies, and that five facts that close the same gap finds can come from inside a single photograph in several ways.
The estimator is a design decision, not an implementation detail. Five reasonable averages of the same ratios gave leans from −12 to +29 per cent, and the difference between two of them was one index. A pipeline that reports its scale drift without saying how it averages its ties has not described the quantity its drift is mostly made of.
What the measurement does not recommend is simply choosing the estimator that leaned least here. The geometric mean and the median leaned least on these streets, at these distances, with this error model; whether they lean least with a matcher’s errors, which are not a rounding, is not measured. The robust conclusion is the structural one: the random part of scale error cancels along a chain, so the systematic part is what accumulates, and it is the part to measure.
What this does not settle
No adjustment. Every chain here is a chain and nothing more. The track and the scene together and its neighbours adjust all cameras and points at once, and a sliding-window adjustment over a few links would re-estimate the reconstructions the ties are built from. How much of the lean survives that is not measured.
The correlation is stronger than a pure hand-back. If each tie’s error were exactly one link’s size error handed on to the next tie, consecutive ties would correlate at −0.5. They correlate at −0.62. Something more is shared between neighbouring ties than the middle link’s size — perhaps because neighbouring ties share many of the same points — and what it is has not been isolated.
One kind of street. A sideways-looking camera, a metre between pictures, facades 5 to 14 m off and thirty-odd shared points per tie. A forward-looking camera, whose baseline points at the scene, reconstructs depth far less well and is exactly the case in which monocular scale drift is most often reported; it was not run.
Rounding. The marks’ error is rounding to a grid, bounded and independent, as throughout this field.
Still open: whether closing a loop removes a lean
A lean is a scale that is consistently wrong in one direction, and the one thing that makes it visible without outside measurement is coming back to a place already seen. A vehicle that drives round a block and returns to its start finds, if its scale has leaned, that its reconstructed return point is not where it began, and the discrepancy is a direct measurement of what the chain accumulated.
The question that leaves is what closing the loop repairs. Closing a loop mends its ends measures a walk once round a ring of walls, left open and closed, and asks where along the walk the recovered positions become better determined — whether the closure improves the whole loop evenly, as a lean spread back round it would suggest, or mostly the cameras near where the walk began and ended.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The workshop that throws drawings away — both name bias, drift, random walk
- What survives being copied — both name bias, drift, random walk
- The wall under the paint — both name scale ambiguity, triangulation
- Two pictures of a ball — both name scale ambiguity, triangulation
Named objects
A flat tag is an object no other essay names yet.
BiasDriftrandom walkRelative posescale ambiguityTriangulation