Many pictures at once

A scale chain leans rather than wanders

A camera driven 200 m along a street, each pair's scale handed to the next through the points they share. Across sixteen streets the scale does not spread the way a random walk would — 7.3 per cent after 199 links, where independent ties would give 108 — and its average leans by an amount the choice of average decides, from +28.6 per cent to −12.2.

Worth reading first: A chain and an adjustment · The one thing a single view cannot give.

A chain and an adjustment measured a chain of pairwise camera poses and found that its rotation does not drift away from the truth but wanders about it, because consecutive links have no reason to err in the same direction. It handed every link its true baseline length, which kept scale out of the measurement, and then said in prose what scale would do: each link’s length is a factor times the previous one, so the logarithm of the scale performs a random walk, its spread grows as the square root of the chain’s length, and a few hundred frames can end up substantially shrunk or grown.

That paragraph contains a prediction, and it can be measured. Measured, it is wrong in both its halves. The spread does not grow. Something else does.

Scale carried along 200 links of a street, on 16 different streetsA camera driven 200 m along a street, one pair of pictures a metre, every mark read to 1 px, and each pair's scale handed to the next through the points they share, tied by the mean of ratios. Each thin line is one of 16 streets; the heavy line is their average. After 199 links the average is 28.6 % from the true scale, with a standard error of 1.8 %, while the streets spread about it by 7.3 % — against 9.5 % at link 50. The spread does not grow with the length; the average moves.-200204060050100150link along the streetscale error relative to the first link (%)average 28.6 %16 streets · ties by mean of ratios · marks read to 1 pxspread after 199 links ±7.3 %
Fig. 1 Scale carried along 200 links of a street, each pair’s scale handed to the next by the mean of its shared points’ distance ratios, marks read to a whole pixel. Each thin line is one of 16 streets and the heavy line is their average. After 199 links the average is 28.6 % from the true scale, with a standard error of 1.8 %; the streets spread about it by 7.3 %, against 9.5 % at link 50.

Handing a scale from one pair to the next

A camera is driven along a street, one pair of pictures every metre, looking sideways at facades between 5 and 14 m away. Each street is drawn from a seeded generator, so the sixteen streets are sixteen genuinely different scenes, and every number here is reproducible.

Each pair of consecutive pictures gives a relative pose, and two views give shape and no size is exact about what that pose contains: a rotation and the direction of the step between the cameras, never its length. So every pair is reconstructed at a baseline of one unit, and the points it sees are placed at distances measured in that unit.

The next pair sees many of the same points. It too reconstructs them at a baseline of one of its own units. The ratio of a shared point’s distance from the middle camera, as the first pair puts it, to its distance as the second pair puts it, is the ratio of the second pair’s unit to the first pair’s. Every shared point gives one such ratio; the tie between the two pairs is some average of them; and multiplying the ties along the chain carries the first pair’s scale to every later pair.

That is the standard way a single moving camera keeps a consistent scale, and with the marks exact it keeps it exactly. With the marks read to a whole pixel, every reconstructed distance carries an error and so does every ratio.

The walk that does not happen

Each tie’s error is easy to measure directly.

The spread a chain of scale ties does not haveEach tie alone is uncertain by 7.6 % in log-scale, across 16 streets and 199 ties each. Independent errors of that size would spread the chains' scale by 107.7 % after 199 links, growing as the square root of the length. The measured spread after 199 links is 7.3 %, and 9.5 % a quarter of the way in. Consecutive ties are correlated at -0.62: each pair's own error enters one tie on top and the next underneath, and cancels.0255075100050100150link along the streetspread of the log-scale across streets (%)if ties were independentmeasuredconsecutive ties correlated at -0.62ties by mean of ratios
Fig. 2 The spread of the log-scale across the 16 streets, link by link, against the spread that independent ties would produce. Each tie alone is uncertain by 7.6 % in log-scale; independent ties of that size would spread the chains by 107.7 % after 199 links. The measured spread is 7.3 % after 199 links and 9.5 % a quarter of the way in.

Across sixteen streets and 199 ties on each, a single tie’s log-scale error has a standard deviation of 7.6 per cent. If those errors were independent from one tie to the next, the chain’s log-scale would be their sum, its standard deviation would be 7.6 per cent times the square root of the number of ties, and after 199 links it would be 107.7 per cent — a spread in which a quarter of streets end up with scales more than twice or less than half the truth.

The measured spread after 199 links is 7.3 per cent. A quarter of the way along, at link 50, it is 9.5. It is not growing at all, and it is fifteen times smaller than the random walk’s.

The prediction had the right arithmetic for its assumption. The assumption is what fails.

Why each tie pays back the last

The ties are not independent, and one scatter plot says how.

Each scale tie against the nextPairs of consecutive ties from 16 streets, every third pair drawn, marks read to 1 px, tied by the mean of ratios. The cloud leans from upper left to lower right: the two are correlated at -0.62. A link whose reconstruction came out too small makes the tie into it read high and the tie out of it read low, so its error is paid back one link later.-0.500-0.25000.2500.500-0.400-0.20000.2000.400one tie's log-scale errorthe next tie's log-scale error1056 consecutive pairs · correlation -0.62ties by mean of ratios
Fig. 3 Consecutive scale ties plotted against each other, every third pair from the 16 streets, marks read to a whole pixel. The cloud leans from upper left to lower right: consecutive ties are correlated at −0.62. A link whose reconstruction came out too small makes the tie into it read high and the tie out of it read low.

Plotting each tie’s error against the next one’s gives a cloud leaning from upper left to lower right, a correlation of −0.62. A tie that reads high is followed by one that reads low.

The mechanism is in how a tie is built. The tie between links k1k-1 and kk divides link k1k-1’s distances by link kk’s. The tie between links kk and k+1k+1 divides link kk’s distances by link k+1k+1’s. Link kk’s reconstruction appears in both — underneath in the first, on top in the second. Whatever error link kk made in its own overall size, it makes the first tie too large by a factor and the second too small by the same factor, and when the two ties are multiplied along the chain that factor cancels exactly.

Written as a sum of logarithms, the chain’s scale after nn links is

logsn=k(logaklogbk),\log s_n = \sum_{k} \left(\log a_k - \log b_k\right),

where aka_k is link kk’s reconstructed distance to the points it shares with the link behind and bkb_k its distance to the points it shares with the link ahead. An error common to both — a mistake in link kk’s size — adds to logak\log a_k and logbk\log b_k equally and vanishes. What survives is only the part of link kk’s error that treats its backward points differently from its forward ones: the inconsistency inside each reconstruction, rather than its scale.

That surviving part is much smaller than a tie’s error, which is why the spread sits at a few per cent. It is also still there, and it has a property the random part does not have.

What does build up

The random part cancels. A systematic part does not, because a bias that makes every tie slightly too large makes every tie slightly too large, and a product of factors each slightly too large grows geometrically.

The heavy line in the first figure is that growth. The sixteen streets’ average scale leans steadily away from the truth, reaching 28.6 per cent after 199 links, with a standard error across the streets of 1.8 per cent — sixteen standard errors from zero. The spread stays put and the whole bundle of streets moves.

Where the bias comes from shows up by changing nothing but the average.

One set of shared points, four ways to tie a scale, four driftsThe same 16 streets and the same reconstructed points, with each pair's scale handed on by five different averages of the same ratios. After 199 links the average scale error is 28.6 % by the mean of ratios, 25.1 % by the upper median, 5.9 % by the median ratio, 6.3 % by the geometric mean, −12.2 % by the harmonic mean. The ratios were identical; the only thing chosen was how to average them, and it chose the direction of the drift.020050100150link along the streetaverage scale error over 16 streets (%)mean of ratiosupper medianmedian ratiogeometric meanharmonic meanafter 199 links: 28.6 % to −12.2 %marks read to 1 px
Fig. 4 The same 16 streets and the same reconstructed points, each tie formed by five different averages of the same distance ratios. After 199 links the average scale error is 28.6 % by the mean of ratios, 25.1 % by the upper median, 5.9 % by the median, 6.3 % by the geometric mean and −12.2 % by the harmonic mean.

The ratios are identical in all five curves. Only the rule that turns a tie’s thirty-odd ratios into one number differs, and the rule decides the drift: +28.6 per cent by the arithmetic mean of the ratios, +25.1 per cent by the upper median, +5.9 per cent by the median, +6.3 per cent by the geometric mean, and −12.2 per cent by the harmonic mean.

The ordering is not an accident of these streets. A ratio of two noisy positive numbers is skewed: an error that shrinks the denominator raises the ratio by more than the same error in the other direction lowers it. The arithmetic mean of skewed ratios is pulled toward the long tail and reads high. The harmonic mean is the reciprocal of the mean of the reciprocal ratios, which are skewed the other way, and it reads low by the mirror-image argument. The geometric mean averages logarithms, in which a ratio and its reciprocal are symmetric, and it sits between the two. This is the same asymmetry depth is a reciprocal found in a single stereo reading, whose depth interval is not centred on the answer, now compounded two hundred times.

The upper median needs a sentence of its own, because it is what a quick implementation produces. The median of an even number of values is conventionally the mean of the two middle ones. Taking the element at position n/2n/2 instead gives the upper of the two, and roughly half the ties in a chain have an even count of shared points. That one indexing choice turns a lean of 5.9 per cent into 25.1, on the same street and the same ratios.

The median’s lean, at sixteen streets

The two estimators that lean least deserve a closer look, because a lean of a few per cent is where the measurement’s own resolution starts to matter.

Scale carried along 200 links of a street, on 16 different streetsA camera driven 200 m along a street, one pair of pictures a metre, every mark read to 1 px, and each pair's scale handed to the next through the points they share, tied by the median ratio. Each thin line is one of 16 streets; the heavy line is their average. After 199 links the average is 5.9 % from the true scale, with a standard error of 2.5 %, while the streets spread about it by 10.1 % — against 10.7 % at link 50. The spread does not grow with the length; the average moves.-2002040050100150link along the streetscale error relative to the first link (%)average 5.9 %16 streets · ties by median ratio · marks read to 1 pxspread after 199 links ±10.1 %
Fig. 5 The same 16 streets with each tie formed by the median ratio. After 199 links the average is 5.9 % from the true scale with a standard error of 2.5 %, and the streets spread about it by 10.1 %, against 10.7 % at link 50.

Tied by the median, the average after 199 links is 5.9 per cent from the true scale and the standard error across the sixteen streets is 2.5 per cent. That is a lean of 2.4 standard errors: suggestive, and not something sixteen streets resolve. The geometric mean’s 6.3 per cent is in the same position. What the figure does resolve is the same flat spread as before, 10.1 per cent after 199 links against 10.7 at link 50, so the median tie shares the cancellation even if its bias is small enough to be uncertain.

The honest summary is not that the median is unbiased. It is that the median and the geometric mean lean by an amount this measurement cannot distinguish from a few per cent either way over two hundred links, while the arithmetic mean and the harmonic mean lean by amounts it measures to within a tenth of themselves, in opposite directions.

Finer marks, smaller leans

A bias that comes from noise should shrink as the noise does, and it does.

How much each link leans, against how finely the marks are readThe average scale error each link adds, over 16 streets of 200 links, at three reading precisions and five ways of tying. By the mean of ratios it is 0.030 %, 0.127 %, 0.662 %; by the median ratio 0.001 % at 0.5 px, 0.029 % at 1 px, 0.221 % at 2 px; by the harmonic mean −0.018 %, −0.065 %, −0.209 %. Finer marks shrink every lean without removing any of them, and no reading precision drawn here makes the mean of ratios and the harmonic mean agree in sign.-0.25000.2500.5000.50012how finely each mark is read (px)average scale error added per link (%)mean of ratiosupper medianmedian ratiogeometric meanharmonic mean16 streets, 200 links eachper link, by the median: 0.001 → 0.029 → 0.221 %
Fig. 6 The average scale error each link adds, over 16 streets of 200 links, at marks read to half a pixel, one and two, for the five ways of tying. Every lean shrinks as the marks get finer, and at every precision drawn the mean of ratios leans outward and the harmonic mean inward.

Read per link, the median’s lean is 0.001 per cent at half-pixel marks, 0.029 per cent at a pixel and 0.221 per cent at two; the harmonic mean’s is −0.018, −0.065 and −0.209. Every estimator’s lean shrinks as the marks get finer, and none changes sign: at every precision drawn, the arithmetic mean of ratios leans outward and the harmonic mean inward.

Per link, those are small numbers. The point of a chain is that it multiplies them, and two hundred links at a pixel turn a lean of a tenth of a per cent a link into a scale error of a quarter. A thousand links, which is a short drive, would compound it further, and the flat spread would not hide it: every street would be wrong in the same direction.

How the lean scales with the reading error

The rate at which each lean shrinks is itself informative, because a bias that comes from the variance of the noise should shrink with the square of the reading error: four times less for marks read twice as finely.

The arithmetic mean of ratios comes close to that. Its lean per link is 0.030 per cent at half a pixel, 0.127 at one and 0.662 at two, falling by a factor of 4.2 from two pixels to one and 5.2 from one to a half. That is the behaviour of a bias produced by averaging a skewed quantity, whose skew grows with the spread of what is averaged.

The harmonic mean falls more slowly, by 3.2 and then 3.6, and the median does something else again: 0.221, 0.029 and 0.001 per cent, a factor of 7.6 and then of 29. A median does not average the ratios at all; it picks one, and which one it picks depends on how the few ratios near the middle happen to be arranged, so its lean need not follow any power of the noise. The measurement does not say what law it follows, only that it is not the variance law and that it falls fastest of the five.

So the five estimators differ not only in the size and sign of their lean at a given precision but in how much a finer instrument helps them. Improving the marks from one pixel to a half cuts the arithmetic mean’s lean fivefold and the harmonic mean’s threefold. A choice of estimator made at one precision does not stay the best choice at another.

What a lean does to a street

A scale error is abstract until it is carried into the positions it produces, and the arithmetic is short enough to do in full.

Each link’s recovered length is its true metre times the scale carried to it. With a lean that grows the log-scale by a fixed amount λ\lambda a link, the recovered distance to link kk is the sum j<keλj\sum_{j<k} e^{\lambda j}, which is (eλk1)/λ(e^{\lambda k} - 1)/\lambda. At the arithmetic mean’s lean of 0.127 per cent a link, the 199 m of street to the last camera comes out at about 227 m. The street is 28 m too long, but the error is not spread evenly along it: the first 50 m are recovered to within a metre and a half, and most of the excess accumulates in the last hundred.

That is an idealisation — it treats every link as leaning by the average amount, which the streets individually do not — and it is quoted as arithmetic on the measured average rather than as a measurement of positions. What it does show is the shape a lean gives a reconstruction: good near where the chain started, increasingly stretched along its length, and consistent enough from run to run that repeating the drive would not reveal it. A reconstruction that looks the same every time is precisely the one the spread-based check passes.

Which arithmetic was wrong

The random-walk account in a chain and an adjustment was correct arithmetic applied to independent ties, and it predicted the one thing a chain built this way does not do. It also missed the one thing the chain does, because a random walk has no direction and the chain has one.

Two consequences follow for anyone carrying scale along a sequence.

Spread across runs is the wrong diagnostic. A set of repeated runs whose scales agree with each other to a few per cent looks consistent, and it can be consistently wrong by a quarter. The agreement is the cancellation at work, and it says nothing about the bias. The number that would catch the bias is a comparison with a length known from outside — the one length that the eighth held number bends the scene counts as one of the seven numbers a survey supplies, and that five facts that close the same gap finds can come from inside a single photograph in several ways.

The estimator is a design decision, not an implementation detail. Five reasonable averages of the same ratios gave leans from −12 to +29 per cent, and the difference between two of them was one index. A pipeline that reports its scale drift without saying how it averages its ties has not described the quantity its drift is mostly made of.

What the measurement does not recommend is simply choosing the estimator that leaned least here. The geometric mean and the median leaned least on these streets, at these distances, with this error model; whether they lean least with a matcher’s errors, which are not a rounding, is not measured. The robust conclusion is the structural one: the random part of scale error cancels along a chain, so the systematic part is what accumulates, and it is the part to measure.

What this does not settle

No adjustment. Every chain here is a chain and nothing more. The track and the scene together and its neighbours adjust all cameras and points at once, and a sliding-window adjustment over a few links would re-estimate the reconstructions the ties are built from. How much of the lean survives that is not measured.

The correlation is stronger than a pure hand-back. If each tie’s error were exactly one link’s size error handed on to the next tie, consecutive ties would correlate at −0.5. They correlate at −0.62. Something more is shared between neighbouring ties than the middle link’s size — perhaps because neighbouring ties share many of the same points — and what it is has not been isolated.

One kind of street. A sideways-looking camera, a metre between pictures, facades 5 to 14 m off and thirty-odd shared points per tie. A forward-looking camera, whose baseline points at the scene, reconstructs depth far less well and is exactly the case in which monocular scale drift is most often reported; it was not run.

Rounding. The marks’ error is rounding to a grid, bounded and independent, as throughout this field.

Still open: whether closing a loop removes a lean

A lean is a scale that is consistently wrong in one direction, and the one thing that makes it visible without outside measurement is coming back to a place already seen. A vehicle that drives round a block and returns to its start finds, if its scale has leaned, that its reconstructed return point is not where it began, and the discrepancy is a direct measurement of what the chain accumulated.

The question that leaves is what closing the loop repairs. Closing a loop mends its ends measures a walk once round a ring of walls, left open and closed, and asks where along the walk the recovered positions become better determined — whether the closure improves the whole loop evenly, as a lean spread back round it would suggest, or mostly the cameras near where the walk began and ended.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasDriftrandom walkRelative posescale ambiguityTriangulation