A chain and an adjustment
Worth reading first: The track and the scene together · Four cameras fit, and one of them can see.
There is an obvious way to get a camera track out of a sequence of photographs. Take the first two, recover their relative pose. Take the second and third, recover theirs. Compose along the sequence, and every camera has a pose relative to the first.
It works, it is what a sequential pipeline does to get started, and every account of it says the same thing about its weakness: the errors accumulate, so the track drifts away from the truth, and the further along the sequence the worse it gets.
That story is repeated widely enough to be worth measuring, and measuring it does not produce it.
What was expected
The figure was built to draw the standard picture. Per-link errors along the bottom, small and roughly equal. A cumulative error climbing steadily away from zero. A caption pointing out that each step is fine and the end is not.
The first run of it, over five links with the marks read to a whole pixel, gave a worst individual link error of 0.876° and a composed error at the end of the chain of 0.264°.
The end of the chain was three times better than its own worst link.
What is actually happening
Consecutive links have no reason to err in the same direction.
Each link’s rotation error comes from the quantisation of that pair’s marks, which is unrelated to the quantisation of the next pair’s. So composing them is a walk with steps of random sign, not a sum of terms with a common bias. A walk of steps of typical size ends up about from the start — which for five steps is barely more than one step, and on any particular walk can be less.
Measured across four reading precisions on the same six views, the ratio of end-of-chain error to worst-link error runs from 0.30 to 1.12 with no trend in it — and the coarsest reading, which has the largest link errors by some way, is not the one that ends furthest from the truth. That is the signature of a random variable being sampled, not of a systematic effect being observed.
So the honest statement is that the composed error is not bounded by the worst link in either direction, and nothing in the chain pulls it back toward the truth. The claim “errors accumulate” is true of the variance and false of any particular chain, and a figure showing one chain and calling it drift is showing one sample of a random variable and naming it after a mechanism.
How the wrong number nearly shipped
The first version of the cumulative measurement reported 176.4°.
That is a number that should have been read as a bug and very nearly was not, because it agreed with the story the figure was built to tell. Errors compose; the end of the chain is bad; here is a large number at the end of the chain. Every part of that reasoning felt like confirmation.
The cause was a double-multiplication. The chained poses are already expressed relative to the first view, and the truth they were being compared against was multiplied by the first view’s rotation a second time. Both sides were wrong in different ways and the difference between them was meaningless.
Correcting it produced 0.264°, which contradicted the story, which is how the story came to be measured properly.
The lesson is not about matrix conventions. It is that a wrong number that agrees with the expected conclusion is far more dangerous than one that does not, because the check a reader applies — does this look about right? — is exactly the check it passes. This site has recorded the same shape before: the caption quoting 906 px for a disagreement that could not exceed 420 on its own canvas survived two phases and every gate, because 906 looked like the sort of number that figure ought to produce.
The four chains, drawn together
The corrected figure draws four chains rather than one, and the choice is the substance of the correction rather than a presentational tweak.
Each curve is the error of the composed pose, view by view, for one reading precision. They rise and fall. None of them is monotone. They cross each other. The chain read most coarsely is not reliably the worst at every view, and the chain read most finely does not stay best.
That is what a walk looks like, and it is what a single curve cannot show. One curve rising would have been read as drift; one curve falling would have been read as a mistake; four curves doing both is the actual behaviour.
There is a modest dishonesty available here that is worth naming, because the figure could have committed it easily. Four chains differing only in reading precision are not four independent samples of the same random variable — they are four correlated samples, since a mark rounded to two pixels and the same mark rounded to one are not independent. A proper statistical statement about the distribution would need genuinely independent draws, which would need a random number generator, which would need a seed, which would make the figure’s numbers depend on a choice nobody can inspect.
So the figure does not make a statistical statement. It makes the weaker and checkable one: here are four chains, they do not behave monotonically, and the ratio of end to worst link varies by a factor of nearly four across them, and never once reaches the sum of the link errors. That is enough to refute “the error grows along the chain” without pretending to have characterised what it does instead.
Where the adjustment’s gain really comes from
If the chain does not drift, the comparison with a bundle adjustment needs a different explanation — and there is a real gain to explain.
Four to nine times better, on identical observations, with no extra input of any kind.
The explanation is in the last sentence of the caption. A chain uses each observation once. When the pose of view 3 relative to view 2 is recovered, only the correspondences between views 2 and 3 take part. The fact that many of those same points are also visible in views 1, 4, 5 and 6 is not used at all — and once the link is computed, it is frozen; nothing later in the sequence can revise it.
An adjustment uses every observation against every parameter simultaneously. A point seen in five views constrains all five of those cameras, and each of those cameras is constrained by every point it sees. The information is the same information; what changes is how much of it each parameter is exposed to.
So the gain is not about correcting drift. It is about redundancy being spent rather than discarded, and it would be there even if every link were perfect in expectation — which, as the previous section established, they roughly are.
How large the gain is, and why that is the wrong question
The obvious next question is how the advantage behaves — whether it grows with the coarseness of the reading, as the redundancy argument suggests it should.
Measured at four, five, six and seven views across four reading precisions, the gain runs from 1.2× to 21× and is not a monotone function of anything. It does not grow with the reading error. It does not grow with the number of views. Consecutive measurements on the same sequence jump by factors of five in both directions.
The reason is the same reason the chain does not drift. Both quantities being divided — the chain’s track error and the adjustment’s — are single draws from short random walks, so their ratio is a ratio of draws, and a ratio of draws has a great deal of scatter and very little shape.
Two earlier versions of the figure’s assertion were about exactly that shape. The first demanded that the chain’s error grow threefold across the reading precisions; the second demanded that the advantage widen with them. Both held at six views and failed at four, which is how they came to be replaced by the claim that does hold everywhere: the adjustment is better at every configuration measured, and not marginally.
That is a weaker sentence than “the gain grows as the data worsens” and it is the one the measurements support. The mechanism behind the stronger sentence is real — an adjustment averages over many more observations than a chain does, and averaging helps most when there is most to average away — but a mechanism is a statement about expectations, and sixteen configurations of a short sequence do not sample an expectation well enough to see it.
The practical reading survives the weakening intact. A chain leaves information unused; an adjustment uses it; and on every sequence measured here that was worth between a small factor and a large one, with no way to predict in advance which.
The chain’s real weakness
Two things a chain genuinely does badly, both of which the drift story obscures rather than describes.
Its errors are frozen. A link recovered early is never revisited, so an unusually bad pair — a moment when two consecutive views happened to have a short baseline, or a stretch where the visible points were nearly coplanar — poisons everything downstream and cannot be repaired by anything later. That is a genuine and serious failure mode, and it is about irreversibility rather than accumulation.
And it has no closure. If the sequence returns to where it started — a camera walked round a building and back — the chain has no mechanism to notice that view twelve is looking at the same thing as view one. An adjustment does, provided the correspondences are there, and enforcing that agreement is a strong constraint that removes a large amount of error at once. The loop-closure step in a real pipeline is exactly this, and it is impossible in a chain by construction.
Both of those are better arguments for adjusting than “errors accumulate”, and both survive the measurement that “errors accumulate” did not.
The variance is real
None of this says the chain is fine. It says the effect is statistical rather than systematic, and the statistical version has real consequences.
The expected error at the end of a chain does grow, as the square root of the length. Over five links that is a factor of about 2.2, which is small enough to be swamped by which particular errors happened to occur. Over five hundred links it is a factor of 22, and no individual chain escapes it by luck.
So the story is right at long sequence lengths and wrong at the length most figures are drawn at, which is a common enough shape for a piece of received wisdom. The correct statement is about the growth of the error’s distribution, and a figure that measures a single chain cannot see a distribution — which is why this one draws four, at different reading precisions, and reads their spread rather than their trend.
What “drift” describes correctly
It is worth separating the phenomenon from the word, because drift is a real thing in sequential reconstruction and this essay has been arguing only that rotation error over five links is not an instance of it.
Where drift genuinely bites is in scale. A chain of pairwise poses recovers each baseline’s direction and not its length, so the lengths have to be tied together — usually by triangulating points shared between consecutive pairs and requiring them to agree. Each of those ties has an error, and the errors multiply rather than add, because each link’s scale is expressed as a factor times the previous one.
A multiplicative chain of factors each with a small random error does not perform a random walk about a constant. Its logarithm performs a random walk, so the scale itself wanders multiplicatively, and a sequence a few hundred frames long can end up with a scale substantially different from where it started. That is the effect that makes a long visual-odometry track shrink or grow, and it is a genuine accumulation because the quantity being composed is a product rather than a sum of signed errors.
This figure does not measure it, because its chain is supplied with each link’s true baseline length — a piece of scaffolding named in the previous essay. So the measurement here is specifically about rotation error, and it says that rotation error over a short chain wanders. Scale drift over a long chain is a different mechanism with a different arithmetic, and lumping the two under one word is part of why the rotation story goes unchecked.
What the chain is still for
Nothing here argues against building the chain. The track and the scene essay uses one as its starting point, and a non-linear optimiser needs a starting point in the right basin far more than it needs a good one.
That is the chain’s role: it is cheap, it is closed-form at every step, and it produces a complete reconstruction that is wrong by a few pixels rather than by a hundred. An adjustment started from it converges in a handful of iterations. An adjustment started from nothing does not converge at all.
The mistake is treating the chain’s output as an answer rather than as a starting point, and the reason that mistake is easy is that the chain’s output looks like an answer. It is a complete camera track and a complete point cloud, drawn as convincingly as the adjusted one, and the difference between them is a factor of four to nine in a quantity nobody computes unless a ground truth is available.
Which is the recurring shape of this whole phase. The short baseline returns a confident wrong reconstruction. The outlier returns a confident wrong geometry with a small residual. The chain returns a confident reconstruction that is four to nine times worse than the same data supports. In every case the failure is not an error, and the only defence is to compute the quantity that would distinguish them.
There is one difference here worth noting, and it is in the chain’s favour. The other two failures need a ground truth to detect: without knowing where the cameras really were, nothing distinguishes a short-baseline reconstruction from a good one. The chain’s shortfall does not. Running the adjustment and watching the reprojection error fall from several pixels to a fraction of one is a measurement made entirely from the data, and it says the chain was leaving information unused without anybody having to know the answer.
That is the practical form of the recommendation, and it costs one solve: build the chain, adjust from it, and compare. If the adjustment barely moves, the chain was already using what there was. If it falls by an order of magnitude, it was not — and the amount it falls is a measurement of how much the sequence’s redundancy was worth.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two views give shape and no size — both name baseline, relative pose
Named objects
A flat tag is an object no other essay names yet.
Baselinebundle adjustmentcamera trackDriftrandom walkRelative posestructure from motion