The cameras decide where a narrow view settles
Worth reading first: The track and the scene together · Seven numbers no picture can name.
A start needs the sign of its depths, not their size moved a narrow-field bundle adjustment’s starting point along a straight path from the true reconstruction of a courtyard to its inside-out twin, and found a band. Every start less than 42 per cent of the way returned to the truth; every start past 56 per cent fell into the twin; between them neighbouring starts went different ways. Read in terms of relief, that looked like a rule a pipeline could use — get the depths the right way round, even at a sixth of their true size, and the adjustment does the rest.
A real initialisation is usually built the way the track and the scene together describes — pairs of pictures chained into a first guess at every camera and point — and a chain and an adjustment found that such a chain’s errors wander rather than accumulate. The earlier essay ended by asking whether the rule survives a start like that. A path through the parameters moves everything together, so every start on it is the true reconstruction scaled in depth. A real initialisation is scaled and scattered: each point’s depth off by its own amount. If scatter moves the band, the sign of the average depth is not enough; if it does not, the rule is general.
Neither of those happened, and what did happen overturns the rule rather than qualifying it.
Scatter turns the switch into a lottery
The measurement is the one the question asked for. Take a start a stated fraction of the way to the twin, add to every point’s depth an independent error of stated size — measured in units of the courtyard’s own spread of depths — and run the adjustment. Twelve seeded starts per cell, the standing six cameras on a 60° arc, a 3° field, every mark exact.
With no scatter every cell goes one way, and the switch sits where the earlier measurement put it: the last start to return is 42 per cent of the way and the first to fall in is 46. Scatter the depths and the switch becomes a band of cells in which some starts return and some fall in — from 42 to 54 per cent of the way at a scatter of 0.2, 38 to 54 at 0.5, and 38 to 58 at 0.8. The band widens about where the switch was; it does not move toward the truth. The flat start halfway along, which fell into the twin every time without scatter, comes home seven times in twelve at scatters of 0.2 and 0.5 and five times at 0.8.
That is the answer to the question as asked, and on its face it supports the rule. The band stays centred; scatter only blurs it. A start whose depths are right on average and scattered would seem to be safe as long as it sits outside the blur.
But a blur is a statement about many starts. Each start has its own realised depths — the blend plus that start’s particular scatter — and the rule is a claim about those.
Each start’s own depths do not predict it
So look at the starts one at a time.
Take every scattered start blended between 42 and 58 per cent of the way, and measure the relief its depths actually have: the slope of its depths on the true ones, along the axis the reversal turns over. A positive slope is a start with the right depth order on average; a negative one has it inside out. If the rule held, the positive starts would return and the negative ones fall in.
They do not sort. Of 80 starts whose depths had the right sign, 26 fell into the twin. Of 100 whose depths were reversed, 43 came home. Starts with a relief of +0.2 — a fifth of the true relief, correctly ordered — fell in; starts at −0.3 returned. The realised sign of a start’s depths predicts where it settles only a little better than a coin.
Something other than the depths is deciding, and on this path there is only one other thing in the start.
Split the start in two
A blended start moves two things: the points and the cameras. The earlier path moved them by the same fraction, so it could never say which of the two was doing the work. The way to find out is to move them by different fractions.
The picture is almost entirely horizontal stripes. With the cameras true, every start returns — including the one whose points are the twin’s own, the courtyard turned completely inside out. With the cameras on the twin’s side, every start falls in — including the one whose points are exactly right. Only in the rows where the cameras are 30 to 70 per cent of the way do the points change the answer.
So the rule is refused, not refined. What a narrow-field start needs to get right is not the sign of its depths; the adjustment repairs any depths at all, including reversed ones, if the cameras are right. And the diagonal the earlier measurement walked along crosses the stripes at the only place where the points matter, which is why a rule about depths looked as if it fitted.
Three rows of that grid make the point on their own. Moving only the points, all the way to the twin’s, never loses the scene. Moving only the cameras returns up to 30 per cent of the way and loses the scene from 40 — sooner than the combined blend, which still returns at 40. The depths are the half of an initialisation the adjustment can repair. The cameras are the half it keeps.
The same grid at a 25° field has the same stripes, narrower in the middle — the points matter only when the cameras are 40 to 60 per cent of the way — so the earlier essay’s other finding survives: the field of view does not change which starts are safe. It changes only how badly the twin fits once reached.
Where a blended camera actually stands
That leaves the middle rows, where the points do vote — and vote in a direction that looks perverse. In the row where the cameras are half of the way, the start returns to the scene when its points are more than half reversed and falls in when they are nearer the truth. The explanation is in where those cameras are, which the word “blend” does not suggest.
Blending a camera’s orientation from the true one to the twin’s does not slide it across the arc toward the mirror position. The twin’s cameras are the true ones reflected in depth, and the straightest path between the two orientations turns each camera over the top of the courtyard, 128 metres out and looking at it all the way. A quarter of the way along, the middle camera looks straight down from 127 metres above. Halfway, it looks back at the courtyard from 128 metres behind it. At the far end it is in front again, with the scene the other way out.
That is the ambiguity a narrow field leaves, drawn as a place. To a nearly parallel projection, a scene reversed in depth and seen from the front is the same picture as the scene seen from behind — a narrow view keeps a second answer, inside out found the two differing only by how much the pinhole departs from parallel, 0.95 pixels at this field. A camera at the halfway point is behind the courtyard. A start with its camera behind the courtyard and its points reversed is the truth seen from the other side, and the adjustment settles in the truth. A start with its camera behind and its points true is the twin seen from the other side, and it settles in the twin.
So in the middle rows the points are not voting against their own sign. The vote is the product of two signs — which side the cameras are on, and which way out the points are — and the points only get to cast it when the cameras have swung far enough round to be behind. With the points left true, the cameras are brought home from every start up to 30 per cent of the way, already swung 109 degrees over and past straight overhead, and lost from 40 per cent, 145 degrees over.
Why the other side of the courtyard is a place the solver can reach
It is worth saying why a camera behind the courtyard is a legitimate configuration at all rather than a numerical accident, because it is the whole of the explanation.
Seven numbers no picture can name found that a set of pictures fixes a reconstruction only up to where the scene is, how it is turned and how big it is. Turning the whole reconstruction — every camera and every point together — by half a turn about the vertical changes no picture at all. So “the cameras behind the courtyard with the points reversed” is not a new answer; it is the twin, turned round, and its reversed points turned round are the scene’s own depth order again. The solver is free to arrive at the true reconstruction in that orientation, and the classifier that reads its depths in the first camera’s own frame correctly calls it the scene.
Far enough away, a pair is one eye measured what a pair loses as its scene recedes; the narrow field here is the same loss taken to the point where front and back are nearly the same picture. That is why the product rule is the right description: two signs, one for the cameras’ side and one for the points’ order, and a reconstruction is the truth when they agree and the twin when they do not.
What the band was measuring
This rereads the band the earlier essay found. Along the path that moves both halves together, a start 42 per cent of the way has its cameras swung about 150 degrees over the courtyard, well behind it, and its points at a sixth of their relief the right way out. By the product rule that is a twin-ward camera with scene-ward points, and which of the two wins depends on how far over the camera has gone. The band where neighbouring starts disagreed is the stretch where the camera is close enough to straight behind that the points’ small, flattened relief is enough to tip it.
The earlier reading — the sign of the depths is what an initialisation needs — was a correct description of that path and the wrong explanation of it. The diagonal of the split grid is the only line on which the depths and the cameras move together, and on that line the depths appear to decide because they are the thing that changes sign at halfway. Off it, they decide almost nothing.
It also explains the lottery. A scattered start near halfway has cameras swung almost exactly behind the courtyard, where they have no preference of their own, and points whose realised relief is small and noisy. The outcome is then set by a few points’ scatter in the particular configuration the solver descends through, which is exactly what the one-start-at-a-time figure shows: a sign that predicts the answer only weakly. And whichever answer the solver reaches, it reports an uncertainty for that valley alone, because an uncertainty is quoted from something and the other valley is not where it was quoted from.
What an initialisation has to get right
For a pipeline the practical reading is simpler than the rule it replaces, and it is about a quantity that is easier to get right.
Get the cameras on the right side. In every start measured, cameras swung up to 20 per cent of the way — about seventy degrees over the top — brought the truth home whatever the depths looked like: flattened, scattered or reversed. At 30 per cent, 109 degrees over, one depth setting in eleven was lost. That is a statement about orientation, not position: every blended camera here stays at its true distance from the target.
Do not rely on the depths to save a bad camera. A start whose cameras are on the twin’s side falls in even from exactly correct points. A relative orientation from a pair of pictures that has picked the wrong member of a narrow-field ambiguity is an initialisation that no quality of depth recovery will rescue.
And check the camera’s side, not the points’ order. The earlier essay recommended reading an initialisation’s depth order before any adjustment runs. The split grid says the thing to read is the cameras’: whether the recovered relative rotation between two views turns the way the camera actually moved. That is often known independently — a vehicle’s heading, a gyroscope, the order in which the pictures were taken along a walk — where the depth order of a flat, distant scene usually is not.
That places this beside four cameras fit, and one can see, which found four camera pairs that reproject two pictures exactly and a test that picks the right one. Here the two answers are not separated by any such test — every camera in both is in front of every point — but the same lesson holds in a weaker form: the discrete choice about the cameras is the one that decides, and the continuous part of the reconstruction is what an adjustment can be trusted with.
Where this reading stops
One path through the cameras. Every camera fraction blends the rotation vectors linearly, which carries a camera over the top. A different interpolation — across the arc in plan, say — would pass through different intermediate cameras and could put the boundary somewhere else. What would not change is the two ends: true cameras bring any depths home, and twin cameras take any depths in.
Only the points were scattered. The scatter here is on the points’ depths. A real initialisation also turns each camera by its own small error, and on the evidence of the split grid that is the scatter that matters; a camera turned a few degrees is far from the hundred-degree swing the boundary sits at, but no scattered-camera map was run.
Exact marks, one courtyard. As in the essays before it, every mark is exact and the scene is the standing courtyard; where the adjustment stops is what happens to a solve once marks are read to a pixel. Marks read to a pixel raise the twin’s floor and could move the boundary between the two cameras’ sides; they were not run.
Twelve starts a cell. The shares in the scatter figure are counts of twelve, and the difference between seven and five returns in twelve is not a finding.
Cameras first
Scattering a narrow-field adjustment’s starting depths blurs the switch between returning and falling into the inside-out twin — a clean change between 42 and 46 per cent of the way becomes a band of mixed outcomes from 38 to 58 — and leaves it centred where it was. But no start’s own depths predict it: of 80 scattered starts with correctly signed relief, 26 fell in, and of 100 reversed ones, 43 came home.
Split the start into its cameras and its points and the reason is plain. With the cameras true, every start returns, the twin’s own inside-out points included; with the cameras on the twin’s side, every start falls in, the true points included. The points matter only in the middle rows, where a blended camera has swung over the top of the courtyard to look back from behind — and there the vote is the product of which side the cameras are on and which way out the points are, because a reversed scene seen from the front is, to a narrow field, the scene seen from behind. What an initialisation needs is its cameras on the right side of the scene.
Still open: how far a start’s cameras can be turned before the side is lost
The boundary measured here is on one path, which turns every camera over the top by the same fraction. Real initialisations err differently: each camera by its own small rotation, and more often about the vertical — a mis-estimated heading — than over the top.
That splits the question into two measurements. The first scatters each camera’s orientation independently about the truth, with the points left true, and finds the scatter at which starts begin to fall in: if it is tens of degrees, any pairwise initialisation that picked the right member of the ambiguity is safe; if it is a few degrees, it is not. The second turns the cameras about the vertical rather than over the top — squeezing or widening the arc — and asks whether an arc squeezed flat, with every camera looking from the same direction, has lost its side the way a camera halfway over the top has.
If both boundaries sit far from anything a pairwise start produces, then the whole of a narrow-field initialisation’s risk is in one discrete choice per pair of pictures, and that choice is the thing to check.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A mirror ball is an equal-area fisheye — both name field of view, orthographic
- A survey is trusted at its own accuracy, unless its error has a shape — both name bundle adjustment, reprojection error
- Any three lines you draw are a cube — both name depth reversal, orthographic
- The eighth held number bends the scene — both name bundle adjustment, reprojection error
- What the removed roof buys — both name field of view, orthographic
Named objects
A flat tag is an object no other essay names yet.
bundle adjustmentConvergenceDepth reversalfield of viewLevenberg–MarquardtOrthographicReprojection errorstructure from motion