The second eye

How flat is flat enough

With exact marks the transition has no width at all — 84° of pose error at exactly coplanar and 0.000° at eight parts in ten thousand of relief. Put three tenths of a pixel of reading error in and the same sweep becomes a slope three decades wide, crossing into usefulness when the out-of-plane parallax reaches about ten times the marking error.

Worth reading first: A flat scene fixes no second eye.

A flat scene fixes no second eye established that marks on one plane leave a two-parameter family of fundamental matrices, and it ended on a question it could not answer: a real wall is not a plane, a real facade has mullions and sills, and a real car park has kerbs. How much relief does a scene need before the family collapses to an answer? The ladder of assumptions is a ladder of conditioning is the shape of the answer and not the answer itself.

The answer turns out to have two halves that disagree completely, and the disagreement is the essay.

With exact marks there is no transition

Sweep the out-of-plane spread from zero to half the scene’s own extent, with marks read exactly, and the recovered translation direction is 84.4 degrees wrong at exactly coplanar and 0.000 degrees wrong at every non-zero relief measured — including eight parts in ten thousand, which puts the furthest mark two hundredths of a millimetre off the plane on a scene metres across.

That is a step function. A scene that is not coplanar is not degenerate, however nearly coplanar it is, and the algebra finds the answer with no difficulty whatever. The design matrix’s second smallest singular value lifts off the floor the instant the relief does, and once it is off the floor there is a unique nullspace and the nullspace is right.

In exact arithmetic the transition has no width; at 0.05 px of reading error it is three decades wideTwo curves against how far the scene departs from being flat. With exact marks the recovery is 84° wrong at exactly flat and 1e-5° wrong at eight parts in ten thousand of relief — a step function, because a scene that is not coplanar is not degenerate however nearly coplanar it is. With 0.05 pixels of reading error the same sweep is a slope: the recovery is useless until the out-of-plane parallax the arrangement offers rises past about ten times the reading error, which here is at a relief of about 0.012 and a parallax of 0.4 pixels. **The width of the transition is the reading error, not the geometry.**0306090-3-2-1out-of-plane spread, as a fraction of the scene's own extent (powers of ten)error in the recovered translation direction, in degreeswith 0.05 px of reading errorexact marksthe geometry is a step and the measurement is a slope9 reliefs
Fig. 1 The two curves this essay is about. The lower one is exact marks — a step at zero and nothing after it.

With a hand in it there is nothing but transition

Put three tenths of a pixel of reading error into every mark, take the median over fifteen draws so the curve is not a picture of one seed, and the same sweep is a slope three decades wide:

  • exactly flat: 60.0 degrees wrong;
  • two parts in a thousand of relief: 65.3 degrees;
  • twelve parts in a thousand: 40.8 degrees;
  • three per cent: 23.1 degrees;
  • eight per cent: 7.5 degrees;
  • twenty per cent: 2.8 degrees;
  • half: 1.1 degrees.

Nothing about the geometry changed between the two sweeps. The width of the transition is the reading error.

The quantity the transition is actually against

Relief measured as a fraction of the scene’s extent is a convenient axis and it is not the axis the geometry sees. What a solver has is two photographs, and what a nearly-flat scene offers it is the amount by which a single homography fails to explain the marks — the out-of-plane parallax, in pixels.

That number is plotted beside the error, and it is the one that makes the crossing legible. At exactly flat it is 0.65 pixels, which is entirely the reading error being fitted. At three per cent relief it is 1.07 pixels, at eight per cent 2.85, at half 14.3 — and the pose error falls below ten degrees between the third and fourth of those, which is where the parallax reaches about ten times the marking error.

So the rule that falls out is not about relief at all: a two-view recovery needs out-of-plane parallax of roughly ten times the marking error, and how much relief that takes depends on the baseline, the distance and the focal length as well as on the scene.

A flat scene has two vanishing singular values and a solid one has oneThe nine singular values of the eight-point design matrix, on a logarithmic scale against the largest. For a scene with depth in it, one of them vanishes: that single null direction is the fundamental matrix, and having exactly one is what having an answer means. For the same marks pressed flat onto a plane, THREE vanish — the second is at 3.4e-9 of the largest against 7.7e-4 for the solid scene — so the nullspace has three dimensions and every matrix in it fits. The gap between the last two curves is the whole difference between a measurement and a family of them, and it is visible in the spectrum without knowing the answer.0-4-8-12-1613579singular value, largest to smallestsize against the largest, in powers of tena scene with depththe same marks, flatthe second smallest is the diagnostic3.4e-9 against 7.7e-4
Fig. 2 The spectrum at three per cent relief, where the second smallest singular value has lifted and the answer exists but is not yet worth much.

Why the noisy curve is not monotone at the flat end

The first four points of the noisy sweep sit between 59 and 65 degrees with no order to them, and that is worth explaining rather than smoothing.

At and near exactly flat, the answer is not determined and the recovered pose is essentially the reading error’s own choice among the family. The median over fifteen draws is a median of samples from a distribution with no centre near the truth, so it wanders. It is not measuring how wrong the answer is; it is measuring where the noise happened to land in a two-parameter family.

That is why the parallax curve is the better instrument at the flat end: it is monotone throughout, it has a meaning at zero relief, and it says how much genuine evidence the arrangement contains rather than how far one draw of the noise landed from the truth.

The same shape, met before

A turn of the head is not a step sideways measured the other degeneracy of a two-view pair — the baseline going to zero — and found the identical structure: the algebra does not notice, a two-millimetre baseline recovers a courtyard to nine digits from exact marks, and what fails is the ratio of baseline to reading error.

Two independent degeneracies, two independent sweeps, one conclusion: a degeneracy is a fact about the geometry, and the width of its neighbourhood is a fact about the instrument. That is worth carrying because it decides what to do about a bad reconstruction. Nothing about the scene needs changing if the marks can be read more carefully; nothing about the marks helps if the arrangement is exactly degenerate.

What “enough relief” costs to arrange

Turning the rule into something a person with a camera can act on takes one more step, because relief is not free to add.

The parallax a point at depth Z₀ + δ offers over a plane at depth Z₀ is roughly f·b·δ/Z₀², with b the baseline and f the focal length in pixels. At f = 620, b = 0.30 metres and a scene at six metres that is 5.2 pixels for every metre of depth difference, so ten times a half-pixel hand needs about a metre of relief — a projecting bay, a parked car, a doorway a person could stand in. A kerb will not do it.

Which is why the failure is common rather than exotic, and why it gets worse so fast with distance: the requirement is δ=10σZ02/fb\delta = 10\sigma Z_{0}^{2}/fb, quadratic in the distance to the subject. The same rig needs 24 cm of relief at three metres, a metre at six, and 3.9 metres at twelve — so a facade photographed from the far pavement needs relief on the scale of a whole projecting wing, while the same facade at arm’s length is settled by a window reveal. A photograph taken square-on from across the street has plenty of texture and almost no relief at the scale the arithmetic needs, and the marks are all beautifully sharp — which is the arrangement flattening a facade is designed for and a two-view solver is not. Sharpness is not the resource being spent.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 4.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.010203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 4.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 3 The relation the arithmetic above is one line of, from the depth field: what a baseline buys at a given depth.

The three quantities, and the one product they make

The closing section says the relief, the baseline and the marking error enter as one product. It is worth writing that product down, because the sweep’s own numbers determine its constant and the constant makes the rule usable on an arrangement that is not this one.

The parallax a mark of relief δ\delta offers over a plane at Z0Z_{0}, from a baseline bb at focal length ff pixels, is fbδ/Z02f b \delta / Z_{0}^{2}. Writing the relief as a fraction ρ\rho of the scene’s own extent LL makes it fbLρ/Z02f b L \rho / Z_{0}^{2} — linear in ρ\rho, with a constant fixed by the arrangement.

The sweep supplies that constant. Its parallax readings are 1.07, 2.85 and 14.3 pixels at three per cent, eight per cent and half, against 0.65 at exactly flat where the whole reading is the noise being fitted. Removing the noise floor in quadrature leaves 0.85, 2.78 and 14.29, which divided by the relief fraction give 28.3, 34.7 and 28.6 — one constant, about 30 pixels of parallax per unit of relief fraction, across a factor of seventeen in the answer.

That closes the arithmetic. The recovery becomes useful at ten times the marking error, so the crossing is at ρ10σ/30\rho \approx 10\sigma / 30: three per cent of relief at a tenth of a pixel, ten per cent at three tenths, sixty-six per cent at two pixels. Those are the three settings the slider draws, and the crossing moves exactly where the rule says.

Two readings follow, and they are the practical content of the whole sweep.

The three quantities are interchangeable, exactly. Halving the marking error, doubling the baseline and doubling the relief all move the crossing by the same factor, because all three enter the product once. A photographer who cannot add relief can step further sideways, and one who can do neither can mark more carefully, and the arithmetic does not care which.

The distance is not interchangeable with any of them, because it enters squared. Doubling the distance to the subject costs a factor of four in relief, or four in baseline, or four in marking precision. That asymmetry is why the failure is so strongly associated with photographing large flat things: a facade is large, so it is photographed from far away, and far away is the one variable that hurts twice.

It also says which lever is cheapest, and it is the surprising one. Stepping two metres sideways rather than one costs nothing at all and doubles the evidence; reading marks to a fifteenth of a pixel rather than a tenth is hard work; adding a metre of relief to a wall is not available. The baseline is the free variable in almost every arrangement, and it is the one people are most reluctant to spend, because a wide step feels like it will make the two pictures harder to match — which it does, and by an amount this collection does not measure.

Why more marks do not move the crossing

The obvious response to a noisy sweep is to use more marks, and it is worth saying what that buys because the answer is not what a reader expects.

More marks reduce the variance of the fit — the spread of answers over draws of the noise — and they do that on both sides of the crossing. What they do not do is create parallax. At exactly flat, forty marks fit the family exactly as well as eight do, and the median error falls from 85 degrees to 48 rather than toward zero: the sequence converges, and it converges on a wrong answer. On a scene with relief the same increase takes the error from 19.8 degrees to under a degree.

That is the distinction an ambiguity is not an uncertainty is written about, and it is why the axis of this essay is relief rather than mark count. Marks buy precision about whatever the evidence determines. Relief decides what the evidence determines.

More marks cure an uncertainty and leave an ambiguity where it wasTwo sequences at the same reading error. On a scene with depth, going from eight marks to forty takes the error from 19.83° to 1.78° — evidence accumulating in the ordinary way. On the same marks pressed flat, the sequence also settles, and it settles at 48°: the reading error lifts each mark a little off the plane, so the fit is not exactly degenerate and does converge, onto an answer that is nowhere near the right one. That is the practical difference between the two failures — one of them is fixed by taking more care and the other is not fixed by anything but moving.00.50011.50210203040how many marks the fit was givenerror in the recovered translation direction, in degrees (powers of ten)flatwith depththe same reading error in both5 counts
Fig. 4 Eight marks to forty, on both kinds of scene. One sequence goes to the answer and the other goes somewhere and stays.

What a reader with a camera should take from the sweep

Three things, in the order they matter.

Photograph the relief, not the texture. A wall covered in detail and a wall with a doorway in it are the same to a feature detector and completely different to a two-view solver. Include something at a visibly different depth, and include it near the middle of the frame rather than at a corner where it will be cropped or vignetted.

Move further than feels necessary. Parallax is the baseline over the depth, so doubling the step doubles the evidence, and a step that looks large in the viewfinder is usually a few per cent of the distance to the subject. This is the same advice another picture of the same sweep arrives at from the many-views side — angular spread is what a reconstruction is short of, not photographs.

And read the residual rather than the reconstruction. A recovered pose always looks like an answer. The homography’s residual in pixels, next to an honest estimate of the marking error, says whether it is one.

Another picture of the same sweep buys nothingTwo ways of adding views to the courtyard, every mark read to 1 px. Filling in a fixed 60° sweep leaves the worst camera-centre error at 2.0e-3 where 3 views gave 1.1e-3. Widening the sweep by 12° per view improves it from 3.2e-3 to 1.3e-3 and then flattens as well. What the reconstruction is short of is angular spread, not pictures. At every point on both curves the Jacobian has exactly 7 flat directions.0.0020.00334567number of viewsworst camera-centre error (fraction of the track's mean radius, log scale)7 flat7 flat7 flat7 flat7 flat60° sweep, filled in12° per view, wideningfilled in: 1.1e-3 → 2.0e-3widened: 3.2e-3 → 1.8e-3
Fig. 5 The many-views version of the second point: what widening the arrangement buys and what filling it in does not.

The instrument a solver can actually use

None of the numbers above are available to a program with two photographs and no truth. What is available is the design matrix’s spectrum and the homography’s residual, and both work.

The spectrum says whether the second smallest singular value is negligible against the largest — 3.4 × 10⁻⁹ at exactly flat against 5.2 × 10⁻³ with depth — and it is computed from the marks alone.

The homography residual says how much of the disagreement between the pictures a single plane map cannot explain, in pixels, and it is directly comparable with the expected marking error. That is the more useful of the two because it has units a reader understands and because it is the quantity the geometry actually cares about.

Fitting both models and comparing is therefore not a heuristic. It is the measurement this essay is about, run at solve time — the same discipline recovering the camera applies to a single view, where the residual of a fit is the evidence about the fit.

A flat scene says everything about the plane and nothing about the eyeFour bars, on a scale of digits. Fitted to four of the marks and tested on all 24, the homography relating the two photographs of the FLAT scene reproduces every mark to 5.9e-13 pixels — the whole relationship between the pictures, from four correspondences, exactly. Fitted to the same four marks of the scene WITH depth it is out by 29.6 pixels, because there is no such map. The fundamental matrix runs the other way: undetermined on the plane, where three combinations of the nullspace are 1.41 apart and all fit, and determined off it. Nothing is lost by the flatness; a different question is being answered.flat — the plane's own map5.9e-13flat — a second eye1.4e+0with depth — the plane's map3.0e+1with depth — a second eye1.4e+1digits, on the same two photographsfour readings
Fig. 6 The two models on the two scenes, which is the comparison a solver makes and the diagnosis it can make from photographs alone.

What this does not settle

Two things, stated because the sweep looks more general than it is.

The sweep uses one arrangement of marks. Twenty-four points, one plane, one baseline, one focal length, one distance. The shape of the result — a step in exact arithmetic, a slope of the marking error’s width in practice — is the claim, and the numbers on the axes belong to this arrangement.

And “relief as a fraction of extent” is a proxy. The geometry cares about parallax in pixels, and that is why parallax is plotted beside it — the same substitution depth is a reciprocal makes when it stops measuring depth and starts measuring disparity. A scene of the same relief photographed from twice as far offers half the parallax, which is far enough away, a pair is one eye approaching the same boundary from the other side.

The two failures on one axis

Putting the exact and the noisy sweeps on the same plot is the point of the hero figure, and it is worth saying what a reader should look at rather than at.

The vertical gap at the left-hand end is the whole story: at exactly flat both sweeps are wrong, and everywhere else the exact one is right and the noisy one is not. That gap is the reading error, drawn.

The horizontal position of the noisy curve’s fall is what a reader can move. Halving the marking error moves it left by a factor of two in relief, which is visible in the slider: at a tenth of a pixel the crossing sits at about a per cent of relief and at two pixels it sits past twenty per cent.

And the height of the plateau at the left is not a measurement of anything. It is where a two-parameter family and a noise draw happened to intersect, and reading it as “the error at flat” would be reading the random number generator — which is why the parallax curve is drawn beside it.

In exact arithmetic the transition has no width; at 0.3 px of reading error it is three decades wideTwo curves against how far the scene departs from being flat. With exact marks the recovery is 84° wrong at exactly flat and 1e-5° wrong at eight parts in ten thousand of relief — a step function, because a scene that is not coplanar is not degenerate however nearly coplanar it is. With 0.3 pixels of reading error the same sweep is a slope: the recovery is useless until the out-of-plane parallax the arrangement offers rises past about ten times the reading error, which here is at a relief of about 0.08 and a parallax of 2.9 pixels. **The width of the transition is the reading error, not the geometry.**0306090-3-2-1out-of-plane spread, as a fraction of the scene's own extent (powers of ten)error in the recovered translation direction, in degreeswith 0.3 px of reading errorexact marksthe geometry is a step and the measurement is a slope9 reliefs
Fig. 7 The hero figure again at the setting the numbers in this essay are quoted from, for the comparison the paragraph above describes.

Why a fraction of the extent is the axis, and not a length

A note on the sweep’s horizontal axis, because the choice of it is doing work.

Relief could have been plotted as a length in metres, and that would have made the curve a fact about one scene at one distance. Plotted as a fraction of the scene’s own extent it is dimensionless, so the same curve describes a two-metre table at a metre and a fifty-metre facade at twenty-five: the arrangement is scaled, the picture is identical, and the parallax in pixels is the same.

That is the same scaling argument the collection makes everywhere about pictures — nothing in a photograph carries a size, so every claim about one is a ratio — and it is why the axis is written the way it is.

What the scaling does not absorb is the baseline. Halving the baseline halves the parallax at every relief, so it moves the whole curve right by a factor of two, exactly as halving the marking error moves it left. The three quantities enter as one product, and the sweep is a slice through it at one setting of the other two.

The short version

A scene that is not coplanar is not degenerate, and in exact arithmetic the transition from a family of answers to an answer has no width at all — 84 degrees of error at exactly flat, zero at eight parts in ten thousand of relief.

With three tenths of a pixel of reading error the same sweep is a slope three decades wide, and the recovery becomes useful when the out-of-plane parallax reaches about ten times the marking error, which here is around three pixels. The width of the transition is therefore a property of the hand and not of the geometry, and the quantity to watch is a residual in pixels rather than a relief in metres.

In exact arithmetic the transition has no width; at 0.8 px of reading error it is three decades wideTwo curves against how far the scene departs from being flat. With exact marks the recovery is 84° wrong at exactly flat and 1e-5° wrong at eight parts in ten thousand of relief — a step function, because a scene that is not coplanar is not degenerate however nearly coplanar it is. With 0.8 pixels of reading error the same sweep is a slope: the recovery is useless until the out-of-plane parallax the arrangement offers rises past about ten times the reading error, which here is at a relief of about 0.2 and a parallax of 6.6 pixels. **The width of the transition is the reading error, not the geometry.**0306090-3-2-1out-of-plane spread, as a fraction of the scene's own extent (powers of ten)error in the recovered translation direction, in degreeswith 0.8 px of reading errorexact marksthe geometry is a step and the measurement is a slope9 reliefs
Fig. 8 Eight tenths of a pixel of hand, where the crossing moves to the right by exactly the factor the rule predicts.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

ConditioningCritical surfaceDegeneracyDegenerate configurationEpipolar geometryerror propagationFundamental matrixHomographyinstrument limitParallaxsingular values