How flat is flat enough
Worth reading first: A flat scene fixes no second eye.
A flat scene fixes no second eye established that marks on one plane leave a two-parameter family of fundamental matrices, and it ended on a question it could not answer: a real wall is not a plane, a real facade has mullions and sills, and a real car park has kerbs. How much relief does a scene need before the family collapses to an answer? The ladder of assumptions is a ladder of conditioning is the shape of the answer and not the answer itself.
The answer turns out to have two halves that disagree completely, and the disagreement is the essay.
With exact marks there is no transition
Sweep the out-of-plane spread from zero to half the scene’s own extent, with marks read exactly, and the recovered translation direction is 84.4 degrees wrong at exactly coplanar and 0.000 degrees wrong at every non-zero relief measured — including eight parts in ten thousand, which puts the furthest mark two hundredths of a millimetre off the plane on a scene metres across.
That is a step function. A scene that is not coplanar is not degenerate, however nearly coplanar it is, and the algebra finds the answer with no difficulty whatever. The design matrix’s second smallest singular value lifts off the floor the instant the relief does, and once it is off the floor there is a unique nullspace and the nullspace is right.
With a hand in it there is nothing but transition
Put three tenths of a pixel of reading error into every mark, take the median over fifteen draws so the curve is not a picture of one seed, and the same sweep is a slope three decades wide:
- exactly flat: 60.0 degrees wrong;
- two parts in a thousand of relief: 65.3 degrees;
- twelve parts in a thousand: 40.8 degrees;
- three per cent: 23.1 degrees;
- eight per cent: 7.5 degrees;
- twenty per cent: 2.8 degrees;
- half: 1.1 degrees.
Nothing about the geometry changed between the two sweeps. The width of the transition is the reading error.
The quantity the transition is actually against
Relief measured as a fraction of the scene’s extent is a convenient axis and it is not the axis the geometry sees. What a solver has is two photographs, and what a nearly-flat scene offers it is the amount by which a single homography fails to explain the marks — the out-of-plane parallax, in pixels.
That number is plotted beside the error, and it is the one that makes the crossing legible. At exactly flat it is 0.65 pixels, which is entirely the reading error being fitted. At three per cent relief it is 1.07 pixels, at eight per cent 2.85, at half 14.3 — and the pose error falls below ten degrees between the third and fourth of those, which is where the parallax reaches about ten times the marking error.
So the rule that falls out is not about relief at all: a two-view recovery needs out-of-plane parallax of roughly ten times the marking error, and how much relief that takes depends on the baseline, the distance and the focal length as well as on the scene.
Why the noisy curve is not monotone at the flat end
The first four points of the noisy sweep sit between 59 and 65 degrees with no order to them, and that is worth explaining rather than smoothing.
At and near exactly flat, the answer is not determined and the recovered pose is essentially the reading error’s own choice among the family. The median over fifteen draws is a median of samples from a distribution with no centre near the truth, so it wanders. It is not measuring how wrong the answer is; it is measuring where the noise happened to land in a two-parameter family.
That is why the parallax curve is the better instrument at the flat end: it is monotone throughout, it has a meaning at zero relief, and it says how much genuine evidence the arrangement contains rather than how far one draw of the noise landed from the truth.
The same shape, met before
A turn of the head is not a step sideways measured the other degeneracy of a two-view pair — the baseline going to zero — and found the identical structure: the algebra does not notice, a two-millimetre baseline recovers a courtyard to nine digits from exact marks, and what fails is the ratio of baseline to reading error.
Two independent degeneracies, two independent sweeps, one conclusion: a degeneracy is a fact about the geometry, and the width of its neighbourhood is a fact about the instrument. That is worth carrying because it decides what to do about a bad reconstruction. Nothing about the scene needs changing if the marks can be read more carefully; nothing about the marks helps if the arrangement is exactly degenerate.
What “enough relief” costs to arrange
Turning the rule into something a person with a camera can act on takes one more step, because relief is not free to add.
The parallax a point at depth Z₀ + δ offers over a plane at depth Z₀ is roughly f·b·δ/Z₀², with b the baseline and f the focal length in pixels. At f = 620, b = 0.30 metres and a scene at six metres that is 5.2 pixels for every metre of depth difference, so ten times a half-pixel hand needs about a metre of relief — a projecting bay, a parked car, a doorway a person could stand in. A kerb will not do it.
Which is why the failure is common rather than exotic, and why it gets worse so fast with distance: the requirement is , quadratic in the distance to the subject. The same rig needs 24 cm of relief at three metres, a metre at six, and 3.9 metres at twelve — so a facade photographed from the far pavement needs relief on the scale of a whole projecting wing, while the same facade at arm’s length is settled by a window reveal. A photograph taken square-on from across the street has plenty of texture and almost no relief at the scale the arithmetic needs, and the marks are all beautifully sharp — which is the arrangement flattening a facade is designed for and a two-view solver is not. Sharpness is not the resource being spent.
The three quantities, and the one product they make
The closing section says the relief, the baseline and the marking error enter as one product. It is worth writing that product down, because the sweep’s own numbers determine its constant and the constant makes the rule usable on an arrangement that is not this one.
The parallax a mark of relief offers over a plane at , from a baseline at focal length pixels, is . Writing the relief as a fraction of the scene’s own extent makes it — linear in , with a constant fixed by the arrangement.
The sweep supplies that constant. Its parallax readings are 1.07, 2.85 and 14.3 pixels at three per cent, eight per cent and half, against 0.65 at exactly flat where the whole reading is the noise being fitted. Removing the noise floor in quadrature leaves 0.85, 2.78 and 14.29, which divided by the relief fraction give 28.3, 34.7 and 28.6 — one constant, about 30 pixels of parallax per unit of relief fraction, across a factor of seventeen in the answer.
That closes the arithmetic. The recovery becomes useful at ten times the marking error, so the crossing is at : three per cent of relief at a tenth of a pixel, ten per cent at three tenths, sixty-six per cent at two pixels. Those are the three settings the slider draws, and the crossing moves exactly where the rule says.
Two readings follow, and they are the practical content of the whole sweep.
The three quantities are interchangeable, exactly. Halving the marking error, doubling the baseline and doubling the relief all move the crossing by the same factor, because all three enter the product once. A photographer who cannot add relief can step further sideways, and one who can do neither can mark more carefully, and the arithmetic does not care which.
The distance is not interchangeable with any of them, because it enters squared. Doubling the distance to the subject costs a factor of four in relief, or four in baseline, or four in marking precision. That asymmetry is why the failure is so strongly associated with photographing large flat things: a facade is large, so it is photographed from far away, and far away is the one variable that hurts twice.
It also says which lever is cheapest, and it is the surprising one. Stepping two metres sideways rather than one costs nothing at all and doubles the evidence; reading marks to a fifteenth of a pixel rather than a tenth is hard work; adding a metre of relief to a wall is not available. The baseline is the free variable in almost every arrangement, and it is the one people are most reluctant to spend, because a wide step feels like it will make the two pictures harder to match — which it does, and by an amount this collection does not measure.
Why more marks do not move the crossing
The obvious response to a noisy sweep is to use more marks, and it is worth saying what that buys because the answer is not what a reader expects.
More marks reduce the variance of the fit — the spread of answers over draws of the noise — and they do that on both sides of the crossing. What they do not do is create parallax. At exactly flat, forty marks fit the family exactly as well as eight do, and the median error falls from 85 degrees to 48 rather than toward zero: the sequence converges, and it converges on a wrong answer. On a scene with relief the same increase takes the error from 19.8 degrees to under a degree.
That is the distinction an ambiguity is not an uncertainty is written about, and it is why the axis of this essay is relief rather than mark count. Marks buy precision about whatever the evidence determines. Relief decides what the evidence determines.
What a reader with a camera should take from the sweep
Three things, in the order they matter.
Photograph the relief, not the texture. A wall covered in detail and a wall with a doorway in it are the same to a feature detector and completely different to a two-view solver. Include something at a visibly different depth, and include it near the middle of the frame rather than at a corner where it will be cropped or vignetted.
Move further than feels necessary. Parallax is the baseline over the depth, so doubling the step doubles the evidence, and a step that looks large in the viewfinder is usually a few per cent of the distance to the subject. This is the same advice another picture of the same sweep arrives at from the many-views side — angular spread is what a reconstruction is short of, not photographs.
And read the residual rather than the reconstruction. A recovered pose always looks like an answer. The homography’s residual in pixels, next to an honest estimate of the marking error, says whether it is one.
The instrument a solver can actually use
None of the numbers above are available to a program with two photographs and no truth. What is available is the design matrix’s spectrum and the homography’s residual, and both work.
The spectrum says whether the second smallest singular value is negligible against the largest — 3.4 × 10⁻⁹ at exactly flat against 5.2 × 10⁻³ with depth — and it is computed from the marks alone.
The homography residual says how much of the disagreement between the pictures a single plane map cannot explain, in pixels, and it is directly comparable with the expected marking error. That is the more useful of the two because it has units a reader understands and because it is the quantity the geometry actually cares about.
Fitting both models and comparing is therefore not a heuristic. It is the measurement this essay is about, run at solve time — the same discipline recovering the camera applies to a single view, where the residual of a fit is the evidence about the fit.
What this does not settle
Two things, stated because the sweep looks more general than it is.
The sweep uses one arrangement of marks. Twenty-four points, one plane, one baseline, one focal length, one distance. The shape of the result — a step in exact arithmetic, a slope of the marking error’s width in practice — is the claim, and the numbers on the axes belong to this arrangement.
And “relief as a fraction of extent” is a proxy. The geometry cares about parallax in pixels, and that is why parallax is plotted beside it — the same substitution depth is a reciprocal makes when it stops measuring depth and starts measuring disparity. A scene of the same relief photographed from twice as far offers half the parallax, which is far enough away, a pair is one eye approaching the same boundary from the other side.
The two failures on one axis
Putting the exact and the noisy sweeps on the same plot is the point of the hero figure, and it is worth saying what a reader should look at rather than at.
The vertical gap at the left-hand end is the whole story: at exactly flat both sweeps are wrong, and everywhere else the exact one is right and the noisy one is not. That gap is the reading error, drawn.
The horizontal position of the noisy curve’s fall is what a reader can move. Halving the marking error moves it left by a factor of two in relief, which is visible in the slider: at a tenth of a pixel the crossing sits at about a per cent of relief and at two pixels it sits past twenty per cent.
And the height of the plateau at the left is not a measurement of anything. It is where a two-parameter family and a noise draw happened to intersect, and reading it as “the error at flat” would be reading the random number generator — which is why the parallax curve is drawn beside it.
Why a fraction of the extent is the axis, and not a length
A note on the sweep’s horizontal axis, because the choice of it is doing work.
Relief could have been plotted as a length in metres, and that would have made the curve a fact about one scene at one distance. Plotted as a fraction of the scene’s own extent it is dimensionless, so the same curve describes a two-metre table at a metre and a fifty-metre facade at twenty-five: the arrangement is scaled, the picture is identical, and the parallax in pixels is the same.
That is the same scaling argument the collection makes everywhere about pictures — nothing in a photograph carries a size, so every claim about one is a ratio — and it is why the axis is written the way it is.
What the scaling does not absorb is the baseline. Halving the baseline halves the parallax at every relief, so it moves the whole curve right by a factor of two, exactly as halving the marking error moves it left. The three quantities enter as one product, and the sweep is a slice through it at one setting of the other two.
The short version
A scene that is not coplanar is not degenerate, and in exact arithmetic the transition from a family of answers to an answer has no width at all — 84 degrees of error at exactly flat, zero at eight parts in ten thousand of relief.
With three tenths of a pixel of reading error the same sweep is a slope three decades wide, and the recovery becomes useful when the out-of-plane parallax reaches about ten times the marking error, which here is around three pixels. The width of the transition is therefore a property of the hand and not of the geometry, and the quantity to watch is a residual in pixels rather than a relief in metres.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A light far enough away — both name conditioning, error propagation, instrument limit
- A picture with two eyes in it — both name degeneracy, error propagation, parallax
- An area, out of one photograph — both name conditioning, error propagation, homography
- Eight points and the basis they are read in — both name conditioning, fundamental matrix, singular values
- How well the floor has to be known — both name error propagation, homography, instrument limit
- Seven marks, three answers — both name critical surface, epipolar geometry, fundamental matrix
Named objects
A flat tag is an object no other essay names yet.
ConditioningCritical surfaceDegeneracyDegenerate configurationEpipolar geometryerror propagationFundamental matrixHomographyinstrument limitParallaxsingular values