The second eye

An ambiguity is not an uncertainty

Eight marks to forty cuts a solid scene's pose error from 19.8° to 0.7° and leaves a flat one at 48°. The two failures look identical from inside — a confident answer, a residual at the floor — and they respond to opposite remedies, so telling them apart is worth more than either measurement.

Worth reading first: The surface two pictures cannot separate · A flat scene fixes no second eye.

This round measured four ways for a two-view recovery to fail, and they sort into two kinds. The sorting is the useful part, because the two kinds respond to opposite remedies and are indistinguishable from inside a solver.

An uncertainty is a direction in which the answer is poorly determined: a small singular value, a wide covariance, a number that moves when a pixel moves. More evidence shrinks it.

An ambiguity is a direction in which the answer is not determined at all: an exact zero in the spectrum, a nullspace with dimension, a finite set of alternatives. More evidence does not shrink it, because every alternative already explains every measurement exactly.

The measurement that separates them

The same marks, the same reading error, the same solver. On a scene with depth, going from eight marks to forty takes the pose error from 19.8 degrees to 0.68, with a dip and a small rise at the end that is the noise draw. On the same marks pressed flat, it goes from 85.0 to 47.7 — and stays there.

The flat sequence does improve, which is the honest part and the part a simpler story would have got wrong. Reading error lifts every mark a little off the plane, so a noisy coplanar scene is not exactly degenerate and the fit does tighten. What it tightens onto is not the answer. It converges, and it converges forty-eight degrees away.

At forty marks the two are a factor of twenty-seven apart. That is the number worth carrying: not “one improves and one does not”, but “both converge, and only one converges on the truth”.

More marks cure an uncertainty and leave an ambiguity where it wasTwo sequences at the same reading error. On a scene with depth, going from eight marks to forty takes the error from 19.83° to 1.78° — evidence accumulating in the ordinary way. On the same marks pressed flat, the sequence also settles, and it settles at 48°: the reading error lifts each mark a little off the plane, so the fit is not exactly degenerate and does converge, onto an answer that is nowhere near the right one. That is the practical difference between the two failures — one of them is fixed by taking more care and the other is not fixed by anything but moving.00.50011.50210203040how many marks the fit was givenerror in the recovered translation direction, in degrees (powers of ten)flatwith depththe same reading error in both5 counts
Fig. 1 Eight marks to forty on both kinds of scene, at the same reading error. Both sequences settle; they settle in different places.

Why the two look identical from inside

A solver has the marks, the fitted matrix and the residual. All three are the same in both cases.

The residual is at the floor for the flat scene, because every member of the family satisfies every mark. The fitted matrix is a perfectly good rank-two matrix. The reprojection error of the reconstruction is at the floor as well, because the points were triangulated against the pose that produced them.

There is no quantity computed by an ordinary pipeline that distinguishes the two. That is not a criticism of pipelines; it is the definition of an ambiguity. The evidence is consistent with several answers, so nothing computed from the evidence can prefer one.

One flat scene, three epipolar geometries, all exact to 1.6e-6 pxTwo photographs of a scene whose points all lie on one plane, with the marks shown in both. Over the right-hand picture are drawn the epipolar lines of 5 of the marks under three different fundamental matrices — three members of a two-parameter family that all satisfy the epipolar constraint at every mark to 1.6e-6 pixels. They are genuinely different matrices, 1.41 apart after normalisation, and they put the second eye in three different places. No measurement on these photographs can choose between them.left pictureright pictureevery one of them fits every mark1.6e-6 px
Fig. 2 Three answers over one pair of photographs, each with a residual at the arithmetic floor.

What does distinguish them

Three things, in order of cost.

The spectrum, where it applies. A plane’s degeneracy shows as a second vanishing singular value — 3.4 × 10⁻⁹ against 5.2 × 10⁻³ — and that is computable from the marks alone. It catches the plane. On the surface two pictures cannot separate it fires as well, and reports two vanishing values rather than three — a pencil rather than a continuum — which is a signal that a solver has no remedy for, since every further mark is on the surface too.

A model comparison. Fit a homography as well as a fundamental matrix and see which explains the marks better. That catches every planar case, including the ones the spectrum would report as merely poorly conditioned.

And evidence the alternatives disagree about. A withheld mark separates the three answers of seven marks, three answers by an order of magnitude. It separates the flat scene’s family by nothing at all, because the withheld mark is on the plane too. So the test is not “hold something back” — it is “hold back something the alternatives disagree about”, which is a statement about geometry rather than about statistics.

A flat scene has two vanishing singular values and a solid one has oneThe nine singular values of the eight-point design matrix, on a logarithmic scale against the largest. For a scene with depth in it, one of them vanishes: that single null direction is the fundamental matrix, and having exactly one is what having an answer means. For the same marks pressed flat onto a plane, THREE vanish — the second is at 3.4e-9 of the largest against 5.1e-3 for the solid scene — so the nullspace has three dimensions and every matrix in it fits. The gap between the last two curves is the whole difference between a measurement and a family of them, and it is visible in the spectrum without knowing the answer.0-4-8-12-1613579singular value, largest to smallestsize against the largest, in powers of tena scene with depththe same marks, flatthe second smallest is the diagnostic3.4e-9 against 5.1e-3
Fig. 3 The cheapest of the three, and the one that works on the commonest case.

The remedies are opposite

This is why the distinction is worth the essay rather than the paragraph.

For an uncertainty: read the marks more carefully, use more of them, use a better detector, widen the baseline. Every one of those buys something, and the amount it buys is calculable in advance from the error propagation.

For an ambiguity: none of those buys anything. What buys the answer is changing the arrangement — a camera centre somewhere else, a point off the plane, a third view, a known length spanning the scene. The remedy is geometric rather than statistical.

A team that misdiagnoses an ambiguity as an uncertainty spends its effort on better features and better optimisation and gets a tighter wrong answer — the shape a chain and an adjustment measures at the level of a whole sequence. A team that misdiagnoses an uncertainty as an ambiguity rebuilds its rig when it needed a sharper lens. Both are expensive and both are avoidable by one eigenvalue.

Where else the pair appears in this collection

The distinction is not a two-view fact, and it is worth listing the other instances because they were all written before the vocabulary existed.

Scale from one view. The one thing a single view cannot give is an ambiguity — one exact parameter, no amount of photography touches it — and how wrong a measurement can be is the uncertainty in everything else.

The seven free directions of a bundle adjustment. Seven numbers no picture can name is an ambiguity of exactly seven dimensions, and the adjustment’s convergence rate is an uncertainty.

And the taught constructions. The rule that draws another room produces a correct perspective of the wrong room, which is an ambiguity from the reader’s side: the drawing is consistent with two scenes and nothing on it says which.

Naming them together is the contribution here. Every one had been measured; none had been sorted.

The one case where an ambiguity becomes an uncertainty

Nothing above is exactly true of a real measurement, and the reason is worth stating carefully because it is where the two categories touch.

An exactly coplanar scene has an exact family. A nearly coplanar scene has a narrow determination — one direction in which the answer is very poorly conditioned rather than undetermined. So in practice every ambiguity is approached through a region of extreme uncertainty, and the boundary between the two is where the determination falls below the reading error.

That is precisely what how flat is flat enough measures: with exact marks the transition is a step, and with a hand in it the step becomes a slope three decades wide. The categories are exact and the boundary between them is not, and the width of the boundary is a property of the instrument.

Which means the practical question is never “is this arrangement degenerate”. It is “does this arrangement offer more parallax than my marking error”, and that has a number.

In exact arithmetic the transition has no width; at 0.3 px of reading error it is three decades wideTwo curves against how far the scene departs from being flat. With exact marks the recovery is 84° wrong at exactly flat and 1e-5° wrong at eight parts in ten thousand of relief — a step function, because a scene that is not coplanar is not degenerate however nearly coplanar it is. With 0.3 pixels of reading error the same sweep is a slope: the recovery is useless until the out-of-plane parallax the arrangement offers rises past about ten times the reading error, which here is at a relief of about 0.08 and a parallax of 2.9 pixels. **The width of the transition is the reading error, not the geometry.**0306090-3-2-1out-of-plane spread, as a fraction of the scene's own extent (powers of ten)error in the recovered translation direction, in degreeswith 0.3 px of reading errorexact marksthe geometry is a step and the measurement is a slope9 reliefs
Fig. 4 The boundary, drawn: a step in exact arithmetic and a slope once a hand is admitted.

The four failures of this round, sorted

The row this essay closes measured four arrangements, and putting them in one table is the point of having measured all four.

A plane. A two-parameter continuum, visible in the spectrum, and every alternative predicts every mark on the plane. Ambiguity. Removed by relief.

A nearly-flat scene. One badly determined direction whose determination is below the reading error. Uncertainty, in the exact sense — but an uncertainty large enough to be useless, which is why it feels like the first case. Removed by more relief or better marks, and the exchange rate between the two is measurable.

Seven marks. Three exact alternatives that disagree about everything except the seven. Ambiguity, and the cheapest one to remove: a single extra mark.

A critical surface. At most three exact alternatives, reported by the spectrum as a two-dimensional nullspace and separated by no residual the pair supports. Ambiguity, removed only by a camera centre or a scene point off the surface — the one case where the diagnostic fires and the remedy does not exist.

Three ambiguities and one uncertainty, and the three ambiguities differ enormously in how expensive they are to remove — which is the second axis the word on its own does not supply.

Two scenes 11.4 m apart, from one pair of photographsA plan. The two eyes are at the bottom; the marks above them are 19 scene points, drawn twice. One set is where the true motion puts them, the other where a second, genuinely different motion puts them — and both reproject onto the same marks in both photographs to the arithmetic floor. They differ by 11.4 metres at the median point and 20.9 at the worst. The scene lies on a quadric surface through both eyes, which is the arrangement on which two motions are indistinguishable; the line joining each pair of copies shows how far apart the two readings of the same photograph are.both explain both photographs exactly11.4 m apart, median
Fig. 5 The most expensive of the four: two scenes metres apart, with nothing in either photograph to choose between them.

The two sequences fall at different rates, and neither falls at the statistical one

The two curves are read for where they end. Reading them for how fast they get there adds something, because neither rate is the one an estimator is supposed to have.

The solid scene falls far too fast. Eight marks to forty is a factor of five, and averaging independent noise would buy 5=2.24\sqrt{5} = 2.24. The error falls by a factor of 29, which is N2.1N^{-2.1} rather than N1/2N^{-1/2}. That is not a better estimator; it is the eight-point algorithm leaving its minimal case. At eight marks the design matrix has exactly enough rows, so the fit interpolates the reading error exactly and carries all of it into the answer. The first few extra marks buy redundancy rather than averaging, and redundancy is worth far more per mark than averaging is. Somewhere past that the curve should turn over into the ordinary square root, and forty marks is not far enough along to see it.

The flat scene falls far too slowly. The same factor of five buys a factor of 1.78, which is N0.36N^{-0.36}, and the exponent is falling as the count rises — the signature of a sequence approaching a non-zero limit rather than approaching zero. It is converging on the member of the family the noise happened to select, and the closer it gets to that member the less each further mark moves it.

So the two curves differ in shape as well as in destination, and it is tempting to make the shape a diagnostic. It is not one, for the reason this whole essay is about: reading either curve requires knowing the truth, and a solver does not.

What is available without the truth is the spread across resamplings — refit on random subsets of the marks, and report how much the answer moves. That is the modern form of an uncertainty estimate and it is exactly the wrong instrument here, for a sharper reason than the covariance is.

A bootstrap measures variance, and on the flat scene the variance behaves impeccably. Every subset of the marks lies on the same plane, so every refit lands near the same noise-selected member of the family, and the spread across subsets shrinks as 1/N1/\sqrt{N} exactly as it should. The bootstrap therefore reports a tight, well-behaved, honestly-computed uncertainty around an answer that is forty-eight degrees wrong — and it reports it with the authority of an empirical measurement rather than of a linearisation.

That is worse than the covariance’s failure, which at least tends to produce a conspicuous number. A resampling estimate on a degenerate arrangement produces a plausible one, because resampling can only ever probe the directions the data vary in, and the family is precisely the direction the data do not vary in at all. The rank is still the only honest report, and no amount of resampling computes it.

What a covariance says and does not say

A reconstruction pipeline that reports uncertainty usually reports a covariance, and a covariance is exactly the wrong instrument for this.

A covariance is computed by linearising around the answer. Around a member of a degenerate family, the linearisation has a null direction, and what a numerical implementation reports there depends on how it regularises — often a very large number, sometimes a very small one, occasionally a failure. None of those is a statement about the geometry.

The honest report is a rank, not a variance: how many directions are determined at all. That is one eigendecomposition of the same matrix the covariance was going to come from, and it answers the question the covariance cannot.

A cylinder's outline is two lines, and the duals name only where they crossThe contour generator of a cylinder, traced on the surface and projected. It is two straight rulings, and every traced point lies on one of them to 5.9e-14 px. The dual formula returns a rank-1 matrix — its second eigenvalue is 1.5e-18 of its first — which as a dual conic is a single point: the image of the axis direction, off the canvas. It names the point the two lines pass through and forgets which two they are.correct from 17 cm, at 160 mm widerank 1 dual · on the lines to 5.9e-14 px
Fig. 6 Rank as this collection measures it elsewhere: the spectrum that says whether a fit has an answer before it says how good the answer is.

Why “identifiability” is the right word

The statistical literature has a name for the distinction and it is worth adopting: a parameter is identifiable when different values of it produce different distributions of the data. Ambiguity is non-identifiability; uncertainty is identifiability with a large variance.

Adopting the word buys one thing. It makes the test explicit — do two candidate answers predict different measurements? — and that test is checkable without any statistics at all. Take two members of the family, predict the marks under each, and compare. If the predictions agree, no estimator of any kind can distinguish them, which rules out a whole category of proposed fixes before any of them is implemented.

That is the argument this essay is finally about. The measurement at the top is one arrangement; the test is general.

One flat scene, one epipolar geometry, all exact to 1.6e-6 pxTwo photographs of a scene whose points all lie on one plane, with the marks shown in both. Over the right-hand picture are drawn the epipolar lines of 5 of the marks under one fundamental matrix — three members of a two-parameter family that all satisfy the epipolar constraint at every mark to 1.6e-6 pixels. They are genuinely different matrices, 1.41 apart after normalisation, and they put the second eye in three different places. No measurement on these photographs can choose between them.left pictureright pictureevery one of them fits every mark1.6e-6 px
Fig. 7 Two members of a family predicting the same marks, which is the test stated as a picture.

Reading a residual honestly

One habit falls out of all of this and it is worth stating on its own, because it changes what a residual is for.

A residual at the arithmetic floor is not evidence that a fit is right. It is evidence that the fit explains the data it was given, which on a degenerate arrangement every wrong answer also does. The residual becomes evidence only when the data contain something the wrong answers would have failed on.

That is the withheld-mark discipline this collection applies everywhere — five marks and the sixth fits on five and predicts the sixth for exactly this reason — and it is the same rule as the one about assertions: a test that has never rejected anything proves nothing. A residual computed only on the marks used in the fit has never had the opportunity to reject anything.

So the honest report of a two-view recovery has three numbers in it rather than one: the residual on the marks fitted, the residual on marks withheld, and the rank of the design matrix. The first says the arithmetic ran. The second says the answer predicts. The third says whether there was an answer to find.

Where the same distinction runs in the other fields

Two more instances, from parts of this collection that have nothing to do with two views, because a distinction that only applies to one field is a technique rather than an idea.

The anamorph’s paintable set. What an eye can paint measures which faces of an object a design can reach, and a face that no ray reaches is not poorly painted — it is unpaintable, and no amount of care with the design changes it. Moving the eye does. That is the same pair: an exact non-availability, remedied by changing the arrangement rather than the execution.

And the screen that names a seat. The screen that names the seat recovers a viewing position from a curved screen’s picture and fails on a flat one — not badly, but exactly, because a plane’s picture contains no information about where it is being viewed from. A better recovery does not help; a curved screen does.

Every field in this collection has at least one of each, and the vocabulary is the same everywhere: does the alternative predict a different measurement? If not, the answer is not in the data and no instrument reads it out.

The vocabulary this round settles on

Three words, used consistently from here on, because the round found that the collection had been using them loosely.

Ambiguity: two or more answers that predict every available measurement identically. Not resolvable by more of the same evidence.

Uncertainty: one answer, determined, with a spread that shrinks as evidence accumulates. Resolvable by more or better measurements.

Degeneracy: the property of an arrangement that produces an ambiguity. A plane is degenerate for a two-view solver; a quadric through both eyes is degenerate; a scene with depth is not.

The three are not synonyms and they answer three different questions — what is true of the answer, what is true of the estimate, and what is true of the geometry. Keeping them apart is the whole of this essay’s contribution, and it is worth the pedantry because the remedy differs at every level: change the arrangement, change the measurement, or accept the answer.

The short version

Two failures that look identical from inside a solver. On a scene with depth, eight marks to forty takes the pose error from 19.8 degrees to 0.68; on the same marks pressed flat, from 85.0 to 47.7 — a factor of twenty-seven apart at the same count, both sequences converged.

The distinguishing question is whether the alternatives predict different measurements. If they do, more evidence resolves them and the failure is an uncertainty. If they do not, no estimator can, and the remedy is a different arrangement rather than better data. The categories are exact; the boundary between them is the width of the reading error.

More marks cure an uncertainty and leave an ambiguity where it wasTwo sequences at the same reading error. On a scene with depth, going from eight marks to forty takes the error from 19.83° to 1.78° — evidence accumulating in the ordinary way. On the same marks pressed flat, the sequence also settles, and it settles at 48°: the reading error lifts each mark a little off the plane, so the fit is not exactly degenerate and does converge, onto an answer that is nowhere near the right one. That is the practical difference between the two failures — one of them is fixed by taking more care and the other is not fixed by anything but moving.00.50011.50210203040how many marks the fit was givenerror in the recovered translation direction, in degrees (powers of ten)flatwith depththe same reading error in both5 counts
Fig. 8 The measurement once more: both sequences settling, twenty-seven times apart.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AmbiguityConditioningCritical surfaceDegeneracyerror propagationFree parameterIdentifiabilityinstrument limitleast squaresMinimal solverreconstruction ambiguitysingular values