The second eye

The surface two pictures cannot separate

There is a quadric through both camera centres on which two genuinely different motions draw identical pictures. Built explicitly, forty-two marks satisfy both epipolar geometries to 2 × 10⁻¹³ pixels, and the two scenes they reconstruct place the same mark at 15.6 metres and 35.1. The design matrix's nullspace has two dimensions rather than one, which is the seven-point situation arrived at from the other side.

Worth reading first: Eight points and the basis they are read in · A flat scene fixes no second eye.

A flat scene is visibly degenerate: a reader looking at the photographs can see that everything is on one plane, and a solver can see it in a spectrum. What makes the arrangement in this essay worth a rung of its own is that nothing about it looks wrong. The scene has depth. The marks are spread across the frame. The design matrix is healthy. And there are two exactly correct answers.

Constructing it rather than describing it

The classical statement is that the ambiguous configurations are ruled quadrics through both camera centres, and the classical presentation is a proof that they are. That is a description; this collection’s habit is to build the object and photograph it.

The construction is one line. Take the true camera pair, and take any second relative pose. Its fundamental matrix F′ gives a second epipolar condition, and the set of world points whose images satisfy both is

Q(X)=(P2X)TF(P1X)=0Q(X) = (\mathbf{P}_2 X)^{\mathsf T}\,\mathbf{F}'\,(\mathbf{P}_1 X) = 0

which is a quadratic form in X — a quadric surface. It passes through both camera centres without being asked to, because P₁C₁ = 0 and P₂C₂ = 0, which is the classical statement recovered as a side effect of writing the definition down.

A scene made only of points on that surface therefore produces photographs consistent with both motions, by construction.

Two epipolar geometries over one pair of pictures: 1.9e-13 and 2.1e-13 pxThe photographs the previous plan came from. Over the right-hand one are the epipolar lines of 5 marks under the true motion and under the second one, drawn together. Neither is an approximation to the other: they cross the picture at different angles and meet at different epipoles, and every mark satisfies both to 2.1e-13 pixels. That is what a critical surface is — not a poorly conditioned answer but two exactly correct ones.left pictureright pictureboth motions satisfy every mark2.1e-13 px
Fig. 1 The two photographs, with both epipolar geometries drawn over the right-hand one. Every mark lies on its line in both.

The measurement

Forty-two marks, found by shooting rays through a grid in the left picture, solving the quadratic along each one, and keeping the roots that land between three and thirty metres in front of both eyes.

Both fundamental matrices satisfy every mark: 1.9 × 10⁻¹³ pixels for the true motion and 2.1 × 10⁻¹³ for the other. Both reconstructions reproject onto both photographs at the same order. And the two scenes are genuinely different — the same six marks are placed at 15.6 metres by one reading and at 35.1 by the other, with a median disagreement of 11.4 metres over the nineteen points both readings put inside sixty metres.

Two worlds, one pair of photographs, and no reprojection residual anywhere that says which is which.

Two scenes 11.4 m apart, from one pair of photographsA plan. The two eyes are at the bottom; the marks above them are 19 scene points, drawn twice. One set is where the true motion puts them, the other where a second, genuinely different motion puts them — and both reproject onto the same marks in both photographs to the arithmetic floor. They differ by 11.4 metres at the median point and 20.9 at the worst. The scene lies on a quadric surface through both eyes, which is the arrangement on which two motions are indistinguishable; the line joining each pair of copies shows how far apart the two readings of the same photograph are.both explain both photographs exactly11.4 m apart, median
Fig. 2 A plan of both reconstructions, with the two eyes below and a line joining each point to its counterpart.

The control

An ambiguity demonstrated on a scene chosen to have it needs a scene chosen not to, or the demonstration is about the two motions being similar rather than about the surface.

So the same second motion is measured against an ordinary scene with depth in it, off the surface: it is rejected by 14.8 pixels, which is not a subtle disagreement. The two motions are genuinely different — the second camera moved twenty centimetres across and twenty-five along — and their pictures of a general scene are visibly different.

The agreement belongs to the surface. That is the claim, and the control is what turns it into one.

Why a plane is a special case of this

A plane through both camera centres is a degenerate quadric — a rank-two one — so the planar degeneracy of a flat scene fixes no second eye is a member of this family rather than a separate phenomenon.

The relationship runs the other way as well and is worth having. A plane that does not pass through the camera centres is still degenerate for the eight-point algorithm, because its own homography absorbs everything; that is a different mechanism and gives a continuum rather than a pair. So “plane” appears twice in the catalogue of degeneracies, for two different reasons, and only one of them is a critical surface.

What the spectrum says, and what it took to get it right

The design matrix’s spectrum is the diagnostic the plane’s degeneracy has, and the obvious expectation is that a critical surface — which looks nothing like a plane — leaves no trace in it. That expectation is wrong, and finding out was the round’s own correction.

Measured, the singular values relative to the largest run 1, 0.65, 0.58, 0.53, 0.46, 0.19, 8.7 × 10⁻³, 0, 0. Two of them vanish, not one and not three. So the linear step returns a pencil of candidates rather than an answer, exactly as seven marks do, and the requirement that a fundamental matrix be singular cuts the pencil to at most three exact solutions — of which the two motions here are two.

That is a tidier result than the one this essay was drafted with, and it unifies the row: seven marks, three answers is the same cubic, reached by withholding evidence rather than by arranging a scene. A plane leaves three dimensions and a continuum; a critical surface leaves two and a cubic; a general scene leaves one and an answer.

The correction is worth recording rather than quietly making. The first construction of this surface verged the same two eyes instead of moving one of them, and the quadric it produced was rank two — a pair of planes — so the “critical” scene was a plane and the demonstration was the planar degeneracy wearing a different name. Every result looked right: both motions fitted, the two reconstructions differed, the control rejected. What gave it away was the spectrum showing three vanishing values where a critical surface gives two, which is a measurement of the object rather than of the result.

A flat scene has two vanishing singular values and a solid one has oneThe nine singular values of the eight-point design matrix, on a logarithmic scale against the largest. For a scene with depth in it, one of them vanishes: that single null direction is the fundamental matrix, and having exactly one is what having an answer means. For the same marks pressed flat onto a plane, THREE vanish — the second is at 3.4e-9 of the largest against 8.0e-3 for the solid scene — so the nullspace has three dimensions and every matrix in it fits. The gap between the last two curves is the whole difference between a measurement and a family of them, and it is visible in the spectrum without knowing the answer.0-4-8-12-1613579singular value, largest to smallestsize against the largest, in powers of tena scene with depththe same marks, flatthe second smallest is the diagnostic3.4e-9 against 8.0e-3
Fig. 3 The spectrum on a scene with plenty of relief, where one value vanishes and there is an answer — the third of the three cases.

Why it is still the hard case

Two vanishing values is a signal, so why is this harder than a plane?

Because the signal is the same one seven marks give, and seven marks is not a pathology. A solver that sees a two-dimensional nullspace on forty-two correspondences has learned something real, and the standard response — impose the determinant condition, take the roots, score them against the remaining marks — does not help, because on a critical surface every remaining mark is on the surface and every root explains it.

So the diagnostic fires and the standard remedy fails. That is a worse position than a plane’s, where the diagnostic fires and the remedy — fit a homography instead — works, and it is a much worse position than seven marks’, where one more correspondence separates the roots by an order of magnitude.

What does work is a third camera centre off the quadric, a scene point off it, or prior knowledge that the scene is not a quadric — the same three remedies seven numbers no picture can name reaches for when a many-view reconstruction runs out of constraints.

Two epipolar geometries over one pair of pictures: 1.9e-13 and 2.1e-13 pxThe photographs the previous plan came from. Over the right-hand one are the epipolar lines of 5 marks under the true motion and under the second one, drawn together. Neither is an approximation to the other: they cross the picture at different angles and meet at different epipoles, and every mark satisfies both to 2.1e-13 pixels. That is what a critical surface is — not a poorly conditioned answer but two exactly correct ones.left pictureright pictureboth motions satisfy every mark2.1e-13 px
Fig. 4 The two geometries again, at the setting the numbers are quoted from: different epipoles, different angles, one set of marks.

What kind of surfaces these are

A ruled quadric through both centres: a hyperboloid of one sheet, a hyperbolic paraboloid, a cone, a pair of planes, a plane through both centres.

The reason to name them is that two of them are things people photograph. A cone through both camera centres is not a scene anybody builds, but a plane through both centres is what a photograph of the ground taken from a moving vehicle at ground level very nearly is — the arrangement the plan hidden in the photograph works in, seen from a solver’s side rather than a surveyor’s — and a hyperboloid is close enough to a valley, a cutting, a curved facade or a section of tunnel to matter.

“Close enough” is doing the same work here that “nearly flat” does in how flat is flat enough: the exact configuration is a measure-zero set and the neighbourhood of it is where reconstructions actually degrade. A scene near a critical surface gives an F that is determined and a pose that is not well determined, with no spectrum to say so.

Two scenes, and what they have in common

The two reconstructions are 11.4 metres apart at the median point and up to 21 at the worst, and they are not random with respect to each other. Both lie on the same quadric — that is what makes both consistent — and the map between them is a projective transformation of space that fixes the surface.

That is why the disagreement is largest where the surface is most curved, and why the two scenes agree exactly at the camera centres. A reader looking at the plan can follow the joining lines and see that they grow from nothing near the eyes to metres at the far edge, which is the shape of the transformation rather than a scatter of errors.

It also says which measurements would not notice. A distance between two nearby marks on the surface is nearly the same in both scenes; a distance between two far-apart marks is not. So even a known length in the scene resolves the ambiguity only if it spans enough of the surface, which is a sharper statement than “a known length fixes the scale”.

Two scenes 11.4 m apart, from one pair of photographsA plan. The two eyes are at the bottom; the marks above them are 19 scene points, drawn twice. One set is where the true motion puts them, the other where a second, genuinely different motion puts them — and both reproject onto the same marks in both photographs to the arithmetic floor. They differ by 11.4 metres at the median point and 20.9 at the worst. The scene lies on a quadric surface through both eyes, which is the arrangement on which two motions are indistinguishable; the line joining each pair of copies shows how far apart the two readings of the same photograph are.both explain both photographs exactly11.4 m apart, median
Fig. 5 The plan again, where the joining lines are shortest near the eyes and longest at the far edge.

How to know it has happened

Nothing in the two photographs says so, which is the difficulty. What can be done is all indirect and all cheap.

Take a third view from somewhere else. A camera centre off the quadric breaks it immediately, and the failure shows as a bundle adjustment that will not settle — which is where the adjustment stops measuring the same thing from the many-view side.

Include something that is obviously not on a smooth surface. A lamp post in front of a curved facade, a person in a valley, anything with its own depth.

Or check the recovered surface. Fit a quadric to the reconstructed points and see whether the two camera centres lie on it. That is a direct test of the exact condition, it costs one small eigenproblem, and nothing else in a standard pipeline computes it.

Finding the points on the surface

A note on how the marks were placed, because it decides what the figure can show.

Sampling the quadric directly in space is the obvious approach and it puts most of the points where neither camera can see them: a hyperboloid through two eyes extends behind both of them and off to infinity in four directions. So the sampling is done in the picture instead. A ray is shot through each of a grid of marks in the left photograph, the quadratic along that ray is solved, and the roots that lie in front of both cameras are kept.

That has two consequences worth having. The marks are spread evenly over the frame rather than over the surface, which is what a feature detector would produce. And the ones that survive are exactly the visible part of the surface, so the figure shows the arrangement a photographer could actually be in rather than a mathematical object with most of itself out of frame.

Sixty-three rays give forty-two usable marks inside the depth window, which is five times what the eight-point algorithm needs and is the reason the residuals above are quoted over forty-two rather than over eight.

Two rays, 5.22 mm apart, in the plane that contains bothThe ray from the left eye through its mark and the ray from the right eye through its. With the marks placed exactly they meet, to 8.0e-15 m. With the same marks read to 1 px they miss by 5.22 mm at a range of 7.22 m. Triangulation is not an intersection; the reported point is a choice about what to minimise, and the gap is the part a residual alone will not tell you.midpoint — 5.22 mm gapfrom the left eyefrom the right eyegap 5.22 mm at 7.22 mexact marks: 8.0e-15 m
Fig. 6 The ray-based construction’s own arithmetic, from the field that measures how far two rays from two eyes miss.

Why this is the strongest ambiguity in the collection

Worth putting beside the others, because the collection now has several and they are not equally serious.

Two views give shape and no size is one parameter, known in advance, fixed by any known length. Seven marks, three answers is a finite set, removed by one more mark. A flat scene fixes no second eye is a continuum, visible in a spectrum, removed by relief.

This one is a set of at most three alternatives that the spectrum reports as a pencil, that no reprojection residual separates, and that no quantity of marks on the scene removes — because every mark on the scene is on the surface, and the surface is what the alternatives agree about. It is the case where the phrase an ambiguity is not an uncertainty has to be taken completely literally: the diagnostic works, the remedy does not.

One mark more, and two of the three miss it by 16 pxThe three matrices from the previous figure, each measured against an eighth correspondence that was withheld from the fit. All three explain the seven they were fitted to, to 9e-9 pixels; only one explains the eighth, at 1.3e-10 pixels against 16.2 and 17.9. That is the difference between this ambiguity and the flat scene's: here one more measurement resolves it, because the alternatives disagree about something a measurement can reach. On a plane or a critical surface no further mark helps, because every alternative predicts every mark.matrix 1, α = -0.53516.17 pxmatrix 2, α = -0.54717.95 pxmatrix 3, α = -0.7680.00 pxthe withheld mark, in pixels3 candidates
Fig. 7 The resolution that works for a finite ambiguity and does not work here, because there is no mark off the surface to withhold.

What a practitioner should carry away

Three sentences, and none of them is “avoid quadrics”.

A healthy-looking spectrum is not evidence that the pose is determined. The design matrix’s rank is a statement about F, and the ambiguity here is downstream of F. Every diagnostic that reads the linear fit is silent.

Camera centres matter as much as scene shape. The condition involves both eyes lying on the surface, so it is as much about where the photographs were taken from as about what was photographed. Two shots along a curved wall, both close to the wall, is the arrangement to be suspicious of; the same wall from two points well off it is fine.

And the test is one eigenproblem. Fit a quadric to the reconstructed points, evaluate it at the two recovered camera centres, and see whether both come out near zero. A pipeline that has already reconstructed a scene has everything it needs, and nothing standard does it.

What the quadric test costs and when it says anything

The test proposed above — fit a quadric to the reconstructed points and evaluate it at the two recovered camera centres — needs one qualification before it is usable, and the qualification is a count.

A quadric in space has ten coefficients and is defined up to scale, so nine degrees of freedom. Nine points in general position lie on a quadric; so do any nine points at all. A pipeline that reconstructs nine marks, fits a quadric and finds it fits perfectly has learned nothing whatever, because there was no way for the fit to fail.

The test acquires power at the tenth point and gains it linearly after that. Forty-two marks give thirty-three excess conditions, so a quadric that still fits at the noise level is saying something strong: thirty-three coincidences, or a scene that really is a quadric. The number to report alongside the residual is therefore N9N - 9, and a pipeline reconstructing a dozen points should be told that its quadric test is nearly vacuous.

The second half of the test is the discriminating half and it is cheap. A scene lying on a quadric is common and harmless — a curved wall is one, and photographed from two points well off it there is no ambiguity at all. What makes the configuration critical is that both camera centres lie on the same quadric, which is two further conditions the fit was never given, so their satisfaction is genuine evidence rather than a fitting artefact. Evaluating the fitted form at the two recovered centres and comparing against its typical magnitude over the scene is the whole of it.

So the honest version of the test has three numbers in it: the quadric’s residual over the marks, the excess count N9N - 9 that says whether that residual means anything, and the form’s value at the two centres against its scale. The first two say the scene is a quadric; the third says and the eyes are on it, which is the condition.

And as everywhere else in this row, the exact condition has a neighbourhood whose width is the marking error. A scene near a critical quadric gives a pose that is determined and badly conditioned, with the spectrum reporting one vanishing value and nothing else amiss — so the continuous form of the test is the quadric residual measured against the reading error, exactly as the plane’s is the homography residual measured against it. The two tests differ only in which surface is fitted, which is the tidiest statement of what a critical surface is.

How large the family of ambiguous motions is

For a fixed camera pair, how many second motions have a critical surface at all? The answer is “essentially all of them”, and that is worth knowing because it changes what the result means.

Pick any second relative pose whatever. The quadric it defines exists, passes through both centres, and is real over some region — so for every alternative motion there is a scene that would make it indistinguishable from the true one. What is not true is the converse: for a given scene, only motions whose quadric contains that scene are ambiguous with the truth, and a general scene of a dozen points lies on no quadric at all.

So the risk is not that ambiguity is common. It is that the scenes which are quadrics — a smooth curved wall, a ground plane, a valley, a cutting — are exactly the scenes a photographer is most likely to be standing in front of when there is nothing else to photograph. The failure selects for featureless environments, which are also the environments where marks are scarce and every other part of the pipeline is already struggling.

The short version

Take the true camera pair and any second motion; the world points whose images satisfy both epipolar constraints form a quadric through both eyes, and a scene on it photographs identically under both motions. Built and measured: forty-two marks satisfying both geometries to 2 × 10⁻¹³ pixels, both reconstructions reprojecting at the floor, and the same marks placed at 15.6 metres by one reading and 35.1 by the other. The second motion is rejected by 14.8 pixels on a scene off the surface, which is what says the ambiguity belongs to the surface.

The design matrix’s nullspace has two dimensions — a pencil, cut by the determinant condition to at most three exact answers — which is the seven-point situation reached by arranging a scene rather than by withholding marks. The diagnostic therefore fires; the remedy that works for a plane does not, because every further mark is on the surface too.

Seven marks, one cubic, three answersSeven correspondences leave a two-dimensional nullspace, so the candidate matrices are a one-parameter line through it and the requirement that a fundamental matrix be singular is a cubic in that parameter. Here it crosses zero three times, at α = -0.768, -0.547, -0.535. Each crossing is a fundamental matrix that satisfies all seven marks — the worst residual across all three is 8.9e-9 pixels. This is the cleanest ambiguity in the collection: not an answer that is poorly determined, a finite set of exactly correct ones.-2e-7-1e-70-0.800-0.700-0.600-0.500α, the mix of the two nullspace directionsthe determinant that a fundamental matrix must make zeroa cubic with three real roots3 matrices, all exact
Fig. 8 The other exact ambiguity this round measured, whose alternatives at least disagree about something a mark can reach.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AmbiguityCritical surfaceDegeneracyEpipolar geometryEssential matrixFundamental matrixMinimal solverQuadricReconstructionreconstruction ambiguityTriangulation