The second eye

Seven marks, three answers

Seven correspondences leave a two-dimensional nullspace, and the requirement that a fundamental matrix be singular is a cubic in the mix — one or three real roots. Here it has three, and all three satisfy every one of the seven marks to 8.9 × 10⁻⁹ pixels. The eighth mark, withheld, separates them by more than an order of magnitude.

Worth reading first: Eight points and the basis they are read in · A flat scene fixes no second eye.

A fundamental matrix has nine entries, is defined up to scale, and must be singular. That is nine minus one minus one: seven degrees of freedom, so seven correspondences ought to be enough.

Eight points and the basis they are read in uses eight because eight makes the problem linear, and a point is a line over there is where the constraint each mark imposes was established. Seven does not, and what happens when the count is dropped to the minimum is the subject here — because the answer is not one matrix and it is not a family. It is three.

Where the cubic comes from

Seven correspondences give seven linear conditions on nine entries, so the nullspace has dimension two: every candidate is F = αF₁ + (1−α)F₂ for two basis matrices and one free number.

The singularity requirement has not been used yet. Imposing it means det(αF₁ + (1−α)F₂) = 0, and a determinant of a three-by-three whose entries are linear in α is a cubic in α. A cubic with real coefficients has one real root or three.

So the minimal problem has one solution or three, and which it has is a property of the seven marks rather than of the algorithm.

Seven marks, one cubic, three answersSeven correspondences leave a two-dimensional nullspace, so the candidate matrices are a one-parameter line through it and the requirement that a fundamental matrix be singular is a cubic in that parameter. Here it crosses zero three times, at α = -0.768, -0.547, -0.535. Each crossing is a fundamental matrix that satisfies all seven marks — the worst residual across all three is 8.9e-9 pixels. This is the cleanest ambiguity in the collection: not an answer that is poorly determined, a finite set of exactly correct ones.-2e-7-1e-70-0.800-0.700-0.600-0.500α, the mix of the two nullspace directionsthe determinant that a fundamental matrix must make zeroa cubic with three real roots3 matrices, all exact
Fig. 1 The cubic, drawn against the mix parameter. It crosses zero three times, and each crossing is a fundamental matrix.

Three answers, all exact

On the arrangement here it has three real roots, and each of them is a matrix that satisfies every one of the seven marks to 8.9 × 10⁻⁹ pixels.

That number is the point. These are not three near-solutions clustered around a true one, and they are not a numerically ill-conditioned single answer split by rounding. They are three distinct matrices — over 0.05 apart after normalisation, with different epipoles — each of which explains the evidence to the last bit of the arithmetic.

The eighth mark chooses

Add one correspondence that was not used in the fit and the three separate immediately: 0.00 pixels for one of them, 16.2 and 17.9 for the other two.

That is an order of magnitude and a half of separation from a single extra measurement, and it is why the seven-point algorithm is used in practice rather than avoided. Inside a robust estimator, seven marks is a cheaper hypothesis than eight and generating three hypotheses per sample rather than one is an advantage: more of the hypothesis space is covered per random draw, and the scoring step against all the other marks resolves them for free.

The other eye, found in the picture — 0.4 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 6.6e-7 px on the left and 9.4e-7 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 6.6e-7 px0.40 m between the eyes
Fig. 2 The mark the three disagree about most visibly, from the field that made it an object.
A point on the left is a line on the rightFive points in the left picture. Each one fixes a line in the right picture — the image of the ray it came along — and its match lies on that line to 2.8e-13 px. The five lines meet at the epipole, 2.9e-10 px from concurrent. Knowing where a point is in one picture does not say where it is in the other; it removes one of the two degrees of freedom.1234512345a point is chosen hereand must be on this linematch to its own epipolar line: 2.8e-13 px5 of 44 correspondences drawn
Fig. 3 And the constraint the eighth mark is scored against: a point in one picture is a line in the other.

Why this ambiguity is not the plane’s

The two failures in this round’s row look similar and are opposites, and the difference is exactly what one more mark does.

A plane leaves a two-parameter continuum of fundamental matrices, and every one of them predicts every mark on the plane. Adding marks buys nothing, because the extra marks are on the plane too and every candidate explains them. A flat scene fixes no second eye is that case.

Seven marks leave a finite set of three, and the three disagree about points not among the seven. One more mark almost anywhere separates them. The ambiguity is real, exact, and cheap to remove.

So the useful distinction is not between “one answer” and “more than one”. It is between an ambiguity whose alternatives disagree about something measurable and one whose alternatives agree about everything the arrangement can measure — which is what an ambiguity is not an uncertainty is written to make precise.

What the three answers look like

Their epipoles are in different parts of the picture and their epipolar bundles run at different angles, so the three are visibly different geometries rather than three versions of one.

What they have in common is the seven marks, and it is worth being precise about how strong that is. Seven marks in two pictures is fourteen numbers. Each of the three matrices reproduces all fourteen relationships exactly. If those seven were the only marks that existed, no amount of care, no better solver and no better camera would distinguish them.

That is the shape of every genuine ambiguity in this collection: two views give shape and no size is one continuous parameter that no pair of pictures reaches, and this is three discrete alternatives that no seven marks reach.

How the cubic is actually solved

A note on method, because the obvious approach is worse than it looks.

Expanding det(αF₁ + (1−α)F₂) symbolically gives four coefficients and a page of algebra with plenty of places to make a sign error. Sampling the determinant at four values of α and solving the resulting four-by-four Vandermonde system is exact for a cubic, is four determinant evaluations and one small solve, and cannot be got wrong in a way that passes a test.

The roots then come from the trigonometric form of the cubic rather than from an iteration. That matters here: an iterative root-finder started near one root returns that root, and the whole content of this essay is that there may be three. A closed form returns all of them or says there is one, and the discriminant says which case it is.

The site's camera, written as the matrix multi-view geometry needsK holds the focal length and the principal point; R's rows are the camera basis — right, down, forward — and t is −R·eye. Projecting all 44 scene points through P = K[R|t] and through the camera itself gives the same picture to 2.0e-13 px. Everything in this field rests on the two being one camera, so it is measured rather than assumed.K — focal length and principal point739.90345.00739.9200.0001.0000R — right, down, forward0.970100.2425-0.0220-0.99590.08790.2415-0.0906-0.9662t = −R·eye00.99596.7148focal 739.85 px · 50.0° acrossP projects 44 points where the camera does, to 2.0e-13 pxcorrect from 17 cm, at 160 mm wide50° across
Fig. 4 The matrix being solved for, from the field that built it and measured its agreement with the constructive camera.
What the eight-point algorithm's basis is worth, as the reading coarsensThe same correspondences at a 2600 mm separation, rounded to five reading precisions and solved twice each. In raw pixel coordinates the design matrix's determined directions span a factor of 1.4e+5; centred and scaled they span 27. At a quarter-pixel reading the two answers differ by 2.5×; at 3 px, by 5.1×. The normalisation does not buy a fixed factor — it buys a floor that does not fall away.0.5120.10.313how finely each point is read off the picture (px, log scale)worst epipolar error against the exact geometry (px, log scale)raw pixelscentred and scaled2.5× apart at 0.25 px, 5.1× at 3 pxspread 27 against 1.4e+5
Fig. 5 And the normalisation that decides how well the linear half of the problem is conditioned.

When there is one root and when there are three

The count is a property of the arrangement, and finding an arrangement with three took a search: consecutive subsets of seven were tried until one gave a discriminant of the right sign.

That is worth reporting as a search rather than presenting as a construction, because it says something about how common the case is. Most sevens on this scene give one root. Some give three. Nothing about a set of seven marks looks different in advance, and the discriminant is the only thing that says which case is in hand.

A practical consequence: an implementation that assumes one root will silently take whichever the solver returns first, and will be wrong on some fraction of its samples. Inside a robust estimator that fraction is absorbed — a wrong hypothesis simply scores badly — but in a pipeline that trusts a single minimal solve it is a source of failure with no symptom, which is the shape where the adjustment stops measures at the other end of the same pipeline.

One correspondence moved 30 px, and where the damage wentThe clean fit is exact to 4.3e-13 px. Moving correspondence 9 by 30 px leaves every other point wrong too — the typical one by 0.49 px and the worst by 1.6 px — because a least-squares fit has nowhere to put a bad row except across all of them. Here the largest residual does fall on the culprit; it is not obliged to.00.50011.50010203040correspondenceepipolar error at every OTHER point, after one match is movedmedian 0.49 pxthe moved oneevery point wrong: median 0.49 px, worst 1.6 pxclean fit 4.3e-13 px
Fig. 6 The reason minimal solvers exist at all, from the depth field: one wrong correspondence is not a small error.

What the three do to a reconstruction

Choosing the wrong root does not produce a slightly wrong scene. It produces a scene that reprojects perfectly onto the seven marks and is otherwise unrelated to the one in front of the camera.

That is the same property the plane’s family has, and it is worth stating twice because it is counterintuitive. A wrong fundamental matrix is not a noisy fundamental matrix. Triangulating against it gives points that satisfy their own epipolar constraint by construction — the constraint the wrong matrix imposes — so every consistency check that reads the seven marks passes.

The checks that fail are the ones that read something else: the eighth mark, the visibility of the reconstructed points, the agreement with a third view. All three are cheap, and a pipeline that runs none of them will report a confident wrong answer with no complaint anywhere. That is the failure mode this collection keeps meeting under different names — a wrong match is not a small error is the same shape at the level of a single correspondence.

The count of solutions, across the field’s minimal problems

Three is not a special number and it is worth putting beside its neighbours, because a reader meeting minimal solvers for the first time will meet several counts.

Seven points and an uncalibrated pair: three solutions, from a cubic. Five points and a calibrated pair: ten solutions, from a tenth-degree polynomial. Three points and a known scene: four solutions, of which visibility usually removes two. Six points and a general camera: three again.

Every one of them is the same phenomenon — the minimum evidence for a problem determines it only up to a finite set — and every one is resolved the same way, by scoring the candidates against evidence that was withheld. Four cameras fit and one can see is this collection’s other instance, where the finite set is four poses and the withheld evidence is which of them puts the scene in front of both cameras.

Why the minimum is worth using

Robust estimation is the reason, and the arithmetic is short.

A sample of n marks is uncontaminated with probability (1−ε)ⁿ if a fraction ε are wrong. At thirty per cent contamination, seven marks are clean 8.2 per cent of the time and eight marks 5.8 per cent — so seven needs about seventy per cent as many random samples to find a clean one. Add that each clean sample of seven yields up to three hypotheses rather than one and the advantage compounds.

The cost is the cubic, the possibility of three roots, and the need to score every root rather than the first. All three are cheap, and none of them is optional: an implementation that skips the scoring throws away most of the advantage it went to the trouble of getting.

One correspondence moved 9 px, and where the damage wentThe clean fit is exact to 2.0e-13 px. Moving correspondence 4 by 9 px leaves every other point wrong too — the typical one by 0.10 px and the worst by 0.7 px — because a least-squares fit has nowhere to put a bad row except across all of them. Here the largest residual does fall on the culprit; it is not obliged to.00.2000.4000.600010203040correspondenceepipolar error at every OTHER point, after one match is movedmedian 0.10 pxthe moved oneevery point wrong: median 0.10 px, worst 0.7 pxclean fit 2.0e-13 px
Fig. 7 The contamination the argument above is about, and what a single instance of it does to a fit.

The sample count is exponential in the sample size

The arithmetic above compares seven marks with eight and finds seventy per cent. Carrying it one step further turns a comparison into a law, and the law is what explains why minimal solvers are pursued so hard.

To be confident to probability pp of drawing at least one clean sample of nn marks from data with a contaminated fraction ε\varepsilon, the number of random samples needed is

N=log(1p)log ⁣(1(1ε)n)    log(1p)(1ε)nN = \frac{\log(1-p)}{\log\!\left(1 - (1-\varepsilon)^{n}\right)} \;\approx\; -\log(1-p)\,(1-\varepsilon)^{-n}

whenever the clean probability is small. NN is exponential in nn, with base 1/(1ε)1/(1-\varepsilon) — so every mark removed from the minimal sample divides the work by a constant factor, and the factor is worse the dirtier the data.

At thirty per cent contamination and ninety-nine per cent confidence: eight marks need 78 samples, seven need 54, and the calibrated five-point solver needs 25. The ratios are 1/0.71/0.7 per mark dropped, which is 1.43 for the first and 2.04 for the pair. At fifty per cent contamination the same counts are 1,178, 588 and 145 — a factor of two per mark, and a solver that uses three fewer marks does eight times less searching.

Against that, the hypothesis counts run the other way: one root for eight marks, three for seven, ten for five. Each hypothesis has to be scored against every other mark, so the scoring work per sample grows with the count. But it grows linearly while the sample count falls exponentially, and a linear cost against an exponential saving has only one outcome once the contamination is appreciable.

That is the whole argument for minimal solvers, and it is worth having as an arithmetic rather than as a preference, because it also says when the argument stops applying. At five per cent contamination the base is 1/0.95=1.051/0.95 = 1.05, so a mark dropped saves five per cent of the sampling and the five-point solver’s ten hypotheses cost ten times the scoring. On clean data the elaborate minimal solver is a bad trade, and the plain eight-point fit over all the marks is better than any of them — which is the arrangement most of this collection’s figures are in, and the reason they use it.

The exponent also explains a practical rule that otherwise looks like superstition: it is worth spending real effort to lower ε\varepsilon before spending any on the solver. Halving the contamination from 50% to 25% takes the five-point count from 145 samples to 17, a factor of eight, which is more than the entire gap between the eight-point solver and the five-point one at the higher rate. A better matcher is worth more than a better minimal solver, and the arithmetic says by how much.

Where the seven-point solver is genuinely the right tool

Two arrangements, and both are common.

A pair of photographs with a lot of wrong matches in them. The arithmetic above is the whole argument: fewer marks per sample means more clean samples, and three hypotheses per sample means more coverage. Every practical two-view pipeline runs the seven-point solver inside a robust estimator for this reason.

And a pair with barely eight matches at all. A textureless scene, a night photograph, a pair taken through a doorway: sometimes there are seven good correspondences and the eighth is a guess. Using seven and scoring the three roots against the doubtful eighth is a better use of it than putting it into the fit, because a bad mark in a fit contaminates the answer and a bad mark in a score only weakens the separation.

The one place it is the wrong tool is the one this round is about. If the seven marks are coplanar, the design matrix loses another rank, the nullspace has three dimensions instead of two, and the cubic construction has nothing to stand on. A flat scene fixes no second eye is where that is measured, and it is one more reason a planar check belongs upstream of any minimal solver rather than downstream of it.

A flat scene says everything about the plane and nothing about the eyeFour bars, on a scale of digits. Fitted to four of the marks and tested on all 24, the homography relating the two photographs of the FLAT scene reproduces every mark to 5.9e-13 pixels — the whole relationship between the pictures, from four correspondences, exactly. Fitted to the same four marks of the scene WITH depth it is out by 29.6 pixels, because there is no such map. The fundamental matrix runs the other way: undetermined on the plane, where three combinations of the nullspace are 1.41 apart and all fit, and determined off it. Nothing is lost by the flatness; a different question is being answered.flat — the plane's own map5.9e-13flat — a second eye1.4e+0with depth — the plane's map3.0e+1with depth — a second eye1.4e+1digits, on the same two photographsfour readings
Fig. 8 The check that belongs upstream, in the form a solver can run from photographs alone.

Why seven and not six

The degree-of-freedom count says seven, and it is worth checking that the count is the right one rather than a coincidence of three numbers subtracting to seven.

A fundamental matrix is a rank-two three-by-three up to scale. The rank-two three-by-threes form a seven-dimensional variety inside the eight-dimensional projective space of all three-by-threes — one dimension lost to the determinant vanishing — and each correspondence imposes one linear condition. Seven conditions cut a seven-dimensional variety in a finite set, and the number of points in that set is the variety’s degree, which is three.

So the three is not an accident of the algebra either. It is the degree of the determinantal variety, and the cubic is that degree showing up in coordinates. The same reasoning gives ten for the five-point calibrated problem, where the variety is the essential matrices and its degree is ten.

That is a satisfying place to leave the count, because it says the number of answers is a property of the problem rather than of the method. No cleverer solver returns fewer.

What the two basis matrices are

The nullspace’s two directions are not meaningful individually, and it is worth saying so because a reader implementing this will want to interpret them.

They come out of the singular value decomposition as the two right singular vectors with the smallest singular values, and any two independent combinations of them span the same pencil. Neither is a fundamental matrix — neither is singular, in general — so drawing the epipolar geometry of F₁ alone is drawing something that is not an epipolar geometry at all.

What is meaningful is the pencil: the one-parameter line through the nullspace, and the three points on it where the determinant vanishes. That is why the cubic is solved rather than the basis inspected, and why the figure plots the determinant against the mix rather than showing the two matrices.

The same caution applies to the flat scene’s three-dimensional nullspace one essay over. Its basis vectors are equally arbitrary; what is meaningful there is the whole two-parameter family after the rank-two projection, and the three members drawn are three points chosen from it to be visibly different rather than three canonical answers.

The short version

Seven correspondences leave a two-parameter nullspace, and the singularity requirement turns it into a cubic in one parameter. The cubic has one real root or three, which is a property of the marks; here it has three, and all three satisfy every one of the seven to 8.9 × 10⁻⁹ pixels.

The eighth mark separates them by 16 pixels against 0.00. That is what makes this ambiguity different in kind from a plane’s: the alternatives disagree about something a measurement can reach, so one more measurement removes them. A plane’s alternatives agree about everything the plane contains, and nothing removes them.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

AmbiguityCorrespondenceCritical surfacedegrees of freedomEpipolar geometryEpipoleFundamental matrixMinimal solverNullspaceRankreconstruction ambiguity