The second eye

Two marks off a known plane find the other eye

Map a courtyard's ground from one picture into the other, and every raised mark lands somewhere the map did not send it — displaced along a line through the image of the other camera, to a fifth of a billionth of a pixel. Two such marks put that image where it is, and with it the whole epipolar geometry.

Worth reading first: A flat scene fixes no second eye · The image of the other eye.

A flat scene is a trap for two-view geometry, and a flat scene fixes no second eye measured exactly how deep: eight marks on one plane leave a two-parameter family of fundamental matrices, all fitting to a millionth of a pixel and putting the other camera in different places. The same essay recorded the consolation. The plane’s own map between the two pictures — the homography that carries every point of the plane in one picture to the same point in the other — is determined perfectly, from four marks, to 5.9×10135.9 \times 10^{-13} px.

How flat is flat enough then asked how much relief a scene needs before the family collapses, and answered in parallax: the pictures need out-of-plane displacement of about ten times the marking error, and a solver can measure that displacement as the residual a single homography leaves.

That residual was treated as a diagnostic. It is more than that. The displacement each raised mark shows after the plane’s map is applied is not a scatter of errors. It has a direction, the direction is the same for every raised mark in a precise sense, and reading it recovers the thing the flat scene could not give.

Where the plane's map leaves every raised mark, with the second camera stepped forwardThe courtyard's ground plane mapped from the first picture into the second by the map fitted to its 22 ground marks. Each of the 22 raised marks lands somewhere the map did not send it, by 18 to 180 px, and every one of those displacements lies on a line through the epipole: the worst misses it by 1.7e-10 px. Here the epipole is inside the picture.the second pictureepipole22 parallax lines miss the epipole by at most 1.7e-10 pxsecond camera stepped forward
Fig. 1 The courtyard’s ground mapped from the first picture into the second by the map fitted to its 22 ground marks, drawn over the second picture. Each of the 22 raised marks lands away from where the map sends it, by 18 to 180 px, and every displacement lies on a line through the epipole — the image of the first camera’s eye — which is inside this picture. The worst line misses it by 1.7e-10 px.

What the plane’s map does to a point that is not on the plane

The homography of a plane is exact for the plane and wrong for everything else, and the way it is wrong is geometric rather than random.

Take a raised point — the top corner of one of the courtyard’s blocks. The first camera sees it along a ray. That ray crosses the ground plane somewhere, at a point directly behind or in front of the corner from the first camera’s point of view, and in the first picture the corner and that ground point coincide. The plane’s map knows nothing about height, so it sends the corner’s mark to wherever the ground point appears in the second picture.

The corner itself appears somewhere else in the second picture. The difference between where the map sent the mark and where the corner actually is — its parallax relative to the plane — is the difference between two points on the first camera’s ray, one on the ground and one at the corner, seen from the second camera.

Two points on one ray from the first eye, seen from the second eye, lie on the image of that ray. And the image of any ray from the first eye passes through the image of the first eye itself: the epipole. So the parallax of every raised mark lies along a line through the epipole. The ground point and the corner are two pages of the same book, and the book’s spine is the line between the eyes.

That is the whole derivation, and it says something strong: the direction of each raised mark’s displacement points at the other camera’s image, whatever the height of the mark and whatever its depth. Only the length of the displacement depends on those.

The lines meet where the eye is

The first figure measures it on the courtyard. The plane’s map is fitted from the 22 ground marks alone — the block bases and the row of ground marks — and applied to the 22 raised marks. Their displacements range from 18 px for marks barely off the ground to 180 px for the tops of the tallest block, pointing in every direction across the picture.

Extended, all twenty-two lines pass within 1.7×10101.7 \times 10^{-10} px of one point, and the point is the image of the first camera’s eye in the second picture: the epipole the image of the other eye computed from forty-four matched marks and a fitted fundamental matrix. Here it is found without any fundamental matrix at all, from a plane and the way everything off the plane disagrees with it.

Nothing about the scene’s structure beyond the plane was used. The raised marks were not triangulated, their heights were not known, and the two cameras’ positions never entered. The plane’s map and a list of raised marks determine a point, and that point is where the other camera is.

Two are enough

If every displacement line passes through the epipole, two of them fix it.

Two marks off a known plane, and the point their parallax lines crossThe courtyard's ground plane mapped from the first picture into the second by the plane's own map, fitted to its 22 ground marks. Two raised marks are chosen whose parallax lines cross at 89.9°. The lines meet 1.7e-10 px from the image of the first camera's eye — the epipole, found from one plane and two points.the second pictureepipoletwo lines meet 1.7e-10 px from the epipolesecond camera stepped forward
Fig. 2 Only two raised marks, chosen so their parallax lines cross at 89.9°. The plane’s map is fitted to the ground marks as before; the two lines meet 1.7e-10 px from the image of the first camera’s eye.

Two raised marks, their two displacements extended, and the crossing is the epipole to 1.7×10101.7 \times 10^{-10} px — the same arithmetic floor as all twenty-two.

The count is worth setting beside the counts the field has met before. Eight points and the basis they are read in needs eight correspondences for a fundamental matrix in general position; seven marks, three answers gets it down to seven at the price of up to three solutions. Plane plus parallax uses four marks on a plane and two off it — six in all — and has one answer, because the plane’s map has already accounted for eight of the fundamental matrix’s degrees of freedom in a form the raised marks cannot disturb, and the epipole’s two coordinates are all that remain.

The saving is not the point, though. The point is the division of labour. The plane does the work of relating the two pictures, which it can do exactly; the raised marks do only the work the plane cannot, which is to say where the other eye is. A flat scene fails because it has only the first kind of mark. A scene with a plane in it and anything at all off the plane has both.

And the whole geometry follows

The epipole and the plane’s map together are not a partial answer. They are the fundamental matrix.

The whole epipolar geometry, from one plane and the crossing of two linesThe fundamental matrix built as the epipole's cross-product matrix times the plane's map — nothing fitted to the raised marks except the one crossing. With exact marks it puts every one of the courtyard's 44 true matches on its own epipolar line to 2.0e-12 px, and differs from the cameras' own matrix by 1.2e-15 after normalisation. With marks read to 1 px, a median pair of raised marks and the plane's map give a matrix whose worst epipolar line misses its true match by 1.2 px.exact marks: worst match off its line2.0e-12 pxread to 1 px: worst match off its line1.17 pxeach true match's distance from its epipolar line, worst over the courtyard (log scale)matrix from a plane and an epipole: 2.0e-12 px exactsecond camera stepped forward
Fig. 3 The fundamental matrix built from the epipole found by two raised marks and the plane’s map, with nothing else fitted. With exact marks every one of the courtyard’s 44 true matches lies on its epipolar line to 2.0e-12 px, and the matrix equals the cameras’ own to 1.2e-15 after normalisation. From marks read to a whole pixel, a median pair gives a matrix whose worst epipolar line misses its match by 1.2 px.

Written out, F=[e]×HF = [\mathbf{e}']_\times H: the cross-product matrix of the epipole times the plane’s homography. The reason is short. The epipolar line of a mark x\mathbf{x} must pass through the epipole e\mathbf{e}', and it must also pass through HxH\mathbf{x}, the place the plane’s point on that mark’s ray appears — because that point is on the ray. A line through two points is their cross product, so the epipolar line is e×Hx\mathbf{e}' \times H\mathbf{x}, which is [e]×Hx[\mathbf{e}']_\times H\,\mathbf{x}.

Built that way from the plane and the crossing of two lines, the matrix puts all 44 of the courtyard’s true matches on their epipolar lines to 2.0×10122.0 \times 10^{-12} px and equals the cameras’ own fundamental matrix to 1.2×10151.2 \times 10^{-15} after normalisation. With marks read to a whole pixel, a median pair of raised marks and a plane’s map fitted to rounded ground marks give a matrix whose worst epipolar line misses its true match by 1.2 px — about what reading to a pixel should cost.

Everything a point is a line over there builds on the fundamental matrix therefore follows from a plane and two points off it: the pencil of epipolar lines, the search along a line, and through the calibration the four candidate poses and the one that can see.

Where the other eye’s image falls decides how well

Exact marks make every arrangement look alike. Rounded marks do not, and the difference is instructive.

Where the plane's map leaves every raised mark, with the cameras side by sideThe courtyard's ground plane mapped from the first picture into the second by the map fitted to its 22 ground marks. Each of the 22 raised marks lands somewhere the map did not send it, by 124 to 617 px, and every one of those displacements lies on a line through the epipole: the worst misses it by 1.2e-10 px. Here the epipole is at (-3313, 132), far outside it, and the lines run almost parallel.the second picture22 parallax lines miss the epipole by at most 1.2e-10 pxcameras side by side
Fig. 4 The same construction with the standing side-by-side pair of cameras. The raised marks’ displacements, 124 to 617 px, still lie on lines through the epipole to 1.2e-10 px — but the epipole is at (−3313, 132), far outside the picture, and the lines run almost parallel.

With the cameras side by side the construction is exactly as valid: twenty-two displacements, twenty-two lines through the epipole to 1.2×10101.2 \times 10^{-10} px. But the epipole is 3,313 px to the left of the picture, and every line through it and a mark in the picture points nearly the same way. Two nearly parallel lines have a crossing that the slightest error moves a long way.

How far two marks put the epipole off, against the angle their lines cross atEvery pair of raised marks, with every mark read to 1 px, in both arrangements. Stepped forward, the parallax lines cross at up to 89° and pairs crossing at more than 20° put the epipole a median 2.9 px from the truth. Side by side, no two lines cross at more than 2.3° and the median pair misses by 419 px. The same plane and the same marks; what decides the answer is where the other eye's image is.0246020406080angle between the two parallax lines (degrees)distance of their crossing from the epipole (px, log scale)stepped forwardside by sideforward, lines over 20° apart: 2.9 px · aside: 419 pxmarks read to 1 px
Fig. 5 Every pair of raised marks, marks read to a whole pixel, in both arrangements, plotted as the distance of the pair’s crossing from the true epipole against the angle between the pair’s two lines. Stepped forward, lines cross at up to 89° and pairs crossing at more than 20° miss the epipole by a median 2.9 px. Side by side, no two lines cross at more than 2.3° and the median pair misses by 419 px.

The scatter makes the dependence plain. With the second camera stepped forward, the epipole is inside the picture, lines through it fan out in every direction, and pairs of raised marks whose lines cross at more than 20° find it to a median 2.9 px from marks read to a whole pixel. Side by side, no two lines in the whole courtyard cross at more than 2.3°, and the median pair’s crossing is 419 px from the epipole.

Nothing about the plane, the marks or the construction differed between the two arrangements. Where the other eye’s image sits relative to the picture decided the precision, through the angles it allows the lines to take.

That is the same geometry as a mismatch on its own line needs a third eye, where two nearly coincident epipolar lines crossed at a place too unstable to report. It is also worth setting against the side-by-side pair’s strengths elsewhere: the arrangement that makes plane plus parallax worst is the rectified-like one that makes depth from disparity simplest, and the arrangement that makes it best puts the epipole inside the frame, which is exactly where depth from two views stops working.

How the reading enters

The last measurement fixes the arrangement that works and varies the one thing a photographer controls, how carefully the marks are read.

The epipole from pairs of raised marks, against how finely the marks are readThe second camera stepped forward, and every pair of raised marks whose parallax lines cross at more than 20°. The median crossing lies 0.52 px from the epipole at 0.25 px, 1.93 px from the epipole at 0.5 px, 2.88 px from the epipole at 1 px, 8.32 px from the epipole at 2 px. The plane's map is refitted from the same rounded marks each time, so the error in it and the error in the raised marks both grow with the reading.024680.2500.50012how finely each mark is read (px)median distance of the crossing from the epipole (px)0.52 px at 0.25 px → 8.32 px at 2 pxstepped forward
Fig. 6 The second camera stepped forward, and every pair of raised marks whose lines cross at more than 20°. The median crossing lies 0.52 px from the epipole with marks read to a quarter of a pixel, 1.93 px at half a pixel, 2.88 px at one and 8.32 px at two.

Read to a quarter of a pixel, the median pair of well-crossed lines puts the epipole 0.52 px from the truth; to half a pixel, 1.93; to one, 2.88; to two, 8.32. The plane’s map is refitted from the rounded ground marks each time, so two errors are growing together — the map’s and the raised marks’ — and the relation is not a clean proportion. It is still the right shape for an instrument: finer reading, better answer, with no floor in sight above the arithmetic’s.

A single pair was used for every number here. With twenty-two raised marks there are 231 pairs, and a least-squares crossing of all twenty-two lines would do better than any median pair; it was not computed, because the claim is what two marks can do.

Why this is the natural repair for a flat scene

The flat-scene failure and this construction are two readings of one fact.

A fundamental matrix fitted to marks on a plane has two degrees of freedom the marks cannot fix, and an ambiguity is not an uncertainty was right that no quantity of marks on the plane will fix them. What those two degrees of freedom are is now clear: they are the epipole’s two coordinates. F=[e]×HF = [\mathbf{e}']_\times H fits every mark on the plane for any e\mathbf{e}', because for a mark on the plane HxH\mathbf{x} is the match itself and the epipolar constraint is satisfied whatever line through it is drawn. The two-parameter family is the family of possible epipoles, each paired with the one correct homography.

So the repair is not “add relief and hope”. It is “add two marks whose position the plane’s map gets wrong”. Each such mark contributes one line through the epipole, which is one equation on its two coordinates, and two marks close the family exactly. The ten-times-the-marking-error rule of the previous essay is the same statement made statistically: a raised mark whose parallax is smaller than the reading error contributes a line whose direction is set by the error rather than by the epipole.

It also makes one practical recommendation more specific than “photograph the relief”. The marks that help most are raised marks whose parallax lines cross at large angles, which means raised marks in different directions from the epipole — and that is decided by where the second camera was, not only by the scene.

The comparison a solver makes, turned into a construction

How flat is flat enough ended with a recommendation for a solver holding two photographs and no truth: fit a homography and a fundamental matrix to the same marks, and read the homography’s residual in pixels against the expected marking error. If the residual is comparable to the error, the scene is effectively flat and the fundamental matrix is not to be trusted; if it is many times larger, there is relief enough.

Plane plus parallax is that comparison taken one step further. The residual the homography leaves is not merely a size to be compared with a threshold. It is a set of vectors, and the vectors are a construction: each one is a line through the answer the fundamental-matrix fit was trying to find. A solver that discards the residual after reading its size has thrown away the part of the data that locates the other camera most directly.

It also inverts the usual order of estimation. The general method fits all seven degrees of freedom of the fundamental matrix at once, from marks that know nothing about which of them lie on which surface. Plane plus parallax fits eight numbers of a homography from marks known to share a plane, which is a better-conditioned problem because a plane’s map is exact for them, and then fits two numbers from marks known not to. Each stage uses the marks that carry its information and nothing else. That is the same discipline the measured cost of each added assumption found worthwhile for a reconstruction: bring in what is known in the order that keeps each step well posed.

When the plane is at infinity

One plane is always available, whatever the scene: the plane at infinity. Its map between two pictures is the one that sends every vanishing point in the first picture to the same vanishing point in the second, and it depends only on the cameras’ rotation and lenses, not on where they stand.

If that is the plane, every mark in the scene is off it, and each mark’s parallax against it is the displacement produced by the cameras’ step alone — the part of the motion that is not rotation. The displacements still point at the epipole, by exactly the argument above, and two of them still fix it.

This is the reading under which far enough away, a pair is one eye is the same result as a flat scene. As the scene recedes, every point approaches the plane at infinity, every mark’s parallax against it shrinks as one over the distance, and the lines through the epipole are defined by displacements smaller and smaller against the reading error. A distant scene is a scene lying on a plane — the one at infinity — and it loses the other eye for the same reason a wall does: there is nothing sufficiently off the plane to point at it.

What this does not settle

The plane is given. Every figure here knows which marks are on the ground. In a real photograph that has to be decided, and a mark wrongly assigned to the plane corrupts the map every raised mark is measured against. Finding the plane is a separate problem, which flattening a facade touches from the single-view side, and it is not attempted.

One pair of cameras for each arrangement. The stepped-forward arrangement was chosen to put the epipole in the frame; any arrangement that does so will fan the lines out, and how far outside the frame the epipole can be before the angles collapse was not swept.

The error model is rounding. As throughout this field, marks are read to a grid, which is bounded and independent. A matcher’s errors near the boundary of a plane, where a patch straddles two surfaces, are neither.

Still open: what the length of each displacement says

Every result here used only the direction of each raised mark’s parallax, which points at the epipole whatever the mark’s height and depth. The length has been ignored, and it is not noise.

For a raised point, the length of its displacement is proportional to its distance from the epipole in the picture, scaled by a factor that depends only on the point’s height above the plane divided by its depth — the same for every picture pair that shares the plane. So the ratio of each mark’s parallax length to its distance from the epipole is a number that belongs to the point, not to the cameras, and it ranks every raised point by how high it stands relative to how far away it is, before any camera is calibrated or any depth is recovered.

The question that leaves is whether that ratio behaves as the algebra says: whether the courtyard’s raised marks, read off both arrangements of cameras, give the same ordering and the same ratios of those numbers; how many pixels of reading error it takes before two blocks of different heights can no longer be told apart by them; and whether a third picture, taking the same plane’s map from a different position, reproduces the same number for every point, as a property of the scene should.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Degenerate configurationEpipolar geometryEpipoleFundamental matrixHomographyParallax