A parallax length is a height over a depth
Worth reading first: A flat scene fixes no second eye · The image of the other eye.
Two marks off a known plane find the other eye took a courtyard’s ground plane, mapped it from one photograph into another with the plane’s own homography, and found that every raised mark lands somewhere the map did not send it, displaced along a line through the epipole — the image of the first camera’s eye in the second picture. Two raised marks fix the epipole to a fifth of a billionth of a pixel, and with it the whole of the geometry between the two pictures.
All of that used only the direction of each displacement. The essay ended by pointing out that the length had been ignored and was not noise: a raised point’s displacement grows with how high it stands above the plane relative to how far away it is. It proposed a specific reading — the ratio of each mark’s displacement to its distance from the epipole, as a number belonging to the point rather than to the cameras — and asked whether that number behaves as the algebra says, how much reading error it survives, and whether a third picture reproduces it.
The length does carry the point’s height over depth. The proposed ratio is not how it carries it, and the number is not quite the point’s alone.
The coefficient on the epipole
Start from what the displacement is. The plane’s map sends a raised mark to the place where the ground point on the same ray from the first eye appears; the mark itself appears somewhere else on the line from there to the epipole. In the homogeneous coordinates in which a homography is a matrix, that is a sum: the mark’s place in the second picture is the plane’s map of its place in the first, plus some multiple of the epipole.
The multiple k is fixed by the two marks and the map, and it is what the geometry says it should be: the point’s height above the plane divided by its depth from the first camera, times a number that belongs to the pair of cameras and not to the point.
On the courtyard, with the second camera stepped forward from the first, the twenty-two raised marks give coefficients from 0.36 to 1.67 once scaled to average one, and the ratio of each to the mark’s own height over its depth from the first camera is the same to 4.5 × 10⁻¹⁴ across all twenty-two. The tops of the tall blocks nearest the first camera are the largest numbers; low marks far off are the smallest. Nothing about the cameras’ positions was used to get there: only the two pictures, the plane’s map fitted from the ground marks, and the epipole found from the raised marks’ own parallax lines.
The ratio of lengths bends away from it
The proposed reading was a ratio of lengths: how far the mark landed from where the map sent it, over how far that place is from the epipole. It is natural, it is measurable with a ruler on the picture, and it is not the coefficient.
Against the point’s own height over depth, the ratio of lengths departs from strict proportion by up to 44.8 per cent. It is not noise — the marks are exact — and it is not a failure of the geometry. Writing out the sum above in ordinary pixel coordinates, the displacement is the coefficient divided by itself plus a quantity that belongs to the pixel: the third coordinate of the plane’s map at that mark, which is how much the map magnifies the picture there. So the ratio of lengths is k/(p + k), with p varying across the picture. For small displacements it is nearly proportional to k; for large ones, and wherever the map stretches or shrinks the picture, it is not.
The algebra the earlier essay leant on — “proportional to its distance from the epipole, scaled by a factor that depends only on the height over the depth” — is right in homogeneous coordinates and wrong on the page. The division that turns a homogeneous point into a pixel is exactly where the extra term comes in. This is the same place a point is a line over there and the rest of this field keep finding the difference between an exact statement about rays and an approximate one about pixels.
What the fundamental matrix throws away
The coefficient is worth setting beside the matrix the earlier essay built from the same ingredients. That essay wrote the fundamental matrix as the epipole’s cross-product matrix times the plane’s map, and showed that it puts every one of the courtyard’s matches on its epipolar line to two trillionths of a pixel. The epipolar line of a mark is the line through the epipole and the mark’s mapped place; the fundamental matrix says the mark is somewhere on that line.
The coefficient says where on it. A mark at its mapped place has a coefficient of zero and is on the plane; a mark at the epipole itself has an infinite one and would sit at zero depth from the first camera; every place between is a different height over depth. So the coefficient is exactly the one number per point that the fundamental matrix, by construction, cannot see. Eight points and the basis they are read in and seven marks, three answers fit the fundamental matrix from correspondences and then have to triangulate to learn anything about the points; plane plus parallax gets the geometry and a ranking of the points in the same step, because it never discarded the position along the line.
That also explains why the ratio of lengths looked right. Along one epipolar line, near the mapped place, the coefficient and the ratio of lengths are nearly proportional — the extra term in the denominator is small next to the plane’s own magnification. Across different lines, and further from the mapped places, the magnification differs and the proportion drifts. A reading that is exact along one line and approximate across lines is the pixel version of a statement that is exact for rays.
Why the second camera cancels
The agreement across three second pictures is not a coincidence of these three positions, and the reason is short enough to state. Moving the second camera changes the plane’s map and changes the epipole, and it changes both in a way tied to the same plane: the map is what the plane looks like from the new position, and the epipole is where the first eye appears from it. The coefficient that reconciles them for a given point changes by one factor that depends on how far the second camera moved relative to the plane — the same factor for every point.
Dividing every coefficient by one of them, or by their average as the figures do, removes that factor. What is left depends on where the point is relative to the plane and to the first eye, and on nothing about the second. That is what “belongs to the point” can honestly mean here: to the point and the first eye, with every second picture reporting the same numbers up to one scale.
Whose number it is
The second claim was that the number belongs to the point and not to the cameras, so that a third picture, taking the plane’s map from a different position, reproduces it. That can be tested directly by reading the same first picture against three different second ones.
With the second camera stepped forward, set to the side, and raised and moved round, the scaled coefficients agree mark by mark to 4.2 × 10⁻¹⁴. The ratios of lengths from the same three pairs disagree by up to 33 per cent. So the coefficient is independent of the second picture, as the earlier essay hoped, and the ratio of lengths is not.
It is not independent of the first. Read the same two photographs the other way round — the camera that stepped forward as the first picture — and the numbers change.
Swapped, the scaled coefficients disagree with the originals by up to 37.5 per cent, and three of the 231 pairs of marks change order. That is exactly what height over depth should do: the two cameras stand at different distances from different marks, so a mark near one camera and far from the other ranks differently depending on which one the depth is measured from. The number belongs to the point and the first camera together — a property of the scene as seen from a particular place, not of the scene alone. It ranks the raised points by how high they stand relative to how far they are from that eye.
How much reading error it survives
With the right reading established on exact marks, the remaining question is how it survives marks read the way a camera reads them.
Read with half a pixel of error in each coordinate, the scaled coefficients are good to 1.34 per cent at the median with the second camera stepped forward, and to 0.48 and 0.50 per cent with it side by side or above. The worst single mark, 4.7 per cent, is one with 21 px of parallax in the stepped-forward pair. That pair is the hard one for a reason the earlier essay found for the direction: its epipole is inside the picture, so the displacements near it are short, and a short displacement has little length for a reading error to be small against.
The ordering is the practical test, since ranking raised points is what the number was proposed for. At a quarter of a pixel of reading error, the stepped-forward pair misorders 0.51 per cent of pairs of marks and the other two none; at half a pixel, 1.30 per cent against 0.01 and 0.03; at a whole pixel, 2.97 per cent against 0.29 and 0.38; at four pixels, 11.3 per cent against 3.1 and 3.4. So the ranking is reliable to a pixel of error when the second camera stands to the side or above, and loses a few per cent of its pairs when it steps toward the scene. The pairs it loses are, as they must be, the pairs whose heights over depth were nearly equal.
What the number is for
A number that ranks raised points by height over depth from one camera, before any camera is calibrated and without any depth being recovered, is worth having, and it helps to be precise about what it gives.
It is a relief map from one viewpoint. Divide every raised mark’s coefficient by one chosen mark’s and the result is each point’s height over depth relative to that mark’s, from the first camera. It does not say how high anything is or how far away; it says which things stand higher relative to their distance, which is the quantity that decides, among other things, what occludes what as the first camera moves.
It needs the plane and the epipole and nothing else — which is why a flat scene cannot supply it. A flat scene fixes no second eye found a two-parameter family of epipoles when every mark lies on the plane; with no raised mark there is no displacement and no coefficient to read. The image of the other eye needed a fitted fundamental matrix to find the epipole; here the epipole comes from the parallax lines themselves, and the plane’s map from the ground marks. Six marks — four on the plane, two off it — give the first coefficients, and each further raised mark adds its own.
It degrades exactly where the direction did. How flat is flat enough found that the parallax must be several times the reading error before a plane’s residual says anything. The coefficient inherits that: a raised mark with a short displacement, near the epipole or barely off the plane, has a coefficient no better than its displacement is long.
And the choice of first camera is part of the answer. Two photographs give two such relief maps, one from each eye, and they disagree about the order of some points. That is not an error in either; it is two viewpoints.
Where this reading stops
One plane, one scene. The courtyard’s ground is the reference plane throughout, and every raised mark stands on it. A plane that is not the ground — a wall, a tabletop — makes the coefficient height off that plane over depth, and a point on the far side of it gets a negative number; nothing here was tested on such a plane.
The epipole is found from the same marks. The coefficient uses the epipole the raised marks’ own lines give. With reading error, the epipole is itself uncertain, and part of each coefficient’s error is shared through it; the sixty-trial figures include that, but do not separate it.
The pinhole is exact. Every picture is a perfect pinhole projection. A lens’s distortion bends the displacements, and fitting a lens from straightness alone is the reminder that a bent line and a raised point can look alike; nothing here separates them.
And the third picture shares the first. Every agreement across pairs above uses the same first picture. Three pictures with no picture in common would give three numbers from three viewpoints and nothing to compare them against without the depths.
The length, read properly
After a known plane’s map, a raised mark’s place in the second picture is the map of its place in the first plus a multiple of the epipole, and that multiple is the mark’s height above the plane over its depth from the first camera, times one number for the pair: to 4.5 × 10⁻¹⁴ on the courtyard’s twenty-two raised marks. The ratio of lengths the earlier essay proposed is that multiple divided by itself plus the map’s own magnification at the mark, and departs from proportion by up to 44.8 per cent.
The coefficient does not depend on the second picture — three second pictures against one first agree to 4.2 × 10⁻¹⁴ — and does depend on the first: read the other way round, the numbers move by up to 37.5 per cent and three of 231 pairs of marks change order. Read to half a pixel, it is good to 0.5 to 1.3 per cent at the median; at a pixel of reading error it misorders 0.3 per cent of pairs of marks when the cameras stand side by side and 3 per cent when one steps toward the scene.
Still open: whether the coefficient reaches past the plane’s own marks
Every raised mark here is a corner of a block standing on the reference plane, and its coefficient is read against a map fitted from marks on that plane. A photograph of a real scene has one plane that is known — a floor, a road — and a great many points on other surfaces, some far above the plane, some below it, some beyond its visible edge.
The coefficient’s algebra does not care where a point is, only that its ray from the first eye crosses the plane somewhere. That suggests a sweep: move a point steadily up, far above the plane, and down, below it, and out past where the plane’s marks were fitted, and measure how well its coefficient is recovered at a pixel of reading error. If the error stays at the half-per-cent level, a single floor is enough to rank everything in a room by height over depth. If it grows quickly beyond the fitted marks’ extent — as a homography extrapolated from a small patch usually does — then the number is local to the plane’s neighbourhood, and the relief map it gives is only as wide as the floor that was measured.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A turn of the head is not a step sideways — both name epipole, homography, parallax, reconstruction ambiguity
- Far enough away, a pair is one eye — both name homography, parallax, reconstruction ambiguity
- A turning frame can be straightened; a travelling one cannot — both name homography, parallax
- An epipole in the picture leaves a blind disc — both name epipole, parallax
- One shutter, two views — both name epipolar geometry, epipole
- Square to the camera is the worst mirror — both name degenerate configuration, epipole
Named objects
A flat tag is an object no other essay names yet.
Degenerate configurationEpipolar geometryEpipoleHomographyParallaxreconstruction ambiguity