Rectifying a pair spends what its epipolar lines lean
Worth reading first: Depth is a reciprocal · Two rays that do not meet.
The second disparity cuts cells counted what a verged pair of eyes can tell apart when it reads the vertical disparity as well as the horizontal one, and arrived at a gain of forty-three. It then left a question it could not answer from its own count: what happens when the same pair is rectified — its two pictures warped so that every match falls on a shared row and the vertical disparity is zero by construction. Rectification is a family, not an operation had found that every member of that family keeps every ray. So either rectifying spends most of a pair’s resolving power for the convenience of a scanline, or the second coordinate was never worth what the count said.
It is the second, and the way to see it runs through a contradiction. Rectifying the verged pair at its own focal length, turned to look straight ahead, produces exactly the pair the earlier count used as its control — the one looking straight ahead, which returned the same number of cells as shells. Taken at face value, that count says a warp that loses no ray has destroyed nearly every distinguishable place in the room. A bijection cannot do that. What it can do is expose a count that was measuring something other than it claimed.
A count that moves with its lattice
The first thing to check about any count of distinguishable places is whether it belongs to the room or to the way the room was sampled. The earlier count laid a lattice of 86,400 points through a room 2.4 m across, 1.2 m tall and half a metre to twelve deep, rounded both disparities at each, and counted the distinct results. The same room on a finer lattice should give the same answer, if the answer is about the room.
It does not. The pair looking straight ahead reports 64 shells on the original lattice, then 113, then 113. The first lattice has only ninety ranges in it, and ninety ranges cannot land on 113 different whole-pixel disparities; the finer two reach the 113 that whole pixels cut space into shells had already found for the same span of depths, and then stop, which is what a count of the room does. The verged pair behaves differently. Its horizontal count runs 179, 252, 282, and its count of pairs of disparities runs 7,663, then 27,695, then 43,851, and it has not settled at eleven million samples.
That correction belongs in the earlier essay itself, and it has been made there: its 7,663 cells and its 64 shells straight ahead are counts of an 86,400-point lattice and not of the room. Its measurements made from gradients — the constant ratio of 4.80 between the two readings, the tube a cell forms, the angle between the two families of slabs — are local rates, involve no lattice at all, and stand. Its counts do not.
But the failure to converge is more interesting than a sampling error, because a count of places ought to converge — the room is finite and the readings are bounded. Something in the verged pair’s labels goes on distinguishing points more and more finely as the lattice tightens, and it cannot be depth, because the pair looking straight ahead reads depth just as finely and stops.
The picture position was in the count all along
The thing that keeps distinguishing points is their position across the picture. A verged pair’s horizontal disparity is not a function of depth alone: its surfaces of constant disparity are circles through the two eyes, as vergence moves the shells and does not respace them found, so two points at the same range but different bearings round to different integers. And its vertical disparity, as the earlier essay itself established, is a reading of lateral position rather than of height. Both of the verged pair’s numbers are partly reports of where in the picture a point is.
A matcher never lacks that information. It found the match at a place in the left picture — a column and a row — and those are two readings as well, far sharper than anything a disparity says about lateral position. At two metres and 25 cm above the plane of the eyes, a metre of sideways motion moves a point 450 px across the picture and changes its vertical disparity by 1.36 px. Counting the disparities without the place is therefore counting a coarse copy of the place and calling it a gain.
So the honest count uses every reading a matcher returns: left column, left row, horizontal disparity, vertical disparity. Four readings of a point that has three coordinates are not four independent numbers, and the first thing to establish is which one is redundant.
The fourth reading is the epipolar constraint
Given the left column, the left row and the horizontal disparity, the point is fixed. The left pixel names a ray; the horizontal disparity names the column of the right picture the match lies in, which is a plane through the right eye; the ray meets the plane once. Project that point into the right picture and its row is determined, and so is its vertical disparity. Nothing is left over for the fourth number to say.
That is the epipolar constraint — a point is a line over there — written as a formula for the vertical disparity, and it is exact: putting a point back from three readings and projecting it reproduces the fourth to the precision of the arithmetic. What rounding the fourth reading adds is therefore not information about the point that the other three lack. It is the extra whole numbers a quantity crosses when the other three are held to one cell each.
The picture shows the mechanism at the one term that can be drawn. Hold the left pixel fixed and let the depth run: the match slides along a line in the right picture, and on a verged pair that line leans. Read only the column and the line is cut into one place per column. Read the row as well and it is cut into every pixel it passes through, which for a line leaning 17.2 per cent across 27 columns is 32 pixels rather than 27. The rows buy the places where the line steps between them, and nothing else.
There are two more terms of the same kind, one for each direction the left pixel can move, and the total is a closed form. If the vertical disparity changes by g across one cell of the other three readings — the sum of how fast it changes in each of the three — then rounding it splits the average cell into 1 + g labels. The whole value of the second disparity, once the picture position is counted, is that one number.
Four per cent, not forty-three times
For the eyes the earlier essay measured, verged on a point 1.2 m away, that number is small everywhere in the picture.
At a point 1.5 m out it adds 0.6 per cent near the middle of the picture and 7.9 per cent at the corners. The pattern is the pattern of the epipolar lines themselves: through the middle of a verged pair’s pictures they are nearly level, and toward the corners they fan, because every one of them passes through the image of the other eye, which sits far off to the side. Where the lines are level the row says nothing the column did not; where they lean it says a little.
A formula is only as good as the check it is answerable to, and this one is answerable to a count that samples no room at all. Visit the whole-number cells of the three readings directly, spread across the picture, work out exactly which whole-number vertical disparities each cell’s corners imply, and count.
For eyes verged at 1.2 m, over 54,351 cells, the count is 3.9 per cent and the formula says 4.1. Across four other pairs, turned in further and further, the two stay within a point and a tenth of each other: 7.9 against 8.0 for eyes verged at 0.6 m, 10.2 against 10.6 at 0.3 m, 21.3 against 21.8 for a 30 cm rig turned in on a point a metre away, and 41.1 against 42.2 for the same rig on a point half a metre away. The count has no lattice to depend on, so the agreement licenses the formula anywhere.
Averaged over everything both pictures see between 0.6 m and 9 m, the second disparity is worth 4.1 per cent to eyes verged at 1.2 m. The earlier count’s forty-three was a count of the picture position leaking through two coarse readings, and its failure to converge was the leak: the finer the lattice, the more positions it had to distinguish.
What a rectification merges
With the right count in hand, the question the earlier essay asked can be answered in its own terms. A rectified pair returns three readings — the left column, the row the two images now share, and the disparity — and nothing else, because its vertical disparity is zero everywhere. How many labels it returns per cubic metre, against how many the verged pair returned with all four readings, is what the rectification spends.
The ratio is a local one, because a rectification is a warp: it turns each picture about its own centre, and a turned picture is magnified on one side and shrunk on the other. Where the rectified grid is coarser than the original, places the pair could tell apart have been merged into one label; where it is finer, the warp has resampled the picture onto more pixels than it had, which buys a finer step rather than a better reading — the distinction whole pixels cut space into shells drew between refining the reading and refining the matcher.
The rig in the figure is turned in 8.5 degrees each side, which is enough to make the effect plain. Rectified at its own focal length of 900 px, it keeps between 73 and 91 per cent of its labels along the row drawn, and less at the edges than in the middle. Every member of the rectification family differs from the others in focal length, among other things, and that freedom is exactly the lever that matters here: raising the focal length magnifies both pictures evenly and scales the labels per cubic metre as its cube. At 958 px the rectified pair is even with the original on average over everything both pictures see, coarser in some places and finer in others; at 1,042 px it is coarser nowhere, and the rectified pictures are 16 per cent larger each way.
So the choice among the family’s members is a resolution decision after all, which is what the earlier essay suspected, but not for the reason it gave. The decision is not whether to keep a second coordinate worth a factor of forty; it is how much magnification to buy so that the warp and the discarded row between them merge nothing.
Across the range of pairs
The same measurement on seven pairs, from human eyes barely turned in to a wide rig turned in hard, puts the pieces side by side.
For eyes verged at 1.2 m, turned in 1.6 degrees each side, rectifying at the same focal length merges 4.0 per cent of the labels, of which the warp of the pictures accounts for half a per cent and the discarded vertical disparity for the rest. Raising the focal length 3.2 per cent, to 929 px, merges nothing anywhere. For a 30 cm rig turned in 16.7 degrees each side, the vertical disparity is worth 34.7 per cent, rectifying merges 32.1, the warp accounts for 11.6 of that, and merging nothing needs a focal length 29 per cent longer.
Everything rises with the turn, and for a reason that is now visible in the formula. The vertical disparity is worth exactly as much as the epipolar lines lean across one cell of the other readings, and the lines lean because the eyes are turned. Turn them in by a degree and a half and the lines barely lean; turn them in by seventeen and they lean enough to matter. The warp that undoes the turn is small in the first case and large in the second for the same reason. Both costs are set by one angle.
What the reading had been counting
It is worth saying plainly what the earlier count found, since its number was real and its interpretation was not. Two disparities without a picture position cannot fix a point, and two points that share both disparities but differ in bearing were being distinguished only because the verged pair’s disparities depend on bearing at all. A pair looking straight ahead has disparities that depend on depth alone, so its count stopped at the number of depths; a verged pair’s depend on bearing too, so its count went on rising with every finer division of bearing the lattice supplied.
That is also why rectifying looked catastrophic in those terms. Rectification makes the disparities depend on depth alone again — that is the whole of what it does to them — and so a count that ignored the picture position saw the bearing information vanish. It had not vanished. It had moved back into the column and the row, where the matcher was always holding it.
The second disparity is a genuine reading, and the tube it cuts, which the earlier essay measured, is a genuine shape; it is simply not independent of the place the match was found. A matcher that keeps its left-picture position loses, by discarding the vertical disparity, the 4.1 per cent the epipolar lines’ lean was worth, and not the ninety-nine per cent a count of disparities alone would suggest.
Where this reading stops
The labels counted are whole-pixel labels. Everything here rounds each reading to a whole pixel and counts the results, which is the quantity the three essays before it counted, and which depth is a reciprocal showed turns a fixed reading error into a lopsided interval of depth, growing without bound toward the range a pair cannot see past. A matcher that interpolates to a fraction of a pixel has more labels in every frame, and the question of how the lean and the warp scale when the reading step shrinks is a different one — though the formula says the vertical disparity’s share stays the same, since it depends on a ratio of rates and not on the step.
Merging is not the only cost of a warp. A rectification that stretches a picture resamples it, and resampling blurs; the labels counted as kept in the stretched part of a picture are labels a whole-pixel matcher can return, not labels it will return correctly. The count here is geometric, and says how many places the grid can name — not whether the interpolated pixels support naming them.
The room is what both pictures see. The earlier count’s room reached two and a half metres across at half a metre out, which no 1280 × 960 picture at 900 px holds. The measurements here are over the places both pictures actually see, from 0.6 m to 9 m, which is why their averages are over a different volume and should not be compared digit for digit with anything the earlier count reported.
And the eyes still only yaw. As in the essay before it, each eye turns about a vertical axis and not about a horizontal one, and carries no torsion. Real eyes obey a torsion rule as they move, and raise the gaze and the line is gone is where its consequences live; a torsion would tilt the epipolar lines further and raise every percentage here.
The count, corrected
A count of the two disparities alone moves with the lattice it is taken on: 7,663, then 27,695, then 43,851 for the verged pair, while the pair looking straight ahead settles at 113. It moves because a verged pair’s disparities depend on bearing as well as depth, so the count was distinguishing picture positions — which a matcher already holds, far more sharply, as the column and row where it found the match.
Counted with the picture position, the vertical disparity is a function of the other three readings, and rounding it adds exactly as much as it changes across one cell of them: 1 + g, checked cell by cell to within a point and a tenth across five pairs. For eyes verged at 1.2 m that is 0.6 per cent near the middle of the picture, 7.9 at the corners, and 4.1 on average.
Rectifying at the same focal length merges 4.0 per cent of the labels — the vertical disparity’s worth, less a little, plus half a per cent from the warp — and a focal length 3.2 per cent longer merges none. For a rig turned in 16.7 degrees the same figures are 32.1 per cent and 29 per cent. What a rectification spends is set by how far the pictures are turned, which is to say by how far their epipolar lines lean.
Still open: whether the shift of the principal point is a free choice after all
The rectification family has one more free number than this essay used: the principal point of the rectified pictures can be slid sideways, which adds a constant to every disparity. Rectification is a family, not an operation observed that a constant added to disparity changes nothing about where points are, and nothing here changes that — the labels per cubic metre are rates, and a constant shift does not move them.
It does move where the whole numbers fall. A shift of half a pixel puts every shell a whole-pixel reading can land on at a different depth, and a room whose contents cluster at one depth — a wall, a face, a workbench — may sit on a shell boundary under one shift and in the middle of a shell under another. The label count cannot see that, because it averages over the room.
The measurement that settles it puts a surface at a stated depth, sweeps the shift across one pixel, and asks how much the error in that surface’s reported depth changes over the sweep — at the near range, where shells are millimetres apart, and at ten metres, where they are 1.7 m apart and half a pixel moves every one of them by more than 80 cm. If the far-range spread is large, the shift is a free choice only on paper, and choosing it well is a way to put a pair’s shells where the scene actually is.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A scroll through two slits ranges in a straight line — both name depth uncertainty, disparity, instrument limit, stereo pair
- A curved screen tilts a stereo pair both ways at once — both name disparity, stereo pair, vergence
- A scroll can be asked its own radius — both name depth uncertainty, disparity, instrument limit
- A scroll round a bend loses its straight-line depth — both name depth uncertainty, disparity, instrument limit
- A stereo picture is drawn for a level head — both name disparity, stereo pair, vergence
- Both coordinates agree on a circle and a line — both name disparity, stereo pair, vergence
Named objects
A flat tag is an object no other essay names yet.
Depth uncertaintyDisparityFocal lengthinstrument limitQuantisationrectified pairStereo pairVergence