What a pair is for

Rectifying a pair spends what its epipolar lines lean

Counted from the two disparities alone, rectifying a verged pair looks as if it throws away at least 99.7 per cent of what the pair can tell apart. That cannot be true of a warp that loses no ray, and it is not: once the place in the picture is counted, the vertical disparity adds 4.1 per cent for eyes verged at 1.2 m, and rectifying at the same focal length gives back all but 4.0 of it.

Worth reading first: Depth is a reciprocal · Two rays that do not meet.

The second disparity cuts cells counted what a verged pair of eyes can tell apart when it reads the vertical disparity as well as the horizontal one, and arrived at a gain of forty-three. It then left a question it could not answer from its own count: what happens when the same pair is rectified — its two pictures warped so that every match falls on a shared row and the vertical disparity is zero by construction. Rectification is a family, not an operation had found that every member of that family keeps every ray. So either rectifying spends most of a pair’s resolving power for the convenience of a scanline, or the second coordinate was never worth what the count said.

It is the second, and the way to see it runs through a contradiction. Rectifying the verged pair at its own focal length, turned to look straight ahead, produces exactly the pair the earlier count used as its control — the one looking straight ahead, which returned the same number of cells as shells. Taken at face value, that count says a warp that loses no ray has destroyed nearly every distinguishable place in the room. A bijection cannot do that. What it can do is expose a count that was measuring something other than it claimed.

A count that moves with its lattice

The first thing to check about any count of distinguishable places is whether it belongs to the room or to the way the room was sampled. The earlier count laid a lattice of 86,400 points through a room 2.4 m across, 1.2 m tall and half a metre to twelve deep, rounded both disparities at each, and counted the distinct results. The same room on a finer lattice should give the same answer, if the answer is about the room.

The count of cells runs 7,663, 27,695, 43,851 as the room is sampled finer, while the pair looking straight ahead stops at 113The same room — 2.4 m across, 1.2 m tall, 0.5 to 12 m deep — sampled on three lattices of 86,400, 1,382,400, 11,059,200 points, with both disparities read to whole pixels. A pair looking straight ahead reports 64, then 113, then 113 shells: the first lattice has only ninety ranges and cannot reach every whole-pixel disparity, and the finer two reach the 113 that a whole-pixel reading cuts between half a metre and twelve. The verged pair's count of pairs of disparities does not settle at all — 7663, 27695, 43851 — and its horizontal count climbs 179, 252, 282. A count that moves with the lattice is a count of the lattice.864001.38×10⁶1.11×10⁷100300100030001000030000points the room is sampled atdistinct whole-pixel labelsverged, both disparities: 43,851verged, horizontal alone: 282straight ahead: 113one room, three lattices, both readings to whole pixels×5.7 and climbing
Fig. 1 The same room counted on three lattices. The pair looking straight ahead stops at 113, the figure a whole-pixel reading cuts between half a metre and twelve; the verged pair’s count of pairs of disparities is still climbing at eleven million points.

It does not. The pair looking straight ahead reports 64 shells on the original lattice, then 113, then 113. The first lattice has only ninety ranges in it, and ninety ranges cannot land on 113 different whole-pixel disparities; the finer two reach the 113 that whole pixels cut space into shells had already found for the same span of depths, and then stop, which is what a count of the room does. The verged pair behaves differently. Its horizontal count runs 179, 252, 282, and its count of pairs of disparities runs 7,663, then 27,695, then 43,851, and it has not settled at eleven million samples.

That correction belongs in the earlier essay itself, and it has been made there: its 7,663 cells and its 64 shells straight ahead are counts of an 86,400-point lattice and not of the room. Its measurements made from gradients — the constant ratio of 4.80 between the two readings, the tube a cell forms, the angle between the two families of slabs — are local rates, involve no lattice at all, and stand. Its counts do not.

But the failure to converge is more interesting than a sampling error, because a count of places ought to converge — the room is finite and the readings are bounded. Something in the verged pair’s labels goes on distinguishing points more and more finely as the lattice tightens, and it cannot be depth, because the pair looking straight ahead reads depth just as finely and stops.

The picture position was in the count all along

The thing that keeps distinguishing points is their position across the picture. A verged pair’s horizontal disparity is not a function of depth alone: its surfaces of constant disparity are circles through the two eyes, as vergence moves the shells and does not respace them found, so two points at the same range but different bearings round to different integers. And its vertical disparity, as the earlier essay itself established, is a reading of lateral position rather than of height. Both of the verged pair’s numbers are partly reports of where in the picture a point is.

A matcher never lacks that information. It found the match at a place in the left picture — a column and a row — and those are two readings as well, far sharper than anything a disparity says about lateral position. At two metres and 25 cm above the plane of the eyes, a metre of sideways motion moves a point 450 px across the picture and changes its vertical disparity by 1.36 px. Counting the disparities without the place is therefore counting a coarse copy of the place and calling it a gain.

So the honest count uses every reading a matcher returns: left column, left row, horizontal disparity, vertical disparity. Four readings of a point that has three coordinates are not four independent numbers, and the first thing to establish is which one is redundant.

The fourth reading is the epipolar constraint

Given the left column, the left row and the horizontal disparity, the point is fixed. The left pixel names a ray; the horizontal disparity names the column of the right picture the match lies in, which is a plane through the right eye; the ray meets the plane once. Project that point into the right picture and its row is determined, and so is its vertical disparity. Nothing is left over for the fourth number to say.

That is the epipolar constraint — a point is a line over there — written as a formula for the vertical disparity, and it is exact: putting a point back from three readings and projecting it reproduces the fourth to the precision of the arithmetic. What rounding the fourth reading adds is therefore not information about the point that the other three lack. It is the extra whole numbers a quantity crosses when the other three are held to one cell each.

A line leaning 17.2 per cent crosses 32 pixels where it crosses 27 columnsA strip of the right picture of a rig with a 30 cm baseline turned in on a point 0.5 m away, 16.7 degrees each side, around the place where the left picture's pixel at column 560, row 420 must be matched. Every depth that pixel could be at is somewhere on the drawn line, which is its epipolar line and leans 17.2 per cent. Reading only the column of the match names 27 places along it; reading the row as well names the 32 pixels shaded, because a leaning line steps up a row as it crosses columns. The lean is one of three terms in what the vertical disparity adds: it is how fast the row changes as the match slides along the line, 17.2 per cent here, and the other two are how fast it changes as the left pixel itself moves across and down, 41.9 and 43.4 per cent. This is the most turned-in rig measured, at a corner of its picture, chosen so the steps can be seen; for eyes verged at 1.2 m the line at the same corner leans 1.3 per cent and crosses one extra row in about eighty columns.right picture, columns 487–512rows 599–60527 columns crossed, 32 pixels+19% from the rows
Fig. 2 One left pixel’s epipolar line in the right picture of a strongly turned-in rig, drawn over the pixel grid. Every depth the pixel could be at lies on the line. Reading only the column of the match names one place per column; reading the row too names every pixel the line passes through, because a leaning line steps up a row as it crosses columns.

The picture shows the mechanism at the one term that can be drawn. Hold the left pixel fixed and let the depth run: the match slides along a line in the right picture, and on a verged pair that line leans. Read only the column and the line is cut into one place per column. Read the row as well and it is cut into every pixel it passes through, which for a line leaning 17.2 per cent across 27 columns is 32 pixels rather than 27. The rows buy the places where the line steps between them, and nothing else.

There are two more terms of the same kind, one for each direction the left pixel can move, and the total is a closed form. If the vertical disparity changes by g across one cell of the other three readings — the sum of how fast it changes in each of the three — then rounding it splits the average cell into 1 + g labels. The whole value of the second disparity, once the picture position is counted, is that one number.

Four per cent, not forty-three times

For the eyes the earlier essay measured, verged on a point 1.2 m away, that number is small everywhere in the picture.

Over the left picture the vertical disparity adds 0.6 to 7.9 per cent to the labelsThe left picture of two eyes 65 mm apart verged on a point 1.2 m away, divided into 192 blocks, each shaded by how many more labels a matcher returns per cubic metre at 1.5 m when it reads the vertical disparity as well as the picture position and the horizontal one. It is 0.6 per cent at the least and 7.9 at the corners, where the epipolar lines lean most. The slider moves the distance the eyes are verged at, and the factor grows as it comes in. That factor is one plus how fast the vertical disparity changes along the other three readings, and it is never the forty-fold gain a count of the two disparities alone reports, because the picture position already carries almost everything the second disparity was counting.left picture, 1280 × 960 px · +0.6% near the middle, +7.9% at the cornerseyes 65 mm apart, verged at 1.2 m, points at 1.5 m×1.006 to ×1.079
Fig. 3 The left picture of eyes 65 mm apart verged at 1.2 m, divided into blocks and shaded by how many more labels per cubic metre the vertical disparity adds at 1.5 m. It is 0.6 per cent near the middle and 7.9 per cent at the corners. The slider moves the distance the eyes are verged at.

At a point 1.5 m out it adds 0.6 per cent near the middle of the picture and 7.9 per cent at the corners. The pattern is the pattern of the epipolar lines themselves: through the middle of a verged pair’s pictures they are nearly level, and toward the corners they fan, because every one of them passes through the image of the other eye, which sits far off to the side. Where the lines are level the row says nothing the column did not; where they lean it says a little.

A formula is only as good as the check it is answerable to, and this one is answerable to a count that samples no room at all. Visit the whole-number cells of the three readings directly, spread across the picture, work out exactly which whole-number vertical disparities each cell’s corners imply, and count.

Counted cell by cell, the vertical disparity adds 3.9 per cent for eyes verged at 1.2 m and the formula says 4.1For five verged pairs, the whole-pixel cells of the left column, the left row and the horizontal disparity are visited across the picture — 54,351 of them for the first pair — and the vertical disparity each cell implies is computed exactly at its corners by putting the point back. The bars are how many more labels rounding it produces: 3.9 per cent counted against 4.1 predicted for the eyes, verged at 1.2 m; 7.9 per cent counted against 8.0 predicted for the eyes, verged at 0.6 m; 10.2 per cent counted against 10.6 predicted for the eyes, verged at 0.3 m; 21.3 per cent counted against 21.8 predicted for the 30 cm rig, verged at 1 m; 41.1 per cent counted against 42.2 predicted for the 30 cm rig, verged at 0.5 m. Nothing here samples a room, so there is no lattice for the count to depend on, and the agreement is what licenses reading the formula anywhere else.eyes, verged at 1.2 m — counted+3.9%formula+4.1%eyes, verged at 0.6 m — counted+7.9%formula+8.0%eyes, verged at 0.3 m — counted+10.2%formula+10.6%30 cm rig, verged at 1 m — counted+21.3%formula+21.8%30 cm rig, verged at 0.5 m — counted+41.1%formula+42.2%counted over the picture against the density formulaworst gap 1.1 points
Fig. 4 Five verged pairs, the vertical disparity’s contribution counted cell by cell against the formula. For eyes verged at 1.2 m it is 3.9 per cent counted and 4.1 predicted over 54,351 cells; for a 30 cm rig turned in on a point half a metre away it is 41.1 against 42.2.

For eyes verged at 1.2 m, over 54,351 cells, the count is 3.9 per cent and the formula says 4.1. Across four other pairs, turned in further and further, the two stay within a point and a tenth of each other: 7.9 against 8.0 for eyes verged at 0.6 m, 10.2 against 10.6 at 0.3 m, 21.3 against 21.8 for a 30 cm rig turned in on a point a metre away, and 41.1 against 42.2 for the same rig on a point half a metre away. The count has no lattice to depend on, so the agreement licenses the formula anywhere.

Averaged over everything both pictures see between 0.6 m and 9 m, the second disparity is worth 4.1 per cent to eyes verged at 1.2 m. The earlier count’s forty-three was a count of the picture position leaking through two coarse readings, and its failure to converge was the leak: the finer the lattice, the more positions it had to distinguish.

What a rectification merges

With the right count in hand, the question the earlier essay asked can be answered in its own terms. A rectified pair returns three readings — the left column, the row the two images now share, and the disparity — and nothing else, because its vertical disparity is zero everywhere. How many labels it returns per cubic metre, against how many the verged pair returned with all four readings, is what the rectification spends.

The ratio is a local one, because a rectification is a warp: it turns each picture about its own centre, and a turned picture is magnified on one side and shrunk on the other. Where the rectified grid is coarser than the original, places the pair could tell apart have been merged into one label; where it is finer, the warp has resampled the picture onto more pixels than it had, which buys a finer step rather than a better reading — the distinction whole pixels cut space into shells drew between refining the reading and refining the matcher.

Rectified at its own focal length the pair keeps 73 to 91 per cent of its labels across the picture, and at 1042 px it keeps all of themA rig with a 30 cm baseline turned in 8.5 degrees each side on a point 1 m away, rectified by three members of the family that differ only in focal length. Along the row 240 px above the middle of the left picture, at 1.5 m, each curve is the labels per cubic metre the rectified readings return over the labels the verged pair returned with its picture position and both disparities counted. At the pair's own 900 px the rectified grid is coarser than the original all along this row, between 73.1 and 91.0 per cent; at 958 px it is even on average over everything both pictures see, coarser in some places and finer in others; at 1042 px nothing the pair could tell apart is merged anywhere, and the price is pictures 16 per cent larger each way.0.80011.201.40-500-2500250500column of the left picture, pxlabels kept by rectifying / labels the pair hadsame focal length, 900 pxeven on average, 958 pxsqueezes nowhere, 1042 px30 cm rig, 8.5° each side, points at 1.5 mkeep 73–91% at 900 px
Fig. 5 A 30 cm rig turned in 8.5 degrees each side, rectified at three focal lengths, along one row of the left picture at 1.5 m. At its own 900 px it keeps 73 to 91 per cent of its labels; at 958 px it is even on average; at 1042 px it merges nothing anywhere.

The rig in the figure is turned in 8.5 degrees each side, which is enough to make the effect plain. Rectified at its own focal length of 900 px, it keeps between 73 and 91 per cent of its labels along the row drawn, and less at the edges than in the middle. Every member of the rectification family differs from the others in focal length, among other things, and that freedom is exactly the lever that matters here: raising the focal length magnifies both pictures evenly and scales the labels per cubic metre as its cube. At 958 px the rectified pair is even with the original on average over everything both pictures see, coarser in some places and finer in others; at 1,042 px it is coarser nowhere, and the rectified pictures are 16 per cent larger each way.

So the choice among the family’s members is a resolution decision after all, which is what the earlier essay suspected, but not for the reason it gave. The decision is not whether to keep a second coordinate worth a factor of forty; it is how much magnification to buy so that the warp and the discarded row between them merge nothing.

Across the range of pairs

The same measurement on seven pairs, from human eyes barely turned in to a wide rig turned in hard, puts the pieces side by side.

Rectifying eyes verged at 1.2 m merges 4.0 per cent of their labels, and 4.1 per cent is what the vertical disparity was worthSeven verged pairs, from eyes 65 mm apart turned in 1.6 degrees each to a 30 cm rig turned in 16.7, each rectified at its own focal length of 900 px. What the vertical disparity adds to the labels, averaged over the places both pictures see, runs from 4.1 to 34.7 per cent; the labels rectifying merges run from 4.0 to 32.1, of which the warp of the picture itself accounts for 0.5 to 11.6 and the rest is the discarded vertical disparity. A rectification that merges nothing needs its focal length raised by 3.2 per cent for the eyes and 29 for the widest rig.010203051015how far each camera is turned in, degreesper centwhat the vertical disparity addslabels merged by rectifyingof those, from the warpfocal length to merge noneeach pair rectified at its own 900 px4.0% for eyes at 1.2 m
Fig. 6 Seven verged pairs rectified at their own 900 px. For eyes verged at 1.2 m the vertical disparity is worth 4.1 per cent, rectifying merges 4.0, the warp of the pictures accounts for 0.5 of that, and raising the focal length 3.2 per cent merges nothing. At 16.7 degrees each side the same figures are 34.7, 32.1, 11.6 and 29.

For eyes verged at 1.2 m, turned in 1.6 degrees each side, rectifying at the same focal length merges 4.0 per cent of the labels, of which the warp of the pictures accounts for half a per cent and the discarded vertical disparity for the rest. Raising the focal length 3.2 per cent, to 929 px, merges nothing anywhere. For a 30 cm rig turned in 16.7 degrees each side, the vertical disparity is worth 34.7 per cent, rectifying merges 32.1, the warp accounts for 11.6 of that, and merging nothing needs a focal length 29 per cent longer.

Everything rises with the turn, and for a reason that is now visible in the formula. The vertical disparity is worth exactly as much as the epipolar lines lean across one cell of the other readings, and the lines lean because the eyes are turned. Turn them in by a degree and a half and the lines barely lean; turn them in by seventeen and they lean enough to matter. The warp that undoes the turn is small in the first case and large in the second for the same reason. Both costs are set by one angle.

What the reading had been counting

It is worth saying plainly what the earlier count found, since its number was real and its interpretation was not. Two disparities without a picture position cannot fix a point, and two points that share both disparities but differ in bearing were being distinguished only because the verged pair’s disparities depend on bearing at all. A pair looking straight ahead has disparities that depend on depth alone, so its count stopped at the number of depths; a verged pair’s depend on bearing too, so its count went on rising with every finer division of bearing the lattice supplied.

That is also why rectifying looked catastrophic in those terms. Rectification makes the disparities depend on depth alone again — that is the whole of what it does to them — and so a count that ignored the picture position saw the bearing information vanish. It had not vanished. It had moved back into the column and the row, where the matcher was always holding it.

The second disparity is a genuine reading, and the tube it cuts, which the earlier essay measured, is a genuine shape; it is simply not independent of the place the match was found. A matcher that keeps its left-picture position loses, by discarding the vertical disparity, the 4.1 per cent the epipolar lines’ lean was worth, and not the ninety-nine per cent a count of disparities alone would suggest.

Where this reading stops

The labels counted are whole-pixel labels. Everything here rounds each reading to a whole pixel and counts the results, which is the quantity the three essays before it counted, and which depth is a reciprocal showed turns a fixed reading error into a lopsided interval of depth, growing without bound toward the range a pair cannot see past. A matcher that interpolates to a fraction of a pixel has more labels in every frame, and the question of how the lean and the warp scale when the reading step shrinks is a different one — though the formula says the vertical disparity’s share stays the same, since it depends on a ratio of rates and not on the step.

Merging is not the only cost of a warp. A rectification that stretches a picture resamples it, and resampling blurs; the labels counted as kept in the stretched part of a picture are labels a whole-pixel matcher can return, not labels it will return correctly. The count here is geometric, and says how many places the grid can name — not whether the interpolated pixels support naming them.

The room is what both pictures see. The earlier count’s room reached two and a half metres across at half a metre out, which no 1280 × 960 picture at 900 px holds. The measurements here are over the places both pictures actually see, from 0.6 m to 9 m, which is why their averages are over a different volume and should not be compared digit for digit with anything the earlier count reported.

And the eyes still only yaw. As in the essay before it, each eye turns about a vertical axis and not about a horizontal one, and carries no torsion. Real eyes obey a torsion rule as they move, and raise the gaze and the line is gone is where its consequences live; a torsion would tilt the epipolar lines further and raise every percentage here.

The count, corrected

A count of the two disparities alone moves with the lattice it is taken on: 7,663, then 27,695, then 43,851 for the verged pair, while the pair looking straight ahead settles at 113. It moves because a verged pair’s disparities depend on bearing as well as depth, so the count was distinguishing picture positions — which a matcher already holds, far more sharply, as the column and row where it found the match.

Counted with the picture position, the vertical disparity is a function of the other three readings, and rounding it adds exactly as much as it changes across one cell of them: 1 + g, checked cell by cell to within a point and a tenth across five pairs. For eyes verged at 1.2 m that is 0.6 per cent near the middle of the picture, 7.9 at the corners, and 4.1 on average.

Rectifying at the same focal length merges 4.0 per cent of the labels — the vertical disparity’s worth, less a little, plus half a per cent from the warp — and a focal length 3.2 per cent longer merges none. For a rig turned in 16.7 degrees the same figures are 32.1 per cent and 29 per cent. What a rectification spends is set by how far the pictures are turned, which is to say by how far their epipolar lines lean.

Still open: whether the shift of the principal point is a free choice after all

The rectification family has one more free number than this essay used: the principal point of the rectified pictures can be slid sideways, which adds a constant to every disparity. Rectification is a family, not an operation observed that a constant added to disparity changes nothing about where points are, and nothing here changes that — the labels per cubic metre are rates, and a constant shift does not move them.

It does move where the whole numbers fall. A shift of half a pixel puts every shell a whole-pixel reading can land on at a different depth, and a room whose contents cluster at one depth — a wall, a face, a workbench — may sit on a shell boundary under one shift and in the middle of a shell under another. The label count cannot see that, because it averages over the room.

The measurement that settles it puts a surface at a stated depth, sweeps the shift across one pixel, and asks how much the error in that surface’s reported depth changes over the sweep — at the near range, where shells are millimetres apart, and at ten metres, where they are 1.7 m apart and half a pixel moves every one of them by more than 80 cm. If the far-range spread is large, the shift is a free choice only on paper, and choosing it well is a way to put a pair’s shells where the scene actually is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Depth uncertaintyDisparityFocal lengthinstrument limitQuantisationrectified pairStereo pairVergence