What a pair is for

A rectification's free shift is free only near the pair

Sliding a rectified pair's principal points apart adds a constant to every disparity, and a constant changes nothing about where points are — on paper. A whole-pixel reading is not paper. The same slide moves every depth the reading can report, and for a wall ten metres away the choice between the best shift and the worst is 1.71 metres. Near the pair it is millimetres. A matcher with half a pixel of its own error erases the choice, and pays more than the worst shift did.

Worth reading first: Depth is a reciprocal · A point is a line over there.

Rectification is a family, not an operation found that the warp which makes a stereo pair’s epipolar lines horizontal is not unique: every member of the family keeps every ray, and one of its free numbers is the sideways position of the rectified pictures’ principal points. Sliding one picture’s principal point relative to the other’s adds the same constant to every disparity. A constant added to disparity is subtracted again when depth is computed, so it changes nothing about where any point is.

Rectifying a pair spends what its epipolar lines lean left that number alone and closed on the one place it might matter. Whole pixels cut space into shells had found that a pair reading its disparity in whole pixels cannot report any depth it likes: it reports one of the depths fB/kfB/k, the focal length times the baseline over a whole number, and a surface between two of those is reported at one of them. Add a constant that is not a whole number to every disparity, and the whole numbers fall at different disparities. The shells move. The earlier essay asked how far, and whether choosing the shift is a way to put a pair’s finest shells where the scene actually is.

The pair throughout is the one the shells were first measured on: 65 mm between the eyes, a focal length of 900 px, so fB=58.5fB = 58.5 metre-pixels. Every reading is a whole pixel, and every depth is computed after taking the shift back out.

The shells slide past the wall

The hero figure draws the depths a whole-pixel reading can report between seven and fourteen metres, four times over, with the principal points slid apart by nothing, a quarter of a pixel, half a pixel and three quarters.

Slid by a quarter pixel at a time, a pair's shells move past a wall 10 m away, and its reported depth jumps by 1.28 mThe depths a whole-pixel reading can report between 7 m and 14 m, for a pair 65 mm apart at 900 px, with the rectified pictures' principal points slid apart by 0, 0.25, 0.5, 0.75 px — each row one shift, each tick one shell, the reading corrected for the shift afterwards. The vertical line is a wall at 10 m, and the marked tick in each row is the depth it is reported at: 9.75 m at 0 px, 10.17 m at 0.25 px, 10.64 m at 0.5 px, 9.36 m at 0.75 px. The shells near the wall are 1.72 m apart and slide together; a shift of 0.150 px would put one exactly on the wall. Nothing about the points changes from row to row — only the depths a whole-pixel reading is allowed to land on.a wall at 10 m0 px−0.25 m0.25 px+0.17 m0.5 px+0.64 m0.75 px−0.64 m7 m8 m9 m10 m11 m12 m13 m14 mf = 900 px, B = 65 mm, whole pixelsshift 0.150 px puts a shell on the wall
Fig. 1 The depths a whole-pixel reading can report between 7 m and 14 m on a 65 mm pair at 900 px, for shifts of 0, ¼, ½ and ¾ px. A wall at 10 m is reported at 9.75, 10.17, 10.64 and 9.36 m. The shells near it are 1.72 m apart and slide together; a shift of 0.150 px would put one exactly on the wall.

A wall at ten metres has a disparity of 5.85 px. With no shift the reading rounds that to 6 and reports 58.5/6=9.7558.5/6 = 9.75 m. With a quarter of a pixel, the disparity read is 6.10, the reading is 6, the shift is taken back out, and the reported depth is 58.5/5.75=10.1758.5/5.75 = 10.17 m. With half a pixel, 10.64 m; with three quarters, 9.36 m. The wall does not move and the points on it do not move. What moves is the set of depths a whole-pixel reading is permitted to land on, and near ten metres those depths are 1.72 m apart.

That is the whole mechanism, and it is a small one to state. The shift is a translation of the lattice of whole pixels against the lattice of disparities, and the depths a reading can report are the images of that lattice under the reciprocal. Every shell slides, the near ones by millimetres and the far ones by most of their spacing, and each surface sits in whichever one the lattice now gives it.

One pixel of shift walks through one shell

Sweeping the shift continuously across a pixel shows what a single surface experiences.

Over a pixel of shift, a surface's reported depth walks through one whole shell, and the far surface's shell is widestThe error in the depth reported for a surface at 0.8, 2, 5, 10 m, as a percentage of its depth, against the shift of the rectified pictures' principal points across one pixel, with a whole-pixel reading and no other error. Each is a sawtooth: the error falls steadily as the shift slides the nearest shell past the surface, and jumps where the reading moves to the next whole pixel. Its swing over the sweep is 1.36% at 0.8 m, 3.40% at 2 m, 8.52% at 5 m, 17.12% at 10 m — one shell's width over the depth, which grows in proportion to the depth, so the same sweep that moves a surface at 0.8 m by 10.9 mm moves one at 10 m by 1.71 m. At every depth one shift reads the surface exactly.-5051000.2000.4000.6000.8001shift of the rectified pictures' principal points (px)error in the reported depth (% of the depth)0.8 m2 m5 m10 mwhole-pixel reading, no matcher errorswing 17.1% at 10 m
Fig. 2 The error in the depth reported for a surface at 0.8, 2, 5 and 10 m, as a percentage of its depth, against the shift across one pixel. Each is a sawtooth with a swing of 1.36, 3.40, 8.52 and 17.12 per cent — one shell’s width over the depth. At every depth one shift reads the surface exactly.

Each curve is a sawtooth. As the shift grows, the shell that holds the surface slides steadily past it, so the error falls steadily; then the reading ticks over to the next whole pixel, a different shell takes the surface, and the error jumps. The sweep of one pixel carries the surface through exactly one shell, so the swing between the best shift and the worst is the width of the shell at that depth.

Expressed as a fraction of the depth, the swing is 1.36 per cent at 0.8 m and 17.12 per cent at 10 m. Expressed in metres it is 10.9 mm at 0.8 m and 1.71 m at 10 m. And at every depth there is a shift that reads the surface exactly: the one that makes the surface’s disparity plus the shift a whole number. For the wall at ten metres that shift is 0.150 px.

The shape of the sawtooth is not symmetric, and the asymmetry is the reciprocal’s. A shell’s catchment runs half a pixel of disparity either side of its whole number, and half a pixel of disparity is more depth on the far side than on the near side. So the error runs further positive than negative — at ten metres from −0.79 m to +0.88 m. Depth is a reciprocal found the same lean in the interval a fixed reading error maps to, and this is that lean at the scale of one whole pixel.

What the choice is worth, depth by depth

The swing is one shell, and the shells part as the square of the depth. So the value of choosing the shift well should grow as the square of the depth too.

What the choice of shift is worth grows as the square of the depth: 0.4 cm at half a metre, 1.71 m at tenFor a surface at each depth from 0.5 to 30 m, the full swing of its reported depth as the shift runs across a pixel — the difference between the best choice and the worst — and, below it, the root-mean-square error of a shift chosen without regard to the surface, about a third of the swing. The swing's fitted slope on logarithmic axes is 2.010: it is one shell, Z²/fB, 0.4 cm at half a metre and 1.71 m at ten. In disparity every row is the same pixel; in depth the choice is worth millimetres near the pair and more than a metre and a half at ten metres.0.512510200.010.1110depth of the surface (m, log scale)error in its depth (m, log scale)best against worst shifta shift chosen blindfB = 58.5 m·pxslope 2.010
Fig. 3 For a surface at each depth from 0.5 to 30 m, the swing of its reported depth as the shift runs across a pixel, and the root-mean-square error of a shift chosen blind. The swing’s slope on logarithmic axes is 2.010: 0.4 cm at half a metre and 1.71 m at ten.

The fitted slope on logarithmic axes is 2.010. At half a metre the difference between the best shift and the worst is four millimetres; at two metres, 6.8 cm; at ten metres, 1.71 m; at thirty metres, more than fifteen. The lower line is what a shift chosen without looking at the scene costs on average — the root-mean-square error over a shift spread evenly across the pixel — and it is about a third of the swing everywhere, which is what a sawtooth’s spread is.

So the answer to “is the shift a free choice?” depends on where the pair is looking. For a pair ranging a workbench at arm’s length the choice is worth a few millimetres against a shell spacing of the same order, and nobody would notice it. For a pair ranging a room it is worth tens of centimetres, and for a pair on a vehicle reading a building across the street it is worth metres. The range a pair cannot see past found that this pair’s reading has a cliff at 58.5 m beyond which an interval loses its far edge; well inside that cliff, the choice of shift is already worth a whole shell.

A room does not average it away

A single wall is a special case. A room holds surfaces at many depths, and it is tempting to argue that each depth is helped by some shifts and hurt by others, so that over a room the shift should average out. It does not, and the reason is instructive.

Averaged over a room, the shift still matters, and it matters most where the room's disparities span fewest shellsThe root-mean-square error in reported depth over 400 surfaces spread evenly in depth through four rooms — 1 to 3 m, spanning 39.0 px of disparity; 3 to 6 m, spanning 9.8 px of disparity; 6 to 10 m, spanning 3.9 px of disparity; 9.5 to 10.5 m, spanning 0.6 px of disparity — against the shift, each as a percentage of its own mean. The swing is 4.7%, 10.2%, 17.8%, 72.1%. Averaging over a room averages over whole shells, but not over the room's far end: the farthest surfaces sit in the widest shells and dominate the error, and where the last shell falls against them is set by the shift. A room whose far part spans less than a pixel of disparity is effectively one surface.-40-2002000.2000.4000.6000.8001shift of the rectified pictures' principal points (px)a room's RMS depth error, against its own mean (%)1–3 m: 39.0 px deep3–6 m: 9.8 px deep6–10 m: 3.9 px deep9.5–10.5 m: 0.6 px deep400 surfaces a room, whole-pixel readingswing 72% in the narrowest
Fig. 4 The root-mean-square depth error over 400 surfaces spread evenly in depth through four rooms, 1–3, 3–6, 6–10 and 9.5–10.5 m, against the shift, as a percentage of each room’s mean. The swings are 4.7, 10.2, 17.8 and 72.1 per cent.

Four hundred surfaces spread evenly in depth through a room from one to three metres deep give a room-wide error that swings by 4.7 per cent as the shift runs across the pixel. From six to ten metres, 17.8 per cent. A shallow room from 9.5 to 10.5 m swings by 72 per cent. The room from one to three metres spans 39 px of disparity — thirty-nine whole shells — and still does not average the shift out.

The explanation is that a room’s error is dominated by its far end. The farthest surfaces sit in the widest shells, so they carry the largest errors, and they are few shells deep in disparity: a room from six to ten metres spans only 3.9 px, and its last shell — the one that decides the error for everything near the far wall — is placed by the shift. Averaging over a room averages over its near shells, which contribute little, and leaves the far shell, which contributes most, exactly as sensitive to the shift as a single wall. A room whose far part spans less than a pixel of disparity is, for this purpose, one surface.

There is a practical rule in that. A pair should choose its shift for the farthest surface it cares about, because that is where the choice is worth most and where no averaging will help. The near surfaces will be read about as well under any shift.

Half a pixel of matcher error erases the choice

Every figure so far reads disparity as if the only error were the rounding to a whole pixel. A real matcher has its own error before the rounding, and that changes the picture completely.

Half a pixel of matcher error makes the shift irrelevant to a wall at 10 m — by making every shift about 1.0 m wrongThe mean error in the depth reported for a wall at 10 m against the shift, with the matcher's own error Gaussian before the whole-pixel rounding, of 0, 0.15, 0.3, 0.5 px — computed exactly over the whole pixels the reading can land on, not sampled. The swing of the mean error over the shift is 1.696 m at 0 px, 0.718 m at 0.15 px, 0.184 m at 0.3 px, 0.007 m at 0.5 px: the matcher's error dithers the rounding, and by half a pixel the mean no longer depends on where the shells sit. What it leaves is the reciprocal's own lean, 0.097 m too far, and a root-mean-square error of 1.03 m at every shift — larger than the 0.91 m that the worst possible shift costs an exact reading.-0.50000.500100.2000.4000.6000.8001shift of the rectified pictures' principal points (px)mean error in the depth of a wall at 10 m (m)matcher ±0 pxmatcher ±0.15 pxmatcher ±0.3 pxmatcher ±0.5 pxexact over the whole pixelsRMS 1.03 m at ±0.5 px
Fig. 5 The mean error in the depth of a wall at 10 m against the shift, with Gaussian matcher error of 0, 0.15, 0.3 and 0.5 px before the whole-pixel rounding, computed exactly over the pixels the reading can land on. The mean’s swing falls from 1.696 m to 0.007 m; at half a pixel the root-mean-square error is 1.03 m at every shift.

The figure computes the reported depth’s distribution exactly, summing over every whole pixel the reading can land on with its Gaussian probability rather than sampling. With no matcher error the mean error at ten metres swings by 1.696 m across the shift — the sawtooth. With 0.15 px of matcher error, by 0.72 m. With 0.3 px, by 0.18 m. With half a pixel, by seven millimetres. The matcher’s own error dithers the rounding, and by half a pixel the rounding’s position no longer matters to the mean: the probability of each neighbouring whole pixel rises and falls smoothly with the shift in exactly the proportions that keep the mean where it was.

This is the dithering an audio engineer adds deliberately, and it has the same price. The mean at half a pixel is 0.097 m too far at every shift — the reciprocal’s own lean, which no dither removes — and the spread of the readings is 1.03 m root-mean-square. That is worse than the 0.88 m the worst possible shift costs an exact reading, and incomparably worse than the zero a well-chosen shift gives one. The matcher error has made the shift irrelevant by making every shift bad.

So the shift is a free choice in exactly the regime where the choice has stopped buying anything. For a matcher that reads to a tenth of a pixel before rounding, the shift is worth nearly its full swing; for one that reads to half a pixel, it is worth nothing, and the pair’s precision at ten metres is set by the matcher rather than by where the shells sit. The second disparity cuts cells made the same separation between labels a reading can print and depths it can tell apart; the shift moves the labels and leaves the telling apart to the matcher.

A shift chosen blind is a dither the pair knows

There is a third way to treat the shift, between choosing it for one surface and letting a matcher’s error make it irrelevant: choose it at random, and remember what was chosen. That is exactly what the “chosen blind” line in the swing figure computes, and its numbers are worth setting against the dithered matcher’s.

At ten metres a shift drawn evenly across the pixel leaves a mean error of +2.5 cm and a root-mean-square error of 0.50 m. The matcher with half a pixel of its own error left +9.7 cm and 1.03 m. Even a matcher whose Gaussian error has the same variance as the random shift — 0.29 px — leaves 0.65 m. At five metres the three are 0.12, 0.25 and 0.19 m; at two metres, 2.0, 4.0 and 2.8 cm. The random shift wins everywhere, by about a factor of two over the half-pixel matcher.

The reason is the difference between two kinds of dither that signal engineers distinguish carefully. A matcher’s error is noise added before the rounding and never removed: its variance adds to the rounding’s. A shift is added before the rounding and subtracted afterwards, because the pair knows what it was. Known, subtracted dither turns the rounding error into a uniform error of exactly one pixel’s width, independent of the disparity being read, with no bias in disparity at all and nothing added on top. The small positive mean that remains in depth is the reciprocal’s lean acting on that uniform error, not a bias of the reading.

So the shift is worth something even when no surface is singled out. A pair that cannot choose the shift well can still choose it deliberately at random, subtract it, and at ten metres be read better than a matcher carrying even 0.15 px of error of its own — 0.50 m against 0.56 m, averaged over where that matcher’s shells happen to fall — which is how the rounding’s cost is paid most cheaply when the scene is unknown.

One shift, one depth

For a matcher precise enough to care, the shift can be chosen to read one surface exactly. The last figure asks what that choice does for the rest of the scene.

A shift of 0.150 px reads the wall at 10 m exactly, and every other wall only by luckWith the principal points slid by 0.150 px — the fraction of a pixel that makes fB/10 a whole number — a whole-pixel reading reports a wall at 10 m exactly. The same shift reports walls nearer in with the errors listed: +0.6 cm at 1.5 m, against a worst of 1.9 cm; −4.7 cm at 2.5 m, against a worst of 5.4 cm; −6.1 cm at 4 m, against a worst of 14.0 cm; −6.1 cm at 6 m, against a worst of 32.1 cm; +54.0 cm at 8 m, against a worst of 58.4 cm. Each nearer wall would need its own shift, in the last column, and one pair has one shift. A single number puts one depth, and the depths whose disparities differ from it by whole pixels, exactly on a shell; everything else is read as before.wallerror at this shiftworst over any shiftits own best shift10 m (chosen)0.0 cm93.5 cm0.150 px1.5 m+0.6 cm1.9 cm0.000 px2.5 m−4.7 cm5.4 cm0.600 px4 m−6.1 cm14.0 cm0.375 px6 m−6.1 cm32.1 cm0.250 px8 m+54.0 cm58.4 cm0.688 pxone shift, chosen for the wall at 10 m43 exact depths between 1.2 and 12 m
Fig. 6 With the principal points slid by 0.150 px, a whole-pixel reading reports the wall at 10 m exactly. Walls at 1.5, 2.5, 4, 6 and 8 m are then reported +0.6, −4.7, −6.1, −6.1 and +54.0 cm off, against worst cases of 1.9, 5.4, 14.0, 32.1 and 58.4 cm; each would need its own shift, and a pair has one.

The shift of 0.150 px that reads the ten-metre wall exactly reads a wall at eight metres 54 cm too far — nearly its worst case — and walls at four and six metres 6 cm short, about average. Each of those walls has its own best shift, listed in the last column, and they are different numbers, because the fractional parts of 58.5/Z58.5/Z at different depths are unrelated. A single shift puts one depth exactly on a shell, together with every depth whose disparity differs from it by a whole number of pixels: between 1.2 and 12 m that is forty-three depths, the shells themselves, and they are fixed once the first one is chosen.

That is the honest statement of what the free number buys. It is one real parameter, and it can satisfy one real condition: that a chosen depth sit exactly on a shell. A scene with one surface of interest — a wall being measured, a workpiece at a known standoff — gets that surface read to the matcher’s precision rather than to the shell’s. A scene with several gets one of them for free and the rest by luck.

Vergence moves the shells and does not respace them found that turning the eyes inward relabels which rays belong to which whole pixel without making the shells any finer. The shift does the same thing along a single ray: it relabels, sharpening nothing. What it can do that vergence cannot is aim — put the relabelling exactly where one surface is.

Why the constant was never free

The argument that a constant added to disparity changes nothing is correct about rays. Every member of the rectification family keeps every ray, and a disparity plus a constant, with the constant subtracted before depth is computed, gives the same depth for the same ray. What the argument assumed without saying is that disparity is read as a real number.

It is read as an integer, and an integer lattice has an origin. Moving the origin by a fraction of its spacing is invisible to anything computed from real numbers and visible to anything computed from rounded ones. Whole pixels cut space into shells is the statement that rounding makes depth discrete; this essay is the statement that the discrete set has a position as well as a spacing, and that the shift is the knob that sets the position. The spacing is set by fBfB, which is the pair; the position is set by the shift, which is the rectification.

The same reasoning says when the knob stops working. It is a knob on the rounding, and anything that makes the rounding irrelevant — a matcher error comparable to the pixel, a sub-pixel interpolation fine enough that the lattice’s origin no longer shows — disconnects it.

What this does not cover

Sub-pixel matchers. A matcher that fits a parabola to its correlation peak reports fractions of a pixel, but it is known to prefer the whole numbers — its estimates cluster toward integers, with a bias that repeats every pixel. The shift moves that pattern exactly as it moves the shells, and a matcher with that bias is somewhere between the exact whole-pixel reading and the dithered one. Its curve is not drawn here.

Surfaces that are not square to the pair. A wall that recedes spans a range of disparities and is read as a staircase of shells; the shift slides the staircase’s treads along the wall, and its effect on a fitted plane is a separate question from its effect on a single depth.

Anything but the pair measured. Every number scales with fBfB. A pair with ten times the baseline has shells ten times finer at every depth and the whole of this essay moves ten times further out.

Still open: a shift that varies across the picture

A single shift is one number for the whole picture, which is why it can place one depth and no more. The rectification family has more freedom than that: a slight rotation of both rectified pictures about their optical axes, or a slight difference in their vertical scale, makes the effective shift vary from one row to the next while keeping every epipolar line horizontal. The shells then sit at different depths in different rows.

For a surface that crosses many rows — a wall, a floor — that is a dither built into the geometry rather than added by the matcher: each row places the wall in a different part of its shell, and a plane fitted to the wall’s rows averages over positions the way the half-pixel matcher averaged over pixels, without paying the matcher’s spread. The measurement that settles whether it works renders a wall ten metres away through a pair whose rectified shift varies by one pixel across the picture, reads every row to a whole pixel with no matcher error, fits the wall’s depth from all its rows, and asks whether the fit reaches the exactness a well-chosen single shift gives one wall — and whether, unlike the single shift, it does so for walls at every depth at once.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasDepth from disparityDepth uncertaintyDisparityPrincipal pointQuantisationRectificationrectified pair