A rectification's free shift is free only near the pair
Worth reading first: Depth is a reciprocal · A point is a line over there.
Rectification is a family, not an operation found that the warp which makes a stereo pair’s epipolar lines horizontal is not unique: every member of the family keeps every ray, and one of its free numbers is the sideways position of the rectified pictures’ principal points. Sliding one picture’s principal point relative to the other’s adds the same constant to every disparity. A constant added to disparity is subtracted again when depth is computed, so it changes nothing about where any point is.
Rectifying a pair spends what its epipolar lines lean left that number alone and closed on the one place it might matter. Whole pixels cut space into shells had found that a pair reading its disparity in whole pixels cannot report any depth it likes: it reports one of the depths , the focal length times the baseline over a whole number, and a surface between two of those is reported at one of them. Add a constant that is not a whole number to every disparity, and the whole numbers fall at different disparities. The shells move. The earlier essay asked how far, and whether choosing the shift is a way to put a pair’s finest shells where the scene actually is.
The pair throughout is the one the shells were first measured on: 65 mm between the eyes, a focal length of 900 px, so metre-pixels. Every reading is a whole pixel, and every depth is computed after taking the shift back out.
The shells slide past the wall
The hero figure draws the depths a whole-pixel reading can report between seven and fourteen metres, four times over, with the principal points slid apart by nothing, a quarter of a pixel, half a pixel and three quarters.
A wall at ten metres has a disparity of 5.85 px. With no shift the reading rounds that to 6 and reports m. With a quarter of a pixel, the disparity read is 6.10, the reading is 6, the shift is taken back out, and the reported depth is m. With half a pixel, 10.64 m; with three quarters, 9.36 m. The wall does not move and the points on it do not move. What moves is the set of depths a whole-pixel reading is permitted to land on, and near ten metres those depths are 1.72 m apart.
That is the whole mechanism, and it is a small one to state. The shift is a translation of the lattice of whole pixels against the lattice of disparities, and the depths a reading can report are the images of that lattice under the reciprocal. Every shell slides, the near ones by millimetres and the far ones by most of their spacing, and each surface sits in whichever one the lattice now gives it.
One pixel of shift walks through one shell
Sweeping the shift continuously across a pixel shows what a single surface experiences.
Each curve is a sawtooth. As the shift grows, the shell that holds the surface slides steadily past it, so the error falls steadily; then the reading ticks over to the next whole pixel, a different shell takes the surface, and the error jumps. The sweep of one pixel carries the surface through exactly one shell, so the swing between the best shift and the worst is the width of the shell at that depth.
Expressed as a fraction of the depth, the swing is 1.36 per cent at 0.8 m and 17.12 per cent at 10 m. Expressed in metres it is 10.9 mm at 0.8 m and 1.71 m at 10 m. And at every depth there is a shift that reads the surface exactly: the one that makes the surface’s disparity plus the shift a whole number. For the wall at ten metres that shift is 0.150 px.
The shape of the sawtooth is not symmetric, and the asymmetry is the reciprocal’s. A shell’s catchment runs half a pixel of disparity either side of its whole number, and half a pixel of disparity is more depth on the far side than on the near side. So the error runs further positive than negative — at ten metres from −0.79 m to +0.88 m. Depth is a reciprocal found the same lean in the interval a fixed reading error maps to, and this is that lean at the scale of one whole pixel.
What the choice is worth, depth by depth
The swing is one shell, and the shells part as the square of the depth. So the value of choosing the shift well should grow as the square of the depth too.
The fitted slope on logarithmic axes is 2.010. At half a metre the difference between the best shift and the worst is four millimetres; at two metres, 6.8 cm; at ten metres, 1.71 m; at thirty metres, more than fifteen. The lower line is what a shift chosen without looking at the scene costs on average — the root-mean-square error over a shift spread evenly across the pixel — and it is about a third of the swing everywhere, which is what a sawtooth’s spread is.
So the answer to “is the shift a free choice?” depends on where the pair is looking. For a pair ranging a workbench at arm’s length the choice is worth a few millimetres against a shell spacing of the same order, and nobody would notice it. For a pair ranging a room it is worth tens of centimetres, and for a pair on a vehicle reading a building across the street it is worth metres. The range a pair cannot see past found that this pair’s reading has a cliff at 58.5 m beyond which an interval loses its far edge; well inside that cliff, the choice of shift is already worth a whole shell.
A room does not average it away
A single wall is a special case. A room holds surfaces at many depths, and it is tempting to argue that each depth is helped by some shifts and hurt by others, so that over a room the shift should average out. It does not, and the reason is instructive.
Four hundred surfaces spread evenly in depth through a room from one to three metres deep give a room-wide error that swings by 4.7 per cent as the shift runs across the pixel. From six to ten metres, 17.8 per cent. A shallow room from 9.5 to 10.5 m swings by 72 per cent. The room from one to three metres spans 39 px of disparity — thirty-nine whole shells — and still does not average the shift out.
The explanation is that a room’s error is dominated by its far end. The farthest surfaces sit in the widest shells, so they carry the largest errors, and they are few shells deep in disparity: a room from six to ten metres spans only 3.9 px, and its last shell — the one that decides the error for everything near the far wall — is placed by the shift. Averaging over a room averages over its near shells, which contribute little, and leaves the far shell, which contributes most, exactly as sensitive to the shift as a single wall. A room whose far part spans less than a pixel of disparity is, for this purpose, one surface.
There is a practical rule in that. A pair should choose its shift for the farthest surface it cares about, because that is where the choice is worth most and where no averaging will help. The near surfaces will be read about as well under any shift.
Half a pixel of matcher error erases the choice
Every figure so far reads disparity as if the only error were the rounding to a whole pixel. A real matcher has its own error before the rounding, and that changes the picture completely.
The figure computes the reported depth’s distribution exactly, summing over every whole pixel the reading can land on with its Gaussian probability rather than sampling. With no matcher error the mean error at ten metres swings by 1.696 m across the shift — the sawtooth. With 0.15 px of matcher error, by 0.72 m. With 0.3 px, by 0.18 m. With half a pixel, by seven millimetres. The matcher’s own error dithers the rounding, and by half a pixel the rounding’s position no longer matters to the mean: the probability of each neighbouring whole pixel rises and falls smoothly with the shift in exactly the proportions that keep the mean where it was.
This is the dithering an audio engineer adds deliberately, and it has the same price. The mean at half a pixel is 0.097 m too far at every shift — the reciprocal’s own lean, which no dither removes — and the spread of the readings is 1.03 m root-mean-square. That is worse than the 0.88 m the worst possible shift costs an exact reading, and incomparably worse than the zero a well-chosen shift gives one. The matcher error has made the shift irrelevant by making every shift bad.
So the shift is a free choice in exactly the regime where the choice has stopped buying anything. For a matcher that reads to a tenth of a pixel before rounding, the shift is worth nearly its full swing; for one that reads to half a pixel, it is worth nothing, and the pair’s precision at ten metres is set by the matcher rather than by where the shells sit. The second disparity cuts cells made the same separation between labels a reading can print and depths it can tell apart; the shift moves the labels and leaves the telling apart to the matcher.
A shift chosen blind is a dither the pair knows
There is a third way to treat the shift, between choosing it for one surface and letting a matcher’s error make it irrelevant: choose it at random, and remember what was chosen. That is exactly what the “chosen blind” line in the swing figure computes, and its numbers are worth setting against the dithered matcher’s.
At ten metres a shift drawn evenly across the pixel leaves a mean error of +2.5 cm and a root-mean-square error of 0.50 m. The matcher with half a pixel of its own error left +9.7 cm and 1.03 m. Even a matcher whose Gaussian error has the same variance as the random shift — 0.29 px — leaves 0.65 m. At five metres the three are 0.12, 0.25 and 0.19 m; at two metres, 2.0, 4.0 and 2.8 cm. The random shift wins everywhere, by about a factor of two over the half-pixel matcher.
The reason is the difference between two kinds of dither that signal engineers distinguish carefully. A matcher’s error is noise added before the rounding and never removed: its variance adds to the rounding’s. A shift is added before the rounding and subtracted afterwards, because the pair knows what it was. Known, subtracted dither turns the rounding error into a uniform error of exactly one pixel’s width, independent of the disparity being read, with no bias in disparity at all and nothing added on top. The small positive mean that remains in depth is the reciprocal’s lean acting on that uniform error, not a bias of the reading.
So the shift is worth something even when no surface is singled out. A pair that cannot choose the shift well can still choose it deliberately at random, subtract it, and at ten metres be read better than a matcher carrying even 0.15 px of error of its own — 0.50 m against 0.56 m, averaged over where that matcher’s shells happen to fall — which is how the rounding’s cost is paid most cheaply when the scene is unknown.
One shift, one depth
For a matcher precise enough to care, the shift can be chosen to read one surface exactly. The last figure asks what that choice does for the rest of the scene.
The shift of 0.150 px that reads the ten-metre wall exactly reads a wall at eight metres 54 cm too far — nearly its worst case — and walls at four and six metres 6 cm short, about average. Each of those walls has its own best shift, listed in the last column, and they are different numbers, because the fractional parts of at different depths are unrelated. A single shift puts one depth exactly on a shell, together with every depth whose disparity differs from it by a whole number of pixels: between 1.2 and 12 m that is forty-three depths, the shells themselves, and they are fixed once the first one is chosen.
That is the honest statement of what the free number buys. It is one real parameter, and it can satisfy one real condition: that a chosen depth sit exactly on a shell. A scene with one surface of interest — a wall being measured, a workpiece at a known standoff — gets that surface read to the matcher’s precision rather than to the shell’s. A scene with several gets one of them for free and the rest by luck.
Vergence moves the shells and does not respace them found that turning the eyes inward relabels which rays belong to which whole pixel without making the shells any finer. The shift does the same thing along a single ray: it relabels, sharpening nothing. What it can do that vergence cannot is aim — put the relabelling exactly where one surface is.
Why the constant was never free
The argument that a constant added to disparity changes nothing is correct about rays. Every member of the rectification family keeps every ray, and a disparity plus a constant, with the constant subtracted before depth is computed, gives the same depth for the same ray. What the argument assumed without saying is that disparity is read as a real number.
It is read as an integer, and an integer lattice has an origin. Moving the origin by a fraction of its spacing is invisible to anything computed from real numbers and visible to anything computed from rounded ones. Whole pixels cut space into shells is the statement that rounding makes depth discrete; this essay is the statement that the discrete set has a position as well as a spacing, and that the shift is the knob that sets the position. The spacing is set by , which is the pair; the position is set by the shift, which is the rectification.
The same reasoning says when the knob stops working. It is a knob on the rounding, and anything that makes the rounding irrelevant — a matcher error comparable to the pixel, a sub-pixel interpolation fine enough that the lattice’s origin no longer shows — disconnects it.
What this does not cover
Sub-pixel matchers. A matcher that fits a parabola to its correlation peak reports fractions of a pixel, but it is known to prefer the whole numbers — its estimates cluster toward integers, with a bias that repeats every pixel. The shift moves that pattern exactly as it moves the shells, and a matcher with that bias is somewhere between the exact whole-pixel reading and the dithered one. Its curve is not drawn here.
Surfaces that are not square to the pair. A wall that recedes spans a range of disparities and is read as a staircase of shells; the shift slides the staircase’s treads along the wall, and its effect on a fitted plane is a separate question from its effect on a single depth.
Anything but the pair measured. Every number scales with . A pair with ten times the baseline has shells ten times finer at every depth and the whole of this essay moves ten times further out.
Still open: a shift that varies across the picture
A single shift is one number for the whole picture, which is why it can place one depth and no more. The rectification family has more freedom than that: a slight rotation of both rectified pictures about their optical axes, or a slight difference in their vertical scale, makes the effective shift vary from one row to the next while keeping every epipolar line horizontal. The shells then sit at different depths in different rows.
For a surface that crosses many rows — a wall, a floor — that is a dither built into the geometry rather than added by the matcher: each row places the wall in a different part of its shell, and a plane fitted to the wall’s rows averages over positions the way the half-pixel matcher averaged over pixels, without paying the matcher’s spread. The measurement that settles whether it works renders a wall ten metres away through a pair whose rectified shift varies by one pixel across the picture, reads every row to a whole pixel with no matcher error, fits the wall’s depth from all its rows, and asks whether the fit reaches the exactness a well-chosen single shift gives one wall — and whether, unlike the single shift, it does so for walls at every depth at once.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Turning the cameras inwards — both name disparity, principal point, rectified pair
- Which reference to measure from — both name bias, depth uncertainty, rectification
- A pixel is not a point — both name principal point, quantisation
- A rig is right on one surface — both name depth from disparity, disparity
- A scroll can be asked its own radius — both name depth uncertainty, disparity
- A scroll of a climbing road measures its grade — both name depth uncertainty, disparity
Named objects
A flat tag is an object no other essay names yet.
BiasDepth from disparityDepth uncertaintyDisparityPrincipal pointQuantisationRectificationrectified pair