Whole pixels cut space into shells
Worth reading first: Depth is a reciprocal · Two rays that do not meet.
The two essays this one builds on treat a disparity reading as a number with an error bar. Depth is a reciprocal took a reading good to one pixel and showed that the depth interval it implies is not centred on the answer; the range a pair cannot see past found the distance, fB/δ, beyond which that interval has no far edge. Both are statements about a continuous quantity blurred by a continuous error.
A disparity is not usually read that way. A matcher that compares two pictures column by column reports whole columns; one that refines to a sub-pixel estimate reports on a finer grid, but a grid. Before any noise enters, the reading is already a member of a discrete set, and the depths it can report are a discrete set too.
That set has a shape, and the shape is worth drawing, because it says things about a stereo pair that the error-bar account does not.
The depths a reading is allowed to print
A rectified pair with focal length f in pixels and baseline B in metres sees a point at depth Z with disparity fB/Z. Read that disparity in steps of s pixels and the only values available are s, 2s, 3s and so on. Each maps back to one depth, fB/(ks), and no other depth can be reported at all.
So a stereo pair reading to whole pixels does not measure distance on a continuous scale. It chooses among a fixed set of shells — surfaces of constant reported depth, which for a rectified pair are planes parallel to the picture — and every point in the scene is assigned to one of them.
For the pair this field uses throughout — 65 mm apart, 900 px of focal length, so fB is 58.5 — the span between half a metre and twelve metres holds 113 shells. The upper strip of the figure draws one tick at each, and the drawing is less a picture of a scale than of a comb whose teeth have been pushed together at one end.
That crowding is the whole of the argument in one glance, and the count explains it. The number of shells between a near distance and a far one is
which depends on the two distances only through their reciprocals. The shells are evenly spaced in inverse depth. Between half a metre and one metre there are 59; between one metre and twelve, 54; between twelve metres and infinity, four. A pair reading whole pixels spends half of its entire vocabulary on the half-metre nearest the cameras.
Why the far shells are a metre apart
The spacing between neighbouring shells follows from the same reciprocal. Moving the disparity by one step moves the depth by
which is the derivative of fB/d taken at a step of s, and it grows as the square of the distance.
The figure measures it instead of trusting the derivative. It takes the shells the reading actually produces, computes the gap from each to the next, and fits a straight line on logarithmic axes. The whole-pixel slope comes back at 1.997, the quarter-pixel slope at 2.000, and the small departure from two at whole pixels is the difference between a derivative and a finite step at the near end, where a single step is a large fraction of the disparity.
At two metres the neighbouring shells are 6.7 cm apart. At ten metres they are 1.39 m apart. A car parked at ten metres and a car parked eleven metres away can easily be reported at the same depth, while a hand moved three centimetres toward a camera at arm’s length crosses several shells.
This is not a new law. It is the same Z² that depth is a reciprocal found as the growth of an error bar, and the same shape as the lopsided interval there. What the discrete picture adds is that the Z² is not only a statistical widening: it is a literal gap in what the instrument can say, and a surface lying between two shells has no correct report available at all.
A level floor, reconstructed row by row
The consequences are easiest to see on the simplest possible scene: a flat, level floor stretching away from a pair of eyes 1.6 m up.
Every row of the picture below the horizon sees the floor at one depth, and its disparity there is exact — the scene has no texture problem, no occlusion and no mismatch. Rounding each row’s disparity to a whole pixel assigns the row to a shell. The row keeps its own vertical position in the picture, because nothing about the rounding touched that, and the reconstruction places it at the shell’s depth and at whatever height its image row implies at that depth.
The floor does not come back as a staircase of flat treads, which is what the word “terraces” suggests and what most people would sketch. It comes back as thirty-five standing plates.
The reason is that each shell collects a band of image rows, and those rows saw the floor across a range of true depths — the nearer ones slightly in front of the shell, the further ones slightly behind. All of them are reported at the shell’s single depth. A row that truly saw the floor in front of the shell has its ray pushed further out along itself, so it arrives below the floor; a row that truly saw it behind has its ray pulled in, so it arrives above. Within one shell the reconstructed heights therefore run from below the floor to above it, at one depth: a vertical plate.
Near the cameras the plates are short and close together, 3.9 cm apart; the floor looks almost continuous and slightly corrugated. By ten metres they are nearly two metres apart and the tallest spans 13.3 cm either side of the true floor. The spacing of the plates, fitted on logarithmic axes, has slope 1.997 — the same law as the shells, measured this time on a reconstruction rather than on the shells themselves.
Plates, not treads, and why the difference is useful
The shape of the artefact is a diagnostic, and it is worth stating how.
A reconstruction whose errors were ordinary noise would scatter the floor’s points above and below the plane at random. A reconstruction with a bias in depth would tilt or bow the floor smoothly. A stepped disparity does neither: it produces a pattern that is vertical within each band and periodic across bands, with a period that grows as the square of the distance. That combination is specific. A surface that comes back as plates spaced by Z²s/fB has been quantised, whatever else happened to it.
The height excursion within a plate is also predictable. A row that truly saw depth Z but is placed at the shell’s depth lands at height relative to the eyes’ height h. Across a shell of width ΔZ that excursion runs to about either side, and substituting the spacing gives
which grows linearly with depth. At 9.8 m with the figure’s numbers that formula gives 13.4 cm, and the figure measures 13.3. The plates are taller as they are further apart, and both are the same rounding seen through two different coordinates.
A quarter-pixel reading makes the same floor into 137 plates, each a quarter the height and a quarter the spacing of its whole-pixel counterpart at the same depth. The pattern is finer everywhere and has exactly the same growth with distance. A finer step is a uniform improvement in the reciprocal coordinate, and so it is a uniform improvement in no linear one.
Two numbers that separate
The quarter-pixel floor raises the question the rest of this essay is about, because it is where the discrete picture stops agreeing with the continuous one.
A reading that prints disparity to a quarter of a pixel has four times as many labels as one that prints whole pixels. Whether it can tell four times as many depths apart depends on something else: how accurately the disparity was known before it was rounded. Rounding to a grid adds a half-step of uncertainty; a matcher’s own error adds whatever it adds. The interval a reported value stands for has half-width equal to the sum.
With no error other than the rounding, the two coincide: every label marks an interval that does not overlap its neighbours, and the count of separable depths is the count of labels, less the single interval that runs off the end of the range. The quarter-pixel strip is simply the whole-pixel strip four times over.
Add a quarter of a pixel of matcher error — a modest figure for a sub-pixel matcher on good texture — and they part.
The reading still prints 449 distinct depths between half a metre and twelve metres. Each one now stands for an interval three quarters of a pixel wide in disparity, so neighbouring labels overlap, and laying non-overlapping intervals end to end from the far end fits only 149 of them into the range. Two thirds of the labels are resolution the reading does not have.
The far cliff moves the same way. With the rounding as the only error, a quarter-pixel reading’s interval keeps a far edge out to 468 m; with a further quarter pixel of matcher error it loses its far edge at 156 m, which is fB divided by the total half-width of three eighths of a pixel. The cliff is set by the sum, and the sum is dominated by whichever term is larger.
This is the same point the range a pair cannot see past made about averaging and calibration floors, arriving from a different direction. There, averaging frames beat down the independent part of the error and stopped at a floor set by the part that does not average. Here, a finer grid beats down the rounding and stops at a floor set by the matcher. In both cases the improvement is real up to the point where another term dominates, and past that point the number of digits printed keeps rising while the information stays where it was.
A longer baseline and a finer step are different purchases
The count of shells has fB divided by s in front of it, so on paper there are two ways to double it: halve the step, or double the baseline. The range a pair cannot see past listed the same two levers for the far cliff and called them interchangeable, and for a reading whose only uncertainty is its rounding they are. Once the matcher contributes an error of its own, they stop being the same thing, and the shells are where the difference shows.
Doubling the baseline doubles every disparity in the scene. Every label moves, every interval’s half-width in disparity stays what it was, and so every quantity in this essay scales together: twice the labels between two distances, twice the depths that can be separated, twice the cliff, and half the spacing at every depth. The matcher’s quarter pixel is a smaller fraction of a larger disparity, which is exactly what an improvement in the instrument looks like.
Halving the step doubles the labels and nothing else of the matcher’s. The rounding’s share of the half-width halves; the matcher’s share does not move. On the quarter-pixel reading above, the half-width is three eighths of a pixel, of which the step contributes one eighth — so halving the step again, to an eighth of a pixel, takes the half-width from 0.375 px to 0.3125 px. The labels double, from 449 to about 900. The separable depths rise by a fifth, from 149 to about 180, and the cliff by the same fifth. That is the arithmetic of printing more digits of a number whose uncertainty sits in the matcher.
So the two levers are interchangeable only while rounding dominates, and a reading that has already reached the matcher’s floor can gain resolution only from the geometry. The baseline lever carries its own costs — overlap between the two views falls as the cameras separate, and a point near the pair may be seen by only one of them — but it is the lever that buys depths rather than labels, and the distinction is visible in the output only by comparing the two counts this essay’s figures print side by side.
Reconciling the two cliffs
A reader who has come from those two essays may notice that the numbers do not match theirs, and the reason is worth a paragraph because it is a reason and not a discrepancy.
The range a pair cannot see past put the far cliff of the 65 mm pair at 58.5 m, for a reading good to one pixel. This essay puts the whole-pixel cliff at 117 m. The two are the same formula with different half-widths. A reading “good to one pixel” there meant an error of up to a pixel either way; a reading rounded to whole pixels with no other error has an uncertainty of half a pixel either way. Halving the half-width doubles the cliff.
That is not pedantry. It says that a matcher which really does locate the correct column every time, and merely reports it as an integer, has twice the range of one whose integer is itself uncertain by a column. The difference between those two instruments is invisible in their output — both print integers — and it is decisive at the far end.
The same shells inside a renderer
The shells are not peculiar to stereo. Anything that stores a quantity proportional to one over the distance, in steps, lays down the same shells.
The precision a depth buffer has left found it in a graphics pipeline, where depth is stored as an affine function of 1/z in a fixed number of codes: half of the buffer’s codes are spent before the harmonic mean of the near and far planes, and the resolution goes as the square of the distance. That is this essay’s set of shells exactly, with the renderer’s near and far planes in the place of a stereo pair’s fB and a code in the place of a pixel. One plane is nearly free then showed that the near plane dominates the budget, which is the renderer’s version of the observation above that half the shells lie within half a metre.
The coincidence is structural rather than accidental. A stereo pair measures a disparity, a renderer computes a projected depth, and both are the reciprocal of distance because both come from the same division by depth that every perspective projection performs. A turn of the head is not a step sideways found the stereo pair’s short-baseline failure hiding in that same division. Quantise anything downstream of it and the shells appear.
What the shells do not establish
Three limits, each stated because the figures could be read as saying more.
The quantisation here is ideal. Every row’s disparity is computed exactly and then rounded; a real matcher’s integer is not always the nearest integer, and sub-pixel estimators are widely reported to be pulled toward whole-pixel values. How a matcher arrives at its number is correspondence, which is a separate problem, and nothing here measures that pull. What is measured is the structure a perfect rounding imposes, which is the floor under any reading reported on the same grid.
The floor is the kindest scene. A level floor has one depth per image row, so its plates are clean. A slanted wall, a curved surface or a textured object mixes depths within a row, and the pattern becomes a set of iso-depth contours drawn across the surface instead of a stack of plates. The spacing law is the same; the appearance is not.
And shells are planes only for a rectified pair. A pair of cameras that are not parallel reports disparity as a difference of angles, and the surfaces of equal disparity are not planes. The depth a pair calls zero found the zero shell of a verged pair to be a circle through both eyes, which is the first hint of what the others are.
Still open: what shells a verged pair cuts
That last limit is the question left open. For a verged pair — two cameras or two eyes turned inward to a fixation point — the zero-disparity shell is the Vieth–Müller circle, and by the same inscribed-angle theorem every other shell of constant disparity, measured as an angle, is also a circle through both eyes, one for each value of the convergence angle. Quantising the disparity therefore cuts the plane of the eyes into a nested family of circular arcs rather than parallel lines, and what remains is to measure that family: how the arcs’ spacing along the line of sight compares with a rectified pair’s Z²s/fB straight ahead, how it changes off to the side where the arcs curve toward the eyes, and whether a verged pair’s shells are more evenly spread across a wide field than a parallel pair’s — which would be a geometric reason for vergence that has nothing to do with where the zero is.
What links here
Computed from the collection, not written here: the essays that point at this one.
- Vergence moves the shells and does not respace them
- A scroll round a bend loses its straight-line depth
- A scroll through two slits ranges in a straight line
- Rectification is a family, not an operation
- A third eye that lands on the next post
- A third ray is worth what its picture is worth
- A mismatch on its own line needs a third eye
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two pictures on one screen — both name baseline, depth uncertainty, disparity, instrument limit
- A scroll can be asked its own radius — both name depth uncertainty, disparity, instrument limit
- An epipole in the picture leaves a blind disc — both name baseline, depth uncertainty, disparity
- Turning the cameras inwards — both name baseline, disparity, rectified pair
- A barrel model folds at a radius it sets itself — both name focal length, instrument limit
- A curved screen is eight flat ones — both name instrument limit, quantisation
Named objects
A flat tag is an object no other essay names yet.
BaselineDepth bufferDepth uncertaintyDisparityFocal lengthinstrument limitQuantisationrectified pair