Depth is a reciprocal
Worth reading first: A point is a line over there · The one thing a single view cannot give.
Two cameras side by side, facing the same direction, with their picture planes aligned. It is the configuration a stereo camera is built in, the one a pair of eyes approximates, and the one in which all the machinery of the previous field collapses into arithmetic.
The collapse is worth seeing. The epipolar lines of such a pair are horizontal and at the same height in both pictures, so a mark at row 173 in the left picture has its match at row 173 in the right. The only thing left to find is the difference in column. That difference is the disparity, and depth follows from it in one line:
where is the focal length in pixels, the distance between the eyes, and the disparity. Nothing more is needed. Given the disparity, the depth is arithmetic.
The reason this essay is not one paragraph long is the shape of that expression. Depth is not proportional to what is measured; it is proportional to the reciprocal of what is measured. And a reciprocal turns a symmetric error into an asymmetric one, then into an unbounded one.
The exact part
Before the uncertainty, the rule itself.
is not an approximation and it is not a small-angle result. It follows directly from similar triangles in the rectified configuration and it is exact. The figure checks it by projecting a point through both cameras, reading the two column positions, computing the disparity, applying the rule, and comparing with the depth the camera actually has. The agreement is 10⁻¹² m over sixty depths from 0.6 m to 40 m — arithmetic, not geometry.
That check is worth making rather than assuming, because the rule is usually derived from a diagram, and a diagram derivation is exactly where a factor of two or a sign convention goes missing. Here the rule is applied to a camera built from an eye, a target and a focal length, with no shared code between the projection and the rule, and the two agree.
Two properties of the exact rule are worth noting before the errors arrive.
Disparity is a difference of two column positions, so it is measured in pixels and does not care about the vertical. A point higher or lower in the frame has the same disparity for the same depth. That is a genuine simplification, and it is why the rectified configuration is worth engineering.
And disparity goes to zero for a point at infinity. A star has the same position in both pictures. Everything nearer has a positive disparity, growing as things get closer, and a point 0.5 m away on the figure’s rig has a disparity of 117 pixels. So the whole of depth is encoded in a quantity that runs from zero to a couple of hundred, and the near end of the range is where it is spent most generously.
Where the interval comes from
Suppose the disparity is read with an error of one pixel — a generous assumption for a good matcher on textured material, and an optimistic one for anything else.
The reported depth is . The true depth might be anywhere corresponding to a disparity between and . Feeding those into the rule gives two depths, and those two are the interval.
They are not symmetric about . The function is convex, so subtracting one pixel moves the depth further than adding one pixel does, and the asymmetry grows as shrinks — that is, as the point gets further away.
At 6.5 m on the figure’s rig the interval runs from 5.86 m to 7.33 m — lopsided by a factor of 1.25, and describable as roughly ±0.7 m without doing much violence.
At 40 m the disparity is 1.46 pixels, and the interval runs from 23.8 m to 126.5 m. That is 16.2 m nearer and 86.5 m further, and there is no honest way to describe it as a plus-or-minus.
The linearisation, and where it stops
The formula everyone quotes for this is the derivative:
It comes from differentiating , and it is correct as a derivative. The figure keeps it alongside the differenced interval so that the place it stops being useful is visible rather than assumed.
At 6.5 m it is 1.2% away from the half-interval. That is fine — the linearisation is doing its job, the curvature over one pixel is small, and quoting the derivative would mislead nobody.
At 40 m it is 47% away. The half-interval is nearly twice what the derivative claims, because the function’s curvature over a one-pixel step is no longer small when the disparity itself is under a pixel and a half.
And there is a second problem with the derivative that the percentage understates. A derivative is symmetric by construction — it produces a single number that is added and subtracted. At 40 m the far half of the interval is 5.32 times the near half, so a symmetric quantity cannot describe them however large it is made. The derivative is not merely 47% small; it is the wrong shape.
This is the same species of error the height recovery’s conditioning turned out to be. There, the sensitivity was computed correctly and the conclusion drawn from it was false, because the variable being perturbed was not the variable anyone actually reads. Here the derivative is computed correctly and describes a quantity — a symmetric error bar — that does not exist. In both cases the arithmetic is right and the object it describes is not the object in question.
The general habit is the one this site’s figures apply everywhere: a derivative quoted without being differenced is a claim about algebra rather than about the measurement. The figure computes both and shows where they part.
The range that has no upper bound
The interval’s far edge is . When falls to , that is a division by zero, and beyond it the far edge is negative — which in this context means there is no upper bound at all.
The disparity is , so it falls below when
For the figure’s rig — 65 mm apart, 900 pixels of focal length, one pixel of reading error — that is 58.5 m.
Past 58.5 m the pair can still produce a depth estimate, and the estimate can be perfectly reasonable. What it cannot produce is an upper bound. A point measured at 70 m might be at 70 m or at 700 m or at infinity, and the two pictures do not distinguish those cases, because the disparity difference between them is smaller than the reading error.
That number is worth carrying around, because it is a property of the rig alone and can be computed before anything is built. It does not depend on the scene, on the matcher’s cleverness, on the estimator, or on how many frames are averaged. It is the focal length times the baseline divided by the reading precision, and it scales the way one would hope and no better.
The three levers, and what each one costs
Everything about the rig’s performance is in , and it is worth taking the three quantities apart because they are not interchangeable even though they enter the expression the same way.
The baseline is the cheapest to change and the most consequential. Doubling it halves every depth error and doubles the range past which nothing is bounded. It costs nothing optically. What it costs is overlap: two cameras far apart see less of the same scene, and a point visible to one may be hidden from the other by the very geometry that makes the measurement precise. At the extreme the two pictures share almost nothing and there is nothing to measure at all. Every stereo rig is a compromise between precision and how much of the scene it can work on, and the compromise is not resolvable by cleverness.
The focal length enters identically in the arithmetic and differently in practice. A longer lens gives more pixels per degree, so a fixed angular error becomes a smaller fraction of the disparity and the depth improves in exact proportion. What it costs is field of view — a longer lens sees a narrower slice — and, on a real instrument, distortion, because the assumption that the picture is a pinhole projection is what makes the disparity mean what it is supposed to mean. That assumption is worth checking rather than making: an uncorrected radial term moves marks along their own radius, which changes the disparity of anything off the optical axis, and the resulting depth error is not a random scatter but a systematic warp of the whole reconstructed surface.
And the reading precision is the one that is not a design choice. It is set by the matcher, by the texture of the material, and by how much of the picture is in focus. Sub-pixel interpolation genuinely helps — a quarter-pixel matcher quadruples the range — but only where there is something to interpolate between. On a blank wall the disparity is unmeasurable at any precision, and the estimate is whatever the regulariser decided, which is not a measurement.
That last point deserves emphasis because it is the one the arithmetic cannot express. assumes a exists. Where the material has no texture, there is no disparity to read and no error bar to quote; the reported depth is an interpolation from elsewhere in the picture, and its uncertainty is not one pixel of anything. A depth map that reports a confident value on a featureless surface is reporting its own prior, and no formula in this essay applies to it.
What a person’s eyes make of it
The figure’s parameters are chosen to be roughly human, and the resulting numbers are instructive.
Sixty-five millimetres is a typical interocular distance. Nine hundred pixels of focal length is roughly the angular resolution of a good camera and in the region of foveal acuity. One pixel of disparity error is optimistic for a matcher and pessimistic for a visual system that pools over time and across many features.
With those, depth from binocular disparity alone is very good indoors — a centimetre or so at arm’s length, four millimetres at half a metre — degrades to a couple of metres at ten, and stops bounding anything at around sixty.
That matches ordinary experience closely enough to be worth remarking on. Reaching for a cup is a precise operation; judging whether a car two hundred metres down the road is closer than a building behind it is not something binocular vision does at all. Whatever supplies the answer at that range — occlusion, familiar size, motion, perspective — it is not the disparity between two eyes, because the disparity is under a pixel.
This site does not claim anything about perception, and this paragraph is not a claim about it. The geometry says what the disparity is; whether a visual system uses it, pools it, or ignores it is a fact about seeing, and the site’s standing limit applies. What the geometry does establish is that beyond a computable range there is nothing there to use.
Why the band is drawn rather than the numbers quoted
A design note about the figure, because it went through a version that was worse.
The obvious drawing is the curve with an error bar at a couple of depths. That version shows a hyperbola, which everyone already expects, and two bars, which look like ordinary measurement scatter.
The version drawn instead fills the whole region between the near and far edges. What that shows is a band that is nearly a line at close range and opens into a wedge, and the wedge is visibly not centred on the curve — the space above the line is much larger than the space below it, and the disproportion grows across the frame.
That asymmetry is the essay’s content, and it is the thing a pair of error bars hides. Three depths are marked with their intervals printed, so the numbers are available, but the shape is what the figure is for.
The first caption for it compared the derivative against the whole interval rather than against the half-interval, which made the derivative look five times wrong at 10 m when it is within 4%. The comparison was between quantities of different kinds, and it hid the real failure — which happens further out, and is about shape rather than size.
What this does not need
It is worth being clear about how little of the previous field’s machinery this essay used.
No fundamental matrix. No essential matrix. No pose recovery, no decomposition, no cheirality test. The rectified configuration is given, the baseline is given, and the calculation is one division.
That is the trade a stereo rig makes. Engineering the configuration — mounting two cameras rigidly, aligning them, measuring the baseline once — replaces the whole of the recovery with arithmetic, and buys the one thing two arbitrary photographs cannot supply, which is a length. The metres in every number above come from the tape measure applied to the rig, not from the pictures.
What the rig does not buy is precision at range, and no amount of care in the mounting changes the in the sensitivity or the in the horizon. Those come from the reciprocal, and the reciprocal comes from the geometry.
Disparity as the quantity actually recovered
There is a way of reading the reciprocal that makes its consequences feel less like a nuisance, and it is worth stating because it changes what a rig should be thought of as producing.
The quantity a stereo pair measures, directly and with a uniform error bar, is disparity — that is, inverse depth. Depth itself is a derived quantity, and every awkwardness in this essay comes from the derivation rather than from the measurement.
In inverse depth the picture is entirely well behaved. The error is one pixel, everywhere, at every range. There is no asymmetry, no , no horizon past which the bound disappears. A point at 40 m and a point at 4 m are measured equally well; they simply have very different disparities.
What changes when the reciprocal is taken is not the information but its parameterisation. Inverse depth is uniformly spaced in what the instrument reads; depth is not. So a rig that is described as “accurate to a metre at ten metres” is being described in a coordinate that stretches, and the same rig is accurate to a hundred metres at fifty and to nothing at all at seventy.
That observation has a practical form which is standard in the field and which this essay’s arithmetic explains. Reconstructions from image measurements are often carried in inverse depth precisely because the uncertainty is well behaved there — a Gaussian in inverse depth is a sensible model and its image in depth is the lopsided thing the figure draws. Converting to metres is a display choice, and it should be made last.
The remaining question is what happens to the arithmetic when the configuration is not given — when the two eyes are close together, or when the marks fed into it are not the right marks. Both turn out to fail in ways that produce confident wrong answers rather than errors, and both are the subject of the essays that follow.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Named objects
A flat tag is an object no other essay names yet.
BaselineDepth uncertaintyDisparityerror propagationFocal lengthrectified pairTriangulation