What a pair is for

The range a pair cannot see past

A stereo rig has a distance beyond which it cannot say "no further than", and the distance is fixed before anything is built. It is the focal length times the baseline divided by the reading precision, and for a human pair of eyes it is fifty-eight and a half metres.

Worth reading first: Depth is a reciprocal · A turn of the head is not a step sideways.

Most statements about measurement accuracy are statements about effort. Read the marks more carefully, average more frames, model the noise better, and the numbers improve.

This one is not. A stereo pair has a range past which it cannot place an upper bound on distance, the range is decided by three numbers, and no amount of care with the estimator moves it.

Zmax=fBδZ_{\max} = \frac{fB}{\delta}

Focal length in pixels, baseline in metres, reading precision in pixels. For the pair used throughout this field — 65 mm apart, 900 pixels, one pixel of reading error — it is 58.5 m.

Where a stereo pair stops being able to say "no further than"The disparity of a point at range Z is fB/Z, so once it falls below the 1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 900 px it is 58.5 m. Doubling the baseline doubles it; nothing about the estimator does.123-2-10baseline (m, log scale)range past which 1 px of disparity error is unbounded (m, log scale)65 mm → 58.5 mrange = fB/δ — 58.5 m at a 65 mm baseline900 px focal length, 1 px reading
Fig. 1 The disparity of a point at range Z is fB/Z, so once it falls below the 1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 900 px it is 58.5 m. Doubling the baseline doubles it; nothing about the estimator does.

Why it is a hard bound

The argument is three lines and its shortness is the point.

The disparity of a point at range ZZ is fB/ZfB/Z. It falls as the point recedes. When it falls below the precision with which disparity can be read, two things at different distances produce readings that cannot be told apart.

Specifically, when the disparity is at or below δ\delta, the reading is consistent with a disparity of zero — and a disparity of zero is a point at infinity. So the interval of depths consistent with the reading extends to infinity, and the pair has no upper bound to report.

Nothing in that argument mentions an algorithm. It does not assume least squares, or a particular noise model, or an estimator of any kind. It says that two configurations of the world produce measurements that differ by less than the measurement precision, and therefore cannot be distinguished by any procedure whatever operating on those measurements.

That is a different kind of statement from an accuracy figure, and the difference is worth naming. An accuracy figure describes how well a method does. This describes what the data contains. A better method cannot cross it, and a method that appears to cross it is supplying the extra information from somewhere else.

What it scales with, and what it does not

The expression has three levers and each behaves exactly as written.

Doubling the baseline doubles the range. Two cameras 130 mm apart rather than 65 reach 117 m rather than 58.5. This is the strongest lever available and it costs overlap: cameras further apart share less of the scene, and at the extreme they share nothing.

Doubling the focal length doubles the range. A longer lens spends more pixels per degree, so the same angular difference becomes a larger disparity. It costs field of view, and it makes the pinhole assumption harder to sustain on a real instrument.

Halving the reading error doubles the range. A quarter-pixel matcher reaches 234 m where a whole-pixel one reaches 58.5. This is the lever most often pulled, because it is a software change rather than a mechanical one, and it is the one most sharply limited by the material — a sub-pixel reading requires something to interpolate between, and a blank wall supplies nothing at any precision.

The figure sweeps the baseline over three decades and the relationship is exactly linear, which is not a surprise and is worth drawing because it makes the absence of any other dependence visible. There is no scene in the figure, no matcher, no estimator, no noise model. There is a rig, and a range.

Where a stereo pair stops being able to say "no further than"The disparity of a point at range Z is fB/Z, so once it falls below the 0.1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 900 px it is 585.0 m. Doubling the baseline doubles it; nothing about the estimator does.234-2-10baseline (m, log scale)range past which 0.1 px of disparity error is unbounded (m, log scale)65 mm → 585.0 mrange = fB/δ — 585.0 m at a 65 mm baseline900 px focal length, 0.1 px reading
Fig. 2 A tenfold better matcher, and a tenfold longer reach. The angular resolution and the baseline enter only as a ratio, so improving either moves the line by the same factor.

The same limit, seen as an angle

There is a second reading of fB/δfB/\delta that makes it feel less like a formula and more like a fact about triangles, and it is worth having because it transfers to configurations that are not rectified pairs.

A disparity of δ\delta pixels at a focal length of ff pixels is an angle of about δ/f\delta/f radians. So the condition for the far bound to disappear is that the angle subtended by the baseline, seen from the point, has fallen below the angular precision of the instrument.

That is the parallax angle, and it is the quantity every distance measurement by triangulation is really reading — a surveyor’s baseline, a pair of eyes, or the Earth’s orbit measured six months apart. The range at which each of them runs out is the range at which its baseline subtends less than one unit of its own angular resolution, and the formula above is that statement with the units cleared.

Reading it that way explains why the three levers are not really three. Baseline and range enter only through their ratio, because that ratio is the angle; focal length and reading precision enter only through their ratio, because that ratio is the angular resolution. There are two quantities, not four, and the condition is that the first exceeds the second.

It also explains why the bound is indifferent to the estimator in the strong sense claimed above. Two points at different distances whose parallax angles differ by less than the instrument’s angular resolution produce, to that instrument, the same pair of directions. No processing of a measurement can distinguish inputs that are identical.

Some rigs, and where each one stops

Putting numbers into the expression is more instructive than the expression is.

A human pair of eyes, taking 900 pixels as a rough stand-in for foveal angular resolution and one pixel as an optimistic disparity precision, reaches about 58 m. Whatever tells a person that a distant hillside is further away than a nearer one, it is not binocular disparity, because past sixty metres there is under a pixel of it.

Two cameras 12 mm apart on the back of a phone, at 1,200 pixels of focal length, reach about 14 m at one pixel and 58 m at a quarter. That is a rig whose useful range is a room, and the arithmetic says so before anything is photographed.

A survey rig with a metre of baseline at 3,000 pixels reaches 3 km at a pixel. A pair of satellite images taken hundreds of kilometres apart reach across a continent. The expression is scale-free and it does not stop working; it is the same statement at every size.

And the interesting inversion: for any desired range, the required baseline is Zδ/fZ\delta/f. Wanting to bound distance at a kilometre with a 1,000-pixel camera at one pixel needs a metre of baseline. That is a design calculation available at the start, and it does not become available at the end after a matcher has been written.

What happens past the range

It is worth being exact, because “cannot see past” overstates it.

Past fB/δfB/\delta the pair can still report a depth. The disparity is not zero, it is merely under a pixel, and a sub-pixel estimate exists. That estimate can be perfectly reasonable and is often approximately right.

What is gone is the upper bound. The near end of the interval survives — a reading of 0.8 pixels does say the point is at least fB/(0.8+δ)fB/(0.8+\delta) away, which is a real and useful constraint. The far end does not. The point may be at the estimate, or at ten times it, or at infinity, and the two pictures contain nothing that distinguishes them.

So the correct statement about a stereo pair beyond its range is that it measures a minimum distance and not a distance. That is a genuinely useful measurement — knowing that nothing is nearer than thirty metres is worth having — and it is a different measurement from the one usually reported.

The one-sidedness is the part that makes a plus-or-minus impossible, and it is the reason the previous essay’s figure draws two edges rather than a bar. The near edge behaves; the far edge runs away; and averaging them produces a number that describes neither.

Where a stereo pair stops being able to say "no further than"The disparity of a point at range Z is fB/Z, so once it falls below the 1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 2400 px it is 156.0 m. Doubling the baseline doubles it; nothing about the estimator does.123-2-10baseline (m, log scale)range past which 1 px of disparity error is unbounded (m, log scale)65 mm → 156.0 mrange = fB/δ — 156.0 m at a 65 mm baseline2400 px focal length, 1 px reading
Fig. 3 The same sweep at 2400 px of focal length. A longer lens moves the whole line up by the same factor it multiplies the focal length by, which is what “linear in f” looks like when both axes are logarithmic.

Where the reading precision actually comes from

The third lever is the one written as a single symbol and is the least like a single number, so it is worth taking apart.

δ\delta is the precision with which the difference between two column positions can be read. That is not the precision with which a mark can be located, and the distinction matters in both directions.

It can be better than the mark’s own precision. A mark whose absolute position is uncertain by two pixels — because the corner it sits on is soft, or the exposure differs between the two cameras — may still have a disparity good to a quarter of a pixel, if the same error is made in both pictures. Two views of one feature through two nearly identical lenses make correlated errors, and the correlated part cancels in the difference. That is the mechanism a sub-pixel matcher exploits, and it is why sub-pixel disparity on well-textured material is routine rather than optimistic.

And it can be much worse. Anything that differs between the two cameras enters the difference directly rather than cancelling: a difference in focus, a specular highlight that moves with viewpoint, a surface seen at a grazing angle by one camera and squarely by the other. On a slanted surface the same physical patch has different shapes in the two pictures, so a matcher comparing patches is comparing things that genuinely differ, and the resulting bias is not noise — it is a systematic pull whose sign depends on the slant.

That last point is the reason fB/δfB/\delta should be read as an optimistic bound rather than as a specification. It assumes a δ\delta that is uniform and unbiased. Real disparity error is neither: it is small on textured fronto-parallel surfaces, large on slanted or shiny ones, and undefined where there is no texture at all. The range computed from a nominal δ\delta is the range the rig achieves on its best material.

None of that is this site’s subject — how a patch is compared is matching, and matching is explicitly not claimed here. What belongs here is the shape of the dependence: whatever the matcher achieves, the range is linear in it, and a matcher twice as good buys exactly twice the range and not more.

Depth from disparity, with the 1 px the reading is worthZ = fB/d on a 12 mm baseline at 1200 px. The line is exact — it returns the camera's own depth to 1e-12 m. The band is what 1 px of disparity error costs, and it stops being a ±. At 3.5 m it runs 2.82–4.64 m, lopsided by 1.64, and the textbook ±Z²δ/fB is 5.9% out. At 14 m it runs 7.1–928.1 m — 914.0 m beyond the estimate against 7.0 m before it, a lopsidedness of 129.91 — and the same formula is 97% out. Past 14.4 m the far edge is infinity.020405101520true depth (m)depth reported from the disparity, with a 1 px reading error2.82–4.64 m5.66–26.50 m7.14–928.15 mat 14 m: +914.0 m against −7.0 munbounded past 14.4 m
Fig. 4 A pair of cameras 12 mm apart on the back of a phone. The band opens almost immediately and the far edge escapes within a room, which is the arithmetic of this essay applied to a rig most people carry.

The comparison with a single view

There is a temptation to read this range as the point at which a stereo pair becomes as bad as a single photograph. It is worth resisting, because the two limits are of different kinds.

A single view supplies no depth at any range. Not a poor estimate, not a wide interval: no depth. What it supplies is ratios along a line, and every three-dimensional statement made from one picture is bought with an assumption about the scene.

A stereo pair past its range supplies a minimum distance and no maximum. That is strictly more than nothing, and it is arrived at without any scene assumption at all.

The two limits also behave differently under improvement. Nothing improves the single view’s depth, because the information is absent rather than faint — the 137× demonstration shows two worlds producing an identical picture, which is absence. The stereo pair’s range moves with ff, BB and δ\delta, so it is a limit on this instrument rather than on the method.

Which raises the obvious question about the other direction. If a longer baseline is straightforwardly better, what is the shortest baseline that still works? The answer is not what the arithmetic above suggests, and it is the subject of the next essay: the algebra does not degrade at short baselines at all, and what degrades is the ratio of baseline to reading error — which fails without announcing itself.

Two rays, 0.85 mm apart, in the plane that contains bothThe ray from the left eye through its mark and the ray from the right eye through its. With the marks placed exactly they meet, to 6.3e-15 m. With the same marks read to 1 px they miss by 0.85 mm at a range of 7.32 m. Triangulation is not an intersection; the reported point is a choice about what to minimise, and the gap is the part a residual alone will not tell you.midpoint — 0.85 mm gapfrom the left eyefrom the right eyegap 0.85 mm at 7.32 mexact marks: 6.3e-15 m
Fig. 5 What the far bound looks like from the triangulation side. Two rays crossing at a shallow angle miss each other by a real length, and the shallower the crossing the longer the region they could have met in.

A second bound, at the other end

The range has a companion that is easy to overlook because it is set by a different mechanism, and a rig has both.

At the near end, the disparity grows without limit as a point approaches. On the figure’s rig a point at 0.5 m has a disparity of 117 pixels; at 0.2 m it would be nearly 300; at 0.05 m, well over a thousand. Two things go wrong there and neither is about precision.

The first is overlap. A point very close to a rig is seen by the two cameras in very different parts of their frames, and eventually falls outside one of them. There is then no correspondence to read, and the depth is not merely imprecise but absent. The distance at which that happens depends on the field of view and the baseline, and for a wide-angle pair it is closer than one would guess.

The second is that the two views of a nearby surface differ so much that they stop being views of the same thing in any useful sense. A face at fifteen centimetres presents visibly different geometry to two eyes 65 mm apart — one sees a cheek the other does not — and no matcher recovers a correspondence that does not exist.

So a stereo pair works between two distances, not below one. The far bound is fB/δfB/\delta and is about precision; the near bound is about overlap and is about the field of view. The two move in opposite directions with the baseline, which is the compromise at the heart of every rig: widening the baseline pushes the far bound out and pulls the near bound in.

That trade has no optimum in general, because the two ends are not commensurable — a metre of extra range is not exchangeable for a centimetre of extra reach. What the arithmetic supplies is both numbers, before anything is built, so the choice is made against them rather than after.

Depth from disparity, with the 1 px the reading is worthZ = fB/d on a 500 mm baseline at 900 px. The line is exact — it returns the camera's own depth to 1e-12 m. The band is what 1 px of disparity error costs, and it stops being a ±. At 18.5 m it runs 17.78–19.30 m, lopsided by 1.09, and the textbook ±Z²δ/fB is 0.2% out. At 120 m it runs 94.7–163.6 m — 43.6 m beyond the estimate against 25.3 m before it, a lopsidedness of 1.73 — and the same formula is 7% out. Past 450.0 m the far edge is infinity.050100150255075100true depth (m)depth reported from the disparity, with a 1 px reading error17.78–19.30 m48.48–61.79 m94.74–163.64 mat 120 m: +43.6 m against −25.3 munbounded past 450.0 m
Fig. 6 A half-metre baseline, out to 120 m. The same law, the same shape, and a far edge that survives a long way further — which is what the limit above says in advance.

The limit as a thing to state

A closing remark about why this deserves a figure and an essay rather than a formula in a caption.

Almost every practical account of stereo vision quotes the depth error Z2δ/fBZ^2\delta/fB and stops. That expression is a derivative, it is correct where the curvature is small, and it says nothing about the range at which the interval stops being bounded — because a derivative is a local object and cannot report a global failure.

Quoting the derivative alone therefore describes a rig as having a smoothly degrading accuracy, and the honest description is that it has a smoothly degrading accuracy up to a specific distance and then a cliff. The cliff is computable before the rig exists, it is not in the derivative, and it is the number an engineer actually needs.

This site’s habit is that every claim should be given a test it could fail, and the corresponding habit for a limit is that it should be a number rather than a caution. “Stereo degrades with distance” is a caution. “This pair cannot bound distance past 58.5 m, and doubling the baseline moves that to 117” is a limit, and the difference is that the second one can be checked, designed against, and found to be wrong.

The distinction matters more than it may seem, because the two statements fail differently in practice. A caution invites the reader to be careful, and careful is not a quantity. A limit invites the reader to compute whether their case falls inside it, and computing that takes three numbers most of which are stamped on the equipment. Two people disagreeing about whether a rig is adequate for a task can settle it in a line of arithmetic if the limit is stated and cannot if only the caution is.

And a limit can be violated, which is the property that makes it worth having. A depth map reporting confident distances at four times fB/δfB/\delta is reporting something other than disparity — a prior, a smoothness assumption, a learned expectation about how large things usually are. Those may be excellent guesses. They are not measurements of the two pictures, and the only way to notice the substitution is to have written down what the two pictures could contain.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineDepth uncertaintyDisparityFocal lengthinstrument limitrectified pairReference length