The second projection

Two pictures on one screen

A stereoscopic display puts a point where two sightlines cross, so the depicted depth is b·D/(b−d) and the disparity that reaches infinity is exactly the separation of the reader's eyes — 63 mm, at any screen distance whatever. The depth budget is set by the width of a head and by nothing about the scene.

Worth reading first: Depth is a reciprocal · The point you have to stand at.

A stereoscopic display shows each eye a different picture. Everything else about it — the polarisers, the shutters, the lenticular sheets, the headsets — is machinery for arranging that one fact, and none of it changes the geometry.

The geometry is a triangle. The two eyes are at known places; each is shown a mark at a known place on the screen; the depicted point is where the two sightlines cross. One similar-triangle argument gives the whole thing:

z=bDbdz = \frac{b\,D}{b - d}

with bb the separation of the eyes, DD the distance to the screen and dd the horizontal separation between the two images of the point on the glass.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 2.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.010203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 2.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 1 Depicted depth against screen disparity, for eyes 63 mm apart at a screen 2 m away. Zero disparity puts the point on the glass; crossed disparity brings it forward; and at d = 63 mm the point reaches infinity. Past that the display is asking the eyes to diverge, and the wall is drawn where it is.

Three regimes, all visible in one formula

d=0d = 0 puts the point at z=Dz = D — on the screen. So the screen plane is the zero-disparity plane, and it is the one depth a stereoscopic display renders with no burden on the viewer at all. Everything a designer wants to be comfortable goes there.

d<0d < 0 — the left eye’s image to the right of the right eye’s, which is crossed disparity — brings the point in front of the screen. The formula handles it with no special case: a negative dd makes the denominator larger than bb and zz smaller than DD.

dbd \to b sends the denominator to zero and the point to infinity. And d>bd > b makes the denominator negative, which the formula reports as a point behind the viewer and which the eyes report as an instruction to point outward.

That third regime is the one worth dwelling on.

The ceiling is a head

Eyes converge inward easily. They do not diverge — the resting state is very nearly parallel and outward rotation past parallel is not available. So a disparity greater than the interocular distance is a request the visual system cannot fulfil, and what it produces is not a very distant object but a failure to fuse.

The formula says the boundary is at d=bd = b exactly, and it says so with DD absent. The gate asserts that: the same 63 mm disparity reaches infinity at a screen 2 m away and at one 6 m away.

That is a genuinely odd and rather good result. A stereoscopic display’s maximum usable disparity is a property of the reader’s skull, and does not change when the screen moves, or grows, or changes resolution. Every other quantity in display engineering scales with something; this one is 63 mm, on any display, for ever.

Its practical form is the one cinema learned expensively. A film graded for a cinema screen has disparities measured in fractions of the screen width; shown on a television a tenth as wide, those fractions become a tenth as many millimetres and everything moves toward the screen plane. Shown on a screen wider than the one it was graded for, they become more millimetres — and the ones near the ceiling go past it. A stereoscopic film is correct for one screen width and cannot be correct for two, which is a much stronger statement than the one this site makes about ordinary photographs, where the wrong width merely stretches depth.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 5.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.10203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 5.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 2 The same curve for a screen 5 m away. Everything has scaled with D except the wall, which is still at 63 mm because it is the reader’s own eye separation. A display’s depth budget in millimetres is fixed; what changes with the screen is how much of the depicted scene fits inside it.

This is the site’s own stereo geometry, read backwards

The depth field already built this. Depth is a reciprocal takes a stereo camera pair with baseline bb and focal length ff, and derives

z=bfdz = \frac{b\,f}{d}

which is the same triangle with the eyes replaced by cameras and the screen by an image plane.

A stereoscopic display is that construction with the arrows reversed. The camera pair converts a depth into a disparity; the display converts a disparity back into a depth; and the two are exact inverses when the display’s geometry matches the camera pair’s. Match means: the ratio of screen distance to screen width equals the ratio of the cameras’ focal length to their sensor width, and the eye separation equals the camera baseline.

Almost nothing meets both conditions. A stereo pair shot with a 65 mm baseline — human separation, the obvious choice — and shown on a screen the viewer sits closer to than the shot’s station point produces a scene compressed in depth; shown further back, a scene stretched. And a pair shot with a wider baseline for a distant subject produces, on any display, a scene depicted at a smaller scale than life, which is the “miniature” look of long-baseline stereo. Every one of those is the same mismatch this field has measured twice already, with a second projection in place of the first.

Depth from disparity, with the 1 px the reading is worthZ = fB/d on a 65 mm baseline at 900 px. The line is exact — it returns the camera's own depth to 1e-12 m. The band is what 1 px of disparity error costs, and it stops being a ±. At 6.5 m it runs 5.86–7.33 m, lopsided by 1.25, and the textbook ±Z²δ/fB is 1.2% out. At 40 m it runs 23.8–126.5 m — 86.5 m beyond the estimate against 16.2 m before it, a lopsidedness of 5.32 — and the same formula is 47% out. Past 58.5 m the far edge is infinity.025507510010203040true depth (m)depth reported from the disparity, with a 1 px reading error5.86–7.33 m13.96–26.69 m23.76–126.49 mat 40 m: +86.5 m against −16.2 munbounded past 58.5 m
Fig. 3 The camera-side version of the same triangle, from the depth field. Depth against disparity for a 65 mm baseline — the same hyperbola, with the same wall at the end, read in the other direction. A display is a stereo pair run backwards, and the conditions under which the two invert each other exactly are conditions almost nothing meets.

Depth resolution, from the same curve

The curve is a hyperbola, so its slope is not constant, and the slope is what says how much depth a given change in disparity buys.

dzdd=bD(bd)2=z2bD\frac{\mathrm{d}z}{\mathrm{d}d} = \frac{b\,D}{(b-d)^2} = \frac{z^2}{b\,D}

Quadratic in the depicted depth — the same shape as a stereo pair’s depth uncertainty, and for the same reason, since it is the same triangle.

Which has a consequence for a display that the camera-side essay states for a measurement. A display with a finite pixel pitch has a finite disparity resolution, so it has a depth resolution that degrades as the square of the depicted distance. On a screen 2 m away with 0.25 mm pixels, one pixel of disparity distinguishes depths of about 8 mm at the screen plane and about 2 m at a depicted 30 m.

So a stereoscopic display renders depth finely near the glass and coarsely far behind it, and the transition is quadratic. That is the same lopsidedness a depth buffer has, arrived at from a completely different direction — one is a quantisation of 1/z1/z in a register and the other is a quantisation of disparity on a panel, and both come out as z2z^2.

The reason they agree is not a coincidence and is worth stating: both are quantisations of something proportional to 1/z1/z, and differentiating 1/z1/z gives 1/z2-1/z^2 whatever the something is.

Where a stereo pair stops being able to say "no further than"The disparity of a point at range Z is fB/Z, so once it falls below the 1 px the reading is worth, the far edge of the depth interval is infinity. That happens at fB/δ, and for a human pair of eyes at 900 px it is 58.5 m. Doubling the baseline doubles it; nothing about the estimator does.123-2-10baseline (m, log scale)range past which 1 px of disparity error is unbounded (m, log scale)65 mm → 58.5 mrange = fB/δ — 58.5 m at a 65 mm baseline900 px focal length, 1 px reading
Fig. 4 The camera-side version, from the depth field. Depth uncertainty against distance for a stereo pair with one pixel of reading error — quadratic, and eventually unbounded. A display’s depth resolution has the identical shape because it is the identical triangle, with the pixel pitch on the panel in place of the reading error on the sensor.

The other distance, and the split between them

There is a second geometric fact about a stereoscopic display and it is the one everything else about the technology is organised around.

The eyes focus on the screen and converge on the depicted point. Those are two different distances the moment the disparity is not zero, and in ordinary vision they are always the same distance — the mechanism that focuses and the mechanism that converges are yoked, because in a world made of real objects the thing being looked at is at one distance.

The gap is conventionally measured in dioptres, which is reciprocal metres, because that is the unit accommodation works in. With the screen at 2 m the eyes accommodate at 0.50 D whatever is depicted, and converge at 1000/z1000/z mm⁻¹.

disparity depicted at vergence split
−20 mm 1.52 m 0.66 D 0.16 D
0 2.00 m 0.50 D 0
30 mm 3.82 m 0.26 D 0.24 D
63 mm 0 D 0.50 D

The last row is worth reading twice. At a screen 2 m away, the split for something depicted at infinity is exactly 0.50 D — which is 1000/20001000/2000, the screen’s own dioptric distance. That is not a coincidence and it is not an approximation: anything depicted behind the screen has a split bounded by the screen’s own dioptric distance, because vergence for a distant object goes to zero and the focus stays where the glass is.

So a display 2 m away has half a dioptre of headroom for the whole of the space behind it, and a display 400 mm away — a desk monitor — has two and a half dioptres. Which is why near stereoscopic displays are harder than far ones, in a way that has nothing to do with the display.

The eyes focus on the glass and converge somewhere elseWith the screen at 2.0 m the eyes accommodate at 0.50 dioptres whatever is depicted, and converge at the reciprocal of the depicted distance. The gap reaches the industry's half-dioptre rule of thumb at a disparity of about 63 mm. This is the geometry of the split; whether it is comfortable is a question about eyes and this site does not answer it.00.2000.400-2502550disparity on the screen — millimetresfocus and vergence, apart, in dioptreszero disparity — on the glassscreen at 2.0 m · eyes 63 mm apartthe split is geometry; comfort is not
Fig. 5 The split, in dioptres, against disparity. It rises to 0.50 D — the screen’s own dioptric distance — for anything depicted at infinity, and grows without bound in front of the screen. The industry’s rule of thumb is a third to a half of a dioptre; where that line falls is drawn, and whether it is the right line is a question about eyes.

Why the screen plane is a design decision

One consequence of the arithmetic is worth stating as a rule, because it is the one thing a designer of a stereoscopic picture actually controls.

The zero-disparity plane is where the depicted scene meets the glass, and moving it costs nothing geometrically — shifting both images horizontally by equal and opposite amounts moves the whole depicted scene toward or away from the viewer, uniformly in disparity, without changing any shape.

So the whole depicted scene can be slid along the depth axis freely, and the constraint is at the ends: nothing may exceed 63 mm of disparity behind, and nothing may come so far forward that the vergence split becomes unreasonable. The design decision is where inside that window to put the scene, and it is a real choice with no geometric preference.

The convention that emerged in cinema is to put the screen plane near the front of the depicted scene, so that most of the action is behind the glass. That places the burden on the divergence limit rather than on the crossed-disparity one, which is the more forgiving of the two — the window behind the screen is bounded by the head’s width and the one in front is not bounded at all, only made progressively more demanding.

Where this site stops

The comfort question is not a geometric one and this site does not answer it.

What is computed here is a split between two distances, exactly, in a stated unit. Whether a viewer finds a given split comfortable, how quickly they adapt, whether it causes fatigue, and how it varies between people are questions about the visual system, and the site has said since its foundation that it computes the geometry of pictures and says nothing about how one is perceived.

The half-dioptre figure quoted above is the industry’s rule of thumb and is quoted as such — it is where a line is conventionally drawn, and this site has no standing to say whether it is drawn in the right place. What the geometry supplies is the ability to say where any particular disparity falls relative to whatever line somebody else draws, which is the useful half.

That boundary is the same one the site draws everywhere. Wide-angle is not distortion computes the stretch and declines to say whether it looks bad. Standing in the wrong place computes the depth factor and notes that viewers mostly do not notice. Geometry produces numbers; what to do about them is somebody else’s field.

The two projections a stereoscopic display has

Everything in the screen field so far concerns a picture projected twice — once by the camera and once off the glass. A stereoscopic display has the second projection twice over, and the two copies do not have the same station point.

Each eye sees the screen from its own position, so each has its own station distance and its own angle to the glass. For a viewer square to a screen at 2 m with eyes 63 mm apart, the two eyes are 31.5 mm either side of the centre line — a difference in station point of 63 mm on a distance of 2000, which is 1.8°.

That is small and it is not nothing, and it has a consequence with a name. The two pictures must be rendered from two different virtual cameras, and those cameras must be parallel with offset principal points rather than converged. Toe the two cameras inward to meet at the subject and the picture planes are no longer coplanar, so the two images of a point differ vertically as well as horizontally — and a vertical disparity is one the eyes cannot fuse at all, at any size.

The geometry says so directly: two cameras with coplanar image planes give correspondences differing only along the baseline, which is the epipolar geometry of a rectified pair. Toe them in and the epipolar lines are no longer horizontal, and every point off the convergence plane acquires a vertical component.

So the standard advice — never converge the cameras, shift the frusta instead — is not a preference. It is the requirement that the pair be rectified, and rectification is exactly the condition that the two pictures can be fused by two eyes that cannot rotate about the line between them.

A point on the left is a line on the rightFive points in the left picture. Each one fixes a line in the right picture — the image of the ray it came along — and its match lies on that line to 2.8e-13 px. The five lines meet at the epipole, 1.2e-10 px from concurrent. Knowing where a point is in one picture does not say where it is in the other; it removes one of the two degrees of freedom.12341234a point is chosen hereand must be on this linematch to its own epipolar line: 2.8e-13 px4 of 44 correspondences drawn
Fig. 6 The condition a stereoscopic pair has to satisfy, from the twoviews field. With coplanar image planes a point in one picture constrains its match to a horizontal line in the other; toe the cameras in and that line tilts. A tilted epipolar line is a vertical disparity, and vertical disparity is the one thing two eyes cannot fuse.

What two pictures buy, and what they still do not

Ending on the thing this field has in common with the twoviews field, because it is the same limit seen from the display’s side.

Two pictures of a scene determine its shape and not its size — that is the two-view field’s central result, and it is a statement about what a pair of pencils of rays contains.

A stereoscopic display has the same property from the other end. Two pictures shown to two eyes determine a shape; what fixes the size is the eye separation, which the display does not control and cannot know. A viewer with eyes 58 mm apart and one with eyes 70 mm apart, watching the same film from the same seat, are shown two scenes of different sizes — the same shape, scaled by the ratio of their separations, which is about 20%.

That is a genuinely peculiar property of the medium and it falls straight out of the formula: bb appears in the numerator, so zz scales with it. Nothing else in the history of pictures has the property that the depicted object’s size depends on the width of the observer’s head.

One reconstruction, drawn at its own scale and at 3.4×The courtyard recovered from two pictures, seen from above. Every pairwise distance ratio matches the world's to 6.3e-14, so the shape is exact. The size is not determined at all: the right-hand plan is the same reconstruction 3.4 times larger and fits the same two pictures equally well. One measured length in the scene — here 0.800 m — fixes it, and nothing in the pictures can.as recovered0.308 across× 3.4 — same two pictures1.046 acrossworst ratio deviation 6.3e-14size fixed only by a supplied 0.80 m
Fig. 7 The limit from the camera’s side. Two views determine a shape and leave the scale free, so the same pair of pictures is consistent with a scene at any size. A stereoscopic display resolves that freedom with the viewer’s own eye separation — which means the scale is set by a measurement the display never makes.
The degeneracy is in the reading, not in the algebraShape error against baseline, twice. Given exact correspondences the recovery is good to 4.7e-9 at a 2 mm baseline and improves from there — the algebra does not notice a short baseline at all. Read the same points to 1 px and the same geometry comes back 6e+9× further from the truth, and it does not refuse: it returns a confident wrong answer. What fails when a photographer swings a camera instead of stepping sideways is the ratio of baseline to reading error, and nothing in the computation reports it.-10-50-2-10baseline (m, log scale)worst deviation in the recovered shape (log scale)exact pointsread to 1 pxat 2 mm: 4.7e-9 exact, 2.9e+1 read to 1 px6e+9× apart
Fig. 8 And the pairing this all rests on. What a second view supplies depends on the baseline, and the whole of stereoscopic display is an argument about one particular baseline — 63 mm, fixed, unchosen, and the thing that sets both the depth scale and the ceiling.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineDemonstrationDepth cueDepth uncertaintyDisparityinstrument limitreconstruction ambiguityStation pointTriangulationViewing distance