The second projection

Turning the cameras inwards

A stereo pair made by rotating two cameras toward a common point puts the same world point at different heights in the two pictures — up to thirty pixels here, on a frame of four hundred. Two eyes level with each other see every point at the same height, so a pair with vertical difference is a pair of pictures of no scene at all.

Worth reading first: Two pictures on one screen · Four numbers and a window.

Two pictures on one screen works out what a stereoscopic display does with two images: the depicted depth is b·D/(b − d), and the disparity that reaches infinity is exactly the separation of the reader’s eyes, at any screen distance whatever.

That essay assumes the pair. This one is about making it, and about the arrangement that looks right and is not.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 2.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.010203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 2.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 1 What a stereo pair is for: screen disparity to depicted depth, with the divergence ceiling at the reader’s own interocular distance.

Two ways to point two cameras

A stereo pair needs two cameras a baseline apart, and something has to decide where the zero-disparity plane is — the depth at which a point lands in the same place in both pictures and therefore appears to sit on the screen.

Toe-in rotates both cameras to look at a chosen point on the axis. It is the obvious construction, it is what a pair of eyes does, and it is what every rig with two tripod heads makes easy.

Off-axis leaves both optical axes parallel and moves the principal points instead, by ∓ b·f/2D — half the baseline times the focal length, over the convergence distance. That is a frustum with its four sides shifted, which four numbers and a window says is the same camera looking through a window that is not centred on it.

Both put the convergence point at zero disparity, exactly, and that is where their agreement ends.

Two cameras turned inwardsThe left picture's marks, with each point's **vertical** difference from the right picture drawn 8× exaggerated. Two eyes level with each other see every point at the same height, so a pair with vertical difference is a pair of pictures of no scene at all. Here it reaches 30.8 px.vertical difference up to 30.8 px, drawn 8× overbaseline 63 mmconvergence 2.0 m
Fig. 2 The left picture’s marks, with each point’s vertical difference from the right picture drawn eight times over. Two eyes level with each other see every point at the same height.
Two cameras turned inwardsThe left picture's marks, with each point's **vertical** difference from the right picture drawn 8× exaggerated. Two eyes level with each other see every point at the same height, so a pair with vertical difference is a pair of pictures of no scene at all. Here it reaches 51.4 px.vertical difference up to 51.4 px, drawn 8× overbaseline 63 mmconvergence 1.2 m
Fig. 3 A nearer convergence distance. The cameras are turned in further and the vertical difference grows with the rotation.

The difference no eye can act on

Horizontal disparity is what a stereo pair is for. Vertical disparity — a difference in the y coordinate of the same world point between the two pictures — is not a signal at all. Two eyes side by side and level with each other see every point at the same height, always, for every scene; a pair of pictures with vertical differences in it is a pair of pictures of no arrangement.

The toed-in pair has it, and here it reaches about thirty-one pixels on a four-hundred-pixel frame.

The off-axis pair has none. Not a small amount: zero, to the last bit, at every point of every scene tried. The two image planes are the same plane, and moving a principal point within a plane cannot change any point’s height in it.

Two cameras with shifted principal pointsThe same baseline and the same convergence distance, with both axes parallel and the principal points moved ∓ b·f/2D instead. The vertical difference is 0e+0 px — zero, not small — because the two image planes are the same plane and a shift within it cannot move anything vertically.vertical difference 0e+0 pxbaseline 63 mmconvergence 2.0 m
Fig. 4 The same baseline and the same convergence distance with parallel axes and shifted principal points.
Two cameras with shifted principal pointsThe same baseline and the same convergence distance, with both axes parallel and the principal points moved ∓ b·f/2D instead. The vertical difference is 0e+0 px — zero, not small — because the two image planes are the same plane and a shift within it cannot move anything vertically.vertical difference 0e+0 pxbaseline 63 mmconvergence 1.2 m
Fig. 5 The same near convergence with shifted principal points. The shift is larger and the vertical difference is still exactly nothing.

Where the difference is, which is the awkward part

The vertical difference is zero on the optical axis, and it is zero along the horizontal centre line, and it is zero wherever the check is easiest to make.

Zero in the middle and nowhere elseA point at a fixed height and depth, walked across the scene. The toed-in pair puts it at the same height in both pictures on the axis — 0e+0 px — and at up to 31.4 px apart at the edges. The off-axis pair is at 0e+0 px everywhere. A defect that vanishes wherever it is checked is the hardest kind to find.-20020-2-1012the point, moved across the scene (m)vertical difference between the two pictures (px)off-axis frusta — exactly zerotoed inzero on the axis
Fig. 6 A point at a fixed height and depth, walked across the scene, with the vertical difference between the two pictures at each position. The flat line is the off-axis pair.

It grows toward the corners, as the product of how far off-axis the point is horizontally and how far it is vertically. So a person checking a rig by putting a target in the middle of the frame and confirming that both cameras agree about its height has confirmed the one place where the defect is absent.

A defect that vanishes wherever it is looked for is the hardest kind to find, and this collection has met the shape before. Sweep the variable the quantity depends on, not the one that is easy: the eight-point circle rule’s error is flat across a frame to four figures and varies with depth, and it was first swept across the frame.

Two cameras turned inwardsThe left picture's marks, with each point's **vertical** difference from the right picture drawn 8× exaggerated. Two eyes level with each other see every point at the same height, so a pair with vertical difference is a pair of pictures of no scene at all. Here it reaches 58.6 px.vertical difference up to 58.6 px, drawn 8× overbaseline 120 mmconvergence 2.0 m
Fig. 7 A wider baseline than a pair of eyes. Everything about the defect scales with the rotation, and the rotation scales with the baseline over the convergence distance.

What a real pair of eyes does, and why it is not a counter-example

The obvious objection is that eyes toe in. A person looking at something near converges on it, both eyes rotate inward, and nobody complains of vertical disparity.

Two things separate that case from a camera rig, and both are worth stating because the objection is a good one.

First, a pair of eyes rotating inward about vertical axes produces a cyclovergence as well — the eyes counter-roll slightly — and the geometry of the resulting pair is not the geometry of two pinhole cameras rotated about their own vertical axes. What the eyes do is not the toe-in construction above.

Second, and more to the point, the eyes are not making a pair of pictures to be shown to a second pair of eyes. A stereoscopic display asks a viewer’s eyes to fuse two images that are already fixed, and the viewer’s own convergence is a separate degree of freedom on top of them. The pair has to be geometrically consistent on its own, because there is no rotation the viewer can perform that removes a vertical difference varying across a frame.

This collection has nothing to say about what happens perceptually when it is asked to. What it can say is that the toed-in pair is not the pair a pair of eyes would have produced, and that the construction which is is the one with no rotation in it.

What keystone correction actually costsA projector turned 15° from square throws its rectangular panel as a quadrilateral. Correction cannot add light outside it, so it shrinks the picture until it fits — and 18.7% of the projector's pixels are thrown away. The fraction is measured on the panel rather than on the wall, because turning the projector makes the wall picture larger while making the panel usage smaller.15° of yaw, 6° of pitch, 1.50 throw ratio81.3% of the panel reaches the corrected rectangleouter: the thrown quadrilateral · inner: what correction can keep18.7% of the panel discarded
Fig. 8 The rotation’s visible form: a projector turned from square throws a trapezium, and each of two toed-in cameras throws one of these onto the world.

Why the rotation does it

Two pinhole cameras at the same place, differing by a rotation, produce pictures related by a homography — the same relation a tile is an off-centre frustum is about, where aiming a camera at each tile of a picture produces a different picture from cropping.

A toed-in pair is not that, because the two cameras are also a baseline apart. But the rotation part of the difference between them is exactly that homography, and a homography that is not a translation does not preserve horizontal lines. The two pictures are related by a rotation composed with a displacement, and the rotation tilts each picture’s own idea of horizontal by an amount that varies across the frame.

The off-axis pair has no rotation in it at all. The two cameras differ by a pure translation along their common x axis, and a pure translation along the image plane’s own x axis moves points horizontally and nothing else. That is the whole reason its vertical disparity is exactly zero rather than merely small: it is a statement about which coordinate a translation can touch.

Two cameras turned inwardsThe left picture's marks, with each point's **vertical** difference from the right picture drawn 8× exaggerated. Two eyes level with each other see every point at the same height, so a pair with vertical difference is a pair of pictures of no scene at all. Here it reaches 17.6 px.vertical difference up to 17.6 px, drawn 8× overbaseline 63 mmconvergence 3.5 m
Fig. 9 A distant convergence. The rotation is small and the vertical difference with it, which is why long shots hide the defect and close ones do not.

The shift, derived rather than quoted

The number ∓ b·f/2D is worth deriving once, because it is the whole of the construction and because getting its sign backwards doubles the quantity it was supposed to null.

Put a point on the axis at the convergence distance D. In the left camera — eye at −b/2, looking straight ahead — that point sits at x = +b/2 in camera coordinates, so it images at pₓ + f·(b/2)/D. In the right camera it images at pₓ − f·(b/2)/D. Their disparity is f·b/D, and it has to be cancelled by the difference of the two principal points.

So the left camera’s principal point moves left, by f·b/2D, and the right’s moves right by the same. Written the other way round, each shift adds to the disparity instead of cancelling it and the convergence point lands at twice the disparity it started with — which is a picture whose zero-disparity plane is not where it was asked to be, and which looks like a mistake about the convergence distance rather than about a sign.

The check that catches it is the one the construction is for: a point on the axis at D must have zero disparity, horizontal and vertical, in both constructions. It is one assertion, it costs nothing, and it is the only thing that distinguishes a correct shift from a shift of the same size in the wrong direction.

A point on the left is a line on the rightFive points in the left picture. Each one fixes a line in the right picture — the image of the ray it came along — and its match lies on that line to 1.4e-13 px. The five lines meet at the epipole, 3.5e-11 px from concurrent. Knowing where a point is in one picture does not say where it is in the other; it removes one of the two degrees of freedom.1234512345a point is chosen hereand must be on this linematch to its own epipolar line: 1.4e-13 px5 of 44 correspondences drawn
Fig. 10 The epipolar constraint: a point in one picture of a pair constrains its match in the other to a line, and the lines are decided by the two cameras alone.

The same fact, from the two-view field

There is a name for a pair of pictures with no vertical disparity anywhere, and this collection has already used it in a different context.

A point is a line over there is about the epipolar constraint: a point in one picture of a pair constrains its match in the other to a line, and the lines are determined by the two cameras alone. When the two cameras are related by a pure translation along their shared image-plane x axis, those lines are horizontal and at matching heights — the pair is rectified, and a match for a point in row j of the left picture is somewhere in row j of the right.

The off-axis construction produces a rectified pair by construction. The toed-in construction does not, and its epipolar lines are sloped and at varying heights.

That is the same statement as the vertical disparity, and it is worth having both because they are useful in different places. For a display, the question is what a viewer’s eyes are asked to do. For a measurement, the question is where a match can be. They are one geometric condition.

Zero in the middle and nowhere elseA point at a fixed height and depth, walked across the scene. The toed-in pair puts it at the same height in both pictures on the axis — 0e+0 px — and at up to 15.7 px apart at the edges. The off-axis pair is at 0e+0 px everywhere. A defect that vanishes wherever it is checked is the hardest kind to find.-10010-2-1012the point, moved across the scene (m)vertical difference between the two pictures (px)off-axis frusta — exactly zerotoed inzero on the axis
Fig. 11 The sweep at a further convergence distance. The rotation is smaller, the vertical difference is smaller, and it is still zero in the middle and nowhere else.

Keystone, which is the same rotation seen on a wall

The rotation that produces the vertical disparity has a familiar visible form, and naming it makes the effect easier to picture than any plot of it.

Turn a projector away from square to a wall and it throws a trapezium rather than a rectangle — the far side of the picture is larger than the near side, and horizontal lines that were parallel in the panel converge on the wall. A projector is a camera run backwards measures what correcting that costs, which is a sixth of the panel discarded at fifteen degrees.

A toed-in camera does the same thing to the scene it records. Each of the two cameras is a projector run the other way, turned by half the convergence angle, and each throws a keystone onto the world — so the left picture is the correct picture stretched one way and the right picture is it stretched the other. The vertical disparity is what remains after the two keystones are laid on top of each other: they agree along the axis and diverge toward the corners, in opposite directions.

Which is why the effect grows as the product of the two off-axis distances rather than as either one. A keystone stretches vertically by an amount that depends on horizontal position, so its effect on a point’s height is horizontal offset times vertical offset, and the two cameras’ keystones have opposite signs.

The other eye, found in the picture — 2.6 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 1.1e-9 px on the left and 3.7e-10 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 1.1e-9 px2.60 m between the eyes
Fig. 12 The image of one eye in the other’s picture, which is what decides where the epipolar lines run. For a pure sideways translation it is at infinity and the lines are horizontal.
Zero in the middle and nowhere elseA point at a fixed height and depth, walked across the scene. The toed-in pair puts it at the same height in both pictures on the axis — 0e+0 px — and at up to 44.9 px apart at the edges. The off-axis pair is at 0e+0 px everywhere. A defect that vanishes wherever it is checked is the hardest kind to find.-50-2502550-2-1012the point, moved across the scene (m)vertical difference between the two pictures (px)off-axis frusta — exactly zerotoed inzero on the axis
Fig. 13 The sweep at a near convergence, where the vertical difference at the edges is at its largest and still nothing in the middle.

What it costs to fix afterwards

A toed-in pair can be rectified after the fact — warped so that its epipolar lines become horizontal and aligned — and that is what a great deal of software does.

The warp is a homography per picture, so it is exact in the sense that the geometry is recoverable. What it is not is free: a homography applied to a sampled picture has to resample it, and the resampling costs resolution unevenly across the frame, worst where the warp is largest, which is at the corners where the problem was. It also crops, because the warped rectangles do not cover each other and only the common part survives.

So the choice is between a construction that never has the problem and a construction that has it and then pays to remove it — and the first costs nothing at all, because a shifted principal point is four numbers in a matrix.

Two cameras with shifted principal pointsThe same baseline and the same convergence distance, with both axes parallel and the principal points moved ∓ b·f/2D instead. The vertical difference is 0e+0 px — zero, not small — because the two image planes are the same plane and a shift within it cannot move anything vertically.vertical difference 0e+0 pxbaseline 120 mmconvergence 2.0 m
Fig. 14 The wider baseline done correctly. The horizontal disparity has grown as it should and the vertical difference has not appeared.

Measuring it rather than asserting it

The comparison is set up so that neither construction can win by accident, and the assertions say so.

Both pairs are built from the same factory, with the same baseline, the same convergence distance and the same field of view, and both are measured on the same set of scene points. The off-axis pair’s vertical disparity must be zero rather than small — a threshold of a fraction of a pixel would let a construction with a tiny rotation in it through. The toed-in pair’s must exceed a pixel, or there would be nothing to report. And both must put the convergence point at zero disparity in both coordinates, which is what says the two are the same pair set up two ways rather than two different pairs.

That last assertion is the one that found the sign error described above, and it found it immediately: with the shifts the wrong way round the off-axis pair’s convergence point came out at forty-seven pixels of horizontal disparity instead of zero, which is four times the shift and is unmistakable. A check that had only compared the two pairs’ vertical disparities would have passed, because the sign error does not introduce any.

A wide render read from a screen that subtends much lessA 27-inch monitor at 650 mm subtends 49.3°. A picture rendered at 100° is therefore being read from 2.60 times its own station distance, and the viewing field has already established what that does: depth is stretched by exactly that factor and nothing in the picture changes. The two routes to the number — from two angles, and from a focal length and a display width — agree to 1e-9.051050100150field of view the picture was rendered at — degreeshow many times the depicted depth is stretchedthe screen subtends 49.3°100° → depth ×2.6027-inch monitor at 650 mmsubtends 49.3°
Fig. 15 The other half of a stereoscopic display, which this collection has already priced: the distance the picture is correct from, against where anybody is actually sitting.

What this does not say

It does not say that a toed-in pair is uncomfortable to look at, or that a viewer will notice, or that vertical disparity of a given size is worse than some other defect of the same size. Those are questions about eyes, and this collection has said since its foundation that it computes the geometry of pictures and has no standing on how they are perceived.

What it says is narrower and is checkable: the two pictures of a toed-in pair are not two pictures of any scene, in the specific sense that no arrangement of a world and two level eyes produces the vertical differences they contain. That is a fact about the pair of images, and it is true whether or not anybody is looking.

The perceptual literature on the question exists and belongs to whoever writes about seeing. The geometry’s contribution is to say exactly how large the offending quantity is, exactly where in the frame it is largest, and exactly what construction makes it zero.

A frustum is four numbers and a windowThe camera's own section, looking down on it. The two slanted lines are the left and right sides: divide either by the near plane and it is a direction, -0.6285 and 0.3041, which between them carry the focal length 739.9 px and the principal point 465. The two crossbars are the near and far planes, and moving either of them moves no mark in the picture at all.the eyenear 0.90 mfar 4.20 mfocal 740 pxprincipal 465, 210
Fig. 16 What the off-axis construction is asking for: a frustum whose principal point is not in the middle of its own rectangle, which a matrix will supply and a lens will not.
Level, tilted, shiftedTilting the camera up to fit the building in makes the verticals converge by 2.81°. Shifting the lens up instead moves the principal point 95 px off centre, frames the same view, and leaves the verticals parallel to 0e+0° — because the picture plane never tilts.level — the top is cut off0.00° of spreadtilted 8°2.81° of spreadshifted 95 px0.00° of spreada shift moves every point by exactly the shift95.0 px, and no direction at all
Fig. 17 The physical instrument that can do it: a shift lens moves the principal point without tilting the picture plane. Two of them, shifted oppositely, are an off-axis pair.

Why the wrong construction is the default

It is worth asking why the toed-in arrangement is the one that gets built, since the correct one is cheaper in every respect once a matrix is available.

For a physical rig it is not cheaper, and for a long time it was not possible. A real lens’s principal point is where the lens puts it, so an off-axis frustum needs either a shift lens or a sensor that can slide behind a fixed lens — both of which exist, both of which are expensive, and neither of which is on an ordinary camera. Two heads that rotate are on every tripod.

So the physical constraint produced the habit, and the habit outlived the constraint. A synthetic camera has no lens and no sensor; its principal point is two numbers in a matrix, and moving them costs nothing at all. Every argument for toe-in is an argument about hardware, and none of it applies.

That is the same shape as the tiling essay’s own moral, arriving in a different field. The exact construction is available because a computed camera has a freedom a physical one does not, and the mistake is importing the physical camera’s constraint when nothing is imposing it.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 2.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.010203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 2.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 18 The depth budget, set by the width of a head rather than by anything about the scene. Choosing the toed-in construction spends one thing for nothing and leaves this untouched.

And the horizontal half is much the same either way

One thing the comparison should not be read as saying is that the toed-in pair gets the useful signal wrong.

Both constructions put the convergence point at zero disparity, exactly, and both produce horizontal disparities of the same order everywhere else — around thirty pixels at the extremes here for the toed-in pair against nineteen for the off-axis one, on the same scene. The difference between those two is a difference in the depth mapping, which is a design decision a stereographer makes deliberately, and it is not the subject of this essay.

The subject is the quantity that ought to be zero and is not. A pair can be built with the intended depth budget and no vertical disparity at the same time; the two are independent, and choosing the toed-in construction spends one of them for nothing.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineCamera tiltDisparityEpipolar geometryParallaxPrincipal pointProjective maprectified pairView frustumViewing position