Turning the cameras inwards
Worth reading first: Two pictures on one screen · Four numbers and a window.
Two pictures on one screen works out what a stereoscopic display does with two images: the depicted depth is b·D/(b − d), and the disparity that reaches infinity is exactly the separation of the reader’s eyes, at any screen distance whatever.
That essay assumes the pair. This one is about making it, and about the arrangement that looks right and is not.
Two ways to point two cameras
A stereo pair needs two cameras a baseline apart, and something has to decide where the zero-disparity plane is — the depth at which a point lands in the same place in both pictures and therefore appears to sit on the screen.
Toe-in rotates both cameras to look at a chosen point on the axis. It is the obvious construction, it is what a pair of eyes does, and it is what every rig with two tripod heads makes easy.
Off-axis leaves both optical axes parallel and moves the principal points instead, by ∓ b·f/2D — half the baseline times the focal length, over the convergence distance. That is a frustum with its four sides shifted, which four numbers and a window says is the same camera looking through a window that is not centred on it.
Both put the convergence point at zero disparity, exactly, and that is where their agreement ends.
The difference no eye can act on
Horizontal disparity is what a stereo pair is for. Vertical disparity — a difference in the y coordinate of the same world point between the two pictures — is not a signal at all. Two eyes side by side and level with each other see every point at the same height, always, for every scene; a pair of pictures with vertical differences in it is a pair of pictures of no arrangement.
The toed-in pair has it, and here it reaches about thirty-one pixels on a four-hundred-pixel frame.
The off-axis pair has none. Not a small amount: zero, to the last bit, at every point of every scene tried. The two image planes are the same plane, and moving a principal point within a plane cannot change any point’s height in it.
Where the difference is, which is the awkward part
The vertical difference is zero on the optical axis, and it is zero along the horizontal centre line, and it is zero wherever the check is easiest to make.
It grows toward the corners, as the product of how far off-axis the point is horizontally and how far it is vertically. So a person checking a rig by putting a target in the middle of the frame and confirming that both cameras agree about its height has confirmed the one place where the defect is absent.
A defect that vanishes wherever it is looked for is the hardest kind to find, and this collection has met the shape before. Sweep the variable the quantity depends on, not the one that is easy: the eight-point circle rule’s error is flat across a frame to four figures and varies with depth, and it was first swept across the frame.
What a real pair of eyes does, and why it is not a counter-example
The obvious objection is that eyes toe in. A person looking at something near converges on it, both eyes rotate inward, and nobody complains of vertical disparity.
Two things separate that case from a camera rig, and both are worth stating because the objection is a good one.
First, a pair of eyes rotating inward about vertical axes produces a cyclovergence as well — the eyes counter-roll slightly — and the geometry of the resulting pair is not the geometry of two pinhole cameras rotated about their own vertical axes. What the eyes do is not the toe-in construction above.
Second, and more to the point, the eyes are not making a pair of pictures to be shown to a second pair of eyes. A stereoscopic display asks a viewer’s eyes to fuse two images that are already fixed, and the viewer’s own convergence is a separate degree of freedom on top of them. The pair has to be geometrically consistent on its own, because there is no rotation the viewer can perform that removes a vertical difference varying across a frame.
This collection has nothing to say about what happens perceptually when it is asked to. What it can say is that the toed-in pair is not the pair a pair of eyes would have produced, and that the construction which is is the one with no rotation in it.
Why the rotation does it
Two pinhole cameras at the same place, differing by a rotation, produce pictures related by a homography — the same relation a tile is an off-centre frustum is about, where aiming a camera at each tile of a picture produces a different picture from cropping.
A toed-in pair is not that, because the two cameras are also a baseline apart. But the rotation part of the difference between them is exactly that homography, and a homography that is not a translation does not preserve horizontal lines. The two pictures are related by a rotation composed with a displacement, and the rotation tilts each picture’s own idea of horizontal by an amount that varies across the frame.
The off-axis pair has no rotation in it at all. The two cameras differ by a pure translation along their common x axis, and a pure translation along the image plane’s own x axis moves points horizontally and nothing else. That is the whole reason its vertical disparity is exactly zero rather than merely small: it is a statement about which coordinate a translation can touch.
The shift, derived rather than quoted
The number ∓ b·f/2D is worth deriving once, because it is the whole of the construction and because getting its sign backwards doubles the quantity it was supposed to null.
Put a point on the axis at the convergence distance D. In the left camera — eye at −b/2, looking straight ahead — that point sits at x = +b/2 in camera coordinates, so it images at pₓ + f·(b/2)/D. In the right camera it images at pₓ − f·(b/2)/D. Their disparity is f·b/D, and it has to be cancelled by the difference of the two principal points.
So the left camera’s principal point moves left, by f·b/2D, and the right’s moves right by the same. Written the other way round, each shift adds to the disparity instead of cancelling it and the convergence point lands at twice the disparity it started with — which is a picture whose zero-disparity plane is not where it was asked to be, and which looks like a mistake about the convergence distance rather than about a sign.
The check that catches it is the one the construction is for: a point on the axis at D must have zero disparity, horizontal and vertical, in both constructions. It is one assertion, it costs nothing, and it is the only thing that distinguishes a correct shift from a shift of the same size in the wrong direction.
The same fact, from the two-view field
There is a name for a pair of pictures with no vertical disparity anywhere, and this collection has already used it in a different context.
A point is a line over there is about the epipolar constraint: a point in one picture of a pair constrains its match in the other to a line, and the lines are determined by the two cameras alone. When the two cameras are related by a pure translation along their shared image-plane x axis, those lines are horizontal and at matching heights — the pair is rectified, and a match for a point in row j of the left picture is somewhere in row j of the right.
The off-axis construction produces a rectified pair by construction. The toed-in construction does not, and its epipolar lines are sloped and at varying heights.
That is the same statement as the vertical disparity, and it is worth having both because they are useful in different places. For a display, the question is what a viewer’s eyes are asked to do. For a measurement, the question is where a match can be. They are one geometric condition.
Keystone, which is the same rotation seen on a wall
The rotation that produces the vertical disparity has a familiar visible form, and naming it makes the effect easier to picture than any plot of it.
Turn a projector away from square to a wall and it throws a trapezium rather than a rectangle — the far side of the picture is larger than the near side, and horizontal lines that were parallel in the panel converge on the wall. A projector is a camera run backwards measures what correcting that costs, which is a sixth of the panel discarded at fifteen degrees.
A toed-in camera does the same thing to the scene it records. Each of the two cameras is a projector run the other way, turned by half the convergence angle, and each throws a keystone onto the world — so the left picture is the correct picture stretched one way and the right picture is it stretched the other. The vertical disparity is what remains after the two keystones are laid on top of each other: they agree along the axis and diverge toward the corners, in opposite directions.
Which is why the effect grows as the product of the two off-axis distances rather than as either one. A keystone stretches vertically by an amount that depends on horizontal position, so its effect on a point’s height is horizontal offset times vertical offset, and the two cameras’ keystones have opposite signs.
What it costs to fix afterwards
A toed-in pair can be rectified after the fact — warped so that its epipolar lines become horizontal and aligned — and that is what a great deal of software does.
The warp is a homography per picture, so it is exact in the sense that the geometry is recoverable. What it is not is free: a homography applied to a sampled picture has to resample it, and the resampling costs resolution unevenly across the frame, worst where the warp is largest, which is at the corners where the problem was. It also crops, because the warped rectangles do not cover each other and only the common part survives.
So the choice is between a construction that never has the problem and a construction that has it and then pays to remove it — and the first costs nothing at all, because a shifted principal point is four numbers in a matrix.
Measuring it rather than asserting it
The comparison is set up so that neither construction can win by accident, and the assertions say so.
Both pairs are built from the same factory, with the same baseline, the same convergence distance and the same field of view, and both are measured on the same set of scene points. The off-axis pair’s vertical disparity must be zero rather than small — a threshold of a fraction of a pixel would let a construction with a tiny rotation in it through. The toed-in pair’s must exceed a pixel, or there would be nothing to report. And both must put the convergence point at zero disparity in both coordinates, which is what says the two are the same pair set up two ways rather than two different pairs.
That last assertion is the one that found the sign error described above, and it found it immediately: with the shifts the wrong way round the off-axis pair’s convergence point came out at forty-seven pixels of horizontal disparity instead of zero, which is four times the shift and is unmistakable. A check that had only compared the two pairs’ vertical disparities would have passed, because the sign error does not introduce any.
What this does not say
It does not say that a toed-in pair is uncomfortable to look at, or that a viewer will notice, or that vertical disparity of a given size is worse than some other defect of the same size. Those are questions about eyes, and this collection has said since its foundation that it computes the geometry of pictures and has no standing on how they are perceived.
What it says is narrower and is checkable: the two pictures of a toed-in pair are not two pictures of any scene, in the specific sense that no arrangement of a world and two level eyes produces the vertical differences they contain. That is a fact about the pair of images, and it is true whether or not anybody is looking.
The perceptual literature on the question exists and belongs to whoever writes about seeing. The geometry’s contribution is to say exactly how large the offending quantity is, exactly where in the frame it is largest, and exactly what construction makes it zero.
Why the wrong construction is the default
It is worth asking why the toed-in arrangement is the one that gets built, since the correct one is cheaper in every respect once a matrix is available.
For a physical rig it is not cheaper, and for a long time it was not possible. A real lens’s principal point is where the lens puts it, so an off-axis frustum needs either a shift lens or a sensor that can slide behind a fixed lens — both of which exist, both of which are expensive, and neither of which is on an ordinary camera. Two heads that rotate are on every tripod.
So the physical constraint produced the habit, and the habit outlived the constraint. A synthetic camera has no lens and no sensor; its principal point is two numbers in a matrix, and moving them costs nothing at all. Every argument for toe-in is an argument about hardware, and none of it applies.
That is the same shape as the tiling essay’s own moral, arriving in a different field. The exact construction is available because a computed camera has a freedom a physical one does not, and the mistake is importing the physical camera’s constraint when nothing is imposing it.
And the horizontal half is much the same either way
One thing the comparison should not be read as saying is that the toed-in pair gets the useful signal wrong.
Both constructions put the convergence point at zero disparity, exactly, and both produce horizontal disparities of the same order everywhere else — around thirty pixels at the extremes here for the toed-in pair against nineteen for the off-axis one, on the same scene. The difference between those two is a difference in the depth mapping, which is a design decision a stereographer makes deliberately, and it is not the subject of this essay.
The subject is the quantity that ought to be zero and is not. A pair can be built with the intended depth budget and no vertical disparity at the same time; the two are independent, and choosing the toed-in construction spends one of them for nothing.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Depth is a reciprocal — both name baseline, disparity, rectified pair
- The plane is a choice — both name parallax, principal point, projective map
- The range a pair cannot see past — both name baseline, disparity, rectified pair
- A set cut for one eye — both name disparity, parallax
- A turn of the head is not a step sideways — both name baseline, parallax
- An inverse perspective is a leaning plane — both name camera tilt, projective map
Named objects
A flat tag is an object no other essay names yet.
BaselineCamera tiltDisparityEpipolar geometryParallaxPrincipal pointProjective maprectified pairView frustumViewing position