Constructing a view

The plane is a choice

A projection has a centre and a surface, and they move independently. Keep the eye and turn the picture plane and every point of any scene lands where one 3×3 matrix says, to 2.5e-13 px. Move the eye instead and the matrix fitted to four points is exact at those four and out by 32.0 px everywhere else. The first is a homography of the picture; the second is parallax, and nothing about the picture can undo it.

Worth reading first: One, two and three point are one construction · A projection of a projection.

Every construction in this field so far has been written with the picture plane square across the view. The horizon at eye level, the distance point on the horizon, the measuring point on the horizon: all true, and all true because the plane is vertical and the axis horizontal.

That is a choice, and it is a different choice from where the eye is. A projection has two independent pieces — a centre and a surface — and this essay is about what happens when the second moves while the first stays put.

The picture plane tilted 14°Pointing the camera up tilts the picture plane with it, and three things happen at once: the verticals converge — 3.59° between the outer two — the horizon drops 213 px below the middle of the frame, and the vertical vanishing point arrives at 3425 px from the principal point. They are one fact: the product of those two offsets is f².correct from 20 cm, at 160 mm wideverticals converge 3.59° · horizon 213 px off centre
Fig. 1 The camera pointed up by 14°, which tilts the picture plane with it. Three things happen at once: the verticals converge by 3.59°, the horizon drops 213 px below the middle of the frame, and the vertical vanishing point arrives at 3425 px from the principal point.

One eye, two planes, one matrix

The theorem is short and it is the spine of everything below.

An image point is a ray through the eye, written in the picture’s own coordinates. Changing the picture plane changes the coordinates and does not change the ray. So the map from one picture to the other is

H=K2R2R1⊤K1−1,H = K_2 R_2 R_1^{\top} K_1^{-1},

a product of a rotation and two calibrations — a 3×33\times3 matrix acting on homogeneous image coordinates, with nothing about the scene in it.

That last clause is the content. The same matrix moves every point of every scene, so the two pictures are related by one map whatever is in front of the camera: a flat wall, a room, a forest.

Checked on a solid scene rather than on a plane: eighty points spread over three depths, projected through a level camera and through one pitched 18° from the same eye, and the map takes every one of them from the first picture to the second to 2.5e-13 px.

And a moved eye is not

The counter-case is what makes the theorem a statement rather than a definition, and it is the case people assume is covered.

Move the eye 0.9 m sideways and take the same scene again. Fit a homography to four of the scene’s own points — which always exists, since four correspondences determine one — and apply it to the rest.

The four fitted points come back exactly. Everything else is out by up to 32.0 px.

That gap is parallax: the amount by which a nearer thing shifts more than a farther one when the eye moves. It is not a small residual and it is not reducible by a better fit, because no map of the picture can move two points at the same image position by different amounts, and two points at the same image position at different depths are exactly what a moved eye separates.

So the pair of statements is:

Same eye, different plane: a homography, for any scene.

Different eye: a homography for a plane, and nothing for a solid scene.

The second half is the reason rectifying a photographed façade works and the reason a photograph cannot be re-staged from a different viewpoint by any amount of processing.

The degeneracy is in the reading, not in the algebraShape error against baseline, twice. Given exact correspondences the recovery is good to 4.7e-9 at a 2 mm baseline and improves from there — the algebra does not notice a short baseline at all. Read the same points to 0.25 px and the same geometry comes back 3e+9× further from the truth, and it does not refuse: it returns a confident wrong answer. What fails when a photographer swings a camera instead of stepping sideways is the ratio of baseline to reading error, and nothing in the computation reports it.-10-50-2-10baseline (m, log scale)worst deviation in the recovered shape (log scale)exact pointsread to 0.25 pxat 2 mm: 4.7e-9 exact, 1.5e+1 read to 0.25 px3e+9× apart
Fig. 2 The same distinction from the twoviews field, measured as an error rather than as a theorem. A turn of the head supplies no baseline, so a reconstruction from it returns a confident wrong answer — a rotation about the eye changes the plane and not the centre, which is exactly the operation this essay is about.

What the tilt does to the furniture

Three things move when the plane leans, and they are three readings of one number.

The horizon leaves the middle of the frame. It drops by ftan⁡θf\tan\theta — 213 px at 14°, 416 px at 26°. It is still the image of the ground’s points at infinity; it is no longer at the principal point, because the principal point is no longer at eye level in the picture.

The vertical vanishing point arrives from infinity. At f/tan⁡θf/\tan\theta from the principal point — 3425 px at 14°, 1751 px at 26°. With the plane vertical it is infinitely far away, which is what “verticals stay parallel” means.

And the verticals converge. 3.59° at 14°, 6.50° at 26°, and exactly 0e+0° with the plane vertical — whatever the field of view, which is the part worth checking because it is the claim a wide-angle lens is usually blamed for.

The picture plane verticalWith the picture plane vertical the world's verticals stay parallel in the picture — 0e+0° between the outer two — and the horizon sits exactly at the principal point. Neither is a property of the lens or the field of view; both are properties of the plane's orientation.horizoncorrect from 20 cm, at 160 mm wideverticals parallel to 0e+0°
Fig. 3 The plane vertical. The verticals stay parallel to 0e+0° and the horizon sits exactly at the principal point — neither a property of the lens, both properties of the plane’s orientation.
The picture plane tilted 26°Pointing the camera up tilts the picture plane with it, and three things happen at once: the verticals converge — 6.50° between the outer two — the horizon drops 416 px below the middle of the frame, and the vertical vanishing point arrives at 1751 px from the principal point. They are one fact: the product of those two offsets is f².correct from 20 cm, at 160 mm wideverticals converge 6.50° · horizon 416 px off centre
Fig. 4 And at 26°: 6.50° of convergence, the horizon 416 px below centre, the vertical vanishing point in at 1751 px. Every one of those is a consequence of the plane’s orientation and of nothing else.

The identity underneath the three

The two offsets are not independent, and the relation between them is the pleasant surprise of this essay.

(ftan⁡θ) × (f/tan⁡θ)  =  f2.(f\tan\theta)\,\times\,(f/\tan\theta) \;=\; f^2.

The horizon’s offset from the principal point, times the vertical vanishing point’s offset, is the focal length squared — at every tilt, to 1e-15 of itself across the sweep.

That is the foundations field’s pole and polar arriving in a field with no conics in it. The horizon is the polar of the vertical direction’s vanishing point with respect to the absolute conic, and the product being f2f^2 is what a polar relation looks like when both points are on the principal axis.

It is also useful. Given a photograph with converging verticals, the vanishing point and the horizon are both measurable, and their product gives the focal length — a one-line calibration from two things a straightedge finds, needing no known object in the scene.

The horizon leaves, and the vanishing point arrivesTilting the picture plane moves the horizon down the frame by f·tan of the tilt and brings the vertical vanishing point in from infinity to f/tan of it. The two are the same number read twice: their product is f² — 729154 px² here — at every tilt in the sweep, to 1e-15 of itself. Which is the pole and polar of the foundations field arriving where nothing is about conics: the horizon is the polar of the vertical direction.01e+32e+3102030tilt of the picture plane (°)pixels from the principal pointthe horizon's offsetthe vertical vanishing pointtheir product is f² = 729154 px²to 1e-15 at every tilt
Fig. 5 The identity checked across a sweep of tilts: the horizon’s offset and the vertical vanishing point’s offset multiply to f² = 729154 px² at every one of them, to 1e-15 of itself.

The convergence is the third measurement, and it is the one on the page

The identity is useful and it has an awkward practical edge: the vertical vanishing point is at fcot⁡θf\cot\theta, which at 14° is 3,425 pixels — some way outside a 690-pixel frame. A reader with a photograph can measure the horizon’s offset with a straightedge and cannot measure the vanishing point at all without extending lines across the desk.

The convergence angle is the quantity that is on the page, and it can stand in. Two verticals at image positions ±a\pm a both run to a vanishing point fcot⁡θf\cot\theta away, so the angle between them is

β  ≈  2atan⁡θf.\beta \;\approx\; \frac{2a\tan\theta}{f}.

Two consequences, and the second is a calibration.

The convergence is not a property of the tilt alone. It carries aa, the verticals’ own separation on the page, so two edges near the middle of the frame converge less than two at its sides at the same tilt — and a photographer judging “how much is this leaning” is reading a product of the tilt and the framing. The measured 3.59° and 6.50° at 14° and 26° correspond to verticals about 107 pixels either side of centre; the same tilts on verticals at the frame’s edge would give twice as much.

And it completes the identity without leaving the paper. The horizon’s offset is h=ftan⁡θh = f\tan\theta and the convergence gives tan⁡θ=βf/2a\tan\theta = \beta f/2a, so eliminating the tilt:

f2  =  2 a hβ.f^{2} \;=\; \frac{2\,a\,h}{\beta}.

Every quantity on the right is a measurement inside the frame — the half-separation of two drawn verticals, the horizon’s drop below the principal point, and the angle between the verticals. Against this sweep: 2×107.3×213/0.0627=729,5002 \times 107.3 \times 213 / 0.0627 = 729{,}500 against f2=729,154f^{2} = 729{,}154, which is four parts in ten thousand and is the arctangent this expression linearises.

So a photograph of a building with converging verticals states its own focal length three different ways: from the two vanishing points, whose product is f2f^{2} but one of which is off the page; from the horizon and the convergence, which are both on it; and from three perpendicular directions, which is the general recovery. The middle one is the reader’s route, and it is the one no treatment writes down, because the convergence angle is usually treated as a defect to be removed rather than as a measurement.

It also says which photographs refuse. β\beta appears in the denominator, so a nearly-level camera gives a nearly-zero convergence and the recovery becomes ill-conditioned — the same pole the ramp’s own vanishing line meets at zero gradient, and the same answer: the construction is exact everywhere and useful over a range, and the range excludes exactly the pictures that did not need correcting.

There is a matching refusal at the other end, and it is less obvious. As the tilt grows the horizon leaves the frame — ftan⁡θf\tan\theta passes half the frame width at θ\theta equal to the half-field of view — so a steeply tilted photograph has a large convergence and no measurable hh. Between the two refusals lies a working band, and its width is the camera’s own half-field: the recovery works when the tilt is smaller than the half-angle of the lens and larger than about a degree. A wide lens therefore has a wide band and a long one a narrow band, which is the reverse of how the two are usually ranked for architectural work and is a fair summary of what this field keeps finding — the lens that makes the picture harder to look at makes it easier to measure.

Why the scene does not appear in the matrix

The scene-independence is the property everything else depends on, and it is worth one section on why it holds, because the neighbouring case where it fails is only one word away.

A point’s image records the direction from the eye to it and nothing else. Two points on the same ray at different distances have the same image, exactly — which is the whole of what one view supplies — so a picture is a picture of the ray bundle rather than of the scene.

Changing the picture plane re-coordinatises that bundle. Nothing about which rays exist changes, so the map between the two coordinate systems is a map of the bundle, and the bundle is the same for every scene.

Move the eye and the bundle itself changes: a different set of rays, with different points on them. There is no map of the old picture that produces the new one, because the new one contains information — the depth-dependent shift — that the old one does not have.

So the theorem’s scene-independence and the counter-case’s parallax are the same fact told twice. A picture knows the rays and not the points, and an operation that only re-describes the rays is available while an operation that needs the points is not.

Four operations inside one group

Once the plane is a free choice, several separately-named techniques turn out to be the same operation.

Tilting the camera rotates the plane about a horizontal axis. Keystone, three-point perspective, converging verticals.

Shifting the lens moves the principal point without rotating the plane. Verticals stay parallel, and the whole image translates, because cxc_x enters the projection additively.

Cropping takes a rectangle out of the picture the plane already carries. Which is the same operation as a shift, exactly — and the figure below measures it: the shifted frame’s points sit at one constant offset from a wider picture’s, spread 1e-13 px over the whole scene.

And rotating the camera about the eye — panning to stitch a panorama — is a plane change with the axis vertical, which is why a rotation homography stitches and a step sideways does not.

All four are elements of the group of homographies of the picture, and the group is the answer to “what can be done to a photograph without moving the photographer”.

Shift or tilt: two ways to include the topThe wide picture is what the eye sees through a vertical picture plane. Sliding the frame up it — a rising front, a shift lens — gives a picture whose points sit at one constant offset from the wide one, spread 1e-13 px over the whole scene: it is a crop, and its verticals stay parallel to 0e+0°. Turning the plane instead gives a picture that is not a crop of it at all — the same points differ by offsets spreading 87.3 px — and its verticals converge 1.34°.the shifted frame — a crop to 1e-13 pxthe wide picture, from the same eyea tilted frame is not a crop of it — 87.3 px of spreadtilt converges the verticals 1.34°
Fig. 6 Shift against tilt, on one wide picture from one eye. The shifted frame is a crop of it — every point at one constant offset, spread 1e-13 px. The tilted picture is not: the same points differ by offsets spreading 87.3 px, and its verticals converge 1.34°.

The frame is not the plane

A distinction the group makes obvious and ordinary language hides: the picture plane is unbounded, and the frame is a rectangle drawn on it.

Choosing the rectangle is a crop, or a shift, or a change of film format. Nothing about the projection changes; a different part of the same picture is kept.

Choosing the plane is a tilt or a rotation. The projection changes — different coordinates on the same bundle of rays — and every point in the frame moves by an amount depending on where it is.

Both are homographies of the picture, so both are in the group; they are different subgroups of it, and the difference is visible in the measurements. A shift’s offsets are constant over the scene — 1e-13 px of spread, which is the definition of a translation. A tilt’s offsets spread by 25.8 px over the same scene, which is what a projectivity does.

The practical form of the distinction is the one every architectural photographer knows and states as a rule of thumb: shifting costs field and tilting costs shape. Shifting keeps the geometry and throws away the parts of the image circle the frame no longer covers. Tilting keeps the frame and converges the verticals. The rule of thumb is exactly the subgroup statement, and the numbers above are its price list.

It also explains an operation that looks like a third option and is not. Cropping the top out of a tilted picture does not straighten anything — the plane is still tilted, so the verticals still converge, and the crop has merely chosen a smaller rectangle on a leaning plane. To straighten, the plane has to change, which is the next essay.

What a plane change cannot do

The boundary is the reason the theorem is worth stating carefully, and it has two halves.

It cannot move the eye. Every picture in the group is a projection of the scene from the same point. The reader’s viewing position for a corrected picture is computed from its own focal length and displayed width and will generally differ from the original’s — but the scene’s relationship to the observer is fixed, and no processing supplies the view from somewhere else.

And it cannot un-mix the depths. Because the map is scene-independent, it carries no information about depth; because it is a bijection of the picture, it cannot add any. Everything a photograph left out stays left out.

Which is exactly what makes the group useful. A transformation that changes nothing about the geometry’s content is a transformation that can be applied without asking what is in the picture, and that is why keystone correction is a slider in every photo editor and re-staging from a different position is not a feature anywhere.

The same cube turned 24° — a three-point constructionNothing about the construction changed. The number of vanishing points inside any finite distance is 3, and 1 of them fall on the canvas.horizon3 vanishing points at a finite distance2583 px · 6269 px · 549 px
Fig. 7 The named cases, as values of two angles. One, two and three point perspective are one construction with the box turned and the plane tilted — three finite vanishing points here, one of them on the canvas. The names describe where the plane is, not what kind of projection it is.
Straightened — and taken from exactly where it wasThe tilted picture is drawn thin and the corrected one over it. The correction is built from the picture alone: where the imaged verticals meet, and the focal length. What comes out agrees with a level camera at the same eye — one the correction was never shown — to 3e-13 px, and its verticals are parallel to 0e+0°. What has not changed is the eye: the cross-ratio of four points along a ground line reads 1.3333 before and after, so every measurement the original supported the corrected one supports, from the same place and no other.corrected from the picture — 0e+0° between the verticalscorrect from 20 cm, at 160 mm wideagrees with a level camera at the same eye to 3e-13 px
Fig. 8 The group’s most-used member, applied. The picture re-cast onto an upright plane through the same eye, agreeing with a level camera to 3e-13 px — a change of surface with the centre held.

What a reader can check in a photograph

Three things follow that need no computation, and they are the reason this is worth knowing rather than merely true.

Converging verticals mean a tilted plane, not a wide lens. The convergence is 0e+0° with the plane vertical at any field of view, so a wide-angle picture with parallel verticals is perfectly possible and a long-lens picture with converging ones is too. Blaming the lens is blaming the wrong parameter.

The horizon’s position measures the tilt. It sits at the principal point when the plane is vertical and drops by ftan⁡θf\tan\theta as the camera points up, so a horizon low in the frame is a camera pointing up and nothing else. Combined with the vertical vanishing point it gives the focal length.

And a picture with no convergence carries no evidence about the tilt. That is the refusal the machinery makes rather than a caution: with the plane vertical the imaged verticals are parallel and their meeting point does not exist, so there is nothing to compute from. A method that returned a tilt anyway would be reporting the intersection of two nearly-parallel lines, which is the failure mode the projector’s recovery refuses by name.

The other kind of map, for contrast

There is a second family of maps of a picture on this site, and the census that separated them is worth invoking because the difference is exactly the one this essay turns on.

A central collineation — a homology or an elation — has a line of fixed points, and the census found three of them: a shadow, a floor anamorph, a mirror. Each changes where a plane is seen from: the plane stays put and the projection centre moves.

A general projectivity has no such line. A rectification is one: both ends of the map are pictures of a plane rather than the plane and a picture of it.

A plane change is in the second family. Both ends are pictures, taken from the same point on two different surfaces, and there is no line of the world sitting in both to be fixed. Which is a compact way of saying what has and has not changed: the eye has not moved, and neither picture is the world.

Where the plane’s freedom is used

Three places on this site turn out to be this essay’s theorem applied, and none of them announces it.

Keystone correction of a projector. A projector is a camera run backwards, and correcting its keystone re-casts its picture onto the plane the wall actually occupies — at a cost of 18.7% of the panel’s pixels at 15° off square, because correction cannot add light outside the thrown quadrilateral.

Stitching a panorama. Frames from one point differ by plane changes, so they compose exactly; frames from a moving camera do not, and the error is the parallax measured above.

And the standard construction’s own conditions. “The horizon is at eye level” is a statement about a vertical picture plane, and the figure that measures it says so in its title. This essay is the rung where the condition is removed rather than assumed.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

centre of projectionHomographyHorizonParallaxPicture planepole and polarPrincipal pointProjective mapthree-point perspectivevertical vanishing point