Constructing a view

The plane is a choice

A projection has a centre and a surface, and they move independently. Keep the eye and turn the picture plane and every point of any scene lands where one 3×3 matrix says, to 2.5e-13 px. Move the eye instead and the matrix fitted to four points is exact at those four and out by 32.0 px everywhere else. The first is a homography of the picture; the second is parallax, and nothing about the picture can undo it.

Worth reading first: One, two and three point are one construction · A projection of a projection.

Every construction in this field so far has been written with the picture plane square across the view. The horizon at eye level, the distance point on the horizon, the measuring point on the horizon: all true, and all true because the plane is vertical and the axis horizontal.

That is a choice, and it is a different choice from where the eye is. A projection has two independent pieces — a centre and a surface — and this essay is about what happens when the second moves while the first stays put.

The picture plane tilted 14°Pointing the camera up tilts the picture plane with it, and three things happen at once: the verticals converge — 3.59° between the outer two — the horizon drops 213 px below the middle of the frame, and the vertical vanishing point arrives at 3425 px from the principal point. They are one fact: the product of those two offsets is f².correct from 20 cm, at 160 mm wideverticals converge 3.59° · horizon 213 px off centre
Fig. 1 The camera pointed up by 14°, which tilts the picture plane with it. Three things happen at once: the verticals converge by 3.59°, the horizon drops 213 px below the middle of the frame, and the vertical vanishing point arrives at 3425 px from the principal point.

One eye, two planes, one matrix

The theorem is short and it is the spine of everything below.

An image point is a ray through the eye, written in the picture’s own coordinates. Changing the picture plane changes the coordinates and does not change the ray. So the map from one picture to the other is

H=K2R2R1K11,H = K_2 R_2 R_1^{\top} K_1^{-1},

a product of a rotation and two calibrations — a 3×33\times3 matrix acting on homogeneous image coordinates, with nothing about the scene in it.

That last clause is the content. The same matrix moves every point of every scene, so the two pictures are related by one map whatever is in front of the camera: a flat wall, a room, a forest.

Checked on a solid scene rather than on a plane: eighty points spread over three depths, projected through a level camera and through one pitched 18° from the same eye, and the map takes every one of them from the first picture to the second to 2.5e-13 px.

And a moved eye is not

The counter-case is what makes the theorem a statement rather than a definition, and it is the case people assume is covered.

Move the eye 0.9 m sideways and take the same scene again. Fit a homography to four of the scene’s own points — which always exists, since four correspondences determine one — and apply it to the rest.

The four fitted points come back exactly. Everything else is out by up to 32.0 px.

That gap is parallax: the amount by which a nearer thing shifts more than a farther one when the eye moves. It is not a small residual and it is not reducible by a better fit, because no map of the picture can move two points at the same image position by different amounts, and two points at the same image position at different depths are exactly what a moved eye separates.

So the pair of statements is:

Same eye, different plane: a homography, for any scene.

Different eye: a homography for a plane, and nothing for a solid scene.

The second half is the reason rectifying a photographed façade works and the reason a photograph cannot be re-staged from a different viewpoint by any amount of processing.

The degeneracy is in the reading, not in the algebraShape error against baseline, twice. Given exact correspondences the recovery is good to 4.7e-9 at a 2 mm baseline and improves from there — the algebra does not notice a short baseline at all. Read the same points to 0.25 px and the same geometry comes back 3e+9× further from the truth, and it does not refuse: it returns a confident wrong answer. What fails when a photographer swings a camera instead of stepping sideways is the ratio of baseline to reading error, and nothing in the computation reports it.-10-50-2-10baseline (m, log scale)worst deviation in the recovered shape (log scale)exact pointsread to 0.25 pxat 2 mm: 4.7e-9 exact, 1.5e+1 read to 0.25 px3e+9× apart
Fig. 2 The same distinction from the twoviews field, measured as an error rather than as a theorem. A turn of the head supplies no baseline, so a reconstruction from it returns a confident wrong answer — a rotation about the eye changes the plane and not the centre, which is exactly the operation this essay is about.

What the tilt does to the furniture

Three things move when the plane leans, and they are three readings of one number.

The horizon leaves the middle of the frame. It drops by ftanθf\tan\theta — 213 px at 14°, 416 px at 26°. It is still the image of the ground’s points at infinity; it is no longer at the principal point, because the principal point is no longer at eye level in the picture.

The vertical vanishing point arrives from infinity. At f/tanθf/\tan\theta from the principal point — 3425 px at 14°, 1751 px at 26°. With the plane vertical it is infinitely far away, which is what “verticals stay parallel” means.

And the verticals converge. 3.59° at 14°, 6.50° at 26°, and exactly 0e+0° with the plane vertical — whatever the field of view, which is the part worth checking because it is the claim a wide-angle lens is usually blamed for.

The picture plane verticalWith the picture plane vertical the world's verticals stay parallel in the picture — 0e+0° between the outer two — and the horizon sits exactly at the principal point. Neither is a property of the lens or the field of view; both are properties of the plane's orientation.horizoncorrect from 20 cm, at 160 mm wideverticals parallel to 0e+0°
Fig. 3 The plane vertical. The verticals stay parallel to 0e+0° and the horizon sits exactly at the principal point — neither a property of the lens, both properties of the plane’s orientation.
The picture plane tilted 26°Pointing the camera up tilts the picture plane with it, and three things happen at once: the verticals converge — 6.50° between the outer two — the horizon drops 416 px below the middle of the frame, and the vertical vanishing point arrives at 1751 px from the principal point. They are one fact: the product of those two offsets is f².correct from 20 cm, at 160 mm wideverticals converge 6.50° · horizon 416 px off centre
Fig. 4 And at 26°: 6.50° of convergence, the horizon 416 px below centre, the vertical vanishing point in at 1751 px. Every one of those is a consequence of the plane’s orientation and of nothing else.

The identity underneath the three

The two offsets are not independent, and the relation between them is the pleasant surprise of this essay.

(ftanθ)×(f/tanθ)  =  f2.(f\tan\theta)\,\times\,(f/\tan\theta) \;=\; f^2.

The horizon’s offset from the principal point, times the vertical vanishing point’s offset, is the focal length squared — at every tilt, to 1e-15 of itself across the sweep.

That is the foundations field’s pole and polar arriving in a field with no conics in it. The horizon is the polar of the vertical direction’s vanishing point with respect to the absolute conic, and the product being f2f^2 is what a polar relation looks like when both points are on the principal axis.

It is also useful. Given a photograph with converging verticals, the vanishing point and the horizon are both measurable, and their product gives the focal length — a one-line calibration from two things a straightedge finds, needing no known object in the scene.

The horizon leaves, and the vanishing point arrivesTilting the picture plane moves the horizon down the frame by f·tan of the tilt and brings the vertical vanishing point in from infinity to f/tan of it. The two are the same number read twice: their product is f² — 729154 px² here — at every tilt in the sweep, to 1e-15 of itself. Which is the pole and polar of the foundations field arriving where nothing is about conics: the horizon is the polar of the vertical direction.01e+32e+3102030tilt of the picture plane (°)pixels from the principal pointthe horizon's offsetthe vertical vanishing pointtheir product is f² = 729154 px²to 1e-15 at every tilt
Fig. 5 The identity checked across a sweep of tilts: the horizon’s offset and the vertical vanishing point’s offset multiply to f² = 729154 px² at every one of them, to 1e-15 of itself.

Why the scene does not appear in the matrix

The scene-independence is the property everything else depends on, and it is worth one section on why it holds, because the neighbouring case where it fails is only one word away.

A point’s image records the direction from the eye to it and nothing else. Two points on the same ray at different distances have the same image, exactly — which is the whole of what one view supplies — so a picture is a picture of the ray bundle rather than of the scene.

Changing the picture plane re-coordinatises that bundle. Nothing about which rays exist changes, so the map between the two coordinate systems is a map of the bundle, and the bundle is the same for every scene.

Move the eye and the bundle itself changes: a different set of rays, with different points on them. There is no map of the old picture that produces the new one, because the new one contains information — the depth-dependent shift — that the old one does not have.

So the theorem’s scene-independence and the counter-case’s parallax are the same fact told twice. A picture knows the rays and not the points, and an operation that only re-describes the rays is available while an operation that needs the points is not.

Four operations inside one group

Once the plane is a free choice, several separately-named techniques turn out to be the same operation.

Tilting the camera rotates the plane about a horizontal axis. Keystone, three-point perspective, converging verticals.

Shifting the lens moves the principal point without rotating the plane. Verticals stay parallel, and the whole image translates, because cxc_x enters the projection additively.

Cropping takes a rectangle out of the picture the plane already carries. Which is the same operation as a shift, exactly — and the figure below measures it: the shifted frame’s points sit at one constant offset from a wider picture’s, spread 1e-13 px over the whole scene.

And rotating the camera about the eye — panning to stitch a panorama — is a plane change with the axis vertical, which is why a rotation homography stitches and a step sideways does not.

All four are elements of the group of homographies of the picture, and the group is the answer to “what can be done to a photograph without moving the photographer”.

Shift or tilt: two ways to include the topThe wide picture is what the eye sees through a vertical picture plane. Sliding the frame up it — a rising front, a shift lens — gives a picture whose points sit at one constant offset from the wide one, spread 1e-13 px over the whole scene: it is a crop, and its verticals stay parallel to 0e+0°. Turning the plane instead gives a picture that is not a crop of it at all — the same points differ by offsets spreading 25.8 px — and its verticals converge 3.07°.the shifted frame — a crop to 1e-13 pxa tilted picture is not a crop of it — 25.8 px of spreadone eye, one wide picture planetilt converges the verticals 3.07°
Fig. 6 Shift against tilt, on one wide picture from one eye. The shifted frame is a crop of it — every point at one constant offset, spread 1e-13 px. The tilted picture is not: the same points differ by offsets spreading 25.8 px, and its verticals converge 3.07°.
Level, tilted, shiftedTilting the camera up to fit the building in makes the verticals converge by 4.55°. Shifting the lens up instead moves the principal point 95 px off centre, frames the same view, and leaves the verticals parallel to 0e+0° — because the picture plane never tilts.level — the top is cut off0.00° of spreadtilted 13°4.55° of spreadshifted 95 px0.00° of spreada shift moves every point by exactly the shift95.0 px, and no direction at all
Fig. 7 The same choice from the lens field, framed as a photographer’s decision. Tilt up and the verticals converge 4.55°; shift the lens 95 px instead, frame the same view, and they stay parallel to 0e+0°.

The frame is not the plane

A distinction the group makes obvious and ordinary language hides: the picture plane is unbounded, and the frame is a rectangle drawn on it.

Choosing the rectangle is a crop, or a shift, or a change of film format. Nothing about the projection changes; a different part of the same picture is kept.

Choosing the plane is a tilt or a rotation. The projection changes — different coordinates on the same bundle of rays — and every point in the frame moves by an amount depending on where it is.

Both are homographies of the picture, so both are in the group; they are different subgroups of it, and the difference is visible in the measurements. A shift’s offsets are constant over the scene — 1e-13 px of spread, which is the definition of a translation. A tilt’s offsets spread by 25.8 px over the same scene, which is what a projectivity does.

The practical form of the distinction is the one every architectural photographer knows and states as a rule of thumb: shifting costs field and tilting costs shape. Shifting keeps the geometry and throws away the parts of the image circle the frame no longer covers. Tilting keeps the frame and converges the verticals. The rule of thumb is exactly the subgroup statement, and the numbers above are its price list.

It also explains an operation that looks like a third option and is not. Cropping the top out of a tilted picture does not straighten anything — the plane is still tilted, so the verticals still converge, and the crop has merely chosen a smaller rectangle on a leaning plane. To straighten, the plane has to change, which is the next essay.

What a plane change cannot do

The boundary is the reason the theorem is worth stating carefully, and it has two halves.

It cannot move the eye. Every picture in the group is a projection of the scene from the same point. The reader’s viewing position for a corrected picture is computed from its own focal length and displayed width and will generally differ from the original’s — but the scene’s relationship to the observer is fixed, and no processing supplies the view from somewhere else.

And it cannot un-mix the depths. Because the map is scene-independent, it carries no information about depth; because it is a bijection of the picture, it cannot add any. Everything a photograph left out stays left out.

Which is exactly what makes the group useful. A transformation that changes nothing about the geometry’s content is a transformation that can be applied without asking what is in the picture, and that is why keystone correction is a slider in every photo editor and re-staging from a different position is not a feature anywhere.

The same cube turned 24° — a three-point constructionNothing about the construction changed. The number of vanishing points inside any finite distance is 3, and 1 of them fall on the canvas.horizon3 vanishing points at a finite distance2583 px · 6269 px · 549 px
Fig. 8 The named cases, as values of two angles. One, two and three point perspective are one construction with the box turned and the plane tilted — three finite vanishing points here, one of them on the canvas. The names describe where the plane is, not what kind of projection it is.
Straightened — and taken from exactly where it wasThe tilted picture is drawn thin and the corrected one over it. The correction is built from the picture alone: where the imaged verticals meet, and the focal length. What comes out agrees with a level camera at the same eye — one the correction was never shown — to 3e-13 px, and its verticals are parallel to 0e+0°. What has not changed is the eye: the cross-ratio of four points along a ground line reads 1.3333 before and after, so every measurement the original supported the corrected one supports, from the same place and no other.corrected from the picture — 0e+0° between the verticalscorrect from 20 cm, at 160 mm wideagrees with a level camera at the same eye to 3e-13 px
Fig. 9 The group’s most-used member, applied. The picture re-cast onto an upright plane through the same eye, agreeing with a level camera to 3e-13 px — a change of surface with the centre held.

What a reader can check in a photograph

Three things follow that need no computation, and they are the reason this is worth knowing rather than merely true.

Converging verticals mean a tilted plane, not a wide lens. The convergence is 0e+0° with the plane vertical at any field of view, so a wide-angle picture with parallel verticals is perfectly possible and a long-lens picture with converging ones is too. Blaming the lens is blaming the wrong parameter.

The horizon’s position measures the tilt. It sits at the principal point when the plane is vertical and drops by ftanθf\tan\theta as the camera points up, so a horizon low in the frame is a camera pointing up and nothing else. Combined with the vertical vanishing point it gives the focal length.

And a picture with no convergence carries no evidence about the tilt. That is the refusal the machinery makes rather than a caution: with the plane vertical the imaged verticals are parallel and their meeting point does not exist, so there is nothing to compute from. A method that returned a tilt anyway would be reporting the intersection of two nearly-parallel lines, which is the failure mode the projector’s recovery refuses by name.

The other kind of map, for contrast

There is a second family of maps of a picture on this site, and the census that separated them is worth invoking because the difference is exactly the one this essay turns on.

A central collineation — a homology or an elation — has a line of fixed points, and the census found three of them: a shadow, a floor anamorph, a mirror. Each changes where a plane is seen from: the plane stays put and the projection centre moves.

A general projectivity has no such line. A rectification is one: both ends of the map are pictures of a plane rather than the plane and a picture of it.

A plane change is in the second family. Both ends are pictures, taken from the same point on two different surfaces, and there is no line of the world sitting in both to be fixed. Which is a compact way of saying what has and has not changed: the eye has not moved, and neither picture is the world.

Four constructions, three of them the same mapA shadow, a floor anamorph, a mirror and a rectification, each decomposed into its fixed points and lines. Three are central collineations with a line of fixed points; the fourth is not, and that is the difference between changing a picture and changing where it is seen from.constructionfixed structurea shadow, ground to floorhomology · ratio 0.6719a floor anamorphhomology · ratio -1.4815a mirror in a vertical planehomology · ratio -1.0000a rectificationgeneral · three fixed points3 of 4 are centrala line of fixed points is what they share
Fig. 10 The census. Three maps with a line of fixed points — a shadow, an anamorph, a mirror — and one without. A plane change belongs with the fourth: both of its ends are pictures.
The ground, turned into a planFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 1e-15 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifiedacross — 0.6000 widths (true 0.6000)along — 0.7667 widths (true 0.7667)diagonal — 1.2023 widths (true 1.2023)worst error 1e-15 relativethe probes were not used to build the map
Fig. 11 The best-known member of the second family. Four corners of a rectangle of known proportions fix the map, and three lengths it was never given come back to 1e-15 relative — a picture turned into another picture of the same plane, from a viewpoint nobody stood at.

Where the plane’s freedom is used

Three places on this site turn out to be this essay’s theorem applied, and none of them announces it.

Keystone correction of a projector. A projector is a camera run backwards, and correcting its keystone re-casts its picture onto the plane the wall actually occupies — at a cost of 18.7% of the panel’s pixels at 15° off square, because correction cannot add light outside the thrown quadrilateral.

Stitching a panorama. Frames from one point differ by plane changes, so they compose exactly; frames from a moving camera do not, and the error is the parallax measured above.

And the standard construction’s own conditions. “The horizon is at eye level” is a statement about a vertical picture plane, and the figure that measures it says so in its title. This essay is the rung where the condition is removed rather than assumed.

What keystone correction actually costsA projector turned 15° from square throws its rectangular panel as a quadrilateral. Correction cannot add light outside it, so it shrinks the picture until it fits — and 18.7% of the projector's pixels are thrown away. The fraction is measured on the panel rather than on the wall, because turning the projector makes the wall picture larger while making the panel usage smaller.15° of yaw, 6° of pitch, 1.50 throw ratio81.3% of the panel reaches the corrected rectangleouter: the thrown quadrilateral · inner: what correction can keep18.7% of the panel discarded
Fig. 12 The projector’s version of the same correction, and its price. Turning the projector 15° from square costs 18.7% of the panel — correction is a homography of the picture and a homography cannot put light where none was thrown.
Four figures of the same height, camera level at 1.60 mThe horizon cuts every one of them at 89.9% of its height — the eye height over the figure height — however far away it is.horizon = eye level, 1.60 m89.89%correct from 26 cm, at 160 mm widespread 0
Fig. 13 The condition this essay removes, stated in its own figure: with the plane vertical the horizon cuts every upright at the same fraction of its height, 89.9% here, however far away it is.
The three vanishing points of one box, drawn to scale with the boxThe picture is the small rectangle. Two of the three vanishing points fall well outside it, which is why they are computed rather than located by eye.the pictureVP₁VP₂VP₃orthocentrefocal length from the triangle — 853.9 pxspread 1e-14% across three routes
Fig. 14 And where the plane’s tilt puts the third vanishing point. Two of the three fall well outside the frame, which is why they are computed rather than found by eye — and why the identity above is worth having as a calibration.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

centre of projectionHomographyHorizonParallaxPicture planepole and polarPrincipal pointProjective mapthree-point perspectivevertical vanishing point