What survives

A projection of a projection

Photograph a photograph and the composite map is a homography — four marks determine it and every other lands where they say, to 1e-13 px. Roll the print and the same four mispredict the rest by 30.5 px, which is why straightening a page of a thick book never quite works.

Every essay in this field so far has been about a single projection: what it destroys, what survives it, where its centre is, and how to get the centre back. This one is about two of them in a row.

The result is short and it is load-bearing for a good deal of the site. A projection composed with a projection, through a plane, is a projection. So the map from an original picture to a photograph of that picture is a homography, and four correspondences determine it completely.

A print, photographed again — flat and rolledFour marks fix a homography; the other 16 are predicted by it. On a flat print they land where it says to 1e-13 px. Rolled to 1/R = 0.90 per metre the same four predict the same 16 to 30.5 px, because a composition of projections is a projection only if the middle surface is a plane.an anchorcorrect from 19 cm, at 160 mm wideflat 1e-13 px · rolled 30.5 px
Fig. 1 A print, stood in the world and photographed again. Four marks fix a homography; the other sixteen are predicted by it. Flat, they land where it says to 1e-13 px. Rolled to 1/R = 0.90 per metre, the same four mispredict the same sixteen by 30.5 px.

Why the composition closes

The argument is one line if the right form is used.

A projection through a centre, written in homogeneous coordinates, is a linear map: xPx\mathbf{x}' \propto P\mathbf{x}, with the proportionality absorbing the perspective divide. Restricted to a plane in the world, it is a 3×3 matrix acting on the plane’s own homogeneous coordinates — which is what a homography is.

Composing two of them gives the product of two 3×3 matrices, which is a 3×3 matrix. That is the whole proof.

The reason it is worth stating rather than assuming is that the composition is doing real work all over this site, in places where it is not obviously the same operation.

Four things that are this theorem

Rectification. Flattening a photographed façade into an elevation is exactly this composition run backwards: the façade is a plane, the photograph is a projection of it, and the rectifying map is the homography that undoes the projection. It works because the composition of the projection and its undoing is the identity, which is a homography.

Rephotography. Photograph a print, and the map from the original file to the new one is a homography. Which means a rephotograph carries no geometric information the original lacked, and none is lost either — every cross-ratio survives both projections, so anything measurable in the original is measurable in the copy.

Stitching. Two frames from one point are related by a rotation homography, which is this theorem with the plane at infinity.

And the tilt correction. Straightening a photograph of a building by sending the vertical vanishing point back to infinity is a homography of the picture plane, so the corrected picture is a perfectly good projection of the same scene from the same point — taken on a differently oriented picture plane.

Four operations that look like four techniques. One matrix product.

The prediction, and why it is a prediction

The figure could have fitted a homography to all twenty marks by least squares and reported how well it fitted. That would prove much less than it appears to.

A least-squares fit to twenty points with eight parameters will report a small residual for a good many maps that are not quite homographies, and the residual would be partly the fit’s own slack. So instead: four marks are named, the homography they determine exactly is constructed, and the other sixteen are predicted by it.

A prediction that lands to 1e-13 px is a statement about the geometry. A fit that lands to 1e-13 px would be a statement about the fit.

That distinction is the same one the least-squares vanishing point makes in the other direction. There, the fit is over-determined on purpose and the residual is the report; here the map is under-determined on purpose so the remaining points are free to disagree. Both are ways of arranging for a number to mean something.

And the counter-case

An identity with no counter-case is a definition, so the figure rolls the print.

A cylindrical print of radius RR is not a plane, and the composition through it is not a homography. The same four anchors still determine a map exactly — four points always do — and the other sixteen now land somewhere else. At 1/R=0.901/R = 0.90 per metre they are out by 30.5 px.

The slider runs the curvature from almost flat to sharply rolled, and the mispredicton grows smoothly from a fraction of a pixel. There is no threshold: any curvature at all breaks the theorem, and the only question is by how much relative to whatever precision the measurement needs.

Reading the curve of the failure

The slider sweeps the print’s curvature, and the shape of the mispredicton against it is worth a look because it says which term is dominating.

At small curvature the error grows very nearly linearly: 6.6 px at 1/R=0.21/R = 0.2, 16.8 at 0.5, 33.9 at 1.0, 65.1 at 2.0. Doubling the curvature roughly doubles the error, which says the departure is first-order in the surface’s deviation from a plane.

That is what should be expected and it is worth confirming, because the alternative — a quadratic onset, error growing as curvature squared — would mean a range of curvatures where the theorem holds to within any given tolerance. It does not. Any curvature at all costs, in proportion.

The practical reading of that: there is no “flat enough” that can be decided by looking. A page that appears flat, a print taped to a wall, a poster with a slight bow — each of them has a curvature, and the question is not whether it is zero but whether the error it produces is below the precision the measurement needs. That is a number, and it is computable from the sag of the surface and the geometry of the two views.

Which is why straightening a page never quite works

The practical consequence is one everybody has met.

Photograph a page of a thick book and the page is not flat: it curves into the gutter. Every scanning application offers to straighten it, and the good ones work by fitting the curvature explicitly. The naive ones fit a homography to the page’s four corners.

The homography maps the corners exactly, because four points always determine one. The text in the middle of the page stays bent, and the residual is largest exactly where the curvature is — near the gutter, which is where the text is hardest to read anyway.

That failure has the same shape as a stitch registered on the sky and a façade rectified from four corners: a fit that is exact on the data it was given and wrong on the data it was not. The residual on the fitted points is zero in all three cases and says nothing at all about whether the model is right.

It is worth having a name for that failure, because it recurs across this site and it is not the same thing as overfitting. Call it what it is: a model with exactly as many parameters as constraints has no residual to report, so the only evidence about it comes from data it did not see.

A 9 mm gap in a canopy, at various heightsThe patch is the hole's shape near the hole and the sun's shape far from it, and the crossover — where it is half of each — is the hole's width over the sun's angular width: 0.97 m. Above that, every gap in the canopy is a pinhole camera imaging the sun, which is why they all go crescent-shaped together during an eclipse.00.2000.4000.6000.8000246distance from the hole to the ground (m)how much of the patch of light is an image of the suncrossover at 0.97 mthe sun subtends 9.30 mradpatch 65 mm wide at 6 m
Fig. 2 A projection through a hole rather than through a lens, and the distance at which it stops imaging the hole and starts imaging the source. The same composition, with the aperture doing the projecting.

The one that is not a homography

There is an operation on this site that looks as though it belongs in the list above and does not, and the exclusion is instructive.

The catoptric anamorph is a picture on paper, reflected in a cylindrical mirror, seen by an eye. Two stages, like a rephotograph — and the composite is nothing like a homography, because the intermediate surface is a cylinder.

That is the same clause failing, in a case where the failure is the point rather than the problem. The anamorph works because the intermediate surface is curved: a flat mirror would give a homography of the drawing, which is to say a drawing that looks like an ordinary drawing seen from an ordinary place, and the whole effect depends on the map being wildly non-projective.

So the curved-surface case is not merely the theorem’s counter-example; it is a technique. The local area scale of the catoptric smear reaches 7.2 between its most and least crowded parts, which is the same quantity the rolled print’s 30.5 px is, at a scale chosen to make a drawing unreadable rather than to make a measurement slightly wrong.

Read together, the flat print and the cylindrical mirror bracket the theorem: one shows what a planar intermediate surface buys, the other shows what abandoning it buys.

What the flatness is doing

It is worth being precise about which assumption the counter-case breaks, because there are two candidates and only one of them is guilty.

The composition needs the middle surface to be a plane. The first projection’s centre can be anywhere, the second’s can be anywhere, the two picture planes can be at any orientations; none of that matters. What matters is that the intermediate surface — the thing being photographed — is a plane, because that is what makes each projection restrict to a 3×3 map on a common two-dimensional coordinate system.

Roll it and each point of the print is at a different depth, so the second projection’s action on the print’s own coordinates involves that depth, and depth is not a projective function of position on a curved surface.

Which is the same reason a refracted picture is not a projection, arrived at from a different direction: a map that depends on a depth cannot be a map of the picture plane.

A façade, flattened out of the photographFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 4e-16 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifieddoor — 0.5068 widths (true 0.5068)sill — 0.3026 widths (true 0.3026)diagonal — 1.1766 widths (true 1.1766)worst error 4e-16 relativethe probes were not used to build the map
Fig. 3 The theorem’s most useful consequence: a photographed plane flattened by the homography that undoes its projection. The composition of the projection and its undoing is the identity, which is why the rectified lengths come back exact.

The cross-ratio, twice

The invariant gives a second and quite independent route to the same theorem, and it is worth having because it needs no matrices.

The cross-ratio of four collinear points survives a projection. Take four points on a line in the original picture: the first projection carries them to four collinear points on the print with the same cross-ratio, and the second carries those to four collinear points in the rephotograph, again with the same cross-ratio.

So the composite preserves collinearity and cross-ratio — and a map of the plane that does both is projective. That is a theorem (the fundamental theorem of projective geometry, in the form for the real plane) and it identifies the composite as a homography without ever writing one down.

The two routes are worth having side by side because they fail differently on the rolled print. The matrix route fails because there is no common 3×3 map; the cross-ratio route fails because the first projection no longer sends collinear points to collinear points on the curved surface. Both diagnoses are correct and they point at the same clause.

Four points on a line, before and after a projectionLength and the ratio of lengths do not survive the projection; the cross-ratio does, agreeing to 0e+0 relative.horizonABCDon the groundin the picturelength AB1.00011.3930ratio AB:CD0.56670.6837cross-ratio1.31681.3168correct from 26 cm, at 160 mm wide34° across
Fig. 4 The invariant that carries the second proof. It survives the first projection and it survives the second, and a map of the plane that preserves collinearity and cross-ratio is projective.

Composing more than two

The theorem extends without effort and the extension is worth noting because it is what makes a chain of operations safe.

A projection of a projection of a projection is still a projection, provided every intermediate surface is a plane — three 3×3 matrices multiply to a 3×3 matrix. So a photograph of a print of a scan of a photograph of a façade is related to the original façade by a single homography, and four correspondences anywhere in that chain determine the whole of it.

The group structure is the reason. Homographies of the plane form a group under composition: they are closed, every one has an inverse (any invertible 3×3 matrix does), and the identity is one. So a sequence of them collapses to one, and there is no accumulation of anything.

That last point is what makes the result useful in practice rather than merely tidy. It says a chain of projective operations does not degrade geometrically — the errors do not compound, because there are no errors to compound. Whatever precision is lost is lost in the resampling and the marking, not in the geometry.

And it says exactly where a chain does break: at the first non-planar link. One curved intermediate surface in a chain of ten flat ones takes the whole composite out of the group, and no amount of care with the other nine repairs it.

What this says about a copy

A small philosophical dividend, because it is the kind of thing this subject is usually vague about.

A photograph of a photograph is geometrically a perfect copy — not in resolution, but in the sense that every projective property of the original is present in it and recoverable from it. The cross-ratios are identical. The vanishing points are the images of the original’s vanishing points under the homography. A camera recovery run on the copy returns the copy’s camera, and the original’s intrinsics are recoverable from the copy given the homography.

So there is no geometric degradation in copying, and there is no geometric information added either. Whatever a photograph of a scene reports, a photograph of that photograph reports the same thing, up to a projective map computable from four marks.

That is not true of most operations on a picture. A crop moves the principal point; a distortion correction changes the projection; a resample loses detail. Rephotographing a flat print does none of it.

Alberti's construction, with the section that fixes the depthsLeft: the panel, six braccia across, its transversals found where the section's rays cross the picture plane. Right: the section, with the eye at its true distance. The transversals agree with a pinhole camera of the same focal length to 6e-14 px.the panelhorizon — the centric point's heightthe section — the eye, the panel, the ground530 px — the viewing distancethree routes agree to 1e-13 pxsection, distance point, and a pinhole camera
Fig. 5 The oldest implementation of the composition: a plan and an elevation, both orthographic projections, read off where they meet the picture plane. It agrees with a pinhole camera because two projections of a plane compose into one.

The theorem’s oldest use

This is not a modern result and it is worth saying where it came from, because the historical use is the same as the modern one.

Renaissance perspective construction is a composition of exactly this kind. Alberti’s lateral section draws a plan and an elevation — two orthographic projections of the scene — and reads the picture off their intersection with the picture plane. That the result agrees with a pinhole camera to 5.7 × 10⁻¹⁴ px, which this site measures, is a consequence of the composition being projective: two projections of a plane, composed, are one projection of it.

The distance-point method is the same thing arranged differently, which is why it agrees with Alberti’s to the same precision. Both are ways of computing a homography of the ground plane by ruler and compass, and both compute the same one, because there is only one.

The measuring-point construction, the diagonal method, the plan-projection method taught in architecture schools — all of them are constructions of the ground plane’s homography, and their agreement is not a coincidence to be verified case by case. It is this theorem.

That reframing is worth having because it converts a list of recipes into one statement with several implementations, which is what the whole construction field on this site is arguing.

A print, photographed again — flat and rolledFour marks fix a homography; the other 16 are predicted by it. On a flat print they land where it says to 1e-13 px. Rolled to 1/R = 0.25 per metre the same four predict the same 16 to 8.3 px, because a composition of projections is a projection only if the middle surface is a plane.an anchorcorrect from 19 cm, at 160 mm wideflat 1e-13 px · rolled 8.3 px
Fig. 6 A gently rolled print. Any curvature at all breaks the theorem and the mispredicton grows in proportion, so there is no “flat enough” that can be decided by looking.
A façade, flattened out of the photographFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 7e-16 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifieddoor — 0.5068 widths (true 0.5068)sill — 0.3026 widths (true 0.3026)diagonal — 1.1766 widths (true 1.1766)worst error 7e-16 relativethe probes were not used to build the map
Fig. 7 The composition run with the picture plane tilted, which changes nothing that matters: the map is still projective and the recovered lengths still come back exact.

Where it stops

Three limits, stated so the theorem is not carried further than it goes.

It needs the print to be planar, which the counter-case measures.

It needs both stages to be projections. A refracted or distorted picture is not one, so photographing a print through a real lens introduces the lens’s own departure, and the composite is a homography plus a radial polynomial rather than a homography.

And it says nothing about anything but geometry. Tone, resolution, noise and colour all degrade in a copy, and none of them is on this site’s subject. The claim is about where the marks are and nothing else.

Within those limits it is unusually clean for a result about pictures, and it is the reason the site’s rectification machinery is allowed to exist at all. Every time a plane in a photograph is flattened, straightened, measured or compared, this theorem is what says the operation is legitimate — and the rolled print is what says to check the plane first.