Measuring from one picture

Flattening a façade out of the photograph

Four corners of a rectangle whose proportions are known are enough to undo the projection of one plane. After that the plane can be measured with a ruler — lengths, angles, areas, all of it — in units of the rectangle's own width, and lengths the map was never given come back to fifteen digits.

A photograph of a building taken from the pavement is a picture of a plane, seen obliquely. Every window in it is a parallelogram that should be a rectangle, every course of brick converges toward a point off to one side, and nothing in it can be measured with a ruler.

It can be measured after one operation. The projection of a plane to a plane is a homography — eight numbers — and eight numbers are determined by four point correspondences. Give the four corners of anything in the picture whose real proportions are known, solve, and the whole plane comes back flat.

A façade, flattened out of the photographFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 4e-16 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifieddoor — 0.5068 widths (true 0.5068)sill — 0.3026 widths (true 0.3026)diagonal — 1.1766 widths (true 1.1766)worst error 4e-16 relativethe probes were not used to build the map
Fig. 1 Left: the façade as photographed, with three lengths marked. Right: the same plane after the homography. The four corners of the reference rectangle built the map; the three marked lengths were never shown to it, and come back to within 7 × 10⁻¹⁶ relative.

Eight numbers from four points

A homography is a 3 × 3 matrix acting on homogeneous coordinates, defined up to scale, so it has eight free parameters. Each point correspondence gives two equations. Four correspondences give eight, and the system is square.

The solve is ordinary Gaussian elimination with partial pivoting on an 8 × 8 system, and it is worth noting one detail of how it is written here: given a singular column, it asserts rather than returning null. A homography from four points three of which are collinear is not a degenerate answer to be handled downstream; it is a statement that the four points do not describe a quadrilateral, and the figure that asked for it is wrong. Refusing is the right response and it is the site’s habit throughout.

Once solved, applying the map is three multiplications and a division per point — the perspective divide, which is the same operation the camera performs and the reason a homography is a projection rather than an affine warp.

A word drawn to be read from 74° off to the sideStraight strokes stay straight and the cross-ratio along each is preserved, which is what makes this a projection rather than a distortion.eye, 64° offgrey: the word before the projectionblack: the same word, projected
Fig. 2 The same eight numbers running the other way. An anamorph is a homography applied to make a picture correct from one oblique point; a rectification is a homography applied to undo one. The cross-ratio check is identical in both.

The property that says it is a projection

The check that the map is right is not that the four corners land where they were told to. They will, whatever the solve did with the rest of the plane, because they are the eight equations.

The property to check is that the map preserves cross-ratio. Take four collinear points anywhere in the source, map them, and the cross-ratio of the four results must equal the cross-ratio of the four originals. If the 8 × 8 solve or the perspective divide has an error in it, the mapped picture still looks like a plausible warp of the original — cross-ratio is what notices. This site’s homography check puts the agreement at machine precision, and asserts collinearity separately, because a warp that bent lines while preserving some numeric coincidence would be a different failure.

That is the same invariant the height measurement runs on, and it is the same invariant that makes the anamorph on this site a projection rather than a distortion. One quantity, three uses.

Why four points, and not three or five

The count is worth pausing on because it is the same count that decides what a picture can and cannot give up, and it recurs across this whole site.

Three points determine an affine map — six parameters, two per point — which handles rotation, scaling, shear and translation and nothing else. An affine map sends parallel lines to parallel lines, and a photograph of a plane emphatically does not: the whole visible content of an oblique photograph is that the parallels converge. So three points cannot describe what a camera did, and a warp built from three will reproduce them and get everything between wrong.

Four points determine a homography, and a homography is exactly the group of transformations a plane can undergo when photographed. Nothing more general is needed, because a camera is a projection through a centre and a projection through a centre restricted to a plane is a homography. That is a theorem and not an approximation, and it is why the residual above is at the floating-point floor rather than merely small.

Five points are one too many, and what happens with five is instructive. Five correspondences over-determine the system, so in general no homography satisfies all of them, and the least-squares fit has a residual. That residual is a measurement: it says how far the five points are from being five points of one plane, photographed. Given five points on a real façade, a large residual means one of them is not on the façade — the balcony, the sign, the projecting cornice — and the method has found it.

So the fifth point is the one that turns a construction into a check, in exactly the way the third line of a vanishing-point bundle does. Two lines always meet; three only meet if they share a vanishing point. Four points always fix a homography; five only fit one if they are coplanar. The pattern is the same one and it is worth reaching for whenever a construction is exactly determined: adding one more input costs nothing and converts an assumption into a residual.

The rectification returns lengths it was not given

The real test of a rectification is a length it has never seen.

The figure above is built by placing a known rectangle in a known plane, projecting it with a camera, and then handing the four image corners and the rectangle’s aspect ratio to the rectifier. Three further lengths in the same plane — a doorway’s height, a sill’s width, a diagonal across the whole façade — are then measured on the rectified result and compared with their true values.

The worst relative error over the three is 7 × 10⁻¹⁶. That is the floating-point floor, and what it says is that the operation contains no approximation: a plane in a picture really is an exact projective image of the plane in the world, and eight numbers really do invert it.

The diagonal is the probe worth watching, because it crosses the whole figure and is the one that would fail first if the map were an affine approximation rather than a homography. An affine map — the four-parameter warp that scaling and shearing gives — reproduces a rectangle’s corners and gets everything inside it wrong, and the error is largest across the diagonal.

The taught two-point cube, with the two far edges placed 8 points apartThe corner angles are 90° because the method forces them. The side ratio is 0.719, so this picture depicts a box whose depth is 1.39× shallower than its width.horizoncorner angles90.000° — forced by the methoddepicted side ratio0.7195lens this drawing implies60° acrossdrawn exactly as the method prescribesthe free step is where the far edges go
Fig. 3 What a recovery of the shape needs. Given intrinsics, a drawn quadrilateral can be asked what solid it depicts — which is exactly the aspect ratio this essay has to supply from outside, arriving as a measurement instead of an assumption.

The aspect ratio is an input, not a recovery

There is one number in the construction above that did not come from the picture, and it is the one this method is most often got wrong on.

Four corners in an image determine the homography to a known rectangle. They do not determine the rectangle. Any of a one-parameter family of rectangles, of every aspect ratio, has a projection matching those four corners exactly — because the family of homographies taking four points to four points includes one for each target shape.

So the aspect ratio is supplied. In the figure it is stated as 0.62 because that is what the reference actually is. In practice it comes from knowing something: a window is a standard size, a door is 2.0 by 0.8, a paving slab is square. Getting it wrong does not produce a visibly wrong rectification — it produces a plausible one, of a building slightly taller or wider than the real one, and every measurement taken on it is off by the same factor in one direction only.

There is a recovery available if more is known. Given the camera’s focal length, or two vanishing points and the assumption of a square principal point, the rectangle’s aspect ratio can be computed from its image — that is what the site’s depictedRectangle does when it asks what solid a drawn quadrilateral depicts. But that needs intrinsics, and a rectification that needs intrinsics is no longer a method that works on a photograph of unknown provenance.

Stating which of the two situations is in play is the whole of using this responsibly. With intrinsics, the rectangle is recovered and the answer is a measurement. Without them, the aspect is an assumption and the answer is a measurement conditional on it.

A 3.4 m object measured from one picture, 11 m awayThe base, the horizon crossing, the top and the vertical vanishing point have a cross-ratio of 1.9101. With the eye at 1.62 m that gives 3.400 m, against a true 3.4 m. The camera is not consulted.horizon — the eye's own heightbase — 0 mhorizon crossing — 1.62 mtop — 3.40 m recoveredthe vertical vanishing point is 8586 px above this framerecovered 3.400 m · true 3.400 m3.5 cm per pixel of click error
Fig. 4 The other half of the same toolkit. Heights come from a cross-ratio along a vertical and need the horizon and the vertical vanishing point; lengths on a plane come from a homography and need four corners. Between them they cover the two things a photograph of a street is usually asked for.

What is measurable on a rectified plane

Once flat, everything that was destroyed by the projection comes back — for that plane and no other.

Lengths come back as ratios, in units of the reference’s width. Angles come back exactly, because a rectified plane is related to the real one by a similarity — a rotation, a scale and a translation — and similarities preserve angle. Areas come back as ratios of the reference’s area. Parallelism comes back, which is the property whose loss makes the unrectified photograph unreadable.

What does not come back is anything off the plane. A balcony projecting half a metre from the façade is not in the plane and its rectified position is wrong by an amount proportional to how far out it stands. This is the standard failure of the method in the field and it is invisible in the result: a rectified façade looks perfectly flat, including the parts of it that are not.

The check for it is available and cheap: a projecting element rectified from two different photographs lands in two different places, and an element in the plane lands in the same place from both. That is a second view, which is what this whole field is defined by not having — so in the single-view case the honest statement is that the method measures a plane and assumes the object is one.

Rectifying without a rectangle

The four-corner method needs a rectangle, and plenty of façades do not offer one. There are two other routes to the same map and both are worth knowing, because both replace the rectangle with something a photograph of a building nearly always has.

The first uses two vanishing points. The horizontal and vertical directions of the façade each have a vanishing point in the picture, found from the drawn lines the same way this site finds every vanishing point — three or more parallels, a least-squares intersection, and a residual that says whether they really are parallel. The line joining the two vanishing points is the image of the façade’s own line at infinity, and mapping that line back to infinity removes all the convergence. What is left is an affine image of the façade: parallels are parallel, midpoints are midpoints, ratios along any line are correct. It is not yet a rectification, because angles and the two axes’ relative scales are still wrong, but it is enough for a large class of measurements and it needs no known rectangle at all.

The second adds two known angles or one known ratio to that affine image, and finishes the job. A right angle in the façade — a window corner, a door jamb meeting a sill — fixes the remaining shear; a known aspect anywhere fixes the remaining relative scale. Between them the affine image becomes a similarity of the real plane, which is a rectification.

Splitting it into two stages this way is more than a convenience. It separates what the picture supplies from what the reader supplies, and it makes the second dependency explicit. The convergence is removed by information in the photograph: the parallels are visible and the vanishing points are measurable, with residuals. The shear and the scale are removed by information about the building, which is an assumption whether or not it is a good one.

That is the same division the four-corner method has, with the boundary drawn in a different place — and seeing it drawn twice makes it clear that it is not an artefact of either method. A single view of a plane determines the plane’s projective structure and nothing about its metric structure. Every route to a rectification has to import the metric part, and the only question is which sentence it is imported in.

The same operation, three times on this site

A homography has now appeared three times here in three different costumes and it is the same eight numbers each time.

It is an anamorph: a picture drawn on a sheet so as to be correct from one oblique viewpoint is the homographic image of the picture that would be correct from straight ahead. The site’s anamorphic word is built by exactly this map, and checked with exactly the cross-ratio test above.

It is a rectification: the inverse operation, undoing the obliquity of a plane in a photograph.

It is a change of picture plane: the relationship between two photographs taken from the same point in different directions, which is why a panorama can be stitched at all and why a re-projection from an equirectangular file to a flat view is a warp with no depth information in it.

Three subjects, taught in three places, one object. This site’s habit is to notice when that happens, and the reason it is worth noticing is practical: an implementation checked once is checked for all three, and a bug in it shows up in whichever is being used at the time.

Two scenes 137× apart, and the one picture they both makeEverything in the second plan — the room, the eye's distance, the eye's own height — is 137 times the first. Every projected vertex agrees to 1e-13 px. A single photograph has no scale, and this is what that means.a room 2.8 m across, eye 1.6 m up1 mthe same plan, 201.5× bigger201.5 midenticalpicturesthe picture — both scenes, drawn twice, one on top of the otherlargest disagreement 2e-13 px over 8 verticesone length has to come from outside the picture
Fig. 5 The limit that applies to rectification exactly as it applies to heights. Two worlds 137× apart, one picture. The rectification returns every ratio on the plane and no size, and the reference width is the one length that turns the ratios into metres.

What the rectified picture is correct from

There is one question this site asks of every picture and it has an odd answer here.

A rectified façade is a picture. Where is it correct from?

The answer is that it is correct from directly in front, at any distance — which is a different kind of answer from the one this site usually gives, and the difference is worth understanding rather than glossing.

A rectification is not a perspective picture of the façade taken from somewhere better. It is an orthographic view: the projection a camera infinitely far away with an infinitely long lens would produce, which is the same thing as an elevation drawing. That is what makes it measurable — an orthographic projection preserves ratios along the plane, which is exactly the property the parallel drawing systems are built on and exactly the property perspective destroys.

So a rectified façade belongs, geometrically, with the isometric drawing and the elevation rather than with the photograph it came from. It has no station point, no vanishing points, no correct viewing distance. The reason it is useful is the reason all the parallel systems are useful: it answers “how big is that, relative to this” and refuses to answer “what did it look like from over there”.

The conversion between the two is exactly the homography, in both directions. Feed the rectified elevation back through the inverse map and the original photograph returns, viewing distance and all. One matrix separates a picture that shows how something looked from a picture that shows how big it is, and no picture is ever both.

What each parallel system does to the three axesThe smallest axis scale is plotted; the note says how many of the three are equal. Only isometric makes all three the same, and even then not equal to one.elevation0.0000two equalcabinet0.5000two equalcavalier1.0000all three equaldimetric0.4714all three differentisometric0.8165all three equaltrimetric0.5479all three differentsmallest of the three axis scalesmeasured from each projection
Fig. 6 Where a rectified plane belongs. It is an orthographic view, so it preserves the ratio in which a point divides a segment — the property this chart measures across the drawing systems, and the reason a rectified façade can be read with a ruler.

Where the errors actually come from

The floating-point agreement above is a statement about the mathematics and says nothing about a real photograph. The error budget of an actual rectification has three terms and they are not equally sized.

Corner placement. The four corners are clicked by a person, or found by a detector, to within a pixel or two. Their effect on the map is strongly non-uniform: a corner near the far end of the plane, where the projection has compressed everything, moves the recovered geometry much more than one near the front. A rectification built on a well-spread quadrilateral is much better conditioned than one built on a small distant one, and choosing the reference is the largest single decision.

The aspect ratio. Discussed above, and it enters as a pure scale factor in one direction — so it distorts the answer without degrading its consistency, which is the worst kind of error because nothing internal to the measurement complains.

Lens distortion. Real lenses do not produce exact projections. Barrel and pincushion distortion bend straight lines by a fraction of a per cent near the centre and by several per cent at the corners, and a homography fitted to distorted corners is a homography fitted to the wrong points. This is the term that dominates in practice, and it is why serious work undistorts before rectifying and why a rectification from an uncalibrated wide-angle photograph should be treated as approximate however clean the arithmetic looks.

None of the three appears in the 7 × 10⁻¹⁶ above, and that is the correct relationship between a check and a field measurement: the check proves the method is exact, so that anything wrong with an answer is known to be in the inputs.