Measuring from one picture

A height, out of one photograph

Four points on a vertical, one cross-ratio, and the height of something nobody measured. The only metric input is the photographer's own eye height, because the horizon is at eye level and that is the one piece of perspective folklore that is exactly true.

A projection destroys length. That is the first thing this site established and it is not in doubt: two objects of different sizes at different distances image identically, and no amount of looking at the picture separates them.

And yet a single photograph of a street will yield the height of a lamp-post to within a few per cent, using nothing but the picture and one number the photographer already knows. The apparent contradiction is worth resolving carefully, because the resolution is the whole of this field.

A 3.4 m object measured from one picture, 11 m awayThe base, the horizon crossing, the top and the vertical vanishing point have a cross-ratio of 1.9101. With the eye at 1.62 m that gives 3.400 m, against a true 3.4 m. The camera is not consulted.horizon — the eye's own heightbase — 0 mhorizon crossing — 1.62 mtop — 3.40 m recoveredthe vertical vanishing point is 8586 px above this framerecovered 3.400 m · true 3.400 m3.5 cm per pixel of click error
Fig. 1 A 3.4 m object measured from one picture. The three marked points are its base, the place where the horizon crosses its vertical, and its top; the fourth point of the cross-ratio is the vertical vanishing point, thousands of pixels off the frame. The recovery is handed those four points and the eye height and nothing else.

What survives, and what it is worth

The quantity that survives a projection is the cross-ratio of four collinear points. Four points on a line in the world, and their four images, have the same cross-ratio, exactly — and this site’s checks put that agreement at 10⁻¹⁶.

A cross-ratio needs four points, which is why the invariant is usually presented as a curiosity. Most measurements a person wants involve two points: the two ends of a thing. Two points have no cross-ratio, and the invariant appears to be useless for the only task anyone brings to it.

The move that makes it useful is to notice that a vertical object in a scene sitting on the ground comes with two free extra points, both of them supplied by the picture rather than by the object.

The first is the vertical vanishing point: the place where all the scene’s verticals converge. It is the image of the point at infinity in the vertical direction, so on the object’s own vertical line it is the point at world height infinity.

The second is where the horizon crosses the object’s vertical. The horizon is at eye level — genuinely, exactly, for any camera whose picture plane is vertical, and this site has already measured what a tilt does to that. So the crossing point is at world height equal to the eye’s height above the ground.

Four points on one line, with world heights 0, e, Z and ∞, where e is the eye height and Z is what is wanted.

Four figures of the same height, camera level at 1.62 mThe horizon cuts every one of them at 91.0% of its height — the eye height over the figure height — however far away it is.horizon = eye level, 1.62 m91.01%correct from 26 cm, at 160 mm widespread 0
Fig. 2 The fact the whole method rests on. The horizon is at the eye’s own height, so where it crosses a standing object is a point of known world height — the one thing in the construction that ties the picture to a real length.

Four lines of algebra

Cross-ratio survives the projection, so the cross-ratio of the four image points equals the cross-ratio of the four world heights. Write the cross-ratio of four values as

CR(a,b;c,d)=(ca)(db)(cb)(da).\mathrm{CR}(a,b;c,d) = \frac{(c-a)(d-b)}{(c-b)(d-a)}.

Put in the base at 0, the horizon crossing at e, the top at Z and the vanishing point at infinity. The fourth argument appears in both numerator and denominator, and in the limit the factors cancel:

CR=(Z0)(Ze)=ZZe.\mathrm{CR} = \frac{(Z-0)}{(Z-e)} = \frac{Z}{Z-e}.

Invert it and the whole method is one line:

Z=eCRCR1.Z = \frac{e \cdot \mathrm{CR}}{\mathrm{CR} - 1}.

Measure the four points in the picture, take their cross-ratio, supply the eye height, and read off the answer. Nothing about the camera enters — no focal length, no principal point, no orientation. The picture is asked and the picture answers.

The recovery is never shown the camera

That last claim is the one that has to be arranged for rather than asserted, and this site has a standing habit for it.

The figure above builds a scene, projects it with a camera whose focal length and position are known exactly, and then puts the camera away. What is handed to the recovery is four image points and one number. The true height is used at the end, to compare, and never during.

The result is 3.400000000 metres against a true 3.4, a relative error of 2.6 × 10⁻¹⁶ — the last bit of double precision. That is not an impressive result about photography; it is a statement that the derivation above has no approximation in it. A projection really does preserve the cross-ratio, the horizon really is at eye height, and the two facts really do compose into a measurement.

What makes it worth running rather than asserting is the failure modes it excludes. A sign error in the cross-ratio, a confusion between the horizon crossing and the principal point, a vanishing point taken from the wrong direction — every one of those produces a plausible number for a plausible-looking object, and only the comparison with a truth the recovery never saw distinguishes them.

The three things that will be refused

The method has three degenerate cases and all three are refusals rather than wrong answers, which is a decision worth defending.

The top must be above its own base. The object’s vertical in the picture is taken as the line through the base and the vanishing point, not the line through the base and the top. When the drawing is right the two are the same line. When a reader clicks a top belonging to something else, they are not, and the method says so instead of returning a number for an object that does not exist.

The top must not be exactly on the horizon. An object whose top lands on the horizon is exactly eye height, and the formula cannot say so: CR goes to infinity and Z = e·CR/(CR − 1) is 0/0 in the limit. This is the one case where the parameterisation is worse than a reader — anyone can see that a top on the horizon means an object at eye height, instantly, with no arithmetic at all.

The camera must not be exactly level. A perfectly level camera has its vertical vanishing point at infinity in the image, which is to say it does not have one, and the method has no fourth point. This one looked like the method’s real weakness and turned out not to be, for reasons that take their own essay.

The vanishing point runs to infinity and the measurement does not careAs the camera comes level the vertical vanishing point leaves the canvas, the page and eventually the plausible — 7.2 × 10⁹ px at a tilt of one part in eight million. The recovered height stays exact to 2e-16 relative the whole way. At exactly level the method has nothing to work with and refuses.0510-6-4-20how far the camera looks down, over eight metres (metres, log scale)where the vertical vanishing point falls (log₁₀ pixels)the vanishing pointthe error in the recovered heightthe error curve is offset by 17 decades to be visiblea flat line at machine precision
Fig. 3 The vertical vanishing point’s position against the camera’s tilt, with the recovery error beside it. As the camera comes level the vanishing point leaves the canvas, the page and eventually the plausible — 7.2 × 10⁹ px at a tilt of one part in eight million — and the recovered height stays exact the whole way. The failure at exactly level is a cliff, not a slope.
A family of parallel ground lines at 30°, and where they meetAll five lines pass through one point on the horizon, off the edge of the frame at x = 1464.horizon — the image of the line at infinityvanishing point at x = 2279 — off the framecorrect from 26 cm, at 160 mm wide34° across
Fig. 4 How the horizon is obtained when the picture does not show one. A bundle of world-parallel lines meets at a point found by least squares, with the residual reported — and two such points determine the horizon.

The horizon has to be found

The derivation treats the horizon as given, and in a real photograph it is not. The sea has one; a street does not. So the first step of any actual measurement is to recover the horizon from the picture, and that step is where most of the error enters.

The horizon is the image of the ground plane’s line at infinity, so it passes through the vanishing point of every horizontal direction. Two such points determine it. Each is recovered the way this site recovers all of them: take three or more image lines known to be parallel in the world — the top and bottom of a wall, a kerb and the gutter beside it, a run of window sills — and find the point they agree on, in the least-squares sense.

Two segments always meet somewhere, so an intersection proves nothing. Three or more only meet at a point if they really do share a vanishing point, and the number that says whether they do is the residual: the largest distance from the fitted point to any of the lines. That residual is what separates a genuine horizontal bundle from three lines that happened to be drawn near each other, and it is the quantity to report.

There is one degenerate case, and it is the same one this site found in the shadow field. A bundle of world-parallel lines that runs across the view gives a strong, well-conditioned vanishing point. A bundle that runs along the view gives one far off the frame, where a small angular error in any line moves the point a long way. And a bundle exactly parallel to the picture plane has its vanishing point at infinity and no fit at all. Choosing which lines in the photograph to use is therefore not a matter of convenience; it decides the conditioning of everything downstream.

The check that the recovered horizon is right is available and worth running: points marching out along the ground must approach it, and approach it at first order — a hundred times the distance must land a hundred times closer. A ground point at 10⁴ units sits 0.10 px off the horizon and one at 10⁶ sits 0.0010 px off it, a factor of 100. Being near the horizon is satisfied by any line near it; being the limit is not.

The reference does not have to be the eye

The eye height is the most convenient reference because the horizon supplies it for free, and it is not the only one.

Any object of known height in the picture will do. Run the same four-point construction on the reference and on the unknown, both with e set to 1, and divide: the ratio of the two heights comes out with no metric input at all. Multiply by the reference’s true height and the unknown is in metres.

That form is often the practical one, because a photographer’s eye height at the moment of exposure is rarely recorded and a doorway in the frame is a standard 2.0 m more often than not. It also degrades more gracefully: an error in the reference height scales every answer by the same factor, so the relative statements in a set of measurements survive a wrong reference entirely.

The two forms are the same formula and it is worth seeing why. Z = e·CR/(CR − 1) is linear in e, so every height in the picture is proportional to whatever single length is supplied, and the ratios between them do not involve it. There is exactly one degree of freedom in the whole reconstruction and any one length fixes it. That is not a property of this method; it is the scale ambiguity of single views, which has its own essay and its own demonstration.

Every ratio, and no size

The eye height is doing something specific in that formula and it is worth isolating.

Set e = 1 and the answer comes out in units of the eye height: the lamp-post is 2.1 eye heights tall. That number is available from the picture alone, with no external input whatsoever. Every ratio in the scene is likewise available — this object is 1.86 times that one, this doorway is 0.71 of that window — and all of them come out of the picture without anybody supplying anything.

What is not available is the size of any of it. That is not a limitation of this method; it is a fact about single views, and it has a demonstration rather than an argument.

Two scenes 137× apart, and the one picture they both makeEverything in the second plan — the room, the eye's distance, the eye's own height — is 137 times the first. Every projected vertex agrees to 1e-13 px. A single photograph has no scale, and this is what that means.a room 2.8 m across, eye 1.6 m up1 mthe same plan, 201.5× bigger201.5 midenticalpicturesthe picture — both scenes, drawn twice, one on top of the otherlargest disagreement 2e-13 px over 8 verticesone length has to come from outside the picture
Fig. 5 Two scenes, one 137 times the size of the other, photographed from proportionally further away. Every projected vertex agrees to within a few parts in 10¹⁴ of a pixel, which is floating-point noise. A photograph has no scale, and this is what that means.

So the eye height is not a convenience. It is the single length that turns a set of ratios into a set of measurements, and any other known length in the scene would do the same job — a door of standard height, a car of known wheelbase, a person whose height is on record. The method is a machine for propagating one length to everything else in the picture, and its accuracy is inherited from that one length before anything else.

What this is used for

The applications are all versions of the same situation: a picture exists, the scene does not, and somebody needs a number.

A security camera records a robbery and the suspect’s height is wanted. The camera’s mounting height is known from the installation, the floor is flat, and the suspect stands upright for one frame. This is the textbook case and it is the one the method was written up for in the computer-vision literature.

A photograph of a demolished building is the only surviving record of its façade. The eye height of the photographer is a guess, but the ratios are exact, so the drawing can be made in units of any one dimension that is independently known — a surviving stone, a plan, a written specification.

A road-accident photograph has to yield the distance between two marks. The ground is a plane, the camera saw the horizon, and a rectification turns the picture into a plan — which is the same machinery applied to a different plane, and its own essay.

In every one of these the structure is identical: the picture gives the ratios, one external length gives the scale, and the error in the answer is the error in the picture plus the error in that one length. Which of the two dominates is the practical question, and it has a measured answer that runs against expectation.

What one pixel of click error costs, against distanceA 1.83 m object at 3 m is measured to 0.28% per pixel; the same object at 201 m to 18.2% per pixel. The whole object shrinks toward the horizon, so a pixel buys more world.05101550100150200distance from the camera (metres)error in the recovered height, per pixel of click error (%)5 m — 0.46%25 m — 2.27%100 m — 9.04%190 m — 17.16%one pixel, on a 690 px picturelinear in distance
Fig. 6 What one pixel of click error costs, against how far away the object stands. At 6 m a pixel is worth half a per cent of the height; at 120 m it is worth eleven per cent. The whole object shrinks toward the horizon as it recedes, so a pixel buys more world.

The assumption in “the horizon is at eye level”

One step of the derivation deserves more scrutiny than it usually gets, because it is the only place a physical fact enters an otherwise purely projective argument.

The claim is that the horizon crosses a standing object at the height of the eye. It is exactly true when the picture plane is vertical, and this site has already measured what happens when it is not: twelve degrees of tilt spreads the fraction at which the horizon cuts a set of equal-height figures by 1.11 percentage points. So a tilted camera has a horizon that is still the image of the ground’s line at infinity — that part is unconditional — but the identification of the crossing point with the eye’s height acquires an error.

The cross-ratio method quietly avoids the problem, and it is worth seeing how. The four points are read off the picture, and the point at world height e is defined as the intersection of the object’s vertical with the horizon. That intersection is the image of the point on the object’s vertical line that is at the same height as the eye, whatever the camera’s tilt — because the horizon is the image of the ground plane’s infinity, and a ray from the eye to any of those points is horizontal, and a horizontal ray from the eye meets the object’s vertical at the eye’s own height. The tilt changes where the point lands on the picture and not what it is.

What the tilt does affect is the earlier, cruder version of the same measurement: the one that reads a height off the fraction of a figure the horizon cuts. That version needs a vertical picture plane and has no way to notice that it does not have one. The four-point version needs the vertical vanishing point instead, which a tilted camera has and a level camera does not — so the two methods want opposite conditions, and between them cover every camera except a set of measure zero.

Why this is not photogrammetry

There is a large, old and highly developed discipline that measures the world from photographs, and it is worth saying plainly what this is not.

Photogrammetry proper uses two or more views. With two views from known or recoverable positions, every point in the overlap gets a depth, and the scene is reconstructed in three dimensions up to one scale factor — the same scale ambiguity as here, resolved the same way, by one known length. Modern structure-from-motion takes hundreds of views and solves for everything at once.

Single-view metrology is the impoverished cousin, and its constraints are severe. It measures only in planes it can identify: the ground, a façade, a vertical. It needs the horizon, which means it needs either two horizontal vanishing points or a known reference. It cannot measure anything that does not sit on a plane the method has already established.

What it has instead is that it works on the pictures that exist. There is no second view of the robbery, of the demolished building, or of the accident before the vehicles were moved. The method’s whole value is that it takes as input the one thing that is always available and never repeatable, and the arithmetic it needs is four points and a division.