Measuring from one picture

The one thing a single view cannot give

Make the world a hundred and thirty-seven times larger and move the eye a hundred and thirty-seven times further away, and the picture does not change by a measurable amount. Every ratio in a scene is recoverable from one photograph and no size is, and that is not a caveat about the method — it is the shape of the method.

Everything in this field returns a ratio. A height comes out as a multiple of the eye height. A length on a rectified plane comes out in units of the reference rectangle’s width. A plan comes out in units of whatever the control points were given in. Each essay states it and moves on, and stating it repeatedly is not the same as demonstrating it.

The demonstration is short and it settles the matter completely.

Two scenes 137× apart, and the one picture they both makeEverything in the second plan — the room, the eye's distance, the eye's own height — is 137 times the first. Every projected vertex agrees to 1e-13 px. A single photograph has no scale, and this is what that means.a room 2.8 m across, eye 1.6 m up1 mthe same plan, 201.5× bigger201.5 midenticalpicturesthe picture — both scenes, drawn twice, one on top of the otherlargest disagreement 2e-13 px over 8 verticesone length has to come from outside the picture
Fig. 1 Two plans, at their own scales, with the scale bars that make them incommensurable: one room 2.8 m across with the eye 1.6 m up, and the same plan 137 times larger with the eye 137 times further away. Below them, the one picture both produce. Every projected vertex agrees to a few parts in 10¹⁴ of a pixel, which is floating-point noise.

What the demonstration actually does

Take a scene: a set of world points. Take a camera: an eye, a target, a field of view, an image size. Project. Keep the results.

Now multiply every world point by 137. Multiply the eye by 137 and the target by 137, leaving the field of view and the image size alone. Project again.

The two sets of image points are the same. Not similar, not close — the same, to the last bits the arithmetic has, which for the eight vertices of the box in the figure is a disagreement of a few parts in 10¹⁴ of a pixel.

The reason is one line of the projection. A world point p images at

(px+fcxcz, py+fcycz)\left( p_x + f\,\frac{c_x}{c_z},\ p_y + f\,\frac{c_y}{c_z} \right)

where c is the point in camera coordinates and f is the focal length in pixels. Scaling the scene and the eye together multiplies every component of c by the same factor, and the factor cancels in cx/cz and in cy/cz. The focal length is in pixels, so it is untouched. The image is unchanged.

That is the whole argument, and it is worth having it as a computation rather than as an algebraic remark, because the algebraic remark is easy to nod at and hard to feel. Every point in the second scene has moved by hundreds of metres. Not one pixel has moved.

Four points on a line, before and after a projectionLength and the ratio of lengths do not survive the projection; the cross-ratio does, agreeing to 0e+0 relative.horizonABCDon the groundin the picturelength AB1.00011.3930ratio AB:CD0.56670.6837cross-ratio1.31681.3168correct from 26 cm, at 160 mm wide34° across
Fig. 2 What a projection destroys and what it does not. Length, the ratio of lengths and the midpoint all go; the cross-ratio comes through untouched. Scale is on the first list, and one number restores it.

What is recovered instead

The scale ambiguity is usually written down as a limitation, and writing it that way misses that it is a statement of exactly what is recovered.

If two scenes produce the same picture precisely when they differ by a uniform scaling — and they do; the demonstration is one direction and the converse is the standard result — then the picture determines the scene up to scale, and up to scale is a precise thing. It means every ratio of lengths, every angle, every shape, the whole geometry of the scene except one number — which is the positive half of what a projection destroys.

So the correct description of what a single view gives up is not “a photograph loses information about size”. It is: a photograph of a scene is a photograph of every scene similar to it, and of nothing else. The set of scenes consistent with one picture is a one-parameter family, and the parameter is a single scalar.

That is an unusually clean state of ignorance. Most measurement problems leave a fog; this one leaves exactly one unknown number, and one measurement of any length anywhere in the scene fixes it forever.

A 3.4 m object measured from one picture, 11 m awayThe base, the horizon crossing, the top and the vertical vanishing point have a cross-ratio of 1.9101. With the eye at 1.62 m that gives 3.400 m, against a true 3.4 m. The camera is not consulted.horizon — the eye's own heightbase — 0 mhorizon crossing — 1.62 mtop — 3.40 m recoveredthe vertical vanishing point is 8586 px above this framerecovered 3.400 m · true 3.400 m3.5 cm per pixel of click error
Fig. 3 The eye height entering the formula Z = e·CR/(CR − 1) is exactly that one number. Set e = 1 and the answer is in eye heights; set e = 1.62 m and it is in metres. Every height in the picture is proportional to it and no ratio between two of them involves it at all.

The miniature, and why it works

The scale ambiguity has a five-hundred-year-old industry attached to it and the industry is a better argument than the algebra.

A model set photographed with a suitably placed camera is indistinguishable from the full-size scene, and film relied on that for most of a century. What the demonstration above says is that this is not a trick or an approximation: a model at one fortieth scale, photographed from one fortieth of the distance, produces the same picture, exactly, and no amount of scrutiny of that picture recovers the difference.

The places where model shots do fail are worth listing, because every one of them is a failure of something other than projection.

Depth of field. A lens focused at 0.5 m has a much shallower depth of field at the same aperture than one focused at 20 m, so a model shot has a characteristic softness front and back. This is optics, not geometry, and it is cured by stopping down.

Motion. Gravity does not scale. A model building collapsing at one fortieth scale falls in one fortieth of the distance but not one fortieth of the time, so it appears to fall too fast — the standard remedy is to shoot at high frame rate and slow the result by the square root of the scale factor.

Texture and the atmosphere. Grain size, surface finish and haze all have absolute scales, and none of them shrink with the model.

Not one of those is a projective failure. The geometry of a miniature is exactly the geometry of the real thing, and the whole miniature industry is an existence proof of the claim this essay opens with. What gives a model away is never the perspective.

The same fact runs the other way, and photographers use it deliberately. Photograph a real city from a high window with a very shallow depth of field and it reads as a model, because the only cue that was ever distinguishing the two was the optics — the geometry never was.

Where the reference length actually comes from

In practice the one number arrives from one of a small number of places, and the choice matters more than it looks.

The camera’s own geometry. A security camera’s mounting height is in the installation record; a dashcam’s is a property of the vehicle. This is the most common source and it is usually the best one, because it is a fixed quantity that was measured once with a tape.

An object of standard size in the frame. A door leaf, a paving slab, a road marking, a vehicle’s wheelbase. These are specified rather than measured, so they carry a manufacturing tolerance rather than a measurement error — which is often smaller, and is occasionally catastrophic when the object turns out not to be standard.

A subsequent measurement of the scene. Go back with a tape and measure anything visible in the photograph. This is the only source that is a genuine measurement, and it is available exactly when the scene still exists, which is exactly when the photograph was least needed.

A person of known height. Common and worse than it looks, because a person’s height in a photograph is their height including shoes and posture, and because the record of it is usually self-reported.

The error in the answer is the error in the picture geometry plus the error in this one number, and the second term is very often the larger. A height recovered from a well-conditioned photograph to half a per cent, scaled by an eye height known to five, is known to five.

The ratios do not care

There is a compensation, and it is the practical reason this field survives its own limitation.

Because the reference enters as a pure multiplicative factor, every relative statement in a set of measurements is immune to it. If the reference is wrong by ten per cent, every recovered height is wrong by ten per cent, and the statement “this object is 1.86 times that one” is exactly right.

So a set of measurements from one photograph has two very different accuracy figures, and the honest way to report it is to give both. The absolute heights are known to whatever the reference is known to. The ratios are known to the picture geometry alone, which is often two orders of magnitude better.

Questions that turn on a ratio can therefore be answered confidently from a photograph whose absolute scale is a guess. Is this person taller than that one? Is this vehicle longer than that bay? Did this object clear that opening? All of them are ratio questions, all of them are answerable, and none of them needs the reference at all.

A façade, flattened out of the photographFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 4e-16 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifieddoor — 0.5068 widths (true 0.5068)sill — 0.3026 widths (true 0.3026)diagonal — 1.1766 widths (true 1.1766)worst error 4e-16 relativethe probes were not used to build the map
Fig. 4 The same structure on a plane. A rectification returns every length in units of the reference rectangle’s width, so the reference is a single multiplier on everything, and every proportion measured on the rectified plane is free of it.
A parallel projection is a perspective one with the eye taken awayThe four small boxes are photographs from 2.5 m, 6 m, 20 m and 200 m with the lens lengthened to match. The isometric drawing is the limit, and its bundles stay parallel to 0e+0 radians.isometric — the limit2.5 m6 m20 m200 msame box, same drawn sizethe eye recedes
Fig. 5 The other end of the same family. Taking the eye away entirely gives a parallel projection, which has no station point and no scale either — the ambiguity survives every limit of the construction.

Two views, and what changes

The field this belongs to has a well-known escape and it is worth stating what the escape does and does not do.

With two views from different positions, the scene is reconstructed in three dimensions — every point gets a depth, not just points on identified planes — and the reconstruction is again determined up to one scale factor. The ambiguity does not go away. What goes away is the restriction to planes: a second view buys depth for everything, and it costs the same one unknown number as before.

The scale is fixed in a two-view reconstruction by knowing the baseline, the distance between the two camera positions. That is the same kind of external input as the eye height here, applied at a different point. A stereo rig with a measured baseline is a camera that knows its own scale; two photographs taken from unrecorded positions are not.

So the escape from planarity is real and the escape from scale is not. Every reconstruction from images, of any number of views, of any sophistication, is determined up to a similarity — and every one of them imports its scale from a ruler somewhere.

What does not scale, and why it is not an escape

A reader looking for a way out of the ambiguity will land on the same idea everyone lands on: surely something in a photograph betrays the size of what it shows. It is worth going through the candidates, because each one fails in a way that says something about where the boundary is.

Depth of field. A real lens has one, and it depends on the absolute distance to the subject. So in principle the blur gradient in a photograph carries scale information — and in practice it is the cue that gives away a miniature. But it is a property of the lens, not of the projection, and it vanishes entirely for an ideal pinhole. This site’s camera is an ideal pinhole, which is why the demonstration is exact.

Atmospheric haze. Distant things are paler and bluer, by an amount that depends on absolute distance through absolute air. Landscape painters have used it for centuries and it does carry real scale information. It is a property of the medium, not the geometry, and it says nothing in a photograph taken indoors or in clear air over a short distance.

Diffraction and grain. Both have absolute scales and both limit resolution, so a very small object photographed very close is softer than the geometry predicts. Again optics, again absent from a pinhole.

Known objects. A doorway, a person, a brick. This is the one that actually works and it is not an escape from the ambiguity — it is the reference length, arriving under a different name. Recognising a brick is supplying the number from outside the picture.

The pattern is consistent and it is the point. Every cue that carries scale is a cue about the apparatus or the medium or the contents, and none is a cue about the projection. The projection is scale-free, exactly, and everything that appears to escape that is information smuggled in from somewhere else.

Why this is the site’s own premise, turned round

This site’s central claim is that a picture determines the point it is correct from: the focal length scaled to the display width is the distance the reader must stand at, and no other treatment of the subject computes it.

The scale ambiguity is that claim in a mirror. What the picture determines is a direction structure — which rays went where — and directions are scale-free. The viewing distance falls out because the display width is a real length supplied from outside the picture; without it, the picture would determine only that the reader must stand at a distance equal to the focal length in units of the picture’s own width, which is a ratio.

Both statements are the same fact seen from either end. A picture knows all the angles and none of the lengths. Give it one length and everything follows: the display width gives the viewing distance, and the eye height gives every height in the scene. The picture supplies the structure and the world supplies exactly one number, once.

What one pixel of click error costs, against distanceA 1.83 m object at 3 m is measured to 0.28% per pixel; the same object at 201 m to 18.2% per pixel. The whole object shrinks toward the horizon, so a pixel buys more world.05101550100150200distance from the camera (metres)error in the recovered height, per pixel of click error (%)5 m — 0.46%25 m — 2.27%100 m — 9.04%190 m — 17.16%one pixel, on a 690 px picturelinear in distance
Fig. 6 The other half of an error budget. Even with a perfect reference the picture contributes its own error, and it grows with distance: half a per cent per pixel of click error at six metres, eleven per cent at a hundred and twenty.

One number, and where it is spent

It is worth counting the degrees of freedom, because the accounting is unusually tidy and it makes clear that the scale ambiguity is not one problem among several.

A scene has, in principle, an enormous number of parameters: the position of every point. A picture of it constrains all of them except a similarity — a rotation, a translation and a scale — since a rotated, moved or resized scene photographed from a correspondingly rotated, moved or resized viewpoint gives the identical picture.

The rotation and the translation are not usually felt as ambiguities, because nobody expects a photograph to say where in the world its subject is or which way it faces. They are ambiguities all the same, and they are resolved the same way: by a coordinate frame supplied from outside.

The scale is felt, because size is a property of an object rather than of its placement, and it is the one similarity parameter people expect a picture to carry. It does not, and the demonstration above is the reason.

So a single view leaves exactly seven numbers unfixed — three of rotation, three of translation, one of scale — and every reconstruction from images, of any number of views, leaves the same seven. What more views buy is depth for every point rather than depth for points on identified planes, and that is a large gain. It is not a gain against this list.

The check, and why it is a check

The measurement in the hero figure is arranged so that it can fail, which is the only reason it is worth printing.

It would be easy to write a demonstration that could not fail. Scale the scene, scale the eye, and observe that the numbers cancel — that is an algebraic identity dressed up as an experiment, and any implementation of a projection satisfies it whether or not the projection is correct.

What makes this one a test is that the two projections go through the camera-building code independently. The second camera is constructed with a scaled eye and a scaled target, so it recomputes its own basis, its own focal length in pixels, its own principal point. If any of those quantities carried a length that did not scale — a hard-coded distance, a near plane, a clamp in world units — the two pictures would differ, and the difference would be exactly the kind of quiet bug that produces plausible pictures.

They do not differ. And the assertion is written with a tolerance in pixels rather than a comparison to zero, which is the honest form: floating-point arithmetic on numbers a hundred and thirty-seven times apart does not produce identical results, it produces results that differ by rounding, and a check demanding exact equality would be a check demanding that the arithmetic be something it is not.