How wrong a measurement from one picture can be
A measurement without an error bar is an opinion with a decimal point. The height recovery in this field returns sixteen correct digits on synthetic data, which says the derivation is exact and says nothing at all about what happens when a person clicks on a photograph and misses by a pixel.
This essay is about that pixel. It has a surprise in it, and the surprise is not the number — it is that the argument everybody makes about where this method goes wrong, including the argument this site made while writing it, is false.
The argument that looked right
The recovery is
Look at the denominator. For an object much taller than the eye, CR = Z/(Z − e) tends to 1 from above, so CR − 1 tends to zero, and the formula divides by the difference of two nearly equal numbers. Differentiating, dZ/dCR = −e/(CR − 1)² = −Z²/e.
That is a strong statement. It says the sensitivity to an error in the cross-ratio grows as the square of the height being measured, so a thirty-metre building measured from a 1.6 m eye should be about three hundred times worse conditioned than a two-metre person.
The argument is correct as far as it goes, and the conclusion drawn from it is wrong, because it stops one step early.
What the measurement says
Measured — at a fixed distance of fourteen metres, varying only the object’s height:
| height | error per pixel of click |
|---|---|
| 1 m | 2.31% |
| 5 m | 0.48% |
| 15 m | 0.18% |
| 24 m | 0.12% |
The error goes down. A taller object at the same distance is recovered better, not worse, by a factor of nearly twenty across that range.
Why the argument failed
The missing step is that dZ/dCR is not the quantity that matters. What matters is dZ/d(pixel), and that is dZ/dCR multiplied by dCR/d(pixel).
The cross-ratio is taken over four points on the object’s vertical in the picture: the base, the horizon crossing, the top, and the vertical vanishing point. As the object grows taller at a fixed distance, its top moves away from the other three — the base and the horizon crossing stay exactly where they are, and the image separation between them and the top grows in proportion to the height.
A cross-ratio of four points is insensitive to a perturbation of one of them in proportion to how far that point is from the others. So dCR/d(pixel) falls as the object grows, and it falls faster than dZ/dCR rises.
The two effects do not merely cancel. Multiplied out, the sensitivity to a pixel falls roughly as 1/Z, which is what the table shows and what the figure draws. The formula’s apparent instability is real in CR and invisible in the thing anybody actually perturbs.
This is a specific and recurring error in reasoning about conditioning: analysing the sensitivity of an answer to an intermediate quantity and forgetting that the intermediate quantity is not what carries the noise. It is worth naming because the analysis reads as complete, and produces a confident, quotable, wrong statement about when a method can be trusted.
What actually costs
The thing that costs is distance, and it costs linearly.
| distance | error per pixel of click |
|---|---|
| 3 m | 0.28% |
| 25 m | 2.27% |
| 100 m | 9.04% |
| 200 m | 18.06% |
The reason is direct. The whole object — base, top, everything — shrinks toward the horizon as it recedes, so the image separation between the four points shrinks with it, and a pixel is a larger share of a smaller separation. Everything about the picture degrades together.
The practical rule that falls out of the curve is worth stating in its own sentence, because it is the one thing a reader should carry away. A single-view height measurement is a near-field instrument. At six metres it is good to about half a per cent per pixel of click error, which with careful marking is a centimetre on a person. At a hundred metres it is good to about nine per cent per pixel, which on a person is seventeen centimetres per pixel — more than the difference between most people.
The security-camera case that this method is famous for sits at the good end of that curve, and it is why the method has the reputation it does. A camera four metres from a doorway is measuring at four metres.
What the perturbation has to be
There is a decision buried in the sensitivity measurement that changes the answer, and it is worth surfacing because it is the sort of thing that gets made by accident.
A “one-pixel error in the clicked top” is not one thing. A pixel of error can be along the object’s vertical, or across it, or in any direction between. The three behave completely differently.
An error across the vertical is caught. The recovery takes the object’s spine to be the line through the base and the vertical vanishing point, and asserts that the clicked top lies on it. A top displaced sideways fails that assertion and the method refuses — which is the right response, because a top that is not above its own base belongs to a different object.
An error along the vertical is the one that matters, and it is what the curves here measure: the clicked point is slid up or down the object’s own spine by one pixel, and the recovered height is differenced. That is what a misplaced click on a real photograph mostly is, because a person marking the top of a lamp-post is uncertain about where the top is and not about which lamp-post they are marking.
Perturbing in a general direction would have mixed the two, and the mixture would have been dominated by whichever component happened to be larger — so the measurement would have reported a number that depended on an arbitrary choice of direction and looked like a property of the method. Choosing the direction deliberately, and having the machinery refuse the other one, is what makes the curve mean something specific.
The same care is worth applying to the base, and there it changes the accounting. A base error along the spine slides the base; a base error across it changes the spine itself, which moves the horizon crossing as well. So the base has two failure directions where the top has one, and neither is refused. That asymmetry is why the base is the harder point in practice even though the arithmetic treats the two symmetrically.
The other three inputs
A pixel on the top is one of four things that can be wrong, and the others behave differently.
The base. Perturbing the base moves two of the four points at once — the base itself and, through the recomputed vertical, the horizon crossing. In practice a base error is worse than a top error of the same size, and the base is also the point that is hardest to place: the top of an object is usually against the sky and the bottom is usually against clutter, shadow or its own footing.
The horizon. The horizon is recovered from vanishing points, and its error is not a pixel — it is a line, with a position error and an angle error. Its position error moves the horizon crossing directly and is the dominant term in most real measurements. This is the input worth spending effort on, and the one an over-determined vanishing-point fit with a reported residual is for.
The reference length. Enters as a pure multiplier, so it scales every answer by the same factor and leaves every ratio untouched. An eye height known to five per cent gives absolute heights known to five per cent and relative heights known to the picture’s own accuracy — which is why the honest way to report a set of these measurements is with two error figures rather than one.
The vertical vanishing point is not on the list, which is the next surprise.
The vanishing point that does not matter
The obvious remaining worry is the fourth point. As the camera comes level the vertical vanishing point runs off to infinity, and a cross-ratio taken over a point in the billions should collapse into noise.
It does not.
At a downward tilt of one part in eight million, with the vertical vanishing point 7.2 × 10⁹ pixels below the frame, the recovered height is still correct to the last bit.
The reason is in how the cross-ratio is computed. This site’s crossRatio parameterises the four points along the line rather than by a coordinate — a decision originally made so that the exactly vertical lines of a two-point construction would not divide by a difference in x. The consequence here is that the enormous coordinate enters as a ratio of two enormous numbers and cancels cleanly, instead of being subtracted from a small one.
So the failure is not a slope, it is a cliff. Every camera with any tilt at all works, to full precision. A camera with exactly zero tilt has no vertical vanishing point — the direction is parallel to the picture plane and its vanishing point is at infinity in the image, which is to say it does not have one — and the method refuses outright.
And the one camera it refuses is the one the elementary method handles exactly: with a vertical picture plane, the horizon cuts a standing figure at a fraction of its height equal to eyeHeight/Z, so the height comes out of a single division with no vanishing point at all. Two routes, each exact where the other has nothing to say, and the boundary between them is a single camera out of all possible cameras.
What an error bar should say
Putting all of it together, a responsible report of a single-view height has three numbers in it and most published ones have one.
The first is the height, in the units the reference was given in.
The second is the picture-geometry uncertainty: the sensitivity per pixel, multiplied by an honest estimate of how well the points can be placed. That is the number the curve in this essay supplies, and it is dominated by distance.
The third is the reference uncertainty, quoted separately and never combined with the second, because the two behave differently. The reference scales everything; the geometry does not. A reader who wants to compare two objects in the same photograph should use the second and ignore the first entirely.
Reporting them separately also makes it obvious when a measurement is worth improving. A height at four metres from a well-placed set of clicks might carry half a per cent of geometry error and five per cent of reference error, and the only useful thing to do with it is go and measure the camera’s mounting height properly.
The errors this analysis does not contain
Everything above is the error a correct method makes on correct inputs that are slightly misplaced. There is a second category, and in real work it is usually larger.
The object is not vertical. A lamp-post leaning two degrees, a person standing with a hip out, a tree. The method assumes the object’s two ends lie on one world vertical, and a lean puts the top somewhere that is not above the base — which the spine assertion catches only if the lean is across the view, and misses entirely if it is toward or away from the camera. A two-degree lean along the line of sight on a five-metre object displaces the top by 17 cm and nothing complains.
The ground under the object is not the ground under the camera. The eye height that the horizon supplies is the height above the plane the horizon belongs to. An object standing on a kerb 15 cm up is being measured from the wrong datum, and 15 cm is a tenth of a person.
The lens is not a pinhole. Barrel distortion bends the object’s own vertical, so the four points are not collinear in the picture and the spine assertion fires or does not depending on how big the effect is. Near the frame edge on an uncorrected wide lens it is several pixels, which is several times the click error this whole essay is about.
The frame is not the moment. For anything moving, the top and the base are marked on one image and the object was in motion during the exposure. A walking person’s head moves several centimetres in a fiftieth of a second.
None of these is a defect of the derivation, and none of them shows up in a per-pixel sensitivity. They are the reason a measurement good to half a per cent on synthetic data is quoted at a few per cent in a report, and a sensitivity curve presented without them invites the reader to believe the wrong number.
The relationship between the two categories is the useful thing to hold on to. The per-pixel curve says how much the picture can be trusted. The list above says how much the situation can be. A measurement is limited by whichever is worse, and it is almost always the second.
The habit this essay is an instance of
The site’s rule is that every claim gets a test it could fail, and the version of that rule this essay demonstrates is narrower and more uncomfortable: an argument about conditioning is a claim, and it needs a test like any other.
The Z²/e sensitivity was derived correctly, stated confidently, and would have gone into an essay as a warning about measuring tall buildings. It survives the derivation and does not survive the measurement, because it is a sensitivity to a quantity nobody perturbs.
What made the difference was writing the measurement before writing the paragraph. The curve was plotted against height because the argument said it should be interesting, and the curve went the other way. That is the cheapest kind of finding to obtain and the easiest to skip, and the number of confident statements about numerical stability in circulation suggests it is skipped often.
Both curves are kept, and that is deliberate. The site’s habit with a failed check is not to delete it — the cross-ratio test that gives the by-eye depth method a perfect score is still in the source, still asserted to pass, and the working version is asserted to reject what it accepts. The same arrangement is here: the sensitivity against height is plotted and shipped, because a reader who has followed the Z²/e argument will want to see it fail rather than be told it did.
It is the third time this site has caught itself the same way, and the three are worth putting side by side because the pattern is more useful than any of them. A cross-ratio test evaluated on four points that could not fail it. An angle test evaluated along a surface’s own axes, where the cylinder cannot fail it. And a conditioning argument evaluated in a variable nobody perturbs. In every case the analysis is correct and the input is chosen — usually by the same machinery that produced the thing being checked, and never by anyone deciding it should be.
The remedy in all three was the same and it is not sophistication. It is to run the check against something it ought to reject, and see whether it does.