Measuring from one picture

Which reference to measure from

Given four candidate scale bars in one photograph, the best is not the longest and not the nearest — it is the longest in the picture. A five-point-two metre bar near the horizon is the longest thing in the scene and the worst reference in it; a four metre bar close to the camera is the best. Walk one bar outward and the term it controls falls as one over its length in pixels, with a fitted exponent of −1.08.

Worth reading first: The answer is an ellipse · The one thing a single view cannot give.

The one thing a single view cannot give establishes that a photograph supplies every ratio and no size, so one length has to come from outside. A real photograph usually offers several candidates, and the field has never asked which to use.

The intuitions are strong and two of them are wrong.

Four candidates

Put four possible references on one floor: a short bar close to the camera, a long bar close to the camera, a bar of moderate length at the same depth as the thing being measured, and the longest bar in the scene lying near the horizon. Use each in turn to give the same unknown length its metre, and score each by the spread of the answer over a hundred and sixty independent noise draws.

The best reference is 4.0 m; the longest is 5.2 m and the worstFour candidate references in one photograph, each used in turn to give the same unknown length its metre, with the spread over a hundred and sixty noise draws as the score. Shorter bars are better. The five-metre bar lying near the horizon is the longest thing in the scene and the worst reference in it, because what the arithmetic uses is its length in the picture and depth has taken that away. The rule is not "use the longest" and not "use the nearest": it is use the one that is longest where you are actually reading it, which is on the print.near and short3.68%0.5 m, 126 pxnear and long2.84%4.0 m, 1088 pxat the unknown's depth3.95%1.4 m, 154 pxfar and long8.36%5.2 m, 241 pxthe spread of the answer, per candidateshorter is better
Fig. 1 Four candidates, one unknown, one hundred and sixty draws each. Shorter bars are better. The longest bar in the scene is the worst reference in it.

The four-metre bar near the camera wins at 2.84 per cent. The five-point-two-metre bar near the horizon loses at 8.36, which is nearly three times worse despite being thirty per cent longer in the world.

What the arithmetic actually uses

The reason is short. A reference gives the answer its metre through the ratio of two measured lengths — the unknown’s and the reference’s — and both are measured after the marks have been read. So the reference’s contribution to the relative error is the relative error of its own measured length, which is the marking error divided by how long it is in the picture.

Image length is world length times focal over depth. A bar four metres long at three and a half metres deep is 1088 pixels; a bar five point two metres long at fifteen metres deep is 241. The second is a quarter the instrument the first is, and the world lengths are the wrong quantity to compare.

The reference's own term falls as one over its length in pixelsOne reference of fixed length in the world, walked from close to the camera out toward the horizon. Its length in the picture falls as one over its depth, and the spread it puts into the answer falls with it. The lower curve isolates the reference's contribution by reading the unknown's own marks exactly; its fitted exponent is one over the image length, not over the world length and not over the depth. The upper curve is what a reader actually gets, and it flattens at the left because the unknown's own marks are still being read to the same precision — which is the same floor the closure sweep finds from the other side.00.5001100200300the reference's length in the picture, in pixelsthe spread of the measurement, in per centwhat a reader getsthe reference's own termone bar, walked outwardlongest in the picture wins
Fig. 2 One bar of fixed world length walked from close to the camera out toward the horizon. Its contribution to the answer falls as one over its length in pixels, with a fitted exponent of −1.08.

The exponent is fitted rather than asserted, over six positions spanning a sixfold change in image length, and it comes out at −1.079 against a predicted −1.

The two curves, and the floor between them

The walk above is drawn twice, and the difference between the curves is the essay’s second finding.

The lower curve isolates the reference’s contribution by reading the unknown’s own marks exactly and jittering only the reference. That is the term the choice of reference controls, and it obeys the one-over-image-length law cleanly.

The upper curve is what a reader actually gets, with both read imperfectly. It flattens at the left, because past a certain point the reference is no longer the limiting term — the unknown’s own marks are — and improving the reference further buys nothing.

That is the same structure the next rung finds from the other side: a total made of a term a reader can improve and a term a reader cannot, with a crossing between them. Recognising which side of the crossing a measurement is on is most of the value of computing either.

Better marks stop helping at 6.0%, which is the assumption's own errorA length running into the picture, measured with the marks read to the precision on the horizontal axis, twice over. The lower curve has the reference's shape exactly right, and it keeps falling: better marks keep buying a better answer, without limit. The upper curve has the reference's aspect wrong by 6 per cent, and it stops — at 6.0 per cent, which is the assumption's own error and nothing else. The crossing between them is where a reader should stop buying lenses, and it can be computed before the photograph is taken.0102030401234the marking error, in pixelsthe error of the answer, in per centupper curve: the closure wrong · lower: the closure exactfloor 6.0%
Fig. 3 The same structure in the closure sweep of the next essay: a falling term, a fixed term, and a crossing where the first stops mattering.

The two intuitions that fail

“Use the longest reference.” False, and the four-candidate comparison is the counterexample. Length in the world is only useful in proportion to how much of it survives into the picture, and depth takes it away as one over the depth.

“Use a reference at the same depth as the unknown.” Plausible — it sounds like an argument about interpolating rather than extrapolating — and it comes third of four here. The bar at the unknown’s depth is 154 pixels long and scores 3.95 per cent, behind the short near bar’s 126 pixels and 3.68 and well behind the long near bar.

That second failure is worth dwelling on because the intuition behind it is not silly. Extrapolating a rectifier a long way beyond its own marks is worse than interpolating, and that effect is real. It is simply much smaller than the image-length effect, and a rule that trades a fourfold loss of image length for a modest gain in interpolation is trading the wrong way.

The reference nobody chose

Before the four candidates, it is worth naming the reference a great many measurements actually use, which is none of them.

A photograph with no scale bar in it — and the plan hidden in the photograph is built on one — is usually measured by assuming something: that a door is two metres, that a car is four and a half, that a person is one metre seventy-five. Those are references too, and they belong in the comparison on exactly the same footing — an assumed length is a supplied length whose supplier is a population average rather than a tape.

What distinguishes them is not the geometry but the bias. A tape gives a length that is right; an assumed height gives a length that is right on average and wrong for the individual, by whatever the population’s spread is. A person assumed to be 1.75 metres is out by five centimetres or so at one standard deviation, which is three per cent, and three per cent is larger than every spread in the comparison above.

So on a photograph with a real scale bar and an assumed person, the bar wins by a large margin whatever their image lengths — because the person’s contribution is not a spread at all but a bias, and biases do not average away. How wrong a measurement can be prices the same effect for a height recovered from one picture, and the sensitivity there is dominated by the same term.

That comparison is the one place where the essay’s rule is overridden, and stating the override is part of the rule.

Five facts that close the same gap, 10.1× apartA photograph gives every ratio in a scene and no size, so one fact has to come from outside it — and "one fact" is not one option. Five are run here on the same picture, the same unknown length and the same noise draws: a length lying on the ground, the camera's own height, a repeated object of unstated size, a standing object of known height, and the focal length with the horizon. All five give an answer. The spread of those answers runs from 0.36 to 3.60 per cent, a factor of 10.1, and which is available depends on the photograph rather than on the arithmetic.a length on the ground3.22%1.00 m across the referencethe camera's height0.36%1.62 m above the grounda repeated object, size unknown3.60%the answer in units of the repeata standing object of known height0.41%1.75 m, upright, anywhere on the …the focal length and the horizon1.94%700 px, and where the ground's li…the spread of the answer, per closureshorter is better
Fig. 4 The five facts that could close the same gap, on one picture. A photograph gives every ratio in a scene and no size, so one fact has to come from outside it — and the five available here span a factor of 10.1 in what they cost. Choosing a reference is one row of this catalogue; the essay’s question is which row, and this is the table it is a row of.

A bias and a spread are not compared the same way twice

The override at the end of the last section — a real bar beats an assumed person because a bias is not a spread — is right for one measurement and understates the case for many. The two behave differently under repetition, and how a reader is going to use the answer changes the ranking.

For a single length the two combine in quadrature and are nearly matched. A scale bar contributing 2.84 per cent of spread and no bias sits beside an assumed 1.75 m person contributing perhaps 1.2 per cent of spread and 3 per cent of bias: 2.842=2.84\sqrt{2.84^{2}}=2.84 against 1.22+32=3.23\sqrt{1.2^{2}+3^{2}}=3.23. A thirteen per cent difference, which is not the large margin the section above claims.

For many lengths off the same photograph the two separate completely. Every measurement made through one assumed height carries the same wrong factor, so averaging ten of them leaves the three per cent exactly where it was while the bar’s spread falls to 0.90. After ten measurements the bar is three times better; after a hundred, ten times.

And the fix follows from the same arithmetic. The bias is common because the reference is common, so using four different people is worth far more than measuring one person four times: four independent draws from the population’s spread average to 1.5 per cent, while four readings of one person’s marks average to nothing at all. That is the one case in this essay where the advice is to choose a worse reference — a person further away, more obliquely seen — in preference to reading a better one again, and it is worth stating because every instinct about measurement says the opposite.

The general shape is the one the rung before this reaches from the other side: a term a reader can improve and a term a reader cannot, with the crossing between them deciding what effort is worth spending. Here the crossing moves with the number of measurements, which is the only variable in this field that a reader controls for free.

What survives the rectifier

There is a subtlety in the arrangement above that changes the answer’s shape and is worth stating explicitly.

The rectifying homography is built from four marks of a reference quad, which fixes the plane’s shape and orientation. The candidate bars then supply only the scale, as a ratio. So the experiment separates the two closures an angle on the ground distinguishes: shape from the quad, size from the bar, and only the second is varied.

That separation is why the result is clean. Varying both at once would confound a shape error with a scale error, and a bar that happened to lie along a well-determined direction would score well for the wrong reason.

In a photograph where the same object supplies both — a paving slab whose size and squareness are both known — the two contributions are correlated, and the honest treatment is to propagate them together. The result then still favours the object with the largest image, and for the same reason.

A façade, flattened out of the photographFour corners of a rectangle of known proportions fix the homography. Three lengths it was never given come back to 4e-16 relative — so every measurement on that plane is available, in units of the rectangle's own width.the picturethe same plane, rectifieddoor — 0.5068 widths (true 0.5068)sill — 0.3026 widths (true 0.3026)diagonal — 1.1766 widths (true 1.1766)worst error 4e-16 relativethe probes were not used to build the map
Fig. 5 The map whose scale the reference is setting, from the essay that built it: the plane flattened, with shape and size supplied separately.

Orientation, which the walk holds fixed

The walk above moves a reference in depth and holds its orientation fixed, running across the picture. Orientation matters too and it matters in the way the answer is an ellipse predicts.

A bar lying across the picture is measured between two ground points whose uncertainty ellipses are short in that direction, so its own measured length is well determined. A bar lying into the picture is measured along the long axes of both ellipses, and its measured length is worse by the depth-over-height factor — five to one at eight metres, in the arrangement that essay uses.

So the rule has a second clause: the best reference is the one longest in the picture and lying across it rather than into it. A bar running away from the camera is short in the picture and badly measured along its own direction, which is the same penalty applied twice.

That also explains why a person is a poor scale bar and a doorway is a good one. A standing person is a vertical, which images short and is read along a direction with no grazing penalty; a doorway’s width lies across the picture at whatever depth it is at. Neither is as good as a metre rule laid across the foreground, which is what a photographer with a tape does without being told.

Several references at once

A photograph with four candidates should not use one of them. It should use all four, and the weighting is the part worth getting right.

Each candidate gives an estimate of the scale with its own variance, and the best combination weights each by the inverse of its variance — which by the law above means weighting each by the square of its image length. A four-metre bar at 1088 pixels then carries about twenty times the weight of a five-metre bar at 241, which is close enough to using the first alone that the combination’s advantage is small.

That is the usual shape of an inverse-variance combination when one member dominates, and it is why “choose the best” is a reasonable approximation to “combine them properly” here. It stops being reasonable when the candidates are comparable, and then the combination is worth the arithmetic.

The one thing a combination does buy in every case is a check. Four references that disagree by more than their spreads allow are four references of which at least one is not the length it is claimed to be — and that is the failure this whole field is least able to detect from a single measurement.

The instruction, and what it is worth

Collecting the essay into an instruction: put a long reference across the near foreground.

Long, because the contribution falls as one over its image length. Across, because a length measured into the picture is read along the long axes of two uncertainty ellipses. Near, because image length is world length times focal over depth and depth is the term a reader controls by walking.

What it is worth is a factor of about three between the best and worst candidate in a scene that offers four, which is the same as a threefold improvement in the marking precision — and marking precision is expensive while walking two metres closer is free. That comparison is the reason the essay exists: the choice of reference is one of the very few free improvements available in this field, and it is usually made by picking whatever object has a known size.

The same reasoning gives the negative instruction, which is more often useful. Do not close on something near the horizon, however large it is. A building’s known width at two hundred metres is a shorter instrument than a metre rule at three, and using it because it is precisely known confuses a well-known length with a well-measured one.

The boundary, stated

One plane, marks read with independent errors, and references whose stated lengths are correct.

The last is the assumption the whole essay rests on and cannot test. Every number here is a spread — how much the answer moves under reading noise — and a reference whose stated length is simply wrong produces a bias that no amount of averaging removes and that the spread does not report. That is the next rung’s subject entirely.

The independence assumption bites here in the way the answer is an ellipse sets out: marks read by one reader on one print share a bias, and shared biases partly cancel in a ratio. So the spreads quoted are, if anything, conservative for a reader measuring everything the same way.

And the arrangement holds the camera fixed. A reader free to choose where to photograph from has a better move available than choosing a reference — get closer, which lengthens every candidate in the picture at once.

Why this was never asked before

It is worth a paragraph to say why a field that has measured heights, façades, plans, balls and shadows has never asked which reference to close with, because the reason is structural rather than accidental.

Every previous essay in the field studies one arrangement with one stated reference, so there was no choice to make. A height, out of one photograph is given a reference height and asks what the recovery costs. Flattening a façade is given four marks and asks what the rectifier does. In each, the closure is part of the problem statement.

A real photograph is not a problem statement. It contains whatever it contains, and the reader’s first decision — before any construction — is which of the available objects to trust and which to measure. That decision is upstream of everything the field has priced, and it turns out to be worth as much as a threefold change in the marking precision.

Naming the gap is the point of the rung. There is one more decision of the same kind upstream of this one, which is what kind of fact to close with — a length, a camera height, a repeated object — and it is the last essay of this ladder.

What is measured here

Four candidates and a law.

The four-metre bar near the camera, 1088 pixels long, leaves a spread of 2.84 per cent. The short near bar, 126 pixels, leaves 3.68. The moderate bar at the unknown’s own depth, 154 pixels, leaves 3.95. The longest bar in the scene, 5.2 metres and 241 pixels, leaves 8.36 — nearly three times the best, while being the longest thing there.

And walking one bar of fixed world length from close to the camera out toward the horizon, the term it controls falls with a fitted exponent of −1.079 against a predicted −1 over a sixfold change in image length, while the total flattens at the left where the unknown’s own marks take over.

The short version

The reference that closes a measurement should be the one longest in the picture, not the one longest in the world and not the one nearest the thing being measured. Its contribution to the answer falls as one over its image length, fitted at −1.08, and image length is world length times focal over depth.

A five-metre bar near the horizon is the longest object in a scene and the worst reference in it. A four-metre bar in the foreground is three times better. And past a certain point improving the reference buys nothing, because the unknown’s own marks become the limiting term — which is the crossing the next rung measures from the other side.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasConditioningDepth uncertaintyerror propagationHomographyLeverageRectificationReference lengthSensitivitysingle-view metrology