Measuring from one picture

Five facts that close the same gap

The gloss that one length has to come from outside the photograph names a single option, and there are at least five — a length on the ground, the camera's own height, a repeated object of unstated size, a standing object of known height, and the focal length with the horizon. All five give an answer on the same picture with the same noise; their spreads run from 0.36 to 3.60 per cent, a factor of ten, and which one is available depends on the photograph rather than on the arithmetic.

Worth reading first: The one thing a single view cannot give · The answer is an ellipse.

The one thing a single view cannot give is the fact this field is built on: scale the world and the eye together and not a pixel moves, so a photograph carries every ratio and no size. The standard gloss is that one length has to come from outside.

That gloss names one option, and there are at least five. This essay runs all of them on one picture and finds that they are not equally good — a factor of ten between the best and the worst, on the same marks with the same noise.

The five

A length on the ground. A scale bar, a paving slab, a known span between two marks lying in the plane being measured. The rectifier’s four marks are read, the map is built, and the answer is in metres.

The camera’s own height. Knowing how far the lens was above the floor fixes the plane’s distance from the eye, so a ray through a mark meets it at a determined point. No object in the scene is needed at all — the closure is a fact about the photographer rather than about the photograph.

A repeated object of unstated size. Two marks a known number of repeats apart give every length in units of the repeat. That is a ratio closure: the answer is metric up to one factor, which is the repeat’s own size, and is complete for anything that only wants proportions.

A standing object of known height. A person, a door, a sign of standard dimensions. The horizon plus one upright of stated height gives the camera’s height, which then gives the ground its metre. This is a height from one photograph run backwards.

The focal length with the horizon. A calibrated camera and a known vanishing line fix the plane’s orientation, which upgrades the shape completely and still borrows one length for the size — so it is a partial closure that has to be finished by one of the others.

Five facts that close the same gap, 10.1× apartA photograph gives every ratio in a scene and no size, so one fact has to come from outside it — and "one fact" is not one option. Five are run here on the same picture, the same unknown length and the same noise draws: a length lying on the ground, the camera's own height, a repeated object of unstated size, a standing object of known height, and the focal length with the horizon. All five give an answer. The spread of those answers runs from 0.36 to 3.60 per cent, a factor of 10.1, and which is available depends on the photograph rather than on the arithmetic.a length on the ground3.22%1.00 m across the referencethe camera's height0.36%1.62 m above the grounda repeated object, size unknown3.60%the answer in units of the repeata standing object of known height0.41%1.75 m, upright, anywhere on the …the focal length and the horizon1.94%700 px, and where the ground's li…the spread of the answer, per closureshorter is better
Fig. 1 Five closures, one picture, one unknown, one hundred and sixty noise draws each. Shorter bars are better, and the range is a factor of ten.

The ranking, and why it is not obvious

The camera’s height wins at 0.36 per cent. The standing object of known height is second at 0.41. The focal length and horizon is third at 1.94. A length on the ground is fourth at 3.22, and the repeated object last at 3.60.

The reason the ground length loses to the camera height is worth stating, because it is the opposite of what a reader would expect from a field whose central slogan is about supplying a length.

A ground length is used through a rectifier built from four marks, and all four marks are read imperfectly. The camera’s height is used through a direct unprojection of two marks, and only two marks are read. Fewer marks, fewer errors, and the closure itself contributes no reading error at all because it is a tape measurement rather than a picture measurement.

That is a general point and it is easy to miss: a closure’s cost is not only its own uncertainty but the number of picture readings it drags in with it.

Spread and bias are different columns

The bars above are spreads — how much the answer moves under reading noise. The catalogue also carries a bias column, and the two do not rank the same way.

The standing object of known height has the second-best spread and the worst bias, at 0.39 per cent. The bias comes from the camera height being estimated from a noisy horizon reading rather than measured, so the closure is itself a recovered quantity with an error in it, and the error is systematic across every subsequent measurement.

A ground length’s bias is 0.15 per cent and a repeated object’s is 0.03. So a reader who wants a single answer wants the small spread, and one who wants an answer to compare against something else wants the small bias — which is the floor a better camera cannot reach arriving in a table.

Better marks stop helping at 6.0%, which is the assumption's own errorA length running into the picture, measured with the marks read to the precision on the horizontal axis, twice over. The lower curve has the reference's shape exactly right, and it keeps falling: better marks keep buying a better answer, without limit. The upper curve has the reference's aspect wrong by 6 per cent, and it stops — at 6.0 per cent, which is the assumption's own error and nothing else. The crossing between them is where a reader should stop buying lenses, and it can be computed before the photograph is taken.0102030401234the marking error, in pixelsthe error of the answer, in per centupper curve: the closure wrong · lower: the closure exactfloor 6.0%
Fig. 2 Where a bias comes from and what it does, from the rung before: a term that better reading never touches.

Which photograph offers which

The ranking is only useful in combination with availability, and availability is the more practical half.

A photograph of an empty floor with nothing in it offers the camera’s height and nothing else, if the photographer measured it, and nothing at all if they did not. That is why the best closure in the table is so often unavailable: it has to be recorded at the time and cannot be recovered later.

A photograph containing a person offers the standing-object closure, at the population’s spread — which is a bias of a few per cent rather than the 0.39 per cent measured here, because the person’s height is assumed rather than known.

A photograph of a tiled floor, a colonnade, a fence — the arrangement counting is a measurement turns into an area — offers the repeated-object closure and gives every proportion with no metre. A photograph with a scale bar offers the ground length. A photograph from a known camera offers the shape and still wants a size.

And a very great many photographs offer more than one, which is the case the previous rung recommends: use them together, and treat their disagreement as the only available evidence that one of them is wrong.

The best reference is 4.0 m; the longest is 5.2 m and the worstFour candidate references in one photograph, each used in turn to give the same unknown length its metre, with the spread over a hundred and sixty noise draws as the score. Shorter bars are better. The five-metre bar lying near the horizon is the longest thing in the scene and the worst reference in it, because what the arithmetic uses is its length in the picture and depth has taken that away. The rule is not "use the longest" and not "use the nearest": it is use the one that is longest where you are actually reading it, which is on the print.near and short3.68%0.5 m, 126 pxnear and long2.84%4.0 m, 1088 pxat the unknown's depth3.95%1.4 m, 154 pxfar and long8.36%5.2 m, 241 pxthe spread of the answer, per candidateshorter is better
Fig. 3 The within-one-closure version of the same choice, from the rung before: four candidate lengths, ranked by what each leaves in the answer.

What using two of them together is actually worth

The advice at the end of the last section — use every closure a photograph offers — is right, and it is worth putting a number on, because the number says what the extra closures are for and it is not what the advice implies.

Independent estimates combine by inverse-variance weighting, so five closures with spreads sis_i give a combined spread of (si2)1/2\left(\sum s_i^{-2}\right)^{-1/2}. On the table’s own figures the weights are 7.72, 5.95, 0.266, 0.096 and 0.077 — and the combination is 0.266 per cent.

Read that against the parts. The best single closure is 0.36. The best two together give 0.271. Adding the remaining three moves it to 0.266, an improvement of two per cent of a per cent. The three weaker closures carry 3.1 per cent of the information between them, because inverse variance is a square and a factor of nine in spread is a factor of eighty in weight.

So combining closures is not how a photograph is measured more precisely. Two good ones exhaust it, and the third onward is arithmetic that changes nothing.

What the extra closures are for is the other thing entirely, and the table supports it much better. Two closures with spreads 0.36 and 0.41 disagree, under noise alone, by a standard deviation of 0.362+0.412=0.55\sqrt{0.36^{2} + 0.41^{2}} = 0.55 per cent. A disagreement of more than about 1.6 per cent is therefore three sigma and is evidence of a fault — a mismeasured tape, a non-standard door, a control point in the wrong plane — rather than of bad luck. That is a test no single closure can perform at any precision, because a single closure has nothing to disagree with and a wrong one produces a self-consistent answer.

The right summary is that a second closure buys a test and a fifth buys nothing, which is close to the opposite of how a list of five options invites a reader to use it.

The one closure that is not a closure

The repeated object deserves separating out, because it does something different from the other four and the difference is a rung on the stratification ladder rather than a difference of quality.

The other four supply a metre. The repeat supplies a unit, and the unit’s size is unknown — so the answer comes back as “three and a quarter repeats” rather than as “two point one metres”. For a great many questions that is the whole answer: is this doorway wider than that one, is the courtyard square, what fraction of the wall is glazed.

Where it is not the whole answer it is still most of it, because it upgrades the picture from projective to similarity structure, and the remaining freedom is a single number. What one picture determines is the collection’s account of that ladder, and the repeat is the cheapest way up it — no tape, no assumed size, nothing supplied at all except the observation that the units are identical.

That observation is a real assumption and it fails on hand-made things, which is the boundary the counting essay states at length.

Two points, and everything metric followsThe imaged circular points are where the horizon meets the image of any circle in the plane, and they are a conjugate pair — the first coordinate here is 169.5 − 446.0i. A rectification built from them and nothing else returns the world's angles to 1.6e-13° and its length ratios to 2.9e-15, and no length at all.horizonthe horizon does not cut the circle — the pair is complexrectified from the two points aloneangles: 1.6e-13°ratios: 2.9e-15length: —circle of radius 1.05 ma dash is a quantity two points cannot buy
Fig. 4 The rung a repeat reaches, from the foundations field: shape without size, which is two numbers on the vanishing line and no metre.

Where each closure sits on the stratification ladder

The five are not five ways of doing one thing. They enter the ladder at different rungs, and reading them that way says what each leaves free rather than only how well it does.

A picture of a plane, on its own, determines the plane’s projective structure: which points are collinear, and every cross-ratio. Adding the vanishing line gives the affine structure: parallelism, midpoints, ratios of lengths along one direction. Adding two more constraints gives the metric structure up to a scale: angles and ratios between directions. And adding one length gives the Euclidean structure entire.

The focal length with the horizon supplies the affine and metric rungs at once and stops short of the last. A repeated object supplies affine and metric and stops in the same place, by a different route. A length on the ground, a camera height and a standing object all supply the last rung, and each of them assumes the earlier ones are already in hand.

That is why the table’s five rows are not interchangeable even where their spreads are similar: two of them leave a scale free and three do not, and a reader who wants proportions and a reader who wants metres want different rows.

The comparison had to be run rather than argued

There is a reason the five are computed on one picture with one noise stream rather than compared by their formulas.

Each closure has a different route from the marks to the answer, with a different number of intermediate quantities, and the error propagates differently along each. Writing five error formulas and comparing them is possible and is exactly the kind of comparison that goes wrong: a factor dropped in one derivation and not another, an intermediate quantity assumed independent that is not, a linearisation valid over a different range in each.

Running all five on identical draws removes every one of those. The differences between the bars are differences between the procedures, because everything else — the picture, the camera, the unknown, the noise — is held fixed by construction.

That is the same argument three procedures, one panel makes for comparing drawing methods on one drawing rather than on their descriptions, and it is the same argument the mirror field makes for reading one set of marks two ways.

What the essay does not settle

Three things, and naming them is part of the result.

The ranking is arrangement-specific. The camera height wins here partly because the unknown lies at moderate depth in the middle of the frame, where a direct unprojection is well conditioned. A different unknown, further out or nearer the horizon, changes the order — and the previous rung’s law says how.

The biases are modelled optimistically. Every closure here is assumed correct on average, with only the picture readings noisy. In practice an assumed person’s height, an assumed slab size and an assumed lens focal length all carry biases of their own, and those would dominate the table entirely.

Combination is not attempted. Five estimates of one quantity should be combined rather than chosen between, weighted by inverse variance and with their correlations accounted for. The correlations are the hard part: four of the five read some of the same marks, so their errors are not independent and a naive combination understates the result’s spread.

The closure a second view supplies, and why it does not help

For completeness, and because it is the obvious question: a second photograph does not close the gap.

Two views give shape and no size measures exactly that. A pair of views recovers the scene’s shape and the camera’s motion up to one overall scale, and the scale is free for the same reason it is free in one view — enlarge the scene and the baseline together and both pictures are unchanged.

What a second view buys is the shape, which one view of a plane already has once the metric closure is supplied, and which one view of a general scene does not. So for a plane the second view adds nothing this table wants; for a solid it adds everything except the row the table is about.

The exception is the one the mirror field found. One shutter, two views shows that a mirror pair’s baseline is twice the camera’s distance from the glass, which is a length a tape measures — so a mirror in the scene converts a second view into a closure, and it is the same closure as the camera’s height by a different route.

The boundary, stated

One plane, one camera, one unknown, one hundred and sixty draws.

The plane is required by four of the five; the camera-height closure needs the plane’s distance and therefore needs the plane. The unknown is a single length rather than a plan, and a plan’s closure is the same fact used many times, with correlations that a single-length experiment cannot show.

And the whole comparison is about spread under reading noise. It says nothing about which closure is easiest to obtain, which is usually the deciding factor — and on that criterion the ranking inverts almost exactly, since the closure with the best spread is the one that has to be measured at the time of the photograph and is therefore the one a reader most often does not have.

The closure as part of the result

A last recommendation, which follows from every essay in this ladder and is easy to state.

Report the closure with the answer. A measurement from a photograph is a picture reading plus a fact from outside, and the fact is not an implementation detail — it is half the result, it carries its own error, and it is the half a later reader is most likely to be able to improve.

“The courtyard is 8.4 by 11.2 metres” is an incomplete statement. “The courtyard is 8.4 by 11.2 metres, closed on a paving slab taken as 600 mm” is a complete one, and it lets somebody who later measures that slab correct the result without retaking the photograph. It also lets them compute the floor, which the floor a better camera cannot reach shows is the term the reported spread never contains.

This collection’s own essays state their closures for exactly that reason, and the habit is the practical residue of the whole field: the picture supplies ratios, something else supplies the metre, and a result that does not say which was which cannot be checked by anybody.

What is measured here

Five numbers and a range.

On one picture, one unknown length and one hundred and sixty noise draws each: the camera’s own height leaves a spread of 0.36 per cent, a standing object of known height 0.41, the focal length with the horizon 1.94, a length on the ground 3.22, and a repeated object of unstated size 3.60. The range is a factor of ten.

The biases rank differently: 0.004 per cent for the camera height, 0.39 for the standing object, 0.065 for the focal length, 0.15 for the ground length and 0.025 for the repeat. And the reason the ground length loses to the camera height is a count — four marks read against two — rather than anything about the closures themselves.

The short version

A photograph needs one fact from outside it, and “one length” is one of at least five facts that will do: a length on the ground, the camera’s height, a repeated object, a standing object of known height, or the focal length with the horizon.

They are not equally good. On one picture with one set of noise draws their spreads run from 0.36 to 3.60 per cent, and the ordering is decided partly by how many picture readings each one drags in with it rather than by the closure’s own quality. Which is available depends on the photograph, and the best of them is the one that has to be recorded at the time and usually is not.

The reference's own term falls as one over its length in pixelsOne reference of fixed length in the world, walked from close to the camera out toward the horizon. Its length in the picture falls as one over its depth, and the spread it puts into the answer falls with it. The lower curve isolates the reference's contribution by reading the unknown's own marks exactly; its fitted exponent is one over the image length, not over the world length and not over the depth. The upper curve is what a reader actually gets, and it flattens at the left because the unknown's own marks are still being read to the same precision — which is the same floor the closure sweep finds from the other side.00.5001100200300the reference's length in the picture, in pixelsthe spread of the measurement, in per centwhat a reader getsthe reference's own termone bar, walked outwardlongest in the picture wins
Fig. 5 And within any one of them, the choice of which reference remains — which is the law the rung before this one fits.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BiasCamera calibrationConditioningerror propagationMetric upgradeProjective stratificationReference lengthscale ambiguitysingle-view metrologyVanishing line