Measuring from one picture

An area, out of one photograph

A patch of ground comes back at 5.205 m² from one photograph, to 1.8 × 10⁻¹⁴, through a homography built from four marks and their four known positions. What is worth knowing is how it degrades — the patch's extent across the picture is read with an error growing as the depth, its extent into the picture with an error growing as the depth squared, and at twenty-eight metres the two are sixteen times apart — which is the depth divided by the camera's height.

Worth reading first: Flattening a façade out of the photograph · The one thing a single view cannot give.

Everything the metrology field has measured so far is a length. A height out of one photograph, a façade flattened, a plan recovered, a ball’s radius — all of them lengths, and all of them reported as a number with an error beside it.

An area is not a length, and it does not behave like one. This essay recovers an area from a single view, which is straightforward, and then measures how it degrades with distance, which is not.

The recovery

Four marks on the ground, four known positions, and the homography between them. That is the rectifying map flattening a façade out of the photograph builds for a wall, pointed downward at a floor — and once it exists, any picture point on the ground can be pushed back to a world position.

An area then costs nothing extra. Push the patch’s corners back, apply the shoelace formula in world coordinates, and the answer is in square metres.

5.205 m² off one photograph, to 1.8e-14On the left, a four-sided patch of ground photographed from a stated camera, with the reference rectangle whose size is known. On the right, the same patch pushed back through the rectifying homography built from that reference — four marks and their four known positions and nothing else. Its area comes back at 5.2050 square metres against a true 5.2050, which is 1.8e-14 of relative error and is arithmetic rather than a fit. An area is not a length, and it is worth noticing that no length inside the patch was measured on the way — the homography carries the whole plane, and the shoelace formula is applied on the far side of it.the photographthe ground, rectified5.205 m²four marks fix the plane; the shoelace does the rest1.8e-14
Fig. 1 On the left, a four-sided patch photographed with a reference rectangle of known size. On the right, both pushed back through the homography built from the reference alone. The patch’s area comes back at 5.205 m² to 1.8 × 10⁻¹⁴.

Two things are worth noticing about that. The patch is not a rectangle and nothing assumes it is — a routine that could only measure rectangles would be measuring the rectangle assumption. And no length inside the patch is measured on the way: the homography carries the whole plane, and the shoelace runs on the far side of it.

Why the degradation is the interesting part

An exact recovery from exact marks says the arithmetic is right. What a reader needs is what happens when the marks are read to four tenths of a pixel, which is careful work on a print, and how that changes with where the patch is.

The naive expectation is that an area, being two lengths, degrades twice as fast as a length. That is very nearly wrong, and the reason is that the patch’s two directions are not read with the same rule at all.

Across the picture the error grows as the depth; into it, as the squareA patch a metre and a half across, at increasing depth, with its marks read to four tenths of a pixel. Its extent across the picture and its extent into the picture are measured separately, because the ground reads them with different rules: across, a pixel of error is worth d over f metres and the fitted exponent is 1.03; into, the ray is grazing, a pixel is worth d squared over f times the camera's height, and the exponent is 1.93. At twenty-eight metres the two are 16 times apart, which is the depth divided by the camera's height. An area is the product of the two, so its own exponent is a crossover rather than a law.01020510152025the patch's depth, in metresthe relative error of the extent, in per centinto: exponent 1.93across: 1.03one patch, two directions16× at 28 m
Fig. 2 A patch a metre and a half across at increasing depth, with its extent across the picture and its extent into the picture measured separately. Two different laws.

Across the picture, a pixel of marking error is worth d/fd/f metres — depth over focal length — so the error of a fixed-size extent grows as the depth. The fitted exponent is 1.03.

Into the picture, the ray meets the ground at a grazing angle, and a pixel of vertical error slides the ground point along that ray by d2/(fh)d^2/(fh) metres, where hh is the camera’s height. The error grows as the depth squared, and the fitted exponent is 1.93.

At twenty-eight metres from a camera 1.62 metres up, the second is sixteen times the first — and sixteen is 28 divided by 1.62, which is the depth over the camera’s height. That ratio recurs.

The area’s own exponent is a crossover

The area is the product of the two extents, so its relative error is a combination of a first-power term and a second-power term. Which dominates depends on where the patch is, so the area’s exponent is not a law at all.

Fitted over four to twenty-eight metres it is 1.62. Fitted over forty to four hundred it is 2.20. Neither number is a property of areas; both are readings of a crossover taken in different places.

That is worth being blunt about, because quoting “an area’s error grows as the square of the depth” would be quoting the asymptote as though it were the rule, and the range where it holds is the range where the relative error already exceeds a hundred per cent. Over the distances a measurement is actually worth making, the exponent is closer to one and a half.

Where the camera’s height enters

The ratio between the two laws is the depth over the camera’s height, and that is the only place the camera’s height appears in the whole analysis. It is worth drawing out, because it is the quantity a reader controls.

A camera on the ground reads a floor entirely at grazing incidence: hh is near zero, the depth term is enormous, and depth is essentially unmeasurable. A camera directly overhead reads the floor at normal incidence: the grazing factor disappears, both directions obey the same law, and the recovered patch is as good in one direction as in the other.

Everything in between is a mixture, and the mixture is set by d/hd/h. So the single most effective thing a reader can do to measure an area from a photograph is to get higher — not to use a longer lens, which scales both errors equally and changes nothing about their ratio.

The answer is an ellipse, elongated by the depth over the camera's heightFive marks read to four tenths of a pixel, with the region each one could actually be in drawn on the ground in plan. A pencil point has a round error and what comes back is not round at all: the ray is grazing, so a pixel across the picture is worth a little and a pixel up it is worth a great deal, and every ellipse points at the camera. The most elongated of these is 7.2 to one, at 11.6 metres from a camera 1.62 metres up — and the ratio is the depth over the height, which is the same number the two extent laws differ by. The ellipses are drawn at 8 standard deviations so that they can be seen.2.7:13.1:14.5:15.8:17.2:1in plan; ellipses at 8σworst 7.2:1
Fig. 3 The same fact in plan, from the essay that draws it: the region a mark could be in is elongated by the depth over the camera’s height, and it points at the camera.

How much height is worth buying

The advice at the end of the previous section — get higher — is right and it is unbounded as stated, so it is worth saying where it stops paying.

The two contributions to an area’s relative error are εd/fL\varepsilon d/fL across and εd2/fhL\varepsilon d^{2}/fhL into, for a patch of extent LL read to ε\varepsilon pixels. Their ratio is d/hd/h and their combination is

σAA    dL1+(dh)2,\frac{\sigma_A}{A} \;\propto\; \frac{d}{L}\sqrt{1 + \left(\frac{d}{h}\right)^{2}},

which has a floor. As hh grows the second term vanishes and the first does not, so the whole expression falls to d/Ld/L and stays there.

That makes the return on height sharply staged. While hdh \ll d the depth term dominates and the error is very nearly proportional to 1/h1/h: doubling the camera’s height halves the area error, which is as good a return as measurement geometry ever offers. At h=dh = d the two terms are equal and the combination is 2\sqrt{2} times its floor. Past h=dh = d there is at most a factor of 1.41 left to win, however high the camera goes.

For the arrangement measured here — a patch at 28 metres, a camera at 1.62 — the ratio is 17, so the depth term is carrying essentially the whole error and almost all of that factor of seventeen is available. A camera on a three-metre pole halves it; one on a ten-metre mast divides it by six; a drone at twenty-eight metres reaches the floor. And a longer lens, as the section above says, moves nothing: ff divides both terms and cancels out of the ratio entirely.

The same quantity d/hd/h is the aspect ratio of the region a single mark could have come from, which is what the answer is an ellipse draws, and it is cotφ\cot\varphi for a ray arriving at elevation φ\varphi — so all of this is one statement about grazing incidence, wearing three sets of clothes across the field.

Why the two laws had to be measured separately

The two exponents above came out of a computation that had to be corrected once, and the correction is the kind worth recording because the wrong version looked like a measurement.

The first attempt measured each extent as a Euclidean length between two recovered ground points — which is the natural thing to write, and it mixes the laws. The endpoints of the across segment each carry a depth error perpendicular to that segment, so the depth error enters its length quadratically rather than linearly, and the fitted exponent came out at 1.71. That is neither law. It is a number produced by a statistic that does not separate what it was meant to separate.

Measuring each extent on its own axis — the across extent as a difference of world xx, the into extent as a difference of world zz — gives 1.03 and 1.93, which are the two laws and which agree with the algebra.

The general shape of that mistake is one this collection has recorded twice before under other names: a quantity evaluated in a way that cannot distinguish the thing being asked about. Here the failure is milder than the vacuous-test kind, because 1.71 is a true statement about a real quantity — it is just not a statement about either of the two laws, and it would have been reported as one.

What an area is used for

Three uses, and each wants a different part of the analysis.

A scene plan — the area of a room, a courtyard, a field — is measured over a patch at moderate depth spanning most of the frame, so both terms matter and the answer wants the full propagation.

A coverage fraction — how much of a canopy is sky, how much of a wall is damaged — is a ratio of two areas on the same plane, and a ratio is far better behaved: a systematic error in the rectifier scales both areas equally and cancels. Counting cloud by counting pixels is where the curved field measures the same cancellation for a sky camera.

A rate — a quantity an error with two terms prices generally, and here how many square metres per second something is spreading over — is a difference of areas across time, and differences amplify: two areas each 3 per cent uncertain give a small difference that is worthless. That is a general property of differencing and it is not special to pictures.

The closure, again

An area recovered this way is metric because the reference rectangle’s size was supplied from outside the photograph. Without it the map is determined only up to a similarity of the plane, and every area comes back in units of the reference’s own area.

That is the one thing a single view cannot give in its usual form, with one difference worth noting: an area needs the closure squared. A reference length one per cent wrong gives a length one per cent wrong and an area two per cent wrong, because the map’s scale enters an area twice.

So the closure matters more here than for a length, and choosing it well matters more. That is the subject of two later rungs in this ladder.

Two ways the recovery fails quietly

Neither of them announces itself, and both produce a definite number.

The patch is not on the plane. A rectifier built for the ground will happily push back a mark on a step, a kerb or a person’s shoulder, and return a ground position for it — the wrong one, by an amount that grows with the height off the plane and with the depth. Nothing in the arithmetic objects. The check is that the patch’s boundary should be visibly on the floor, which is a judgement rather than a measurement, and it is why a floor with a referent insists on the referent being in the plane.

The reference is read wrong. Reading the reference rectangle one tenth too wide gives an area ten per cent out — measured, 0.09999, which is exactly the fractional error and not a fraction of it. That is the control this essay’s assertion uses, and it is what says the exact recovery is a measurement rather than a property of the solver.

An area is not a length twice over

It is worth collecting the ways an area differs from a length, because they are not obvious and every one of them has a number attached.

The closure enters twice. A reference read one per cent wrong costs a length one per cent and an area two, because the rectifier’s scale appears in both directions of the patch. That is the squaring above, and it means a careless reference is twice as expensive here.

The two directions obey different laws. A length has one law, whichever direction it runs in; an area has both at once, always, and its exponent is a mixture that depends on the patch’s aspect and depth. A patch elongated across the picture behaves almost like a length; one elongated into it behaves almost like a squared law.

A ratio of areas is far better than either. The rectifier’s scale cancels in a ratio, so a coverage fraction is insensitive to the closure entirely — which is why sky-fraction measurements from a fisheye can be made without any reference length at all, and why they are the metrology this collection trusts most on a bad photograph.

And an area cannot be carried. A length can be transferred across a picture by the straightedge construction carrying a height across the room uses, because the cross-ratio survives a projection. An area has no such construction: area is not a projective invariant, there is no ratio of areas that survives a general projection, and every route to one goes through a rectification with a closure attached.

That last one is the structural difference. What a projection destroys lists area among the casualties and the cross-ratio among the survivors, and the whole of this essay is downstream of which side of that list the quantity falls on.

Four points on a line, before and after a projectionLength and the ratio of lengths do not survive the projection; the cross-ratio does, agreeing to 0e+0 relative.horizonABCDon the groundin the picturelength AB1.00011.3930ratio AB:CD0.56670.6837cross-ratio1.31681.3168correct from 26 cm, at 160 mm wide34° across
Fig. 4 The survivor, from the foundations field — which is why a length can be carried across a picture and an area cannot.

The boundary, stated

A plane, four marks whose world positions are known, and a patch lying in that plane.

The plane is required and is often the hard part: a “floor” that slopes, a lawn that is not level, a cobbled yard that dishes toward a drain, are all planes to the eye and not to the arithmetic. The floor that is not a plane is where the light field measures what a curved receiver costs a shadow, and the same departure costs an area the same way.

Four marks is the minimum and more is better in the ordinary way; the homography has eight degrees of freedom and each correspondence gives two equations, so four is exact and five is the first with a residual to look at.

And the whole of it is a geometric measurement. What is inside the patch — whether a region of grass is one region or two, where a stain’s edge is — is a question about the picture’s content rather than its geometry, and this collection does not answer it.

A reader’s rule for the two laws

Collecting the essay’s two exponents into something usable: the ratio of the two error laws is the depth over the camera’s height, so a reader can compute it before making the measurement and decide what to do about it.

At head height, a patch three metres away is read about twice as badly in depth as across; at ten metres, six times; at thirty metres, eighteen. So a patch whose sides run across the picture is a much better shape to measure than one whose sides run into it, and a reader with a choice about how to frame the shot has a real lever.

The lever has two settings. Raise the camera and the ratio falls in direct proportion, which is why an area from a first-floor window is worth several from the pavement. Move closer and the ratio falls too, though at the cost of covering less of the patch — so for a fixed patch there is a distance that minimises the total error, and it is nearer than most photographers stand.

Neither of those is available by processing the picture afterwards, which is the point. The choices that matter for an area measurement are made before the shutter, and the two exponents are what say which choices they are. Standing in the wrong place is the collection’s statement of the same idea from the viewing side: where the eye is decides what the picture can be.

What is measured here

Four numbers.

A four-sided patch of 5.205 square metres comes back from one photograph to 1.8 × 10⁻¹⁴ relative error on exact marks; the same picture read with a reference a tenth too wide returns it 10.0 per cent out. The patch’s extent across the picture degrades with a fitted exponent of 1.03 and its extent into the picture with 1.93, and at twenty-eight metres the two are sixteen times apart — the depth divided by the camera’s height. And the area’s own exponent, fitted over four to twenty-eight metres, is 1.62, rising to 2.20 over forty to four hundred, which is what says it is a crossover rather than a law.

The short version

An area comes out of a single photograph exactly, through the same rectifying homography a length does, and needs a reference in the plane whose area it is quoted in units of — so a closure one per cent wrong costs two per cent rather than one.

Its two directions degrade under different laws: across the picture as the depth, into the picture as the depth squared, sixteen times apart at twenty-eight metres from a camera at head height. The single most effective way to improve an area measured from a photograph is to raise the camera, which changes the ratio; a longer lens scales both terms equally and does not.

21 of 36 tiles are countable — 8.07 m² of 13.84A tiled floor photographed from a stated camera, magnified to fill the frame, with the tiles that can still be told apart drawn solid and the rest drawn faint. Counting is a measurement of area that needs no length at all — multiply the count by the tile, and the answer is exact wherever the count is right. What it does not do is degrade gracefully. The limit is set by the tile's edge running INTO the picture, which foreshortens as one over the depth squared, so a floor of 62 centimetre tiles stops being countable at about 19 metres — where the across edge alone would have allowed 217. The magnification is uniform within each setting, so it changes how large the floor is drawn and not which of its tiles are separable.21 of 36 tiles still countablethe floor starts 18 m away
Fig. 5 And the next rung takes the other route to an area entirely, which needs no length at all and does not degrade until it stops working.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Area scaleConditioningerror propagationGrazing incidenceGround plane rectificationHomographyJacobianRectificationReference lengthsingle-view metrology