Systems that kept the measure

Size that means rank

In a great many pictures the drawn height records importance rather than distance. That is a decision rather than a mistake, and it can be caught with a straightedge by carrying one figure's height across the room by the taught construction and see where it lands. A tenth of rank in the picture already misses the drawn head by 21 px, the miss is exactly linear in how much rank is there, and the whole test needs two references and no arithmetic.

Worth reading first: A picture with no size–distance signal · A height, out of one photograph.

The king is large because he is the king. Assyrian reliefs do it, Egyptian ones do it, Byzantine donor panels do it, and a great many medieval altarpieces do it — the drawn height of a figure records where it stands in a hierarchy rather than where it stands on the floor.

A picture with no size–distance signal establishes the first half of what that means: a picture whose figures do not shrink with depth gives a reader no way to read depth from size, and it has no horizon findable by the usual construction. This essay asks the sharper question. Given a picture, can a reader tell which rule was used — and if so, how much rank has to be in a picture before the test fires?

The picture, as a mixture

Almost no surviving picture uses one rule alone. A donor kneeling in the foreground is drawn larger than the saints behind, and also nearer; a king in a hunting scene is drawn larger than his attendants, some of whom are further away than others. So the honest object to study is a mixture, and the parameter is how much of the drawn height is rank.

Sized by distance: a photographFive figures standing on one ground, drawn with their heights decided partly by how far away they are and partly by how important they are. At the left-hand end of the slider the picture is a photograph; at the right-hand end the drawn height records rank alone and the ground is irrelevant to it. Most surviving pictures that use the convention are somewhere between, which is the reason for the slider: an Assyrian relief, a Byzantine donor panel and a great many altarpieces are not choosing one rule and abandoning the other, and a test that only fired on the pure case would catch none of them.correct from 15 cm, at 160 mm widefive figures, distance only
Fig. 1 Five figures on a ground with their heights decided entirely by distance. This is a photograph, and it is the left-hand end of the slider.

At one end the picture is a photograph. At the other the ground is irrelevant and the heights are ranks alone. A test that only fired at the far end would catch almost nothing that was ever painted.

Sized by rank at 100 per centFive figures standing on one ground, drawn with their heights decided partly by how far away they are and partly by how important they are. At the left-hand end of the slider the picture is a photograph; at the right-hand end the drawn height records rank alone and the ground is irrelevant to it. Most surviving pictures that use the convention are somewhere between, which is the reason for the slider: an Assyrian relief, a Byzantine donor panel and a great many altarpieces are not choosing one rule and abandoning the other, and a test that only fired on the pure case would catch none of them.correct from 15 cm, at 160 mm widefive figures · 100% rank
Fig. 2 The same five figures with their heights decided entirely by rank. The second figure matters most and is drawn tallest, at a depth where the photograph makes it short.

The test a reader can run

The instrument is the taught construction, used as a detector rather than as a drawing aid. Carrying a height across the room sets it out: given one figure of known drawn height and a horizon, the drawn height of a figure standing anywhere else on the same floor is fixed, and it is found by joining and meeting.

Two figures fix a point on the horizon between them. With a level camera the horizon through that point is horizontal, and two figures are therefore enough — which matters, because every statistic that needs three of them is corrupted by the very effect being detected.

With the horizon in hand, each remaining figure’s head is predicted from its own feet, and the miss between the predicted head and the drawn one is the reading.

Carried across, the height lands on the drawn headThe test a sceptical reader runs with a straightedge, drawn on the picture it is run on. Two of the figures fix a point on the horizon; the horizon is then used to carry the first figure's height across to each of the others, and the predicted head is compared with the drawn one. Where the picture is a photograph the two coincide exactly. Where the heights record rank they do not, and the worst miss here is 0.0 pixels. Nothing about the camera is needed and nothing is measured: the whole construction is joins and meets, which is why it survives being carried out on a photograph of the painting.two references, a straightedge, no arithmeticexact
Fig. 3 The construction run on the photograph: two references, a horizon, and a predicted head that lands on the drawn one exactly.

Nothing is measured anywhere in it. The whole construction is joins and meets, so it survives any projection and can be carried out on a photograph of the painting rather than on the painting — which is what a reader actually has.

Linear in the thing it detects

The reading is 8.5 × 10⁻¹⁴ pixels on the photograph and 209 on the pure rank picture, and in between it is a straight line.

A tenth of rank in the picture is already 21 px of missThe miss against the amount of rank in the picture, read from exact marks and from marks read to about a pixel. The clean line is straight, which is worth knowing: the detector is linear in the thing it detects, so a reading of forty pixels means twice what a reading of twenty means. The noisy line lifts off the axis at the left, and where it does is the sensitivity of the whole test — below that, a picture drawn honestly and a picture with a little rank in it cannot be told apart by anybody working from a photograph of a painting.050100150200020406080100how much of the drawn height is rank, in per centhow far the carried height misses, in pixelsexact marksread to a pixelthe detector is linear in what it detects21 px at 10%
Fig. 4 The miss against how much of the drawn height is rank, from exact marks and from marks read to about a pixel. The clean line is straight, which is what makes a reading interpretable.

That linearity is worth more than the endpoints. A detector that responded quadratically, or that saturated, would give a reading with no interpretation: forty pixels would mean “some” and eighty would mean “some more”. Here forty means twice what twenty means, and a reader who measures thirty-five pixels on a real panel can say that about a sixth of the drawn height is doing something other than recording distance.

The statistic that looked right and was not

The first version of this test used a different statistic and it is worth recording, because it failed in a way that is easy to reproduce and hard to notice.

Every pair of figures gives one horizon point, so five figures give ten of them, and the obvious detector is their scatter: they are one point in a photograph and spread out in a rank-scaled picture. That is true, and the scatter is not monotone in the effect. Two uprights whose feet line and heads line are nearly parallel meet a very long way off the paper, so the statistic is heavy-tailed, and a single such pair dominates everything.

Measured on this arrangement, with the rank content raised evenly from none to all, the scatter read 0, 41, 509, 139, 175, 439, 1751. A detector whose reading falls when the effect grows is not a detector.

The carried-height version is bounded, because it is a distance between two marks that are both on the page, and it uses two references rather than all pairs. Both properties come from the same decision: measure in the picture, not in the space of horizon points.

The arrangement that gives nothingThree uprights at one depth: the line through the feet and the line through the heads are parallel in the picture, so they meet nowhere on it. 3 meeting points and none of them within 690×430 px. A photograph of a crowd standing in a row is this arrangement.the camera's horizoncorrect from 19 cm, at 160 mm wide46° across
Fig. 5 Where the discarded statistic comes from, in the field that uses it properly: uprights, their pairwise meets, and a horizon fitted through them.

The threshold

A reader working from a photograph of a painting reads marks to somewhere around a pixel. The detection threshold is where the miss clears that.

At ten per cent of rank the clean miss is 21 pixels, which is far above a pixel of reading error — so the test fires comfortably on a picture with a tenth of its height decided by rank. Below about two per cent it does not, and a painter who enlarged a donor by a fiftieth would not be caught by anything a reader could do from a reproduction.

That is the right shape of answer for this collection. The test has a sensitivity, the sensitivity is a number, and the number is not zero. A test offered without one is offered as a trick.

Where the factor of four in the threshold comes from

Two numbers in the sections above pin the test’s calibration, and dividing them turns the threshold into a statement about the construction rather than about the marks.

The reading is linear with a constant of 209 pixels at full rank, so a tenth of rank gives 21 — which is the quoted figure, and it means a reader can invert directly: rank fraction = miss ÷ 209 on this panel. Nothing else is needed, which is what a linear detector is for.

Now the threshold. If a reader places marks to about a pixel, the test ought to fire wherever the clean miss exceeds a pixel — which is at 1/209=0.48%1/209 = 0.48\% of rank. The measured liftoff is at about two per cent, four times higher.

The factor of four is the construction’s own amplification. The horizon is fitted from two reference figures and then used to predict a third head, so a marking error enters twice: once in placing the horizon, and again in reading off the prediction. Two nearly-horizontal lines meeting at a distant point locate that point poorly — the same 1/sinφ1/\sin\varphi conditioning that makes a mirror pair’s two joining lines want to be widely separated — and the poorly located horizon then propagates into every prediction made from it.

So the honest statement of the sensitivity has three numbers rather than one: a panel constant of 209 pixels, a marking error of about a pixel, and an amplification of about four between them. And the third is the one a reader can improve. Choosing the two reference figures as far apart in depth as the panel allows widens the angle at which their joining lines meet, lowers the amplification, and moves the threshold back toward half a per cent — which is the same instruction the mirror pair’s own conditioning gives, arriving in a completely different subject.

That is worth having because it is the one part of the test a reader controls. The panel’s constant is the painter’s; the marking error is the reproduction’s; the amplification is a choice made while laying the ruler down.

What the test cannot say

Three things, and the third is the one that matters.

It cannot say which figure is out of place. The construction uses two references and tests the rest, so a reading tells a reader that the set is inconsistent rather than which member of it caused the inconsistency. Choosing a different pair of references moves the blame around, and the honest procedure is to run every pair and look at the pattern — which is a robust-estimation problem rather than a construction.

It cannot distinguish rank from a different height. A picture of an adult and three children on a floor fails the test exactly as a rank-scaled picture does, because the test assumes equal heights and children are not adults. That is not a defect to be repaired; it is the assumption the construction is built on, and the same assumption a picture with nothing straight in it has to make when it finds a horizon in a crowd.

And it cannot say the picture is wrong. A rank-scaled picture is a correct rendering of a statement about rank, made by somebody who was not trying to record where anybody was standing. What the test measures is what a reader who brings a distance assumption to it gets, which is a fact about the reader — the same asymmetry an inverse perspective is a leaning plane records for a system that looks wrong until the assumption behind the looking is stated.

A scroll keeps the midpoint along its length and loses it acrossLeft, a segment lying along the roll: the image of its midpoint and the midpoint of its image are the same point to 0e+0 px. Right, a segment running away from the eye: the two are 21.9% of the segment apart. One projection, two answers, because the eye is at infinity in one direction and seven metres away in the other.along the rollthe two midpoints coincide — 0e+0 pxone mark, drawn twiceacross itthey separate by 21.9%the image of the midpointthe midpoint of the imageno single viewpoint — the rays miss by 6.9 m21.9% of the receding segment
Fig. 6 The general form of that last point, from the scroll field: a system that is a map in one direction and a picture in the other, measured on what it does rather than ranked against a photograph.

Why two references and not a fit

There is an obvious alternative to the construction above: fit the best single figure height and camera to all five figures at once, and report the residual. It is a better estimator in every ordinary sense, and it is the wrong instrument here for two reasons.

The first is that it needs a camera. A least-squares fit over five figures has to be parameterised by something — a focal length, a horizon height, a common object height — and every parameterisation smuggles an assumption about the picture that the reader has no independent access to. The straightedge version assumes only that the picture is a projection of equal uprights on a plane, which is the hypothesis under test.

The second is that a fit absorbs the effect. Given five figures with one enlarged, a least-squares horizon moves to split the difference, and every figure ends up with a small residual instead of one figure ending up with a large one. That is exactly the right behaviour when the departures are noise and exactly the wrong behaviour when one of them is the signal — and it is why what a hand leaves has to fit its horizon on a relative objective and then convert back into pixels before the reading means anything.

So: two references, chosen and stated, and every other figure tested against them. It is the procedure a sceptical reader with a ruler would use, and it has the property that its answer can be checked by drawing on the page.

What a viewer sees, which is not what the test measures

A reading of two hundred pixels sounds like something a viewer could not miss, and viewers miss it constantly. That gap is worth a paragraph because it explains why the convention survives.

A picture is read by a viewer who has no horizon marked, no straightedge, and no reason to compare a figure’s height against a construction. What a viewer has is an expectation about how large people are and a rough sense of depth from overlap, from position on the ground and from what the figures are standing next to. A large figure at the back reads as nearer first, and only reads as anomalous if the ground or the overlaps contradict it — which a competent painter arranges that they do not.

So the convention hides inside the slack between the geometry and the perception. This collection can measure the first and says nothing about the second, which is the honest limit it states wherever the two meet: the geometry says where a mark must go for a picture to be a projection, and it says nothing at all about what a viewer will accept.

It is worth naming the other conventions the same reading can come from, because a test that fires does not identify the cause.

Register pictures draw every figure at the same height regardless of band, which is a picture with the depth signal removed rather than replaced; the test fires, and it fires with a signature of its own that the next essay measures.

Two grounds — the carpet drawn from above with the people on it drawn from the side — is a picture with two optical axes, and figures on the tipped-up ground do not obey the standing ground’s construction at all. A carpet and the people on it measures that arrangement and finds no single camera fits both halves.

Continuous narrative with one copy enlarged is the case the same person, twice on one panel uses as its control, and it is the same reading arriving from a panel that is otherwise a photograph.

A tilted picture plane changes the relation between drawn height and depth without any convention being involved, and the horizon is at eye level — if the picture plane is vertical is where the collection prices the assumption this test is quietly making.

So a large reading means the picture is not a photograph of equal figures on one level floor. Which of the four it is takes more than the one number.

Why the convention is worth the reader’s respect

There is a temptation, given a test that catches something, to treat the catching as a verdict. It is worth resisting for a reason this field states everywhere and this essay is a good place to state again.

A picture that scales by rank is answering a question a photograph cannot answer: which of these people matters. That question has no geometric content and no camera can record it, so a painter who wants it in the picture has exactly one channel available — size — and using it means giving up the channel’s other job. The convention is a trade, and the thing traded away is the one this collection can measure, which is why the collection can see the cost and not the benefit.

The same asymmetry runs through the whole field. What perspective gave up is the collection’s statement of it from the other side: the system that keeps the geometry gives up things that other systems keep, and a table of what each keeps is a fairer instrument than a ranking.

The boundary, stated

A level camera, figures standing on one plane, and two of them the same height. Each is required and each is checkable.

The level camera makes the uprights vertical in the picture and the horizon horizontal, which is what lets two references fix it. A pitched camera converges the verticals, the horizon is still a line but is no longer found from one point, and the construction needs three references — at which point the corruption problem returns and the test has to be run on subsets.

The one plane is what makes a horizon exist. And the two equal heights are the reference: the test is a comparison, so it measures every other figure against the pair chosen, and choosing a pair that is itself rank-scaled reports the honest figures as the anomalies.

Carried across, the height misses the drawn head by 73 pxThe test a sceptical reader runs with a straightedge, drawn on the picture it is run on. Two of the figures fix a point on the horizon; the horizon is then used to carry the first figure's height across to each of the others, and the predicted head is compared with the drawn one. Where the picture is a photograph the two coincide exactly. Where the heights record rank they do not, and the worst miss here is 73.3 pixels. Nothing about the camera is needed and nothing is measured: the whole construction is joins and meets, which is why it survives being carried out on a photograph of the painting.two references, a straightedge, no arithmetic73 px
Fig. 7 A third of rank in the picture, where the misses are large enough to read and small enough that a viewer would not notice them.

What is measured here

Four numbers.

Carrying a height across a photograph of five figures lands on the drawn head to 8.5 × 10⁻¹⁴ pixels. On a picture whose heights are ranks alone it misses by 209. In between the miss is linear in the rank content to the precision of the arithmetic — 21 pixels at a tenth, 42 at a fifth, 209 at the whole. And the statistic this replaced, the scatter of the pairwise horizon points, read 0, 41, 509, 139, 175, 439, 1751 over the same evenly-spaced sweep, which is what says a bounded detector was worth building.

The short version

Size as rank and size as distance can be told apart on a print, with a straightedge and two references, by carrying a height across the room and seeing where it lands. The miss is zero on a photograph, two hundred and nine pixels on a picture scaled by rank alone, and linear in between.

A tenth of rank is already twenty-one pixels, which clears a reader’s reading error comfortably; a fiftieth is not. And what the test reports is that the picture is not a photograph of equal figures on a level floor — which convention it is instead takes more than one number.

Sized by rank at 70 per centFive figures standing on one ground, drawn with their heights decided partly by how far away they are and partly by how important they are. At the left-hand end of the slider the picture is a photograph; at the right-hand end the drawn height records rank alone and the ground is irrelevant to it. Most surviving pictures that use the convention are somewhere between, which is the reason for the slider: an Assyrian relief, a Byzantine donor panel and a great many altarpieces are not choosing one rule and abandoning the other, and a test that only fired on the pure case would catch none of them.correct from 15 cm, at 160 mm widefive figures · 70% rank
Fig. 8 Seventy per cent, which is about where a good many surviving panels sit: still readable as a room, and no longer a record of one.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

DetectionDiminutionDrawing conventionForeshorteningGround planeHorizonIdentifiabilitySensitivitySize-distanceStraightedge construction