Higher on the page, and where that stops being true
Worth reading first: A picture with no size–distance signal · A picture in bands.
A picture with no size-distance signal measured what a system with no diminution loses: nothing in an oblique drawing’s drawn size tells a reader how far away anything is, because the size-distance slope is zero to twelve decimal places. Size that means rank then measured a convention that keeps the machinery of diminution and repurposes it, drawing importance rather than distance at the same falling rate a real depth cue would use. Both essays take diminishing size as their subject. This one takes a different, older convention that never draws anything smaller at all: a figure standing further back on a ground plane is simply drawn higher on the page.
That convention is usually presented as an alternative to perspective — a simpler, cruder rule invented before anyone worked out vanishing points. It is not an alternative. It is what a pinhole camera already does, for a particular class of point, before any lens distortion or any deliberate convention is added at all. A point on a ground plane below the eye images higher on the page exactly as it moves further away, because that is what the projection arithmetic returns, and the interesting content of “higher on the page means further away” is not that it is true — it is where it stops being true, and by how much it compresses everything that stays true.
This is worth stating carefully because the ordinary way of introducing the convention gets the causal arrow backwards. A draughtsman working before any construction of a vanishing point existed did not invent a rule and then discover, centuries later, that a camera happens to obey something similar. A pinhole obeys this ordering for any depiction of a ground plane, drawn or photographed, whether or not anyone drawing it knew what a station point was — the ordering is a fact about central projection, not a fact about a particular historical technique for approximating one. The convention, read this way, is not a stage on the way to perspective; it is the single fact about perspective’s own image that survives when every other fact about a specific station point, a specific focal length and a specific picture plane is stripped away and only the ranking of heights on the page is kept.
The picture: rungs receding to the horizon
The claim is easiest to see as a picture before it is measured as a function, so start with an ordinary photograph of a row of posts on the ground.
Two things are true of that image and worth separating immediately, because the rest of the essay is organised around the difference. The ordering — further rungs sit higher — holds for every pair drawn here, with no exception and no reversal anywhere in the picture. The spacing does not hold anything like as cleanly: the gaps between adjacent rungs shrink sharply toward the horizon, so a reader’s eye reading off “roughly how much further” from the drawing alone is reading a quantity the picture actively misrepresents by compressing it. A pinhole gives the order for free and gives the metric a severe discount, and 57% is the size of that discount stated as a single number: better than half of the entire six-to-250-metre range is squeezed into the final tenth of the vertical band the ground occupies on the page.
The map, drawn rather than described
“Higher on the page means further away” is usually stated as a sentence. It is worth having as the function it actually is, because a function can be asked whether it ever turns around, and a sentence cannot.
A curve that falls monotonically and never turns is exactly what “further is always higher, with no exceptions” requires arithmetically, and it is the reason a reader can, within the convention’s own scope, always read depth order off image height with total confidence — there is no depth at which two rungs swap places, because the function that maps one to the other has no local minimum or maximum anywhere on the ground plane below the eye. What the curve does not do is fall at a constant rate: near the camera, moving the object by one metre moves its image by nearly three hundred pixels; far from the camera, moving it by one metre moves the image by less than a hundredth of a pixel. Both facts are needed to state the convention honestly, and most accounts of “higher on the page” state only the first.
The reason the rate itself falls as the square of the depth rather than merely falling is worth having explicitly, because it is the same arithmetic that gives ordinary diminution its own inverse-depth law, one power lower. A point’s height above the horizon is proportional to eye-height-over-depth to begin with, which is why the curve falls at all; the rate of change of that quantity with respect to depth brings down a second factor of one-over-depth from the calculus of the falling curve, and the two together are what put a square in the denominator rather than a single power. That is not a numerical coincidence discovered by fitting a curve to the data above — it is the same projection arithmetic a picture with no size-distance signal already measured for drawn size, differentiated once, because “how high above the horizon” and “how large” are both, at bottom, one length in the world divided by one depth, read off two different parts of the same image.
The same map, drawn from a different eye
The 57% figure above is doing real work in this essay’s argument, so it is worth checking that it is a fact about the shape of the convention rather than an artefact of the one eye height used to draw it.
Both rates move and the fraction does not, which is the signature of a self-similar curve: raising the eye rescales the whole map by a constant factor without moving where, in depth, most of the accumulation has already happened. That is not a coincidence built into this one pair of numbers — the underlying quantity is height over depth, and multiplying the eye height by a constant multiplies the numerator of every term in the curve by that same constant, which cancels out of any question phrased as a fraction of the curve’s own total range. A taller station point changes how large everything is drawn. It does not change how much of the story is being told in the last few metres before the horizon.
Where the ordering inverts
The picture and the map above both stayed below eye level throughout, which was a choice rather than an oversight — the whole convention depends on it, and the essay’s sharpest claim is about exactly where that dependence gives out.
The crossover is not a fitted threshold discovered by sweeping a parameter until something looked close to zero. It is an exact consequence of the projection’s own geometry: the top of an object as tall as the eye sits on the horizon whatever its depth, because a ray from the eye to a point at eye height is horizontal regardless of how far along it the point sits, and a horizontal ray from the eye defines the horizon by construction. Below that height, a nearer object’s top subtends a larger angle below the horizontal than a further object’s does, which draws it lower and gives the familiar ordering; above it, the same geometry runs in reverse, because now the ray climbs toward the horizon rather than falling toward it, and a nearer object’s top climbs further above the horizontal than a further one’s does. The convention is not “higher means further.” It is “closer to the horizon in the direction the eye is looking down means further,” and below eye level that reads as higher on the page purely because down is drawn as up-the-page in a level ground picture. Restated that way, the inversion above the eye is not a failure of the rule — it is the same rule, correctly followed, once the direction being measured turns from downward to upward.
Two consequences follow immediately from stating it this way rather than as “higher is further, except sometimes.” First, the crossover height is not a fixed number carried by the convention itself — it is wherever the eye happens to be, so a picture with a low horizon puts the inversion low in the scene and a picture with a high horizon puts it correspondingly higher, and a reader who does not know the eye height of the camera that made a picture cannot say, from the ordering alone, which side of the crossover any one tall object’s top actually sits on. Second, the spread at the crossover is not merely small, it is exactly nothing — 0.0e+0 px across all five sampled depths — which is a stronger claim than “the convention becomes unreliable near eye level.” Unreliable would mean the ordering still exists but is hard to read; exactly zero spread means the ordering has genuinely stopped existing at that one height, because every depth maps to the identical image position and there is no information left in the picture to order them by even in principle. A convention that degrades gracefully and a convention that hits a true singularity are different findings, and this is the second kind.
The control: an ordering with no metric at all
A register picture is the cleanest control this convention has, because it is built from a different premise and arrives at a related but distinguishable conclusion — an ordering with none of a pinhole’s own metric attached to it.
A register keeps exactly the half of the story the ground-plane convention above also keeps — an order, further along corresponds to a certain arrangement on the page — and drops the half that convention still has, which is a single fixed camera behind the whole drawing. The ground-plane picture above is drawn from one stated eye height and one stated horizon, so its 57% figure is a fact about a fully specified camera; a register supplies neither, so a ratio read off it is not wrong exactly, it is underdetermined, with the 2.58-to-6.40 spread measuring the size of that underdetermination directly rather than measuring any error in the reading.
Fixing one assumption collapses the spread
The 2.5× spread above is a property of leaving the assumption open. Closing it is a single choice, and the effect of that choice is worth seeing as its own number rather than as a footnote to the previous figure.
This is what separates “the register is ambiguous” from “the register is wrong”: choosing a horizon position and an eye height is exactly choosing a camera, and once a camera is chosen a register picture reads precisely, because it is at that point being read as an ordinary ground-plane picture with the bands reinterpreted as points on it. The ambiguity was never in the arithmetic; it was in the missing camera, and the moment a reader — or a modern reconstruction — supplies one, the same nine-fold spread this essay’s own control measured a moment ago collapses to the single number 3.4688. What a register cannot do is supply that camera itself; a reader who wants the ratio has to bring the same “outside length” this collection has already found perspective needing for scale, only here what is needed is not a length but an assumed viewpoint.
A different system, the same accumulation
The 57% figure has been this essay’s own number throughout. It is worth checking it is not special to the particular depth range chosen for these figures by seeing the same accumulation measured on an independently built system, over a different range.
Sixty-three per cent over a 3-to-45-metre range and 57% over the 6-to-250-metre range measured earlier in this essay are not the same measurement and are not meant to agree exactly — different depth ranges, sampled differently, from a different construction entirely — but both numbers are far above the 10% a linear map returns under an identical shading test, and that “far above ten” is the fact the two independent constructions agree on. A pinhole’s accumulation near the horizon is not a property of any one depth range chosen to make the point dramatic. It is a property of the map itself, and any reasonable range shows it. The 10% control in this borrowed figure is doing the same job the linear comparison already did inside a picture with no size-distance signal: establishing what an unaccumulated map would report, so that the pinhole’s own 57-to-63% has something honest to be compared against.
The honest limit
Every number in this essay describes a level camera looking at a flat, unobstructed ground plane, and both restrictions are doing real work that a casual reading of “higher on the page” glosses over.
A tilted camera — one whose optical axis is not horizontal — moves the horizon off the frame’s own half-height and changes every rate in the map above without changing the underlying claim that image height is still monotone in depth on the ground plane it can see; the 57% figure specifically would need to be re-measured for that camera rather than assumed to carry over, since the accumulation fraction is a property of the particular relationship between eye height and sampled depth range rather than a universal constant. Carrying a height across the room and figures on a street that slopes both work with ground planes this essay’s own construction excludes — an inclined floor and a floor interrupted by a step — and the ordering claim above is silent about either, because “the top of an object at eye height sits on the horizon at every depth” is a fact about a level ground and a level camera together, not about height above the eye in some more general sense.
And nothing here says how steeply a reader perceives the compression to be, only how steeply the arithmetic compresses it. The horizon and the fraction is where this collection’s own caution about perceived versus computed effects gets its fullest statement for exactly this convention’s neighbourhood; the geometry computed here is exact and the human reading of “how much closer do these two rungs look” is a separate question this site does not answer.
What this is an instance of
The convention measured here is not a rival to perspective and never was — it is what perspective already does to any point below eye level, isolated and looked at directly rather than folded into a fuller construction alongside vanishing points and measuring lines. The horizon is at eye level establishes the crossover this essay’s own aboveeye figure measures directly: the horizon is exactly where a level camera’s optical axis meets the picture, which is exactly the eye height, which is exactly the height at which “higher means further” runs out of anything to say. A picture in bands is the convention that keeps this essay’s ordering and discards its metric on purpose, rather than merely lacking the metric for want of a stated camera, and comparing the two — an ordering convention that never had a metric to lose, against one that has a metric and spends 57% of it in the last tenth of the page — is the accurate way to place “higher on the page” among the other conventions this field has measured rather than beneath perspective as a lesser version of it.
What perspective gave up reads every convention in this field as a trade rather than a deficiency, and the trade here is stated precisely by the numbers above: a pinhole buys a total ordering of everything on the ground plane below the eye, for the entire depth of the scene, at the cost of spending the majority of the picture’s available height on the nearest tenth of that depth and almost none of it on the rest. A register buys the same ordering for less — no camera has to be specified at all — and spends the ability to say anything about magnitude in exchange. Both are honest accounts of what each convention keeps; neither is what the other one is trying and failing to be.
Read across the three essays on this anchor together, a single line runs through all of them. A picture with no size-distance signal measured a system that carries no diminution at all and asked what has to replace it; the answer there was exactly this convention, height on the page standing in for depth. Size that means rank measured a system that keeps the diminishing rate but repurposes what it stands for, so the same falling curve now reports importance rather than distance and a reader who assumes otherwise is caught by a straightedge test rather than by eye. This essay closes the sequence by measuring the convention those two essays both lean on without quite stating its own arithmetic — where it comes from, how much of the depth range it spends on the nearest tenth of the page, and the one height at which it has nothing left to say. None of the three is the general case the other two approximate; each is a different answer to the same question, what a picture does with depth once size has been asked to say something else, or nothing at all, about it.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A set cut for one eye — both name depth compression, depth cue, diminution, foreshortening
- A frame is an interval — both name depth cue, diminution, foreshortening
- A picture with nothing straight in it — both name eye level, ground plane, horizon
- A wall does not get darker as it goes away — both name depth cue, diminution, foreshortening
- Nothing moves when the object does — both name diminution, foreshortening, size-distance
- Stepping closer is not zooming — both name depth cue, diminution, foreshortening
Named objects
A flat tag is an object no other essay names yet.
Depth compressionDepth cueDiminutionEye levelForeshorteningGround lineGround planeHorizonSize-distance