Three distances in one landscape
Worth reading first: A scroll is a camera that moves · What a 360-degree photograph actually is.
Guo Xi’s eleventh-century essay on landscape names three distances a painter should be able to draw: the towering distance, looking up at a mountain from below; the deep distance, looking into it from the front; and the level distance, looking across from a height. They are described as three things a painter does, and they are usually read as three moods.
This essay reads them as three cameras, because that is what makes the arrangement measurable — and what comes out is a trade with a number on both sides.
Three stations, one strip
Set three cameras at different heights above one landscape: low for the near ground, higher for the middle distance, higher again for the far. Each takes its own band of depth. Then lay the three bands out on one surface, each butted against the one before.
Butting them is not a simplification. Try instead to put all three in one camera’s frame and the bands overlap and invert — the middle band’s near edge lands below the near band’s far edge, so the far ground would be drawn nearer the bottom of the picture than the near ground. A painter does not do that; a painter places the bands, and the placing absorbs the offset.
So the discontinuity a reader actually sees is not an offset. It is a change of rate: how many pixels of picture one metre of depth is given, just before and just after each join.
What the far distance costs a single camera
The reason for the whole arrangement is a fact about a single camera that has nothing to do with China.
The image of a ground point at depth sits at a height above the horizon proportional to , where is the eye’s height. So an interval of depth from to occupies a picture height proportional to : the square of the depth, not the depth. That is the same one-over-depth-squared law depth is a reciprocal measures for a stereo pair, arriving in a landscape.
The consequence is severe. A camera at 1.6 metres gives the 8-to-26-metre band ninety-seven pixels of a four-hundred-and-thirty-pixel picture, the 26-to-70-metre band twenty-seven, and the 70-to-200-metre band ten. The far distance — the mountains, which are the subject — is a smear ten pixels tall at the horizon.
Raising the eye for that band restores it. Eleven metres of elevation gives the same 70-to-200-metre band seventy-two pixels instead of ten, which is 6.9 times as much room, and the mountain becomes a thing with parts in it rather than a band of texture.
What the seam costs
The price is that the rate changes abruptly at each join. Just below the first join a metre of depth is worth a certain number of pixels; just above it, 2.50 times as many. At the second join the factor is 2.75.
A reader’s eye absorbs that, and the reason it absorbs it is worth naming: a rate of depth is not something a viewer measures. What a viewer reads is the ordering, the overlaps and the relative sizes of things known to be similar, and none of those is discontinuous across the join. A tree just below the seam and a tree just above it differ in drawn size by a factor that reads as a difference in distance, which is exactly what it is being used to mean.
The control says that the seam is the convention rather than the layout. Three cameras at the same height, butted the same way, give a rate ratio of exactly 1.000 at both joins — no seam at all, and no gain either.
The gain and the seam are the same number
The three measurements — 2.50 at the first join, 2.75 at the second, 6.9 for the far band — look like three independent readings of the arrangement. They are one reading, and seeing that turns the convention into a conserved trade rather than a set of choices.
A ground point at depth images at a height above the horizon proportional to , so the rate at which depth runs — picture height per metre of depth — is proportional to . At a join, the two bands are evaluated at the same depth, so the cancels and the jump in rate is simply
Not approximately: the join depth drops out, the focal length drops out, the picture’s size drops out. The eye heights are 1.6, 4 and 11 metres, and and — the two measured seams, from the elevations alone.
The same cancellation runs the other way for the gain. A band’s picture height is proportional to its own with everything else fixed, so raising the eye from 1.6 m to 11 m multiplies the far band’s allowance by — the 6.9 the figure reports.
So the gain is the product of the seams, exactly: . That is the whole trade in one line, and it says the convention has no free lunch anywhere in it. A painter who wants the far distance times larger must raise the eye by , and the seams along the way must multiply to whatever else is done with them.
What is free is how the factor is distributed, and that is where the count of bands stops being Guo Xi’s number and starts being a design decision. Two joins carrying a total of 6.875 need each if they are shared evenly. Six joins need each, which is a rate change small enough to sit inside the variation a viewer already tolerates. The total is fixed; the conspicuousness of any one seam is not.
And that identifies the limit the family runs to. Let the number of joins grow without bound and the ratio at each goes to one: the eye rises continuously as the picture recedes, no seam anywhere, and the arrangement becomes a translating camera whose track climbs. That is a handscroll with a rising eye — the pushbroom of this field’s first rung with a height that varies along the track. So the three-distances convention and the scroll are not two neighbouring conventions but the two ends of one family, indexed by how many pieces the elevation change is broken into, and the piecewise version’s only advantage is that a painter can draw three cameras and cannot draw a continuum.
Which is not a panorama, and not a scroll
Three names are close enough to conflate and the field has already separated two of them.
A panorama is one station with the picture surface curved around it, so it is a single projection onto a cylinder and it has a centre. What a 360-degree photograph actually is sets that out. The three-distances arrangement has three centres and is not a projection onto anything.
A handscroll is a translating eye, which a scroll is a camera that moves models as a pushbroom: one station per column of the picture, all at the same height, moving sideways. That gives a continuous family of stations rather than three, and its seams are infinitesimal rather than visible — the price appears as a bend in every receding line instead of as a jump at a join, which is what a straight line in a scroll is a hyperbola measures.
The three-distances arrangement is neither. It is a piecewise picture: finitely many projections, each correct on its own band, joined at edges where nothing is continuous but the subject.
Why the rate and not the offset
The choice of statistic is worth defending, because the obvious one gives a wrong impression and gives it confidently.
Measure the offset — how far apart the two cameras place the same depth in the picture — and the numbers are large: several hundred pixels at each join, on a picture four hundred and thirty pixels tall. That reads as a catastrophic discontinuity, and it is an artefact of pretending the bands share a frame. They do not; a painter places them, and placing them removes the offset by construction. A statistic that measures something the arrangement was never trying to do is measuring the model.
The rate survives the placing. Translating a band up or down the surface does not change how many pixels a metre of depth gets inside it, so the ratio of the rates on either side of a join is invariant under exactly the freedom the painter has. That makes it the right quantity: it is what remains after the arrangement’s own degrees of freedom have been used up.
This is the same move the slip that leaves no trace makes in the construction field, where a residual had to be converted from a ratio into a displacement before it measured the hand rather than the pavement’s length. Choosing a statistic that is invariant under the things that do not matter, and sensitive to the thing that does, is most of the work in both cases.
The three named distances, as three elevations
Guo Xi’s three distances map onto the three stations in the obvious way, and the mapping is worth writing down because it explains the names.
The towering distance is the low station looking up: a near mountain drawn from below, with its base close and its summit high in the frame. That is a camera at 1.6 metres with a steep upward axis.
The deep distance is the middle station looking into the landscape past the near range, which needs enough elevation to see over the first ridge and not so much that the ground flattens.
The level distance is the high station looking across, where the far ranges are laid out one behind another and the ground is nearly a plan. Eleven metres is a modest version of it; a real painting implies a great deal more.
The names are about what the painter is looking at rather than about where the painter is standing, which is the ordinary way a practical vocabulary works. Recovering the station from the name is this essay’s contribution, and it is what turns three moods into three numbers.
What a reader can and cannot recover
The ledger for the arrangement follows from the ledger for each band.
Within a band, everything a single perspective gives is available: the ground is a projection from a point, the horizon of that band is findable, the depths are recoverable up to the usual one length from outside, and the whole of the metrology field applies unchanged.
Across a join, almost nothing. The two bands have different stations, so a length in one and a length in the other are not comparable without knowing both elevations — which the picture does not state and cannot. The ratio of two depths in different bands is not recoverable, and neither is any angle between directions lying in different bands.
That is the same division a picture in bands finds for a register, arrived at from a completely different construction. A register removes the depth signal and gives an ordering; a three-distances landscape keeps the depth signal within each band and gives an ordering between them. The second is strictly more informative and it has the same boundary.
Recovering the stations, and what it would take
A reader given a three-distances painting and asked where the painter stood has a well-posed problem within each band and an under-determined one across them.
Within a band, the horizon of the ground and one known length give the station height, which is the standard single-view recovery. Across bands, the three heights are only comparable if something is common to two bands — an object appearing in both, a ridge line continued, a ground plane shared. A painting with no such object gives three heights in three unrelated units.
That is a genuine gap rather than a difficulty, and it is the same gap the collection meets whenever two projections are joined without a shared feature. A projection of a projection is where the composition is a single homography and everything transfers; here the composition is not a map at all, because the two bands are pictures of different parts of the world from different places.
The convention is a choice of picture surface, one level up
There is a way of reading the arrangement that puts it beside the curved field rather than beside the scroll field, and it is the reading that generalises.
No surface keeps everything establishes that a picture surface is a decision with a measurable price: a plane keeps straight lines and cannot reach a hundred and eighty degrees, a cylinder reaches all the way round and bows every horizontal, stereographic keeps every angle and no scale. Each is a single map from directions to marks, and the field’s whole table is a comparison of what each spends.
A three-distances landscape is a piecewise answer to the same question. Rather than choose one map and pay its price everywhere, it chooses three and pays the price only at the joins. That is a strictly larger design space and it buys exactly what a piecewise construction always buys: better behaviour in the interior, discontinuity at the boundaries.
The same trade is in the cube map, where six flat pictures cover every direction and the price is paid at the twelve edges — six flat pictures of everything measures that. It is in a tiled render, where a curved screen is approximated by eight flat frusta. And it is in every atlas of a curved surface. Naming the family is worth doing because it says what the seams are: not failures of the convention, but the whole of its cost, concentrated where it can be put somewhere unobtrusive.
The boundary, stated
Three level cameras at three heights, with the bands butted. Three simplifications, each of which a real painting departs from.
Real landscapes are painted with pitched axes as well as raised eyes, and the towering distance in particular is defined by looking upward. Adding pitch changes where each band’s horizon falls and therefore how the bands butt, without changing the rate argument, which depends on elevation alone.
Real bands are not rectangular: a painter runs the near ground up one side and the far range down the other, so the joins are irregular curves rather than horizontal rules, and the discontinuity is distributed rather than concentrated. That makes the seam far less visible and does not make it smaller.
And the count is not always three. Some scrolls have one transition and some have half a dozen, and nothing in the geometry prefers three — the number is Guo Xi’s and it is a claim about painting rather than about projection.
What a reader can measure on a real scroll
The arrangement above is built rather than read, so it is fair to ask what a reader with a photograph of a real landscape could actually recover.
Within a band, the horizon is findable wherever the picture contains two objects of equal height standing on the ground — the construction the same person, twice on one panel uses, applied inside one band. Two figures, two trees of a kind, two posts of a fence: any of them gives a horizon point, and three give a line.
Each band’s horizon then sits at that band’s own eye level. So the relative elevations of the three stations are readable directly off the picture, as the heights of the three horizons above their bands’ ground lines — provided something of stated size appears in each band to give it a metre, which is exactly the closure problem the metrology field spends four essays on.
Without such an object, what comes back is the ratio of the three elevations, which is still a real measurement and is the one a reader is most likely to want: it says how much the painter raised the eye between the near ground and the far, in units of itself.
What is measured here
Three numbers and a control.
The far band, 70 to 200 metres, occupies seventy-two pixels under its own station and ten under a single camera at the near station — a gain of 6.9. The rate at which depth runs jumps by 2.50 at the first join and 2.75 at the second. And three cameras at the same height, butted the same way, give a rate ratio of 1.000 at both joins, which is what says the jump is the elevation rather than the layout.
The short version
A landscape assembled from three stations gives its far distance almost seven times the picture a single camera would allow it, because the image of a depth interval falls as one over depth squared and raising the eye is the only thing that gives it back.
The price is a discontinuity in the rate at which depth runs — two and a half times at one join and two and three quarters at the other — which a viewer absorbs because a rate of depth is not something a viewer reads. Within a band the picture is an ordinary perspective and everything applies; across a join, almost nothing does.
What links here
Computed from the collection, not written here: the essays that point at this one.
Reads more easily once this is understood
Essays that name this one as worth reading first.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A set cut for one eye — both name depth compression, foreshortening, free parameter, station point
- Every row is a different camera — both name handscroll, moving viewpoint, station point
- The centre a scroll does not have — both name handscroll, moving viewpoint, station point
- Two grounds, and what the second one costs — both name foreshortening, free parameter, station point
- Two stations in one picture — both name drawing convention, free parameter, station point
- A carpet and the people on it — both name foreshortening, station point
Named objects
A flat tag is an object no other essay names yet.
Depth compressionDrawing conventionElevationForeshorteningFree parameterHandscrollMoving viewpointPiecewise mapSeamStation point