The eye that moves

A scroll is a camera that moves

A Chinese handscroll is not a picture with a wandering viewpoint or a picture with no viewpoint. It is the image of an eye that travels along a track and records one vertical line at a time, and that object has an exact geometry — orthographic along the roll, perspective across it.

Worth reading first: Parallel projection is not primitive perspective · The eye taken to infinity.

A Chinese handscroll is several metres long and is looked at a shoulder-width at a time, unrolled from one hand into the other. The standard account of its geometry is that the viewpoint moves with the reader, and that account is true and stops one step before the interesting part. If the viewpoint moves, it moves along something, and what it moves along is a track.

That is enough to build the thing.

A scroll in plan: the eye travels, and images one line at a timeThe eye runs along the track at the bottom. Each position images the single vertical plane it is level with, so a world point is drawn by exactly one position of the eye — the one at its own x. The paper advances 26 px for every metre of travel whatever the scene does, which is why the roll is a map along its length.the eye's trackthe eye at x = -10.5 mevery point is drawn by the one position of the eye that is level with itplan — the eye's track and the scans it makes28 m of travel
Fig. 1 The arrangement in plan. The eye runs along the track at the bottom; each position images the one vertical plane it is level with; a world point is drawn by exactly the position of the eye that is level with it. Nothing about the scene is assumed — the pavilions sit at whatever depths they sit at, and the paper advances at a constant rate regardless.

The two projections, separated

Give the eye a position that depends on one parameter — it slides along the world’s x axis, at a fixed height, looking sideways at whatever is in front of it. At each position it records a single vertical line of the picture. Then a world point at (X,Y,Z)(X, Y, Z) is imaged by the one position of the eye at x=Xx = X, and its two picture coordinates come out of two different operations:

u=u0+sXv=v0+f(heY)Zzcu = u_0 + s\,X \qquad v = v_0 + \frac{f\,(h_e - Y)}{Z - z_c}

The first has no ZZ in it. Depth does not affect where a point lands along the roll, because the scan that images a point is chosen by its xx alone. The second is the ordinary perspective divide, unchanged from lib/camera.js.

So a handscroll is orthographic along its length and perspective across it. That is not a compromise between the two families this site has been describing; it is one member of each, glued along an axis. The consequences are all downstream of that one sentence and every one of them is measurable.

The same row of 5 pavilions, by a pinhole and by a scrollAbove, one pinhole at 62°: the row is oblique to the eye, so the far pavilion is drawn 2.38× smaller than the near one. Below, the scroll: every pavilion is drawn at the same height to twelve decimal places, because along the roll the eye is at infinity. Neither is a defect — they are answers to different questions.one pinhole, standing backfar end 2.38× smallera scroll, unrolledevery end the same sizeno single viewpoint — the rays miss by 9.4 m across the lower stripthe upper strip is one pinhole at 62°
Fig. 2 The same row of pavilions, twice. Above, a single pinhole standing off to one side, which is how anybody photographing a long façade actually stands: the far end is drawn 2.38× smaller than the near one. Below, the scroll, which draws them at the same height to twelve decimal places. The scroll’s strip states no viewing distance because it has none to state.

The pinhole in that figure stands obliquely rather than square-on, and the reason is worth a sentence because it is a fact about pinholes that the word diminution obscures. Imaged height is fH/zfH/z, where zz is depth along the optical axis — so a row of equal objects all at one depth is drawn at one size by a pinhole too. What a pinhole does to such a row is stretch its ends, which is a different essay entirely. To make a pinhole diminish a row it has to see the row obliquely, and then the far end really is further along the axis. The first version of the figure above asserted a diminution of 1.000× and was right to fail.

Why the scroll’s scale does not run out

The property that makes a handscroll possible as an object is that it can be arbitrarily long without anything degenerating. A perspective picture cannot: extend the depth range far enough and everything past some distance lands within a few pixels of the horizon, which is what a horizon is. Extend a scroll and it simply gets longer, at ss pixels of paper per metre of travel, for ever.

The reason is the missing ZZ. Along the roll, the map from world to paper is a scaling, and a scaling has no accumulation point. A mile of river laid out on eight metres of silk is drawn at one constant scale from end to end, and the scale is a number the painter chooses rather than a consequence of where anybody stands.

This is the same fact that makes a parallel projection unbounded, arriving in one direction only. The scroll takes the eye to infinity along the roll and leaves it seven metres away across it — which is the sentence this whole field turns on.

What each parallel system does to the three axesThe smallest axis scale is plotted; the note says how many of the three coincide and whether the system is an orthographic projection or an oblique construction. cavalier and isometric both make all three the same — but isometric is the only orthographic one that does, at 0.8165, and the oblique route gets there by drawing depth at true length.elevation0.0000two equal, orthographiccabinet0.5000two equal, obliquecavalier1.0000all three equal, oblique ←dimetric0.4714all three different, orthographicisometric0.8165all three equal, orthographictrimetric0.5479all three different, orthographicsmallest of the three axis scalesmeasured from each projection
Fig. 3 The parallel systems, for comparison: every one of them has taken the eye to infinity in all directions at once, which is why every axis scale is a constant. A scroll does it in one direction and keeps a real focal length in the other, so it belongs to neither column of this table.

What the compression is, and what it is not

The number ss — pixels of paper per metre of travel — has no analogue in a pinhole camera, and it is worth being precise about what it does and does not control.

It is not a focal length. A focal length sets an angular scale and is tied to a viewing distance by the arithmetic this whole site is built on: the focal length scaled to the width the figure is displayed at is the distance the reader must be at. The scroll’s compression sets a linear scale, metres of world to millimetres of paper, and there is no distance at which it becomes correct because there is no angle involved.

It is also not free of consequences. Two of them:

The compression sets how much world a section of paper covers, which is the quantity the next two essays measure. A section of paper 60 cm wide at a compression of 26 px per metre covers about 26 metres of the eye’s travel, and that number is what the missing centre is made of.

And it is the one parameter a viewer never renegotiates. The reader of a scroll chooses how fast to unroll and how much to have open, and both of those change the experience considerably; neither changes the geometry. The compression was fixed when the silk was painted.

The rays of 12 m of scroll, and the point they miss by 3.72 mPlan of one section. Each ray leaves the eye at its own column, so the eyes lie along a track rather than at a point. The circle is the least-squares centre drawn at the radius of its own miss — 3.72 m, which the closed form puts at 3.72 m, the standard deviation of a track that long. A single column of the same scroll fits exactly.the best point, missed by 3.72 m12 m of the eye's trackno single viewpoint — the rays miss by 3.72 mthe eyes are a track, not a point
Fig. 4 What a section of the picture is made of: rays that leave the eye at their own columns. The circle is the least-squares point they come closest to, drawn at the radius of its own miss. This is the subject of the third essay in this field, and it is here to make the point that the geometry above is not an approximation to a camera — it is a different object.

The model comes out of how the object is used

Nothing above was assumed. The two coordinates fall out of one fact about handscrolls that has nothing to do with geometry: they are read a section at a time, and the section is about the width of the reader’s outstretched hands.

Consider what that implies. At any moment there is a piece of the picture in front of the reader and a great deal of picture rolled up on either side. The rolled-up parts are not being looked at from anywhere, because they are not being looked at. So whatever geometric relationship holds between the reader and the picture, it holds locally, and the painter’s problem is to make every local section work without any global station point existing.

The pushbroom model is the minimal solution to that problem. Give each column of the picture its own eye; place those eyes on a track; and the section in front of the reader at any moment is drawn by a short arc of that track, which is nearly a point. The shorter the section, the nearer to a point, and the third essay in this field measures exactly how near.

A perspective picture solves the same problem the other way and pays for it. It has one station point, so a reader in the right place sees a correct projection and a reader anywhere else does not — which is the whole of the viewing field and is a real cost on an object several metres long, because there is no single place to stand in front of a five-metre picture.

Where the reader has to be for a 46° picture to be correctShown 160 mm wide, this picture is a correct projection only from 19 cm away. Drawn to scale.the picture, 160 mm wide19 cm46°the eyefocal length 813 px19 cm at 160 mm wide
Fig. 5 The cost a single station point carries, on this site’s own machinery: the picture is a correct projection from one point and from nowhere else. For a picture the size of a page that is a mild constraint. For a picture the size of a room it is the constraint that the handscroll’s geometry exists to avoid.

What the reader can recover, and what is gone

A useful way to characterise any drawing system is to ask what a viewer holding a ruler can get back out of it. For a perspective picture the answer is ratios along a line and nothing metric without a reference length. For a scroll the answer splits, and the split is exactly the split in the model.

Along the roll, everything. Distances along the eye’s track are recoverable up to the one scale factor ss, because the map is a scaling. Two pavilions measured on the silk are in the same ratio as the two pavilions on the ground, whatever their depths. A reader who knows the width of one courtyard knows the length of the whole journey.

Across the roll, the usual perspective answer. Depth is recoverable only up to the usual ambiguities, from the usual invariants, with the usual reference required.

That asymmetry is unusual enough to be worth stating plainly: a handscroll is a map in one direction and a picture in the other, and both halves of that are exact. No other system on this site does two different things to its two axes.

A scroll keeps the midpoint along its length and loses it acrossLeft, a segment lying along the roll: the image of its midpoint and the midpoint of its image are the same point to 0e+0 px. Right, a segment running away from the eye: the two are 21.9% of the segment apart. One projection, two answers, because the eye is at infinity in one direction and seven metres away in the other.along the rollthe two midpoints coincide — 0e+0 pxone mark, drawn twiceacross itthey separate by 21.9%the image of the midpointthe midpoint of the imageno single viewpoint — the rays miss by 6.9 m21.9% of the receding segment
Fig. 6 The asymmetry as a measurement. Above, a segment lying along the roll: the image of its midpoint and the midpoint of its image coincide to arithmetic noise, which is this site’s own test for whether a projection preserves ratio. Below, a segment running away from the eye: the two separate by 21.9% of the drawn segment. The fourth essay in this field is about what that costs and what it buys.

The model has a name, and the name is worth knowing

An eye that translates while imaging one line at a time is a linear pushbroom camera, and it is not a historical curiosity. It is what a satellite is. Earth-observation instruments from Landsat onward image a single across-track line and use the spacecraft’s own motion to build the second dimension, for the same reason a scroll does: the track is free and the sensor is cheaper in one dimension than in two.

That coincidence is not an argument that the painters of the Song dynasty anticipated remote sensing, and nothing here should be read that way. It is an argument about the space of drawing systems: the map from world to page that a handscroll uses is one somebody else arrived at independently, from an engineering constraint rather than a pictorial one, and the geometry does not care which. A system is a rule, and rules can be reached from either end.

What the coincidence does buy is a body of results. The image of a straight line under a pushbroom, the form its epipolar geometry takes, the conditions under which its images can be rectified — all of these are worked out, and the next essay uses the first of them.

One thing a scroll cannot do

Every system in this field buys something and pays for it, and the scroll’s bill is worth naming before the measurements start, because it is not the one usually named.

The usual complaint is that a scroll has no consistent viewpoint, which is true and is the subject of the third essay and is not really a cost — it is what the system is. The actual cost is narrower and harder: a scroll cannot turn a corner.

The eye’s track runs along one direction. A courtyard whose far side faces back toward the near side has two walls the track can only see from one side, and no amount of unrolling reaches the other. A pinhole has the same problem and solves it by turning, which a scroll cannot do without abandoning the constant-scale property that is the whole point. A panorama turns and keeps its centre; a scroll travels and loses it; nothing does both.

What the painters did about this is a question about painting rather than about geometry, and the geometry has nothing to say beyond noting that the problem exists and is structural. The convention of the removed roof, which the next field measures, is one answer to a closely related problem in a different system.

Curving straight lines and having no centre are two different thingsThe rms miss of the best single centre, for two cameras that both draw straight world lines as curves. A rotating eye keeps its centre exactly — 2e-15 m, which is the solver's noise floor. A translating eye has none: 7.97 m over 27 m of track. A panorama is a projection and a scroll is not, and no amount of looking at the curves tells them apart.a rotating eye — the panorama2e-15 ma projectiona translating eye — the scroll7.97 mnot oneboth of these draw a straight world line as a curverms miss of the least-squares centreone of them is a projection
Fig. 7 The two ways an eye can move, on the one question that separates them. Both draw straight world lines as curves. A rotating eye keeps its centre exactly — the number is the solver’s own noise floor — and a translating one has none at all. The last essay in this field is about how thoroughly these two get conflated.

Where the model stops

Three honest limits, because a model of a convention is not a reconstruction of an object.

A real handscroll is not a strip photograph. It has passages that are not one continuous space at all — a mist band, a change of season, a jump of a hundred li — and no single map from a world to a page describes those, because there is no single world. What the model describes is the geometry within a continuous passage, which is where the interesting questions are and is not all of the object.

The track need not be straight, and in a real scroll it is not. The model here runs the eye along a line at constant height. A painter who takes the viewer along a winding river is running the eye along a curve, and the geometry of that is the geometry here with XX replaced by arc length. Nothing in the results below depends on the track being straight; they depend on it being a track rather than a point.

And the model says nothing about why any of this was done. The handscroll’s geometry serves a way of looking that unrolls in time, and time is the half of the object this site cannot measure. What can be said is that the geometry is not a failure to achieve perspective — it is a different rule, consistently applied, whose properties are exactly computable and mostly better than perspective’s on the axis the object cares about.

What each drawing system preservesEvery cell is read out of the system's own map rather than typed: the centre is fitted from its rays, measure is the site's own midpoint test, diminution is the drawn size of a fixed object at two depths, bounded depth asks whether infinite depth lands in finite page, and straightness is the sag of a receding line. Only perspective has a centre, and it is the one system with no true measure. No row has both.a centretrue measurediminutionbounded depthstraight linesperspectivehandscrollisometricdimetrictrimetriccavaliercabinetelevationfilled means the system keeps itmeasured from each system's own projectionno row is filled in both of the first two
Fig. 8 The scroll’s row on the field’s comparison table, with the pinhole’s above it for reference. Every cell is read out of the system’s own map rather than typed: a centre is fitted from its rays, measure is this site’s own midpoint test, and straightness is the sag of a receding line. The scroll is the only row that answers differently in its two directions, and that is what the rest of this field is about.

What the rest of the field does with this

Four consequences, one per essay, and each of them is a measurement rather than a description.

A straight line that changes depth is drawn as a hyperbola, exactly, and the closed form for it agrees with the projected samples to fourteen decimal places. A line that holds its depth is drawn straight. So curvature in a scroll is a depth signal, not a stylistic one.

A section of the picture has no centre of projection, and the miss is not a small residual: it is the standard deviation of the eye’s own track, in metres of the world, and it grows linearly with how much of the scroll is open. There is no width at which the failure begins and none at which it saturates.

Measure survives along the roll and not across it. The midpoint of a segment lying along the length images to the midpoint of its image exactly; the midpoint of a receding one is 22% of the way off. One projection, two answers.

And a scroll is not a panorama, though both draw straight lines as curves. A panorama is an eye that rotates, and a rotating eye keeps its centre exactly however far it turns. That both objects curve their lines and only one of them is a projection is the distinction the two are usually conflated over, and it takes one number to separate them.

Two figures to keep in view

The field’s whole argument is visible in two pictures and it is worth naming them here, at the start, so the measurements that follow have somewhere to land.

A scroll has a centre only for a section of no widthThe least-squares miss against how much of the scroll is fitted at once. It passes through the origin — one column is a pinhole camera and fits to 4e-15 m — and it is straight: 2.000× for twice the section. The closed form (the line) is the standard deviation of the eye's own track, and the measurements sit on it to 4e-15 m.0246801020metres of the eye's track the section coversrms miss of the best centre, mmeasuredthe closed formleast-squares fit against √12 of the track lengththey agree to 1e-9 m
Fig. 9 The first: the miss against how much of the scroll is fitted at once. A straight line through the origin, so there is no width at which a scroll becomes a projection and none at which it stops failing to be one. Everything the third essay in this field says is on this line.
Where a depth range lands on the pageEach rule is one depth, evenly spaced from 3 m to 60 m, drawn at the height its system puts it. The shaded band is the top tenth of each strip. A pinhole files 69% of the whole depth range into it; a linear depth map puts exactly 10% there, because a tenth of a page is a tenth of anything under a linear map. That band is what a horizon is.a pinhole69% of the rangean oblique system10% of the rangeevenly spaced depths, drawn where each system puts themthe shaded band is the top tenth of the stripa horizon is a band, not a line
Fig. 10 And the second, which belongs to the next field and explains what the roll direction is for. A pinhole files 69% of its depth range into the last tenth of the page; a linear depth map puts exactly 10% there. A scroll’s roll direction is the linear case, which is why a journey can be laid out along it at constant scale for as long as the silk lasts.

What links here

Computed from the collection, not written here: the essays that point at this one.

Reads more easily once this is understood

Essays that name this one as worth reading first.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Focal lengthForeshorteningHandscrollMoving viewpointOrthographicParallel projectionPicture planepoint at infinityPushbroomStation point