A carpet and the people on it
Worth reading first: The point you have to stand at · One, two and three point are one construction.
A Persian miniature of a garden shows the carpet, the pool and the flowerbeds spread out as a plan — rectangles drawn as rectangles, the layout legible — and the people standing on them drawn as people are drawn, upright and at full height, with their faces visible.
Both halves are perfectly clear and the combination is impossible. This essay measures how impossible.
Why this is a geometry question and not a taste question
The usual framing of these pictures is that their makers were not interested in illusionistic space, which may well be true and is not what is at issue. The question here is narrower and has a yes-or-no answer: could a single camera have produced both effects?
That question is worth asking because if the answer were yes, the convention would be a choice among available options and the interesting discussion would be about why that option was taken. If the answer is no, then the convention is doing something no projection can do, and the discussion has to start somewhere else.
The answer is no, and the measurement below says how emphatically: the two effects require optical axes that are exactly perpendicular. Not merely different, not different by an amount that a compromise could split — orthogonal.
The two requirements, each satisfiable alone
To draw a horizontal rectangle as a rectangle, a camera must look down its normal. Any other direction images it as a general quadrilateral, because the two horizontal directions have finite vanishing points and the sides converge. This is not approximately true: the control in assertNoCameraHasBothGrounds puts a camera nine metres directly above a three-metre carpet and finds its imaged sides in the ratio 1.000000000000.
To draw a standing figure at its full relative height, a camera must look horizontally. A camera tilted down foreshortens verticals, by a factor that is the cosine of the tilt at the centre and varies across the frame.
Each requirement is satisfiable and each is satisfiable exactly. The optical axis the first demands is straight down; the axis the second demands is horizontal; and the angle between them is 90.000000000°.
That is the measurement, and its being exactly a right angle rather than merely large is what makes the impossibility structural rather than a matter of degree. There is no compromise angle at which both are nearly satisfied — 45° is the worst of both, giving a carpet that is neither a rectangle nor a band and figures at 71% of their height.
What the plan view does to a standing figure
The interesting half of the measurement is not that the plan view shortens people. It is what it replaces their height with.
Consider a vertical segment of height standing at horizontal distance from the point directly under an eye at height . Its two endpoints image at radial distances and , so the drawn length of the segment is
which is proportional to . A figure standing directly under the eye is drawn at exactly zero length, whatever its height. A figure twice as far from that point is drawn twice as long.
So the plan view does not compress height into a small quantity. It converts height into position: the mark a standing figure makes on the page measures how far off-axis it stands, and its actual height enters only as a multiplier shared with everything else at that radius.
Why that makes the convention necessary rather than lazy
Put those two facts together and the convention stops looking like an avoidance of difficulty.
A picture of a garden with people in it has to do two things: show the layout, and show the people. The layout is a planar configuration best shown from its normal. The people are vertical objects best shown from the horizontal. The two directions are orthogonal, and no projection through a single centre supplies both.
A perspective painter facing the same problem chooses, and the choice is visible in what European painting does: the ground is shown obliquely, at a compromise angle, and the layout is inferred rather than shown. Ask what shape a garden in a Renaissance painting is and the answer requires a rectification and four known points. Ask what shape a garden in a Persian miniature is and the answer is on the page.
Neither is a failure. They are different allocations of a resource that cannot be spent twice.
What a compromise angle actually costs
The impossibility is exact, so the interesting question is what the best available compromise gives, and it is worth computing rather than assuming.
A camera tilted down by from horizontal images a horizontal rectangle with its two receding sides converging — the departure from a true rectangle grows as the tilt falls away from vertical — and images a vertical segment at a foreshortening that grows as the tilt rises. At 45°, halfway between the two requirements, a carpet’s near and far edges differ by a factor set by the depth range, and a standing figure is drawn at about seven tenths of the height a level camera would give it.
Neither number is catastrophic and both are wrong, and the picture that results is the one European painting settles on: a floor whose pattern converges, figures somewhat foreshortened, and a layout that has to be inferred rather than read.
That is not a criticism; it is what a projection through a centre offers, and it comes with the station point and the evidential force attached. The point of computing the compromise is to establish that there is no third option — no tilt at which both requirements are met, and no camera parameter that helps. The requirements are orthogonal in the strict sense, and orthogonal requirements do not admit a clever compromise.
The convention’s own map
The map the miniature actually uses is neither of the two cameras above; it is a third thing, and it is worth writing down because it is simpler than either.
Horizontal position goes to horizontal position. Height above the ground goes straight up the page, undiminished. Depth goes up the page too, scaled by the cosine of a tilt angle that is a convention rather than a camera parameter.
That is an oblique parallel projection with the axes assigned differently from the cabinet and cavalier systems — the depth axis runs vertically on the page instead of diagonally — and everything the previous essays measured about oblique projections applies to it unchanged: no diminution, no horizon, uniform depth scale, exact ratios along each axis.
What it adds is the assignment. By sending depth up the page rather than diagonally, it puts the ground plan in its true proportions in the two directions that matter and reserves the vertical for both depth and height, which is the ambiguity the previous essay priced.
Why the ambiguity is smaller than it looks
The vertical axis carrying two quantities sounds fatal and is not, and the reason is worth measuring rather than asserting: the two quantities have very different magnitudes.
A garden is tens of metres deep and its people are under two metres tall. With depth compressed by and height at full scale, a figure’s own height occupies a small and roughly constant part of the vertical extent, while its depth position ranges over the whole picture. So the vertical coordinate is mostly depth, with a small height offset riding on it, and a reader resolves the two by the simplest possible rule: the bottom of a figure is where it stands.
That rule is exact under this map. The foot of a standing figure is at , so its page height is exactly — pure depth, no height contribution at all. The convention of drawing every figure standing on a visible ground, feet showing, is what makes the map readable, and it is a drawing convention doing a job the geometry cannot do for itself.
Perspective needs no such convention because the horizon does the work. That is a fair statement of what a horizon is for, as distinct from what it is: a line whose height is the eye’s own, which separates above-eye-level from below and therefore separates height from depth.
The same convention in a modern document
There is a familiar object with exactly this geometry and it is worth naming, because it shows the arrangement is a solution rather than a period style.
An architectural presentation drawing — the kind with a plan of a garden or a floor and elevations of what stands on it — is this map. The plan is at true proportions; the vertical elements are drawn at true height; the relationship between them is a convention rather than a projection; and a reader reads the layout off the plan and the heights off the elevations without ever supposing the sheet is a picture of anything.
The same is true of a stage plan, a games board with pieces drawn upright on it, a map with buildings drawn in elevation, and an isometric video-game view with characters drawn frontally. The last of those is worth a sentence, because it is the same compromise being made again under different constraints: a tile grid drawn at true proportions so the game state is readable, and characters drawn upright and frontal so their faces work, with the two related by a convention nobody notices.
Four traditions, one geometric problem, one solution. That does not make the solution obvious — it took each of them their own route to it — but it does make it a solution rather than a style.
parallel field already names them. Elevation draws vertical extents at true scale and collapses depth entirely; the plan is the same projection down a different axis. The miniature’s map is the two of them assigned to one page, which is what a presentation drawing does on a larger sheet.Three cameras, three answers, one control
The measurement in assertNoCameraHasBothGrounds is built from three checks and the first is the control that makes the other two mean something.
The overhead camera images the carpet as a true square, ratio 1.000000000000. If that failed, the “impossibility” below would be an artefact of a badly placed camera rather than a property of the requirements.
The same camera images a figure under the eye at zero length, and figures off the axis in exact proportion to their radius — checked by taking two radii and requiring the ratio of drawn lengths to equal the ratio of radii to nine decimal places.
And the level camera flattens the carpet to a band 56 px deep against 271 px wide, while keeping the figures. Both cameras are ordinary pinholes from lib/camera.js; nothing is special-cased.
The right angle between their axes is then computed from the two view directions rather than assumed from the setup, which is a small thing and is the difference between a measurement and a restatement of how the cameras were placed.
One number to carry
If one figure from this essay is worth remembering it is the right angle, and it is worth remembering because of how exact it is.
Two requirements, each satisfiable on its own, whose optical axes are 90.000000000° apart. Not “very different”. Not “hard to reconcile”. Orthogonal, which in this context means that satisfying one of them perfectly satisfies the other not at all, and that every compromise between them is a compromise on both.
A convention that meets both is therefore doing something no camera can do, and the only way to do it is to stop being a single projection. That is what the miniature does, and it is what every other system in this field does, by a different route each time.
The limit of what the geometry can say## The limit of what the geometry can say
A closing caution, because this is the essay in the field where overreach would be easiest.
Nothing above is a claim about why Persian painters drew this way, what the convention meant, or how it was understood. The measurement says that the two effects the pictures exhibit cannot come from one camera, and that the map which produces both is an oblique parallel projection with a particular axis assignment. Everything past that — whether the convention was arrived at as a solution to this problem, whether anybody involved thought of it as a problem, what else the arrangement was doing — is a question about a tradition, and the geometry has no standing to answer it.
What the geometry does establish is the negative, and the negative is worth having because it is so often stated the other way. These pictures are not inconsistent perspective. They are consistent something-else, and the something-else is a projection whose properties are exactly computable and whose two halves are impossible to obtain together by the means the objection assumes.
There is one further thing the measurement is good for, and it is a test rather than a claim. If a picture really is drawn by the map above, then every horizontal rectangle in it should be drawn as a rectangle, every vertical should be drawn at a height proportional to its real height, and every depth interval should occupy page in proportion to its size. Those are three checkable predictions about a picture, and a picture that failed them would be drawn by something else.
That is the form an account of a convention should take: not an interpretation, but a map with consequences that a picture can contradict. The measurements in this field are all of that kind, which is the only reason they are on a site about geometry rather than in an essay about painting.
The two views, once more
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Assembled from several views — both name demonstration, drawing system, foreshortening, occlusion, orthographic
- What the removed roof buys — both name demonstration, drawing system, occlusion, orthographic, station point
- A scroll is a camera that moves — both name foreshortening, orthographic, picture plane, station point
- What perspective gave up — both name demonstration, drawing system, occlusion, station point
- A centre and a measure are exclusive — both name demonstration, drawing system, station point
- An inverse perspective is a leaning plane — both name camera tilt, demonstration, picture plane
Named objects
A flat tag is an object no other essay names yet.
Camera tiltDemonstrationDrawing systemForeshorteningGround plane rectificationOcclusionOrthographicPicture planeSite planStation point