The rectangle behind the lens

Stepping closer is not zooming

Changing the focal length leaves the ratio between any two things in a picture exactly alone — to twelve decimal places, at every focal length there is. Moving changes it. Hold the subject's drawn size across a step from 3 m to 1.5 m and the background halves, which is the whole of the shot everybody knows and nobody derives.

Worth reading first: A focal length is not an angle · The point you have to stand at.

Two things are commonly said about long lenses, and one of them is true.

The true one is that a long lens produces a picture in which near and far things are drawn at similar sizes — the “compressed” look of a telephoto photograph of a street, where cars a hundred metres apart appear stacked.

The false one is that the lens does the compressing. It does not. The compression is entirely a consequence of where the camera was standing, and it would be identical in a picture taken from the same place with any lens at all and then cropped.

The subject held, the background moved: a step is not a zoomLeft, the camera 3 m from the subject at 50 mm. Right, 1.50 m from it at 25.0 mm, chosen to hold the subject at exactly the same drawn size. The subject is unchanged to twelve decimal places and the background is smaller by 0.526×. No focal length alone can do that: the near-to-far ratio is 10.00 at every focal length there is.3 m, 50 mm1.50 m, 25 mmsubject ×1.000000 · background ×0.526 · zoom alone would give ×1 for bothnear-to-far ratio 10.00 → 19.00changing the focal length leaves it at 1.000000000000
Fig. 1 The demonstration. Left, a camera 3 m from the subject at 50 mm. Right, 1.5 m from it at 25 mm, chosen so the subject is drawn at exactly the same size. The subject is unchanged to twelve decimal places and the background has halved. No focal length alone can do that.

The invariance, in one line

A thing of height HH at depth zz is drawn at height fH/zfH/z. Two things, at depths zaz_a and zbz_b:

hahb=fHa/zafHb/zb=HaHbzbza\frac{h_a}{h_b} = \frac{f H_a / z_a}{f H_b / z_b} = \frac{H_a}{H_b}\cdot\frac{z_b}{z_a}

The focal length has cancelled. It appears in both drawn heights identically, so it cannot appear in their ratio.

That is the whole argument and it is worth noticing how little it assumes: nothing about the lens being rectilinear, nothing about the sensor, nothing about the two objects being at the same height or on the axis. It is a consequence of a projection being a projection, and the same statement holds for a shadow from a lamp or a reflection in a mirror, both of which this site has already shown are the same operation with the centre moved.

The gate asserts it at 101210^{-12} rather than merely stating it, and asserts the other half too: a step of half the subject distance changes the same ratio by a factor that is not one. Without the second half the first is satisfied by a function that ignores its arguments.

What a change of focal length actually is

Since it cannot change any ratio in the picture, what does it do?

It changes the field of view, and that is all. A picture at 100 mm from a given place contains exactly what is in the middle of a picture at 50 mm from the same place, magnified. The two are crops of each other — not approximately, exactly, up to the resolution of whatever recorded them.

Which is why the previous essay’s finding matters here. A long lens and a small sensor do the same thing, because both do the only thing available: they take a narrower slice of the same pencil of rays. There is nothing else for either of them to change, because the pencil of rays is decided by where its vertex is.

Where the reader has to be for a 24° picture to be correctShown 160 mm wide, this picture is a correct projection only from 38 cm away. Drawn to scale.the picture, 160 mm wide38 cm24°the eyefocal length 1623 px38 cm at 160 mm wide
Fig. 2 And the consequence for the reader, which is where the “compression” actually lives. A 24° picture is correct from 2.35 of its own width; a 40° one from 1.37. So the long-lens look is a picture whose station point is far behind where anybody stands, and the flattening a viewer sees is the depth compression of reading a picture from too close — an effect of the viewing, not of the taking.

The shot that separates them

Hold the subject’s drawn size fixed and move. That requires changing the focal length in step with the distance — f=fz/zf' = f \cdot z'/z — and it is the whole trick.

With the subject at 3 m and the background at 30 m, stepping in to 1.5 m needs the focal length halved, from 50 mm to 25 mm. Then:

  • the subject is drawn at exactly the size it was, to twelve decimals;
  • the background has moved to 28.5 m and is drawn at ×0.526\times 0.526 of its previous size;
  • and the near-to-far ratio has gone from 10.00 to 19.00.

The last number is the measurement. The gate checks it against a closed form rather than a threshold —

(zbΔ)zszszb\frac{(z_b - \Delta)\,z_s}{z_s' \, z_b}

— and against the control, which is a step of nothing, where the change is exactly 1.000000000000. The first draft asserted “the background changes by more than five per cent”, which failed at exactly the place the claim is trivially true, and reported the control as a defect.

That is the effect every account of cinema attributes to Hitchcock, and it is worth stating exactly what is remarkable about it. It is not that two things change at once. It is that one thing is held perfectly still while the space around it changes, so the viewer has no reference for what moved. Because the subject’s drawn size is the strongest depth cue a picture has, holding it fixed removes the evidence that the camera went anywhere.

The subject held, the background moved: a step is not a zoomLeft, the camera 3 m from the subject at 50 mm. Right, 0.90 m from it at 15.0 mm, chosen to hold the subject at exactly the same drawn size. The subject is unchanged to twelve decimal places and the background is smaller by 0.323×. No focal length alone can do that: the near-to-far ratio is 10.00 at every focal length there is.3 m, 50 mm0.90 m, 15 mmsubject ×1.000000 · background ×0.323 · zoom alone would give ×1 for bothnear-to-far ratio 10.00 → 31.00changing the focal length leaves it at 1.000000000000
Fig. 3 The same shot taken further. From 0.9 m at 15 mm the subject is still drawn at exactly its original size and the background is down to a third of it. There is no limit to this except the camera walking into the subject; the effect is the ratio of the distances, and the ratio is unbounded.

Why the wrong explanation is so durable

It is worth taking the folk account seriously rather than dismissing it, because it is nearly right and it fails for an interesting reason.

The folk account is: long lens, compressed picture. As a correlation this is almost perfect, because a photographer who fits a long lens is a photographer who is about to stand further back. The lens and the distance travel together, in practice, essentially always.

So the belief is well supported by every photograph anybody has taken. What it gets wrong is the direction of the arrow — it takes a reliable correlate for the cause — and the way to break the correlation is exactly the shot above, which is precisely why the shot is startling. It shows the two coming apart.

This is the same shape as a finding the wrong field records about the taught sixty-degree cone of vision: the rule is right and its stated reason is not, and nobody notices because following the rule works. A correlation that holds in every case anybody meets is indistinguishable from a mechanism until somebody constructs the case where they differ.

What actually changes as the camera moves

The invariance above says the focal length does nothing to any ratio. It is worth stating positively what a step does, because “it changes the perspective” is a phrase rather than a quantity.

Stepping from zz to zz' multiplies every object’s drawn size by z/zz/z' — a different factor for every object, since each has its own depth. So the picture is not scaled, it is re-scaled non-uniformly along depth, and the non-uniformity is exactly the thing a focal length cannot produce.

Two consequences worth extracting.

The near-to-far ratio changes by (zbΔ)zs/(zszb)(z_b - \Delta)z_s / (z_s'\,z_b), which the gate asserts as a closed form. Everything about the “compression” of a scene is that one expression: it is greater than one when stepping back and less when stepping in, and its size is governed by the step relative to the subject’s distance rather than by the step in metres.

And the change is dramatic only when the subject is close. Stepping 1.5 m nearer a subject at 3 m nearly doubles the ratio. Stepping 1.5 m nearer a subject at 30 m changes it by 5%. That is why landscape photography is insensitive to a few paces and portraiture is not, and it is a statement about the ratio Δ/zs\Delta/z_s rather than about the subjects.

The midpoint of one segment, under both familiesThe parallel projection places it exactly halfway (0e+0 px out). The perspective projection places it 22 px away from halfway, 7% of the drawn length.halfway along the drawn linethe actual midpointcorrect from 26 cm, at 160 mm wide22 px apart
Fig. 4 The non-uniform re-scaling in its simplest form. A segment’s drawn midpoint is not the midpoint of the real segment, and the drift depends on the depths. Moving the camera changes those depths and therefore changes the drift; changing the focal length scales the whole drawing and leaves it exactly alone.

What a crop can and cannot recover

A practical corollary worth stating because it is regularly disputed.

Cropping a picture is exactly equivalent to having used a longer lens from the same place, in every respect this site measures — the framing, the perspective, the station point of the result, the ratios between everything in it. The only thing lost is resolution.

That equivalence is not an approximation. Both operations take a narrower cone from the same vertex, and a picture is fully determined by which cone it takes and from where. Anything that could distinguish the two would have to be a property the picture does not have.

Cropping cannot recover a photograph taken from the wrong place. No amount of it changes a single ratio, and the ratios are the entire content of the perspective. A photograph taken from too far away is a photograph of a flattened arrangement, permanently.

Which is a useful thing for the site to be able to say cleanly, because it is the practical form of the one thing a single view cannot give: a picture is a complete record of one pencil of rays and contains nothing whatever about any other one.

One reconstruction, drawn at its own scale and at 3.4×The courtyard recovered from two pictures, seen from above. Every pairwise distance ratio matches the world's to 6.3e-14, so the shape is exact. The size is not determined at all: the right-hand plan is the same reconstruction 3.4 times larger and fits the same two pictures equally well. One measured length in the scene — here 0.800 m — fixes it, and nothing in the pictures can.as recovered0.308 across× 3.4 — same two pictures1.046 acrossworst ratio deviation 6.3e-14size fixed only by a supplied 0.80 m
Fig. 5 The general form of the limit, from the two-view field. A pair of pictures determines shape and not size; a single picture determines ratios and not size; and a crop of a single picture determines exactly what the original did. Each is a statement about what a set of rays contains, and none of them is improved by looking harder.

The same invariance, in the other three projections

The site’s habit is to check whether a result about the camera holds for the other three projections from a centre, because when it does the result is about projection rather than about photography.

A shadow. Move the lamp and the shadow’s shape changes; change nothing but the distance to the surface the shadow falls on and every part of the shadow scales together. So the ratio between two parts of a shadow is invariant to the surface’s distance and not to the lamp’s position — the same statement with the lamp playing the camera.

A reflection. The mirror’s size decides how much is visible and not what shape it is; moving the viewer changes the shape. Again the same split.

A parallel projection is the degenerate case and it is instructive. With the centre at infinity there is no distance to change, so nothing at all can alter a ratio: a parallel drawing preserves the midpoint exactly and preserves every other ratio too. The “zoom” of a parallel system is a scale factor and is all it has.

Three checks and one degenerate case, all agreeing. What changes a picture’s internal ratios is the position of the centre, and nothing else in any of the four constructions has any purchase on them.

The one place the folk account is literally right

There is a genuine exception and it is worth conceding, because a claim with no boundary is a claim nobody can test.

If the subject is what is being held fixed — a portrait framed head-and-shoulders, a building filling the frame — then choosing a focal length does determine the distance, because framing plus focal length determines it. In that setting “a long lens flatters a face” is a true statement with a true mechanism: the long lens forces the photographer back, and the distance flattens the face.

So the trade’s rule of thumb is sound advice and its mechanism is one step removed from where it is usually placed. The lens does not flatten the face. It obliges the photographer to stand where faces are flattened, and the obligation is what does the work.

That distinction is not pedantic, because it says what happens when the constraint is removed. A photographer who steps back with a short lens and crops gets the identical flattering perspective at lower resolution — which is exactly what a phone with a “portrait” mode does, and why it works.

A 1.7 m figure, drawn at every distance, by two systemsThe falling curve is a pinhole: f·H/Z, dropping -4.06 px per metre on average and 20.0× overall across the range. The flat one is an oblique system, whose slope is 0e+0 px per metre — zero, not nearly zero. A system with no diminution has no size–distance signal at all, so depth has to be carried by something else, and the something else is where the figure stands on the page.0100200204060distance of the object, metresits drawn height, pxa pinholean oblique systemdrawn height of a fixed object against its distance20.0× against 1.000
Fig. 6 The curve the whole argument runs on. Drawn size against distance for an object of fixed height — one over z, with the focal length as a scale factor that multiplies the whole curve and cannot change its shape. Every ratio taken between two points on this curve is independent of that scale factor, which is the invariance stated as a picture.

The shot works because it removes the evidence

There is a question the arithmetic raises and does not settle, and it is worth putting carefully because it is on the edge of what this site can say.

The dolly zoom is startling. Two other operations produce the same background change and are not: walking toward a subject with a fixed lens, and walking away with one. Everybody has done both, thousands of times, and nobody finds either unsettling.

The geometric difference is that the ordinary versions come with a consistent signal — the subject grows or shrinks along with everything else, and the change in the ratio is accompanied by a change in the subject’s drawn size that agrees with it. The dolly zoom removes the second signal while keeping the first, so the picture asserts one thing about the camera’s motion and another about the scene’s shape.

What this site can say is that the two signals disagree, and by exactly how much: the subject reports no motion at all, to twelve decimals, while the background reports a step of half the subject distance. What it cannot say is why a viewer finds that unsettling rather than merely unusual — that is a question about the visual system, and the site has been explicit since its foundation that pictorial depth cues compete and that this machinery has no standing to adjudicate between them.

The distinction is the same one a picture with no size–distance signal is built on. There the finding is that removing a cue leaves the picture consistent with a range of scenes; here it is that holding one cue fixed while another moves leaves it consistent with none. Both are statements about what a picture contains rather than about what a viewer does with it.

A 2.4 m figure, drawn at every distance, by two systemsThe falling curve is a pinhole: f·H/Z, dropping -5.73 px per metre on average and 20.0× overall across the range. The flat one is an oblique system, whose slope is 0e+0 px per metre — zero, not nearly zero. A system with no diminution has no size–distance signal at all, so depth has to be carried by something else, and the something else is where the figure stands on the page.0100200300204060distance of the object, metresits drawn height, pxa pinholean oblique systemdrawn height of a fixed object against its distance20.0× against 1.000
Fig. 7 The cue the shot holds fixed. Drawn size against distance for an object of known height is the strongest depth signal a picture has, and the dolly zoom is the arrangement that pins it while everything else moves. What the geometry supplies is the disagreement between the two signals; what a viewer does with the disagreement is a different subject.

Why this belongs on this site rather than in a photography book

Two reasons, and the second is the one that makes it a measurement rather than an explanation.

The first is that the site has already built the machinery that makes the claim exact. Every figure is projected from a stated camera, so “the ratio is unchanged” is not an assertion about how pictures look — it is a computation over eight vertices at two focal lengths, and the answer is 1.0000000000001.000000000000.

The second is that the site’s own premise is on the other side of the same coin. A picture states the point it is correct from, and what this essay establishes is that the point it is correct from is the only thing about the camera the picture records. Focal length, sensor, format and crop are all recoverable only as one angle, and the angle is the station distance in disguise.

So the two halves fit together: the picture forgets everything about the equipment except one number, and that number is where to stand.

A box drawn from a known camera, and the camera recovered from the drawingThree vanishing points found from the twelve drawn edges alone give back the focal length to 4e-15 relative.recovered principal pointused to drawrecoveredgapfocal length853.90853.904e-15principal x345.0345.02e-12angle44.0°44.0°correct from 20 cm, at 160 mm wide44° across
Fig. 8 The site’s round trip, which is what “the picture records one number” means operationally. Three vanishing points read off the drawn edges give back a focal length and a principal point — and that is everything the picture contains about the camera. There is no fourth quantity to be found because there is nothing else in there.
The subject held, the background moved: a step is not a zoomLeft, the camera 3 m from the subject at 50 mm. Right, 2.40 m from it at 40.0 mm, chosen to hold the subject at exactly the same drawn size. The subject is unchanged to twelve decimal places and the background is smaller by 0.816×. No focal length alone can do that: the near-to-far ratio is 10.00 at every focal length there is.3 m, 50 mm2.40 m, 40 mmsubject ×1.000000 · background ×0.816 · zoom alone would give ×1 for bothnear-to-far ratio 10.00 → 12.25changing the focal length leaves it at 1.000000000000
Fig. 9 And a small step, which is where the effect is subtle enough to be missed and large enough to matter. From 2.4 m at 40 mm the background is at ×0.80 — a fifth smaller, on a subject drawn at exactly the same size. Nobody watching would call that a dolly zoom; everybody watching would say the picture felt different, which is the effect doing its work below the threshold at which it can be named.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Crop factorDemonstrationDepth cueDiminutionfield of viewFocal lengthForeshorteningStation pointTelephoto compressionViewing distance