The rectangle behind the lens

A pupil sees around an edge

Two backgrounds identical everywhere a pinhole can see, differing only in the strip an occluder hides from it, produce identical pinhole pictures and pupil pictures 42 per cent apart. So no function of the sharp image — no kernel, no depth-dependent kernel, nothing — produces the picture a real lens makes, and the reach behind the edge is R(Z₂/Z₁ − 1), which is 120 mm here.

Worth reading first: The centre has an area.

The centre has an area licensed almost everything: a world point images as a disc whose centre is exactly the pinhole’s mark, so every construction built on point positions survives a finite aperture untouched.

It ended by naming one exception, and this is it. An edge does not blur the way a picture of an edge blurs, and the reason is not a subtlety of the model. It is that a pupil with an area can see part of a scene that a pinhole cannot see at all.

The arrangement

An opaque edge at 1.2 metres, a background at the focus distance of six metres, an 85 mm lens at f/1.4 — a pupil 30 millimetres across.

Take the sensor position where a pinhole at the pupil’s centre says the edge falls. The chief ray grazes the edge. Rays through the upper half of the pupil pass it and land on the background behind the edge; rays through the lower half are stopped.

So light from a strip of background that the pinhole picture does not contain at all reaches the sensor.

A pupil reaches 120 mm behind the edge that a pinhole cannot see pastA plan of the object side of a 85 mm lens at f/1.4. The pupil is on the left, an opaque edge stands 1.2 metres away, and the background is at the focus distance of 6 metres. Five rays leave one point of the sensor and fan out across the pupil; the ones drawn faint are stopped by the edge, and the ones drawn solid get past it and land on the background behind the edge — as far as 120 millimetres behind it, which is R(Z₂/Z₁ − 1) and comes out at 121. A pinhole at the pupil's centre receives the middle ray only, so everything in that strip is, to a pinhole, absent. **The scene behind an edge is in the photograph and not in the pinhole picture of it.** The axis is compressed; the vertical scale is millimetres of the scene.the pupilan opaque edge120 mmf/1.4, focused at 6 m3 of 5 rays get past
Fig. 1 A plan of the object side: five rays from one sensor point, some stopped by the edge and some getting past it onto the hidden strip.

How far behind, exactly

The reach has a closed form and it is one line. A ray leaving the sensor point through the pupil at height a lands at a(1 − Z/Z_f) + X_c·Z/Z_f at depth Z. Setting the occluder condition and the background depth gives the width of the strip as

reach=R(Z2Z11)\text{reach} = R\left(\frac{Z_2}{Z_1} - 1\right)

with R the pupil’s radius. On this arrangement that is 121.43 millimetres, and the measurement — sweeping the sensor across the edge and asking which background positions are hidden from the chief ray and still lit through the pupil — returns 120.44, within one per cent of it.

Twelve centimetres of scene that is in the photograph and is not in the pinhole picture of it.

The impossibility, stated as a measurement

An inequality would be weak — “the blurred picture is not quite the convolution of the sharp one” invites a better kernel. So the claim is made in the form that admits no better kernel.

Two backgrounds are built. They are identical everywhere except inside the hidden strip, where one is dark and the other is bright. A pinhole at the pupil’s centre cannot see inside that strip at all, so the two pinhole pictures are identical to the last bit — measured, and the difference is 0.

Through the pupil they differ by up to 42 per cent of full scale.

Two different outputs, one input. There is therefore no function whatever — not a convolution, not a space-varying convolution, not a depth-dependent kernel, not anything trained on anything — that takes the pinhole picture to the pupil picture. The information is not in the input.

Why this is the strongest form of the claim

Because it needs no assumption about how blur is modelled.

The usual statement — “defocus is not shift-invariant” — is true and weaker: it says one particular family of models fails, and leaves room for a richer family. The usual repair — a per-pixel depth and a per-pixel kernel — is also true and also weaker, because it is defeated by the same example: the per-pixel depth at the sensor point above is the occluder’s, and the occluder’s kernel cannot produce background light that the input does not contain.

The measurement here is about information, not about kernels. Two scenes with one pinhole picture and two pupil pictures rules out every map from the first to the second at once, and the ruling-out is a single equality and a single inequality rather than an argument about model classes.

This is the same move an ambiguity is not an uncertainty makes in the two-view field: the test for whether an answer is available is whether two candidates predict different measurements, and it settles a whole category of proposed remedies before any of them is implemented.

A depth buffer's blur and the real one differ by 70% over 11 pxTwo edge profiles across the same occluding edge. The solid curve is what the pupil actually delivers, integrated over the disc with the occlusion applied ray by ray. The other is what a renderer produces when it keeps one depth per sample and gathers with the blur radius belonging to whatever surface is at each pixel — the near surface's radius here is 33 pixels and the far one's is 0.0. They part company by 70 per cent of full scale over a band 11 pixels wide, and the sign of the difference is always the same: the gather has no way to let the background through, because the sharp image it is gathering from does not contain it.00.2500.5000.7501-10-50510across the edge, in pixels of the picturehow much of the background reaches the sensorwhere a pinhole puts the edgethe pupilone depth per samplethe same edge, two ways of blurring it70% apart
Fig. 2 The consequence for a renderer: the pupil’s own profile against what a per-pixel gather produces.

What a photographer already knows

The effect has a name in practice and it is not usually thought of as geometry.

A wide aperture makes a foreground obstruction — a branch, a fence, a rail — go not merely soft but transparent, so the subject behind it is visible through the blur. A frame is an interval records the same kind of transparency in time, where a moving occluder lets the background through for part of the exposure. That is exactly this measurement: the pupil is seeing around the obstruction, and the fraction of the pupil that gets past it is the fraction of the background that reaches the sensor.

The rule photographers use follows from the formula. The reach is R(Z₂/Z₁ − 1), so it grows with the aperture and with the ratio of the two depths: get close to the fence, keep the subject far, open up. Every part of that advice is a term in one expression.

In the picture, the reach is the blur’s own radius

The reach is quoted in millimetres of scene, which is where it is easiest to picture and not where a reader of the photograph meets it. Converting it into the image gives a statement that ties this rung to the previous one exactly.

The strip is R(Z2/Z11)R(Z_{2}/Z_{1} - 1) wide at depth Z2Z_{2}, so it subtends R(1/Z11/Z2)R(1/Z_{1} - 1/Z_{2}) at the camera and lands on the sensor at

fR(1Z11Z2)=f22N(1Z11Z2),f R \left(\frac{1}{Z_{1}} - \frac{1}{Z_{2}}\right) = \frac{f^{2}}{2N}\left(\frac{1}{Z_{1}} - \frac{1}{Z_{2}}\right),

which is half the occluder’s own blur-disc diameter — the expression the previous rung derived, with the same factor and the same reciprocal difference. On this arrangement the occluder’s disc is 3.49 mm across and the strip is 1.72 mm wide on the sensor: the disc’s radius, to a per cent and a half, the residue being the (Zff)(Z_{f} - f) the exact form carries.

So the see-around is not an additional effect with its own size. The transparent fringe is exactly as wide, in the picture, as the blurred edge’s own spread, and it could not be otherwise: the disc is what the edge’s image is smeared over, half of it falls on the near side and half on the far, and the far half is precisely where light from behind the edge is arriving.

Three things follow that the world-side formula hides.

The two are not independent settings. A photographer cannot buy more transparency without buying more visible blur, or less blur without losing the transparency, because they are one length seen twice. Every arrangement that makes a foreground obstruction disappear makes it conspicuously soft first.

The focal length cancels in the world and does not in the picture. The reach in millimetres of scene depends only on the pupil’s radius and the depth ratio — which is the checkable claim the closing section invites — while the reach in pixels carries f2/Nf^{2}/N, so a long lens spreads the same strip over more of the frame. That is the same split the sharp band finds between the band’s depth and the background’s magnification, arriving one rung later on a different quantity.

And it says where to look for the effect in a photograph. Anywhere the blur of a foreground edge is nn pixels across, exactly n/2n/2 pixels of background that no pinhole could have seen are in the frame. The effect is not rare and not subtle; it is present at every defocused edge in every photograph, at half the width of something already visible.

The relation to a penumbra

This collection has computed the same integral before, with the roles exchanged, and putting the two beside each other is the cheapest way to see the structure.

A shadow from an area source is the same geometry with the source where the pupil is: the penumbra at a point on the floor is the fraction of the source visible from that point past the occluder, and the penumbra is the lamp’s image is where this collection measured it. A defocused edge is the fraction of the pupil visible from a point on the sensor past the occluder.

Same integral, same chord-of-a-disc arithmetic, same closed form. What differs is which end of the chain has the area — the light source, or the eye — and this collection’s habit of treating a shadow as a second projection is exactly what makes the two one calculation.

Which also says something about the reach: a wide aperture’s transparency and a soft shadow’s penumbra are the same phenomenon, and the numbers convert.

A 35 cm source, an edge, and the band betweenThe penumbra is 17.5 cm wide by the projection — the source's width times the receiver-to-occluder distance over the source-to-occluder distance — and 17.4 cm by counting how much of the source each point can see. The two routes share no arithmetic.source, 35 cmthe occluder's edgefraction of the source visiblepenumbra 17.5 cmprojection: 17.50 cmsampled: 17.41 cm
Fig. 3 The same integral with the area at the other end of the chain, from the light field.

What the profile actually looks like

The intensity across a defocused edge is not a smooth ramp and not a Gaussian. It is the chord area of a disc cut by a line, which is what the fraction of an unoccluded pupil is, and that has a definite shape: flat at both ends, steepest in the middle, with a vertical tangent at neither.

Writing it down is one line — the open share is (arccos d − d√(1−d²))/π with d the normalised position of the cut across the pupil — and it is worth having because it is the exact profile a synthetic aperture produces, against which any approximation can be checked.

Two features of it matter. It reaches zero and one at finite distances, so a defocused edge has a definite width rather than a tail; and it is antisymmetric about its midpoint for a uniform background, which is why the naive convolution model looks right in the simplest test anybody runs. Put structure in the background and the antisymmetry goes, which is the whole content of the previous section.

That is why the demonstration here uses two structured backgrounds rather than measuring a profile. A profile against a uniform background is exactly the case where the wrong model is right.

Two scenes a pinhole cannot tell apart, 42% apart through a pupilTwo backgrounds identical everywhere a pinhole at the pupil's centre can see, and different only inside the strip the occluding edge hides from it. Their pinhole pictures are identical to the last bit. Their pupil pictures differ by up to 42 per cent of full scale. So there is no function of the pinhole picture that produces the picture a real lens makes — not a blur kernel, not a depth-dependent kernel, not anything — because two different outputs would have to come from one input. That is the strongest form the claim takes, and it needs no assumption about how the blur is modelled.00.2500.5000.7501-50-2502550across the edge, in pixels of the picturewhat the sensor receivesidentical to a pinhole, different to a pupil42% apart
Fig. 4 The two profiles again, whose difference is the structure a uniform-background test would never have shown.

Where the licence is lost, precisely

The centre has an area drew a boundary and this essay is the other side of it, so it is worth restating the boundary sharply.

Safe: the centroid of an isolated blob. Nothing occludes it, the pupil is symmetric, and the centroid is the pinhole mark exactly.

Not safe: any position read off an edge, a silhouette, or a boundary between two depths. The intensity profile across such a boundary is not centred on the pinhole position, is not symmetric, and depends on what is behind the edge.

That has a practical consequence for calibration that is worth stating because it is the opposite of the usual instinct. A defocused target of dots is fine; a defocused chequerboard is not. A chequerboard corner is a meeting of four edges, and its apparent position under defocus depends on the depths and brightnesses on both sides. A dot’s centroid does not, which is the licence the centre has an area established and the reason the boundary between the two essays is worth drawing sharply.

Seven distances, seven discs, every centre on the pinhole's mark to 2.9e-17 mmThe patch a point on the axis images as, at seven distances from 0.9 to 16 metres, drawn to a scale that is the same at every aperture. The lens is focused at 3 metres, where the patch collapses to a point; either side of it the disc grows, and it grows faster toward the camera than away from it, which is the asymmetry the depth of field inherits. Every disc is centred on the mark a pinhole at the pupil's centre would have made — worst departure 2.9e-17 millimetres across all seven, which is the floor of the arithmetic rather than a tolerance. The number over each disc is its diameter in the pixels of this collection's figures.13.50.9 m6.61.4 m2.52.1 m0.03.0 m1.84.4 m3.37.0 m4.716.0 mf/2.8, focused at 3 mcentres to 2.9e-17 mm
Fig. 5 The case that is safe: isolated blobs, whose centres do not move however large they get.

What can be recovered, and what cannot

The impossibility is about recovering the pupil picture from the pinhole one. The reverse direction is worth a sentence because it is not symmetric.

Going from a pupil picture to a pinhole picture — deblurring — is ill-posed in the ordinary way at every point, and additionally impossible at an edge for a reason of its own: part of the pupil picture there is light from a region the pinhole picture has no pixel for. A perfect deblurring would have to discard information rather than invent it, which is a nicer problem, but the discarding is not a convolution either.

What can be done, and is, is going the other way with more input: render the scene from many points across the pupil and average. That is what distributed ray tracing does, it is exact in the limit, and it is exact precisely because it evaluates the integral rather than approximating it with a kernel. The cost is that the scene has to be evaluated many times, which is the price of the information a single depth buffer does not have — one depth per sample is not enough.

A depth buffer's blur and the real one differ by 87% over 11 pxTwo edge profiles across the same occluding edge. The solid curve is what the pupil actually delivers, integrated over the disc with the occlusion applied ray by ray. The other is what a renderer produces when it keeps one depth per sample and gathers with the blur radius belonging to whatever surface is at each pixel — the near surface's radius here is 17 pixels and the far one's is 0.0. They part company by 87 per cent of full scale over a band 11 pixels wide, and the sign of the difference is always the same: the gather has no way to let the background through, because the sharp image it is gathering from does not contain it.00.2500.5000.7501-10-50510across the edge, in pixels of the picturehow much of the background reaches the sensorwhere a pinhole puts the edgethe pupilone depth per samplethe same edge, two ways of blurring it87% apart
Fig. 6 The gap a gather leaves, at a narrower aperture, where the effect is smaller and has not gone.

Two apertures, and what changes between them

The slider on the section figure walks the f-number, and what it shows is worth naming because the scaling is exact and simple.

The reach is proportional to the pupil’s radius, so it is proportional to 1/N. From f/1.2 to f/8 the pupil shrinks by a factor of nearly seven and the reach with it — from about 140 millimetres of hidden background to about 21. The number of rays in the figure that get past the edge falls in step.

Nothing else in the geometry moves. The depths are the same, the edge is in the same place, the chief ray is unaffected. So the whole of the transparency is bought with aperture, exactly, and the exchange rate is a straight line.

That is the same linearity the blur diameter has, and it is why a single number — the entrance pupil’s diameter, which is f/N — controls both. Two lenses of different focal lengths with the same pupil diameter give the same reach behind the same edge at the same depths, which is a slightly surprising statement and a directly checkable one.

What it means for a matting or a depth estimate

Two downstream consequences, because this is a geometric fact with practical teeth.

Alpha matting. A defocused foreground edge is genuinely partially transparent — not “looks transparent”, but passes a measurable fraction of the background — and the fraction is the open share of the pupil, which is a known function of position. So the alpha channel a matting algorithm is trying to estimate has a geometric meaning at a defocused edge, and it is computable if the depths and the aperture are known.

Depth from defocus. The blur diameter is a function of depth, so measuring the blur gives the depth — except at edges, where the profile is not the blur of anything and reading a diameter off it gives a depth that belongs to neither surface. That is why depth-from-defocus methods report their worst errors exactly at depth discontinuities, which are the places a depth map most needs to be right.

Both failures come from the same place, and both have the same partial remedy: treat the edge as a mixture of two depths rather than as one surface, and the arithmetic becomes tractable again. What no remedy recovers is the part of the background the pinhole never saw, because the sensor received it and the model has nowhere to put it.

A depth buffer's blur and the real one differ by 100% over 11 pxTwo edge profiles across the same occluding edge. The solid curve is what the pupil actually delivers, integrated over the disc with the occlusion applied ray by ray. The other is what a renderer produces when it keeps one depth per sample and gathers with the blur radius belonging to whatever surface is at each pixel — the near surface's radius here is 8 pixels and the far one's is 0.0. They part company by 100 per cent of full scale over a band 11 pixels wide, and the sign of the difference is always the same: the gather has no way to let the background through, because the sharp image it is gathering from does not contain it.00.2500.5000.7501-10-50510across the edge, in pixels of the picturehow much of the background reaches the sensorwhere a pinhole puts the edgethe pupilone depth per samplethe same edge, two ways of blurring it100% apart
Fig. 7 The gap at a narrow aperture, where a depth-from-defocus method would be most confident and is still wrong at the edge.

Why the background is put at the focus distance

A choice in the arrangement that makes the demonstration as clean as it can be, and which a reader reproducing it should keep.

With the background exactly in focus, every ray from a given sensor point converges on a single background point, so the sensor point’s own conjugate is unambiguous and “which background positions are hidden from the chief ray” is a question with a definite answer. Move the background off focus and each sensor point samples a patch of it, and the strip’s edges blur into the measurement.

The result does not depend on the choice — a pupil sees around an edge whatever the background is doing — but the measurement of the reach does, and quoting a number from a blurred boundary would have been quoting the blur. So the arrangement is chosen to make the quantity sharp, and the closed form it is checked against holds generally.

The short version

A finite pupil sees around an occluding edge, as far as R(Z₂/Z₁ − 1) behind it — 120.44 millimetres measured against 121.43 from the closed form, on an 85 mm lens at f/1.4 with an edge at 1.2 metres and a background at six.

Two backgrounds differing only inside that strip give identical pinhole pictures and pupil pictures 42 per cent apart, so no function of the pinhole picture produces the pupil picture. That is an impossibility rather than an inequality, and it rules out every kernel-based model of defocus at once — including a per-pixel depth-dependent one, which is the model every real-time renderer uses.

A pupil reaches 84 mm behind the edge that a pinhole cannot see pastA plan of the object side of a 85 mm lens at f/2. The pupil is on the left, an opaque edge stands 1.2 metres away, and the background is at the focus distance of 6 metres. Five rays leave one point of the sensor and fan out across the pupil; the ones drawn faint are stopped by the edge, and the ones drawn solid get past it and land on the background behind the edge — as far as 84 millimetres behind it, which is R(Z₂/Z₁ − 1) and comes out at 85. A pinhole at the pupil's centre receives the middle ray only, so everything in that strip is, to a pinhole, absent. **The scene behind an edge is in the photograph and not in the pinhole picture of it.** The axis is compressed; the vertical scale is millimetres of the scene.the pupilan opaque edge84 mmf/2, focused at 6 m3 of 5 rays get past
Fig. 8 The reach at f/2, halfway down the sweep, where two of the five rays still get past.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

ApertureCircle of confusionConvolutiondepth of fieldEntrance pupilExit pupilIdentifiabilityOcclusionPinholeSensorSilhouette