A pupil sees around an edge
Worth reading first: The centre has an area.
The centre has an area licensed almost everything: a world point images as a disc whose centre is exactly the pinhole’s mark, so every construction built on point positions survives a finite aperture untouched.
It ended by naming one exception, and this is it. An edge does not blur the way a picture of an edge blurs, and the reason is not a subtlety of the model. It is that a pupil with an area can see part of a scene that a pinhole cannot see at all.
The arrangement
An opaque edge at 1.2 metres, a background at the focus distance of six metres, an 85 mm lens at f/1.4 — a pupil 30 millimetres across.
Take the sensor position where a pinhole at the pupil’s centre says the edge falls. The chief ray grazes the edge. Rays through the upper half of the pupil pass it and land on the background behind the edge; rays through the lower half are stopped.
So light from a strip of background that the pinhole picture does not contain at all reaches the sensor.
How far behind, exactly
The reach has a closed form and it is one line. A ray leaving the sensor point through the pupil at height a lands at a(1 − Z/Z_f) + X_c·Z/Z_f at depth Z. Setting the occluder condition and the background depth gives the width of the strip as
with R the pupil’s radius. On this arrangement that is 121.43 millimetres, and the measurement — sweeping the sensor across the edge and asking which background positions are hidden from the chief ray and still lit through the pupil — returns 120.44, within one per cent of it.
Twelve centimetres of scene that is in the photograph and is not in the pinhole picture of it.
The impossibility, stated as a measurement
An inequality would be weak — “the blurred picture is not quite the convolution of the sharp one” invites a better kernel. So the claim is made in the form that admits no better kernel.
Two backgrounds are built. They are identical everywhere except inside the hidden strip, where one is dark and the other is bright. A pinhole at the pupil’s centre cannot see inside that strip at all, so the two pinhole pictures are identical to the last bit — measured, and the difference is 0.
Through the pupil they differ by up to 42 per cent of full scale.
Two different outputs, one input. There is therefore no function whatever — not a convolution, not a space-varying convolution, not a depth-dependent kernel, not anything trained on anything — that takes the pinhole picture to the pupil picture. The information is not in the input.
Why this is the strongest form of the claim
Because it needs no assumption about how blur is modelled.
The usual statement — “defocus is not shift-invariant” — is true and weaker: it says one particular family of models fails, and leaves room for a richer family. The usual repair — a per-pixel depth and a per-pixel kernel — is also true and also weaker, because it is defeated by the same example: the per-pixel depth at the sensor point above is the occluder’s, and the occluder’s kernel cannot produce background light that the input does not contain.
The measurement here is about information, not about kernels. Two scenes with one pinhole picture and two pupil pictures rules out every map from the first to the second at once, and the ruling-out is a single equality and a single inequality rather than an argument about model classes.
This is the same move an ambiguity is not an uncertainty makes in the two-view field: the test for whether an answer is available is whether two candidates predict different measurements, and it settles a whole category of proposed remedies before any of them is implemented.
What a photographer already knows
The effect has a name in practice and it is not usually thought of as geometry.
A wide aperture makes a foreground obstruction — a branch, a fence, a rail — go not merely soft but transparent, so the subject behind it is visible through the blur. A frame is an interval records the same kind of transparency in time, where a moving occluder lets the background through for part of the exposure. That is exactly this measurement: the pupil is seeing around the obstruction, and the fraction of the pupil that gets past it is the fraction of the background that reaches the sensor.
The rule photographers use follows from the formula. The reach is R(Z₂/Z₁ − 1), so it grows with the aperture and with the ratio of the two depths: get close to the fence, keep the subject far, open up. Every part of that advice is a term in one expression.
In the picture, the reach is the blur’s own radius
The reach is quoted in millimetres of scene, which is where it is easiest to picture and not where a reader of the photograph meets it. Converting it into the image gives a statement that ties this rung to the previous one exactly.
The strip is wide at depth , so it subtends at the camera and lands on the sensor at
which is half the occluder’s own blur-disc diameter — the expression the previous rung derived, with the same factor and the same reciprocal difference. On this arrangement the occluder’s disc is 3.49 mm across and the strip is 1.72 mm wide on the sensor: the disc’s radius, to a per cent and a half, the residue being the the exact form carries.
So the see-around is not an additional effect with its own size. The transparent fringe is exactly as wide, in the picture, as the blurred edge’s own spread, and it could not be otherwise: the disc is what the edge’s image is smeared over, half of it falls on the near side and half on the far, and the far half is precisely where light from behind the edge is arriving.
Three things follow that the world-side formula hides.
The two are not independent settings. A photographer cannot buy more transparency without buying more visible blur, or less blur without losing the transparency, because they are one length seen twice. Every arrangement that makes a foreground obstruction disappear makes it conspicuously soft first.
The focal length cancels in the world and does not in the picture. The reach in millimetres of scene depends only on the pupil’s radius and the depth ratio — which is the checkable claim the closing section invites — while the reach in pixels carries , so a long lens spreads the same strip over more of the frame. That is the same split the sharp band finds between the band’s depth and the background’s magnification, arriving one rung later on a different quantity.
And it says where to look for the effect in a photograph. Anywhere the blur of a foreground edge is pixels across, exactly pixels of background that no pinhole could have seen are in the frame. The effect is not rare and not subtle; it is present at every defocused edge in every photograph, at half the width of something already visible.
The relation to a penumbra
This collection has computed the same integral before, with the roles exchanged, and putting the two beside each other is the cheapest way to see the structure.
A shadow from an area source is the same geometry with the source where the pupil is: the penumbra at a point on the floor is the fraction of the source visible from that point past the occluder, and the penumbra is the lamp’s image is where this collection measured it. A defocused edge is the fraction of the pupil visible from a point on the sensor past the occluder.
Same integral, same chord-of-a-disc arithmetic, same closed form. What differs is which end of the chain has the area — the light source, or the eye — and this collection’s habit of treating a shadow as a second projection is exactly what makes the two one calculation.
Which also says something about the reach: a wide aperture’s transparency and a soft shadow’s penumbra are the same phenomenon, and the numbers convert.
What the profile actually looks like
The intensity across a defocused edge is not a smooth ramp and not a Gaussian. It is the chord area of a disc cut by a line, which is what the fraction of an unoccluded pupil is, and that has a definite shape: flat at both ends, steepest in the middle, with a vertical tangent at neither.
Writing it down is one line — the open share is (arccos d − d√(1−d²))/π with d the normalised position of the cut across the pupil — and it is worth having because it is the exact profile a synthetic aperture produces, against which any approximation can be checked.
Two features of it matter. It reaches zero and one at finite distances, so a defocused edge has a definite width rather than a tail; and it is antisymmetric about its midpoint for a uniform background, which is why the naive convolution model looks right in the simplest test anybody runs. Put structure in the background and the antisymmetry goes, which is the whole content of the previous section.
That is why the demonstration here uses two structured backgrounds rather than measuring a profile. A profile against a uniform background is exactly the case where the wrong model is right.
Where the licence is lost, precisely
The centre has an area drew a boundary and this essay is the other side of it, so it is worth restating the boundary sharply.
Safe: the centroid of an isolated blob. Nothing occludes it, the pupil is symmetric, and the centroid is the pinhole mark exactly.
Not safe: any position read off an edge, a silhouette, or a boundary between two depths. The intensity profile across such a boundary is not centred on the pinhole position, is not symmetric, and depends on what is behind the edge.
That has a practical consequence for calibration that is worth stating because it is the opposite of the usual instinct. A defocused target of dots is fine; a defocused chequerboard is not. A chequerboard corner is a meeting of four edges, and its apparent position under defocus depends on the depths and brightnesses on both sides. A dot’s centroid does not, which is the licence the centre has an area established and the reason the boundary between the two essays is worth drawing sharply.
What can be recovered, and what cannot
The impossibility is about recovering the pupil picture from the pinhole one. The reverse direction is worth a sentence because it is not symmetric.
Going from a pupil picture to a pinhole picture — deblurring — is ill-posed in the ordinary way at every point, and additionally impossible at an edge for a reason of its own: part of the pupil picture there is light from a region the pinhole picture has no pixel for. A perfect deblurring would have to discard information rather than invent it, which is a nicer problem, but the discarding is not a convolution either.
What can be done, and is, is going the other way with more input: render the scene from many points across the pupil and average. That is what distributed ray tracing does, it is exact in the limit, and it is exact precisely because it evaluates the integral rather than approximating it with a kernel. The cost is that the scene has to be evaluated many times, which is the price of the information a single depth buffer does not have — one depth per sample is not enough.
Two apertures, and what changes between them
The slider on the section figure walks the f-number, and what it shows is worth naming because the scaling is exact and simple.
The reach is proportional to the pupil’s radius, so it is proportional to 1/N. From f/1.2 to f/8 the pupil shrinks by a factor of nearly seven and the reach with it — from about 140 millimetres of hidden background to about 21. The number of rays in the figure that get past the edge falls in step.
Nothing else in the geometry moves. The depths are the same, the edge is in the same place, the chief ray is unaffected. So the whole of the transparency is bought with aperture, exactly, and the exchange rate is a straight line.
That is the same linearity the blur diameter has, and it is why a single number — the entrance pupil’s diameter, which is f/N — controls both. Two lenses of different focal lengths with the same pupil diameter give the same reach behind the same edge at the same depths, which is a slightly surprising statement and a directly checkable one.
What it means for a matting or a depth estimate
Two downstream consequences, because this is a geometric fact with practical teeth.
Alpha matting. A defocused foreground edge is genuinely partially transparent — not “looks transparent”, but passes a measurable fraction of the background — and the fraction is the open share of the pupil, which is a known function of position. So the alpha channel a matting algorithm is trying to estimate has a geometric meaning at a defocused edge, and it is computable if the depths and the aperture are known.
Depth from defocus. The blur diameter is a function of depth, so measuring the blur gives the depth — except at edges, where the profile is not the blur of anything and reading a diameter off it gives a depth that belongs to neither surface. That is why depth-from-defocus methods report their worst errors exactly at depth discontinuities, which are the places a depth map most needs to be right.
Both failures come from the same place, and both have the same partial remedy: treat the edge as a mixture of two depths rather than as one surface, and the arithmetic becomes tractable again. What no remedy recovers is the part of the background the pinhole never saw, because the sensor received it and the model has nowhere to put it.
Why the background is put at the focus distance
A choice in the arrangement that makes the demonstration as clean as it can be, and which a reader reproducing it should keep.
With the background exactly in focus, every ray from a given sensor point converges on a single background point, so the sensor point’s own conjugate is unambiguous and “which background positions are hidden from the chief ray” is a question with a definite answer. Move the background off focus and each sensor point samples a patch of it, and the strip’s edges blur into the measurement.
The result does not depend on the choice — a pupil sees around an edge whatever the background is doing — but the measurement of the reach does, and quoting a number from a blurred boundary would have been quoting the blur. So the arrangement is chosen to make the quantity sharp, and the closed form it is checked against holds generally.
The short version
A finite pupil sees around an occluding edge, as far as R(Z₂/Z₁ − 1) behind it — 120.44 millimetres measured against 121.43 from the closed form, on an 85 mm lens at f/1.4 with an edge at 1.2 metres and a background at six.
Two backgrounds differing only inside that strip give identical pinhole pictures and pupil pictures 42 per cent apart, so no function of the pinhole picture produces the pupil picture. That is an impossibility rather than an inequality, and it rules out every kernel-based model of defocus at once — including a per-pixel depth-dependent one, which is the model every real-time renderer uses.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- The corner sees an ellipse — both name aperture, circle of confusion, entrance pupil, sensor
- A close picture carries its own distance — both name circle of confusion, depth of field
- Assembled from several views — both name identifiability, occlusion
- Counting shadows is not counting lamps — both name identifiability, occlusion
- Facing the reader is not being reachable — both name identifiability, occlusion
- The entrance pupil walks with the angle — both name aperture, entrance pupil
Named objects
A flat tag is an object no other essay names yet.
ApertureCircle of confusionConvolutiondepth of fieldEntrance pupilExit pupilIdentifiabilityOcclusionPinholeSensorSilhouette