The lamp is the second eye
Worth reading first: A shadow is a second projection · Two rays that do not meet.
A single photograph fixes each point of the scene on a ray and no further. That is the whole of what one view supplies, and everything the metrology field does with one picture is a way of supplying the missing coordinate from somewhere else — a known height, a known plane, a known ratio.
A lit scene supplies it for free, and the reason is the identification this field is built on. A shadow is a projection from the lamp; a lamp is therefore a second centre of projection; and two centres are a stereo pair.
The construction
Three steps, none of which needs anything the picture does not have except the lamp’s position.
The camera’s ray. A point’s image fixes the direction from the eye, through the calibration: the ray direction is in the camera’s own frame, rotated into the world. The point is somewhere on that line.
The shadow’s position on the floor. The shadow’s image fixes a second ray, and that ray meets the known floor plane at a known place. So the shadow’s world position is available from the picture alone, exactly as a ground plan is — one picture and a known plane give the plane’s points.
The lamp’s ray. The point, its shadow and the lamp are collinear by definition, so the line from the lamp through the recovered shadow position passes through the point.
Two lines in space, and the point is where they meet. The recovery is exact — 9e-16 m from where the point was put — and the closest approach between the two lines, 1e-15 m, is the residual that would report a failure if there were one.
That last quantity is worth its own sentence. Two lines in space generically do not meet, so solving them as a least-squares closest approach rather than as an intersection returns an answer and a number that could be large. An intersection routine returns a point either way, which is the failure the refraction field measures at length — rays that miss their own common centre by millimetres and a solver that reports a centre regardless.
It is a stereo pair, and it degrades like one
The identification is not a metaphor, and the way to check that is to see whether the error behaviour is the stereo behaviour.
It is. A stereo pair’s depth precision depends on the angle subtended at the point by the two eyes — the baseline over the range — and the depth error per pixel of matching error grows as that angle closes. Replace the second eye with a lamp and the same relation holds with the lamp’s offset playing the baseline.
The degenerate configuration has a name in both readings. In stereo it is a point on the line through the two eyes, where the two rays coincide and the triangulation is singular. Here it is a lamp on the camera’s own optical axis: the shadow of every point falls exactly behind it in the picture, the two rays coincide, and there is nothing to intersect.
That is why a camera-mounted flash produces a picture with no shadows visible. The shadows are there, hidden behind their own objects, and the geometry says a picture with no visible shadows carries no depth information from them at all.
What has to be known, and what does not
The accounting of inputs is the part that makes this a technique rather than an observation.
Needed: the camera’s calibration. Focal length and principal point, so an image point becomes a ray. Recoverable from the picture itself in many scenes — three vanishing points give the focal length — so this is not necessarily an external input.
Needed: the lamp’s position. Three numbers, and they can also come out of the picture: the lamp comes out of the shadows by intersecting drawn lines, given at least two objects standing on the known plane.
Needed: the floor plane. One plane, to convert the shadow’s image into a world position.
Not needed: anything about the object. No known height, no assumed shape, no second photograph. Which is the whole gain: the measurements the metrology field makes from one picture all need a scene assumption, and this one needs a lighting assumption instead.
Worth noticing that two of the three inputs are themselves recoverable from the same photograph. A picture with a level ground plane, two uprights and their shadows contains its own calibration and its own lamp — so the depth of everything else in it is available with no external information at all, and the only thing that has to be supplied from outside is one length to fix the scale, which no single view can ever provide.
This is what a shadow scanner is
The technique is not hypothetical. A shadow scanner — a stick waved between a lamp and an object while a camera watches — is exactly this construction repeated for every pixel and every frame, and its accuracy is governed by exactly the curve above.
What the moving stick supplies is the correspondence. In the figures here, a point and its shadow are known to belong together because the figure knows the scene. In a real photograph, matching an object point to its own shadow point is the hard part, and the stick makes it easy: at each moment the shadow’s leading edge is a single curve, so every pixel it crosses is being told which plane through the lamp it lies on.
That is the same trade a structured-light scanner makes and the same one stereo matching makes. The geometry is easy and the correspondence is the work — which is what the twoviews field says about photographs and what a wrong match costs measures in the other direction.
The lamp’s image is the epipole
The stereo reading has a second consequence that is visible in every photograph anyone has ever taken of a lit scene, and naming it makes the identification concrete.
In a stereo pair, the image of one camera’s centre in the other camera’s picture is the epipole, and every point’s match lies on the line joining the epipole to that point’s own image. Replace the second camera with a lamp and the statement survives word for word: the lamp’s image is the epipole, and a point’s shadow lies on the line in the picture joining the lamp’s image to the point’s image.
That is why the shadows in a photograph of a lamp-lit room appear to radiate from the lamp. They are not merely pointing away from it; every object and its own shadow lie on a common line through the lamp’s image, exactly, and the line is the epipolar line of that point.
The sun is the case where the epipole is at infinity. Sunlit shadows are parallel in the world, so their images meet at a vanishing point on the horizon — which is the image of the sun’s direction and is the epipole of a second view whose centre is infinitely far away.
The sun is a second view that is parallel
Following that through gives a genuinely different pair, and the difference is the one the parallel field spends five essays on.
A lamp and a camera are two projections through finite centres — a proper stereo pair, with a baseline and an epipole in the frame. The sun and a camera are a perspective view and a parallel one, and the second view has no centre to be a baseline from.
What survives is that the sun’s rays still fix a direction, so a point and its shadow still determine the point: the ray from the shadow’s world position, in the sun’s known direction, meets the camera’s ray. The construction is identical and the conditioning question changes shape — there is no “lamp too close to the camera” configuration, only a sun direction too close to the camera’s own axis, which is the same degeneracy with the distance taken out of it.
That is worth having because it says which quantity the technique is really sensitive to. Not the lamp’s distance, and not its brightness: the angle between the two rays at the point being measured, which for the sun is a property of the time of day and for a lamp is a property of where it was put.
The angle is the whole of the design
Everything a practitioner would decide follows from the conditioning curve, and it is worth stating as design advice because it is unusually clean.
Put the lamp far off the camera’s axis. The measurement’s precision is set by the angle between the two rays at the point, and nothing else about the lamp matters — not its brightness, not its distance, except as those change the angle.
And do not put it too far. At a large angle the shadow of a point may not be visible from the camera at all: the object hides its own shadow, or the shadow falls outside the frame. The usable range is bounded below by conditioning and above by visibility, which is exactly the trade a stereo rig makes between baseline and overlap.
The floor’s flatness matters as much as the angle. The shadow’s image is converted to a world position using the floor plane, so a floor that is not the plane it is assumed to be feeds an error straight into the answer — and the receiving surface’s shape is not a small correction: a dish costs 5.67 mm and a step 74.95 mm on a scene about a metre across.
The miss is the check, and it is not decoration
The closest-approach distance deserves more than the sentence it got, because it is the only quantity in the construction that could ever report that something is wrong.
Every input to the measurement is an assumption: that the camera’s calibration is right, that the floor is where it is believed to be, that the lamp is where it is believed to be, that the shadow point identified in the image is the shadow of the object point identified in the image. If any of those is wrong, the two rays no longer meet, and the distance between them at their closest approach is how far from meeting they are.
So the number is a consistency check with a physical meaning in metres, available for every point independently, and it costs nothing to compute because the least-squares solve produces it on the way to the answer.
Two failure modes it catches, and they look identical from anywhere else:
A wrong correspondence. A shadow point matched to the wrong object point gives two rays that generally miss by a large margin, so the miss flags it — which is exactly the diagnostic a stereo pipeline uses and the reason a triangulation is written as a closest approach rather than an intersection.
A wrong assumption about the floor. If the floor is not flat where the shadow lands, its world position is misplaced, the lamp’s ray is drawn through the wrong point, and the two rays miss. The miss is then reporting the receiving surface, which is the quantity the essay two doors down spends its length on.
What the miss cannot catch is an error that moves both rays consistently — a scale error, most obviously. Scaling the whole configuration leaves every ray pair meeting exactly, which is the same absence every single-view method has.
Why one picture can do this at all
There is a fair objection to answer, because the site has been emphatic that one view fixes a ray and nothing more.
The objection’s answer is that this is not one projection. The photograph is one projection and the shadow is a second one, from a different centre, and both are recorded in the same image. The picture carries two views because the scene was lit from a place that is not the camera.
Read that way, the strange thing is how much of the second view is available. The lamp’s picture is not photographed — there is no sensor at the lamp — and yet the position of a shadow on a known plane recovers precisely the one number the lamp’s own image would have supplied. The floor is acting as the lamp’s film.
That is the sharpest form of this field’s identification. A shadow on a plane is a photograph taken by the lamp, developed on the floor, and re-photographed by the camera — which is a projection of a projection, and therefore a homography, and therefore invertible.
What this does not give
The boundaries, because a technique’s reach is the honest part of describing it.
No scale. As with every single-view method, the whole reconstruction can be scaled — camera, lamp, scene and floor together — and every image is identical. One measured length fixes it; nothing in the picture does.
Nothing for a point whose shadow is not visible. A point in shadow, a point whose shadow falls behind the object, a point whose shadow leaves the frame: no correspondence, no measurement. The method is defined on exactly the set of points whose shadows can be seen and identified.
And nothing at the degenerate angle. As the lamp approaches the camera the measurement’s cost per pixel grows without bound, and the figure walks it rather than warning about it — 5.9 mm per pixel at 39°, and 1 mm at 15.4° with the trend going the way the geometry says.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Depth is a reciprocal — both name baseline, error propagation, triangulation
- The image of the other eye — both name baseline, correspondence, epipolar geometry
- Two views give shape and no size — both name baseline, correspondence, triangulation
- A lens destroys the invariant — both name error propagation, single-view metrology
- A point is a line over there — both name correspondence, epipolar geometry
- A shadow can be un-cast — both name centre of projection, error propagation
Named objects
A flat tag is an object no other essay names yet.
Baselinecentre of projectionCorrespondenceEpipolar geometryerror propagationShadow projectionsingle-view metrologyskew raysStructured lightTriangulation