Light and mirrors

The lamp is the second eye

One photograph, one lamp whose position is known, and a point's place in space comes back to 9e-16 m — the camera's ray through the point, the lamp's ray through the image of its shadow, and the intersection of two lines. It is triangulation with one of the two eyes replaced by a light, and it degrades exactly like a stereo pair: 5.9 mm of depth per pixel at 39° between the rays, 1 mm at 15.4°.

Worth reading first: A shadow is a second projection · Two rays that do not meet.

A single photograph fixes each point of the scene on a ray and no further. That is the whole of what one view supplies, and everything the metrology field does with one picture is a way of supplying the missing coordinate from somewhere else — a known height, a known plane, a known ratio.

A lit scene supplies it for free, and the reason is the identification this field is built on. A shadow is a projection from the lamp; a lamp is therefore a second centre of projection; and two centres are a stereo pair.

The lamp is the second eyeOne camera, one lamp, one point. The camera's ray through the point's image fixes it on a line; the image of its shadow fixes where the lamp's ray through it meets the floor; and two lines that are not parallel meet. The point comes back at 1e-15 m of closest approach and 9e-16 m from where it was put — depth out of a single photograph, with no second camera and nothing assumed about the object. The rays cross at 28.5°, and that angle is what the measurement is worth.horizonthe pointits shadowcorrect from 22 cm, at 160 mm widerays cross at 24.8° · recovered to 4e-15 m
Fig. 1 One camera, one lamp, one point. The camera’s ray through the point’s image fixes it on a line. The image of its shadow fixes where the lamp’s ray through it meets the floor. Two lines that are not parallel meet — the point comes back at 1e-15 m of closest approach and 9e-16 m from where it was put, with the rays crossing at 28.5°.

The construction

Three steps, none of which needs anything the picture does not have except the lamp’s position.

The camera’s ray. A point’s image fixes the direction from the eye, through the calibration: the ray direction is (ucx,  vcy,  f)(u - c_x,\; v - c_y,\; f) in the camera’s own frame, rotated into the world. The point is somewhere on that line.

The shadow’s position on the floor. The shadow’s image fixes a second ray, and that ray meets the known floor plane at a known place. So the shadow’s world position is available from the picture alone, exactly as a ground plan is — one picture and a known plane give the plane’s points.

The lamp’s ray. The point, its shadow and the lamp are collinear by definition, so the line from the lamp through the recovered shadow position passes through the point.

Two lines in space, and the point is where they meet. The recovery is exact — 9e-16 m from where the point was put — and the closest approach between the two lines, 1e-15 m, is the residual that would report a failure if there were one.

That last quantity is worth its own sentence. Two lines in space generically do not meet, so solving them as a least-squares closest approach rather than as an intersection returns an answer and a number that could be large. An intersection routine returns a point either way, which is the failure the refraction field measures at length — rays that miss their own common centre by millimetres and a solver that reports a centre regardless.

It is a stereo pair, and it degrades like one

The identification is not a metaphor, and the way to check that is to see whether the error behaviour is the stereo behaviour.

It is. A stereo pair’s depth precision depends on the angle subtended at the point by the two eyes — the baseline over the range — and the depth error per pixel of matching error grows as that angle closes. Replace the second eye with a lamp and the same relation holds with the lamp’s offset playing the baseline.

What the angle between the two rays is worthThe recovery is exact at every one of these — 3e-14 m — and that is not the question. What changes is the price of a pixel: 5.9 mm of depth per pixel with the lamp 39° off the camera's ray, and 1 mm at 15.4°. A lamp beside the lens gives a shadow under the object and a measurement worth nothing.02.5057.5020253035angle between the camera's ray and the lamp's (°)millimetres of depth per pixel of error in the shadowexact at every anglethe cost of one pixel is what moves
Fig. 2 The recovery is exact at every one of these — 3e-14 m — and that is not the question. What changes is the price of a pixel: 5.9 mm of depth per pixel of error in the shadow’s image with the lamp 39° off the camera’s ray, and 1 mm at 15.4°. A lamp beside the lens gives a shadow under the object and a measurement worth nothing.

The degenerate configuration has a name in both readings. In stereo it is a point on the line through the two eyes, where the two rays coincide and the triangulation is singular. Here it is a lamp on the camera’s own optical axis: the shadow of every point falls exactly behind it in the picture, the two rays coincide, and there is nothing to intersect.

That is why a camera-mounted flash produces a picture with no shadows visible. The shadows are there, hidden behind their own objects, and the geometry says a picture with no visible shadows carries no depth information from them at all.

The lamp is the second eyeOne camera, one lamp, one point. The camera's ray through the point's image fixes it on a line; the image of its shadow fixes where the lamp's ray through it meets the floor; and two lines that are not parallel meet. The point comes back at 4e-15 m of closest approach and 1e-14 m from where it was put — depth out of a single photograph, with no second camera and nothing assumed about the object. The rays cross at 17.2°, and that angle is what the measurement is worth.horizonthe pointits shadowcorrect from 22 cm, at 160 mm widerays cross at 17.2° · recovered to 1e-14 m
Fig. 3 The lamp brought close to the camera. The rays cross at a shallow angle and the recovery is still exact — this is the configuration where the arithmetic is fine and the measurement is worthless, which is a distinction worth keeping separate.
The lamp is the second eyeOne camera, one lamp, one point. The camera's ray through the point's image fixes it on a line; the image of its shadow fixes where the lamp's ray through it meets the floor; and two lines that are not parallel meet. The point comes back at 1e-15 m of closest approach and 4e-15 m from where it was put — depth out of a single photograph, with no second camera and nothing assumed about the object. The rays cross at 22.1°, and that angle is what the measurement is worth.horizonthe pointits shadowcorrect from 22 cm, at 160 mm widerays cross at 22.1° · recovered to 4e-15 m
Fig. 4 Between the two ends of the slider. The recovery is exact here as everywhere; what the offset changes is the angle the two rays cross at, and therefore what a pixel of error is worth.
Two rays, 2.53 mm apart, in the plane that contains bothThe ray from the left eye through its mark and the ray from the right eye through its. With the marks placed exactly they meet, to 3.3e-14 m. With the same marks read to 1 px they miss by 2.53 mm at a range of 7.45 m. Triangulation is not an intersection; the reported point is a choice about what to minimise, and the gap is the part a residual alone will not tell you.midpoint — 2.53 mm gapfrom the left eyefrom the right eyegap 2.53 mm at 7.45 mexact marks: 3.3e-14 m
Fig. 5 Why the answer is solved as a closest approach rather than as an intersection. Two rays in space generically miss, and the distance between them is the only quantity in the construction that can report a failure.

What has to be known, and what does not

The accounting of inputs is the part that makes this a technique rather than an observation.

Needed: the camera’s calibration. Focal length and principal point, so an image point becomes a ray. Recoverable from the picture itself in many scenes — three vanishing points give the focal length — so this is not necessarily an external input.

Needed: the lamp’s position. Three numbers, and they can also come out of the picture: the lamp comes out of the shadows by intersecting drawn lines, given at least two objects standing on the known plane.

Needed: the floor plane. One plane, to convert the shadow’s image into a world position.

Not needed: anything about the object. No known height, no assumed shape, no second photograph. Which is the whole gain: the measurements the metrology field makes from one picture all need a scene assumption, and this one needs a lighting assumption instead.

Worth noticing that two of the three inputs are themselves recoverable from the same photograph. A picture with a level ground plane, two uprights and their shadows contains its own calibration and its own lamp — so the depth of everything else in it is available with no external information at all, and the only thing that has to be supplied from outside is one length to fix the scale, which no single view can ever provide.

The lamp, from the shadows aloneThe two intersections are the light and the point below it. Nothing about the lamp was given to the construction — it is shown three posts, three shadow tips and the camera's own horizon — and the recovered position is 1e-12 mm from the truth. The light's foot sits 351 px below the horizon, which is what says it is a lamp and not the sun.the shadow lines meet below the horizon — a lamp in the roomhorizoncorrect from 25 cm, at 160 mm wide3 posts · foot 351 px below the horizon
Fig. 6 The input that is not free, recovered from the picture that needs it. Three posts and their shadow tips give the image of the lamp and the image of its foot — nothing about the lamp is supplied — and the recovered position lands 1e-12 mm from the truth, with the foot 351 px below the horizon, which is what says it is a lamp and not the sun.

This is what a shadow scanner is

The technique is not hypothetical. A shadow scanner — a stick waved between a lamp and an object while a camera watches — is exactly this construction repeated for every pixel and every frame, and its accuracy is governed by exactly the curve above.

What the moving stick supplies is the correspondence. In the figures here, a point and its shadow are known to belong together because the figure knows the scene. In a real photograph, matching an object point to its own shadow point is the hard part, and the stick makes it easy: at each moment the shadow’s leading edge is a single curve, so every pixel it crosses is being told which plane through the lamp it lies on.

That is the same trade a structured-light scanner makes and the same one stereo matching makes. The geometry is easy and the correspondence is the work — which is what the twoviews field says about photographs and what a wrong match costs measures in the other direction.

One correspondence moved 30 px, and where the damage wentThe clean fit is exact to 2.3e-13 px. Moving correspondence 9 by 30 px leaves every other point wrong too — the typical one by 0.54 px and the worst by 1.7 px — because a least-squares fit has nowhere to put a bad row except across all of them. Here the largest residual does fall on the culprit; it is not obliged to.00.50011.50010203040correspondenceepipolar error at every OTHER point, after one match is movedmedian 0.54 pxthe moved oneevery point wrong: median 0.54 px, worst 1.7 pxclean fit 2.3e-13 px
Fig. 7 The other half of every triangulation, from the field that owns it. The clean fit is exact to 2.3e-13 px; one correspondence moved by 30 px leaves every other point wrong as well, the typical one by 0.54 px. A shadow scanner’s stick exists to make this half easy.

The lamp’s image is the epipole

The stereo reading has a second consequence that is visible in every photograph anyone has ever taken of a lit scene, and naming it makes the identification concrete.

In a stereo pair, the image of one camera’s centre in the other camera’s picture is the epipole, and every point’s match lies on the line joining the epipole to that point’s own image. Replace the second camera with a lamp and the statement survives word for word: the lamp’s image is the epipole, and a point’s shadow lies on the line in the picture joining the lamp’s image to the point’s image.

That is why the shadows in a photograph of a lamp-lit room appear to radiate from the lamp. They are not merely pointing away from it; every object and its own shadow lie on a common line through the lamp’s image, exactly, and the line is the epipolar line of that point.

The sun is the case where the epipole is at infinity. Sunlit shadows are parallel in the world, so their images meet at a vanishing point on the horizon — which is the image of the sun’s direction and is the epipole of a second view whose centre is infinitely far away.

Six posts in sunlight from 34°The shadows are parallel in the world, so in the picture they meet at one point on the horizon — found from the drawn shadows to 3e-13 px.horizonshadows meet at x = -58, off the frameon the horizon, as it must be
Fig. 8 The epipole, in the case everybody has seen. Six posts in sunlight: the shadows are parallel in the world and meet in the picture at one point, found from the drawn lines to 3e-13 px. That point is the image of the second centre of projection.

The sun is a second view that is parallel

Following that through gives a genuinely different pair, and the difference is the one the parallel field spends five essays on.

A lamp and a camera are two projections through finite centres — a proper stereo pair, with a baseline and an epipole in the frame. The sun and a camera are a perspective view and a parallel one, and the second view has no centre to be a baseline from.

What survives is that the sun’s rays still fix a direction, so a point and its shadow still determine the point: the ray from the shadow’s world position, in the sun’s known direction, meets the camera’s ray. The construction is identical and the conditioning question changes shape — there is no “lamp too close to the camera” configuration, only a sun direction too close to the camera’s own axis, which is the same degeneracy with the distance taken out of it.

That is worth having because it says which quantity the technique is really sensitive to. Not the lamp’s distance, and not its brightness: the angle between the two rays at the point being measured, which for the sun is a property of the time of day and for a lamp is a property of where it was put.

The angle is the whole of the design

Everything a practitioner would decide follows from the conditioning curve, and it is worth stating as design advice because it is unusually clean.

Put the lamp far off the camera’s axis. The measurement’s precision is set by the angle between the two rays at the point, and nothing else about the lamp matters — not its brightness, not its distance, except as those change the angle.

And do not put it too far. At a large angle the shadow of a point may not be visible from the camera at all: the object hides its own shadow, or the shadow falls outside the frame. The usable range is bounded below by conditioning and above by visibility, which is exactly the trade a stereo rig makes between baseline and overlap.

The floor’s flatness matters as much as the angle. The shadow’s image is converted to a world position using the floor plane, so a floor that is not the plane it is assumed to be feeds an error straight into the answer — and the receiving surface’s shape is not a small correction: a dish costs 5.67 mm and a step 74.95 mm on a scene about a metre across.

What the receiving surface costs a four-point fitThe same lamp and the same occluder, cast onto four surfaces. Four marks of the shadow are matched to four known points of the occluder, the map they determine exactly is built, and the other sixty-eight points are predicted by it. On the plane the prediction is right to 3e-13 mm; on the others it is not, and the fit is still exact at the four everywhere — 2e-13 mm — which is what makes the rest a prediction rather than a residual.a flat floor3e-13 mma homology — four points determine ita dished floor5.67 mmnot a homologya ridged floor9.07 mmnot a homologya floor with a step74.95 mmnot a homologyworst error of the predicted pointsexact at the four fitted, on every surface
Fig. 9 The input this technique is most exposed to. The floor is assumed flat when the shadow’s image is turned into a world position, and a floor that is not flat is not a small correction — a step costs an order of magnitude more than a dish.

The miss is the check, and it is not decoration

The closest-approach distance deserves more than the sentence it got, because it is the only quantity in the construction that could ever report that something is wrong.

Every input to the measurement is an assumption: that the camera’s calibration is right, that the floor is where it is believed to be, that the lamp is where it is believed to be, that the shadow point identified in the image is the shadow of the object point identified in the image. If any of those is wrong, the two rays no longer meet, and the distance between them at their closest approach is how far from meeting they are.

So the number is a consistency check with a physical meaning in metres, available for every point independently, and it costs nothing to compute because the least-squares solve produces it on the way to the answer.

Two failure modes it catches, and they look identical from anywhere else:

A wrong correspondence. A shadow point matched to the wrong object point gives two rays that generally miss by a large margin, so the miss flags it — which is exactly the diagnostic a stereo pipeline uses and the reason a triangulation is written as a closest approach rather than an intersection.

A wrong assumption about the floor. If the floor is not flat where the shadow lands, its world position is misplaced, the lamp’s ray is drawn through the wrong point, and the two rays miss. The miss is then reporting the receiving surface, which is the quantity the essay two doors down spends its length on.

What the miss cannot catch is an error that moves both rays consistently — a scale error, most obviously. Scaling the whole configuration leaves every ray pair meeting exactly, which is the same absence every single-view method has.

Why one picture can do this at all

There is a fair objection to answer, because the site has been emphatic that one view fixes a ray and nothing more.

The objection’s answer is that this is not one projection. The photograph is one projection and the shadow is a second one, from a different centre, and both are recorded in the same image. The picture carries two views because the scene was lit from a place that is not the camera.

Read that way, the strange thing is how much of the second view is available. The lamp’s picture is not photographed — there is no sensor at the lamp — and yet the position of a shadow on a known plane recovers precisely the one number the lamp’s own image would have supplied. The floor is acting as the lamp’s film.

That is the sharpest form of this field’s identification. A shadow on a plane is a photograph taken by the lamp, developed on the floor, and re-photographed by the camera — which is a projection of a projection, and therefore a homography, and therefore invertible.

A box and its shadow, both projections from a pointThe rays from the lamp to the corners are the same construction as the rays from the eye to the corners — one operation, two centres.correct from 26 cm, at 160 mm wide34° across
Fig. 10 The identification underneath the whole essay: the rays from the lamp to a box’s corners are the rays from an eye to the same corners, with the centre moved. One machine, two centres, and a picture that contains both.
The other eye, found in the picture — 2.6 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 1.1e-9 px on the left and 3.7e-10 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 1.1e-9 px2.60 m between the eyes
Fig. 11 The epipole in the field that named it: the image of one camera’s centre in the other camera’s picture. Replace the second camera with a lamp and this point is the lamp’s own image, which is where a photograph’s shadows appear to radiate from.
Depth from disparity, with the 1 px the reading is worthZ = fB/d on a 65 mm baseline at 900 px. The line is exact — it returns the camera's own depth to 1e-12 m. The band is what 1 px of disparity error costs, and it stops being a ±. At 6.5 m it runs 5.86–7.33 m, lopsided by 1.25, and the textbook ±Z²δ/fB is 1.2% out. At 40 m it runs 23.8–126.5 m — 86.5 m beyond the estimate against 16.2 m before it, a lopsidedness of 5.32 — and the same formula is 47% out. Past 58.5 m the far edge is infinity.025507510010203040true depth (m)depth reported from the disparity, with a 1 px reading error5.86–7.33 m13.96–26.69 m23.76–126.49 mat 40 m: +86.5 m against −16.2 munbounded past 58.5 m
Fig. 12 And the error law the two share. Depth from a pair is a reciprocal, so a fixed reading error becomes an interval that is not symmetric and eventually is not bounded — the same shape the lamp’s conditioning curve has.

What this does not give

The boundaries, because a technique’s reach is the honest part of describing it.

No scale. As with every single-view method, the whole reconstruction can be scaled — camera, lamp, scene and floor together — and every image is identical. One measured length fixes it; nothing in the picture does.

Nothing for a point whose shadow is not visible. A point in shadow, a point whose shadow falls behind the object, a point whose shadow leaves the frame: no correspondence, no measurement. The method is defined on exactly the set of points whose shadows can be seen and identified.

And nothing at the degenerate angle. As the lamp approaches the camera the measurement’s cost per pixel grows without bound, and the figure walks it rather than warning about it — 5.9 mm per pixel at 39°, and 1 mm at 15.4° with the trend going the way the geometry says.

Two scenes 137× apart, and the one picture they both makeEverything in the second plan — the room, the eye's distance, the eye's own height — is 137 times the first. Every projected vertex agrees to 1e-13 px. A single photograph has no scale, and this is what that means.a room 2.8 m across, eye 1.6 m up1 mthe same plan, 137× bigger137 midenticalpicturesthe picture — both scenes, drawn twice, one on top of the otherlargest disagreement 1e-13 px over 8 verticesone length has to come from outside the picture
Fig. 13 The freedom no lighting arrangement removes. Two scenes 137× apart make pictures identical to 1e-13 px; adding a lamp scales with everything else and changes nothing.
The lamp is the second eyeOne camera, one lamp, one point. The camera's ray through the point's image fixes it on a line; the image of its shadow fixes where the lamp's ray through it meets the floor; and two lines that are not parallel meet. The point comes back at 5e-16 m of closest approach and 2e-15 m from where it was put — depth out of a single photograph, with no second camera and nothing assumed about the object. The rays cross at 38.6°, and that angle is what the measurement is worth.horizonthe pointits shadowcorrect from 22 cm, at 160 mm widerays cross at 38.6° · recovered to 2e-15 m
Fig. 14 And the well-conditioned end of the range: the lamp as far off the axis as this scene allows, with the rays crossing at their widest. The recovery is no more exact than at the near end — it is worth more.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Baselinecentre of projectionCorrespondenceEpipolar geometryerror propagationShadow projectionsingle-view metrologyskew raysStructured lightTriangulation