One homography makes a shadow map the eye's picture
Worth reading first: A shadow map's texels land by two distances and two cosines · The divide is postponed, not avoided.
A shadow map’s texels land by two distances and two cosines worked out how large each texel of a shadow map is by the time it reaches the eye’s picture — a closed form in the two cameras’ focal lengths, their distances to the surface and the two cosines of its tilt to each. One warped shadow map, and what it cannot reach then asked what one map could do for the worst case, a lamp forty metres ahead facing back toward the eye. Drawn square, the map needed 90.8 million texels for none of them to land on the near floor larger than a pixel. Fitted to the floor and with its rows re-spaced by the one projective parameter a perspective shadow map applies, it needed 805 thousand; with its rows re-spaced in any shape at all, 706 thousand — the counts as that essay now gives them, corrected in place after this measurement found its first counts overstated. And nothing can do better than one texel per pixel of the eye’s picture of the floor, 148 thousand.
The factor of 4.8 between the best row warp and that bound was the room left. The essay closed by naming the obvious larger instrument: a full projective map of the whole picture — a homography, eight numbers rather than one — which crowds texels toward a point rather than toward an edge and so can follow a need that varies across the rows as well as down them. It asked how much of the 4.8 such a warp would take back.
All of it, to within five per cent. And the reason it can is not a matter of optimisation at all.
Six ways to spend the same map
The hero figure sets the full warp beside everything the earlier essay measured, for the lamp facing back, on a logarithmic scale.
Two new bars sit between the row warps and the bound. The first is a true one-parameter perspective warp: the same single number a perspective shadow map applies, but applied as a genuine projective map of the picture, so it crowds columns toward its vanishing point as well as rows. It needs 752 thousand texels, fewer than the one-parameter re-spacing of rows at 805 thousand and more than the 706 thousand a free re-spacing reaches by giving every row its own density.
The second is the full projective warp: both of the perspective terms free, and the rotation and shear of the map chosen with them. It needs 154 thousand texels, against the bound’s 148 thousand — 4.2 per cent above it. The shortfall of nearly five has become one of a few per cent.
The number is checked by a route that shares nothing with its search. Every texel of the warped map is carried back through the warp into the lamp’s picture, down onto the floor, and forward into the eye’s picture, and its area measured there. The largest covers 1.0000 px, the median 0.9930. The map is not merely within budget. Very nearly every texel in it lands at a pixel.
Every count in the figure is priced on one footing: the floor’s need taken at every third pixel of the eye’s picture and, as well, at points spread down every row of the lamp’s own picture. That second set is what corrected the earlier essay’s row-warp counts, which were priced on the eye’s points alone and so left most rows of the facing lamp’s picture with no point in them; each empty row had been given its neediest neighbour’s density. The corrected counts are lower, by 1.8 and 1.6 times, and the gap to the bound that was the earlier essay’s room is 4.8 rather than seven and a half.
What one perspective parameter buys, and what it does not
The first new bar deserves a closer look, because it sits in an unexpected place. A one-parameter perspective beats a one-parameter re-spacing of rows — 752 thousand against 805 for the facing lamp — and the reason is the columns. It does not beat a free re-spacing of rows, which reaches 706 thousand, and the reason is that a free re-spacing has a number for every row where the perspective has one for the whole picture.
The facing lamp sees the floor near the eye at the far edge of its picture, crowded toward its horizon, and it sees that floor foreshortened in both directions: rows crowd because the floor recedes from the lamp, and columns crowd because the same distant floor is narrower in the lamp’s picture than it is in the eye’s. A re-spacing of rows can give those rows more texels. It cannot narrow the columns there, because a row warp keeps every column at one width from top to bottom, so it must spend on the whole width of every crowded row what its most demanding point needs. A perspective crowds rows and columns together toward one point, and where it gives a row more texels it also gives that row fewer columns to spread them over. That is the sense in which a perspective shadow map “crowds texels toward a point rather than toward an edge”. With one number it buys seven per cent over the row warp with one number; the free row warp, spending a number on every row, buys fourteen.
For the overhead lamp the comparison is 921 thousand for the one-parameter perspective against 1.01 million for the one-parameter row warp and 896 thousand for the free one. In neither case does one perspective parameter do most of the work. What does is freeing the rest of the projective map at once — the second perspective term, the rotation and the shear — which takes the facing lamp from 752 thousand to 154 and the overhead lamp from 921 thousand to 159. The perspective’s first parameter is the smallest part of what a homography can do; its value is that it is the first step on a road that ends at the bound.
A row warp wastes most of its texels
Sizing a map so that no texel exceeds a pixel says nothing about how small the rest are, and the rest are where the texels go.
Under the best row warp, the band shows each distance’s smallest and largest texel as it lands in the eye’s picture. The largest touch one pixel — that is the condition the map was sized to meet — and the smallest fall to 0.088 of a pixel, eleven to a pixel, with most of the map somewhere between. Every texel below one pixel is resolution spent on detail the eye’s picture will average away. The row warp matches the need along one direction of the lamp’s picture and spends the other direction at the rate its most demanding column requires.
Under the full warp the band collapses to a line at one pixel. There is nowhere left to save, because there is nowhere the map is finer than it needs to be.
The map becomes the eye’s picture of the floor
A map in which every texel lands at exactly one pixel of the eye’s picture is, in effect, the eye’s picture itself, sampled a second time from the lamp. The figure below shows that literally, by drawing the warped map’s own texel grid where it lands.
Every sixteenth row and column of the warped map, carried into the eye’s picture, is a straight line, and together they form a grid of equal squares — the cells differ in area by one and a half per cent. The map’s texel grid and the eye’s pixel grid are, up to a scale and a turn, the same grid.
That is not something the search found by luck, and the reason makes the whole result exact rather than empirical. The eye’s picture of a plane is a homography of the plane: a pinhole maps a flat surface to its picture by a projective transformation. So is the lamp’s. The composition — from the lamp’s picture of the floor to the floor, and from the floor to the eye’s picture — is one homography, the one a plane induces between two views, and two marks off a known plane find the other eye used it to find a second camera. A warp of the lamp’s picture by exactly that homography turns the lamp’s picture of the floor into the eye’s picture of the floor, and a map spread evenly over the result has one texel per pixel everywhere. The search over the warp’s four numbers was a search for the plane’s homography, and its four per cent is the crop — the eye’s picture of the floor out to sixty metres is not a rectangle, and the map is — together with the sampling the bound itself is counted from.
The row warp could not find it because a homography is not a warp of rows. It moves every column differently from every other, which is exactly the freedom a row re-spacing gives up.
Three lamps, one bound
The facing lamp was the hard case. The same measurement for the lamp beside the eye and the lamp thirty metres overhead says whether the result is special to it.
For all three the full warp comes within four per cent of the bound, and for two of them it lands a per cent under it. That is not a map beating one texel per pixel. The bound is counted from samples of the eye’s picture three pixels apart and the warp is checked at the same samples, so each is uncertain by about a per cent, and a warp that meets the bound exactly can appear to undercut it by that much.
The lamp beside the eye had almost nothing to gain: it sees the floor along nearly the same rays as the eye, so its picture of the floor is already nearly the eye’s, and a map fitted to the floor was about five per cent above the bound before any warp. The overhead lamp gains 5.6 times over the best row warp, and the lamp facing back 4.6 times. The earlier essay’s lesson that the lamp’s position decides how much a warp can do still stands; what changes is that a full warp does all of it, for every position.
What no single map reaches
The earlier essay drew one more bar, and it is the one the full warp cannot move: a second surface on the same ray from the lamp.
The full warp reaches the floor’s bound for every lamp. Add a wall through every floor point, turned toward the eye, and the fewest texels any single map can spend rises: 1.12 times the floor’s bound for the lamp beside the eye, 1.70 times for the lamp facing back, and 171 times for the lamp overhead. Along one ray from the lamp, the floor and the wall want different densities — they are at the same place in the lamp’s picture and at different angles to the eye — and a map, however it is warped, gives one ray one density. The homography that is perfect for the floor is, for the wall, a different homography, because the wall is a different plane.
That makes the division of labour between one map and several exact rather than a matter of taste. A single map warped by the right homography is perfect for one plane. Several maps — one per surface, or per slab of distance from the eye, as cascaded shadow maps do — are what the second plane costs. For the overhead lamp the wall’s cost is of another kind: its rays skim every vertical wall, and a grazed surface needs texels without limit, which no number of maps from that lamp can supply.
Why the projective warp was always the right instrument
The earlier essays described a perspective shadow map as a map warped by one projective parameter to crowd its texels toward the lamp picture’s edge nearest the eye. That is a fair description of what it does and a misleading one of why it works. It works because the eye’s picture of the ground is a projective image of the lamp’s picture of the ground, and a projective warp can therefore undo the difference between them exactly. The one parameter is enough to capture the largest part of that difference for a lamp roughly in line with the eye; the full homography captures all of it for a lamp anywhere.
A point at infinity is an ordinary vertex found that homogeneous coordinates let a renderer treat the horizon as just another place. The same coordinates are what make the plane’s homography a matrix like any other, one the renderer can apply to its shadow pass with no more cost than the lamp’s own projection matrix. The divide is postponed is the reason that cost is nothing: a homography followed by the pipeline’s own divide is still one matrix and one divide per vertex.
What a renderer must know to use it
A plane’s homography between two views needs three things: the two cameras and the plane. A renderer has all three every frame. It knows the lamp, since it is about to render from it; it knows the eye; and the ground plane is part of the scene it is drawing. So the warp is not something to be searched for at run time. It is written down: the lamp’s projection, then the plane’s homography into the eye’s picture, then the fit of the eye’s picture of the floor into the map’s rectangle.
That fit is where the remaining freedom lives, and four numbers and a window and a tile is an off-centre frustum are the reason it costs nothing to use: an off-centre frustum is itself a projective map of the picture, and a renderer that can draw into an off-centre frustum can draw into a warped one. The near plane can be any plane found the same kind of freedom hiding in the projection matrix’s third row, and this is its analogue in the fourth: the matrix a shadow pass uses has room for exactly the warp it needs.
The price is that the warp belongs to the eye. When the eye moves, the plane’s homography changes, and the map must be warped again — the perspective shadow map’s familiar habit of re-fitting every frame, carried to its conclusion. A map warped for one eye position is perfect for that eye and slightly imperfect for the next, and the imperfection grows with how far the eye has moved relative to its height above the plane.
What the measurement assumes
One plane is the receiver. The floor is a single flat surface, and the plane’s homography is exact for it alone. Real scenes have terrain, stairs and furniture; for them the full warp is the best warp for whichever plane is chosen and a compromise for the rest, and how good a compromise depends on how far the scene departs from that plane.
The eye’s picture is what is matched. A texel “lands at one pixel” when its area on the eye’s picture is one pixel squared. A texel stretched into a long thin shape covers one pixel’s area and still aliases along its length; the plane’s homography maps squares to squares only where the two views agree in shape as well as area, and the grid figure is the evidence that for the floor they do.
The shadow is on the floor and cast by something near it. A shadow map’s job is to know, at each floor point, what lies between it and the lamp. The occluders do not affect how many texels the floor needs, which is all measured here.
Still open: what the second plane costs a warped map
The full warp is perfect for one plane and imperfect for every other, and a scene with a floor and walls asks how imperfect. The earlier bars measured the fewest texels any single map could spend on a floor and a wall together; they did not ask how close the floor’s own homography comes to that, or whether some other warp, perfect for neither surface, serves both better.
That question has a definite shape. For each lamp and each wall, the plane’s homography for the floor gives the wall texels of the wrong density — too coarse where the wall faces the eye, too fine where it grazes the lamp — and a warp chosen for the pair would trade some of the floor’s perfection for the wall. The measurement that settles it takes a floor and one vertical wall, searches the full projective warp for the fewest texels meeting both surfaces’ need, and compares that with the floor’s homography, with the two-surface bound, and with two maps, one warped for each plane — asking whether one compromise warp is ever within a small factor of two perfect ones, or whether a second plane always wants a second map.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A curved screen is eight flat ones — both name projective map, resolution, sampling grid, view frustum
- A hole is not preserved — both name point light, sampling grid, shadow projection
- A shadow across an edge — both name homography, projective map, shadow projection
- A shadow can be un-cast — both name homography, point light, projective map
- A wire with a corner in its shadow — both name point light, projective map, shadow projection
- Copying square by square — both name homography, projective map, sampling grid
Named objects
A flat tag is an object no other essay names yet.
HomographyPoint lightProjective mapResolutionSampling gridShadow projectionView frustum