What a machine computes
The divide is postponed, not avoided
A renderer does not divide by depth. It multiplies by a four-by-four matrix that carries the depth in a fourth coordinate and divides later, and the postponement is not an optimisation — it is what makes clipping and texture interpolation possible at all. The matrix and this site's pinhole put every point on the same pixel to five parts in a hundred trillion.
What happens behind the eye
A point behind the camera has a perfectly plausible image. Dividing by a negative fourth coordinate flips both signs, so the point lands through the principal point on the far side of the frame, and a segment crossing the eye plane is drawn straight, inside the frame, and running in exactly the opposite direction — a direction cosine of −1.0000.
The precision a depth buffer has left
Depth is stored as an affine function of one over the distance, so half of a buffer's codes are spent before the harmonic mean of the near and far planes — twenty centimetres out of a kilometre. The resolution goes as the square of the distance, and the fix that works is not more bits.
A pixel is not a point
Where the sample sits inside a pixel is a convention, and getting it wrong shifts every mark by half a pixel in each axis. What that costs can be measured by recovering the camera from the picture — the answer is a principal point exactly 0.707 px from the truth with the focal length untouched, and the other half-pixel mistake does precisely the reverse.
A texture does not interpolate on the page
Walking across a drawn surface at a constant rate walks across the real one at a rate that changes, and the worst gap is a closed form in the depth ratio alone — 0.52 at ten to one, more than half the whole range. It is exactly the error a person makes dividing depth by eye, made by a machine, and the fix is the fourth coordinate the pipeline kept.
Four numbers and a window
A projection matrix is built from six numbers and one of them is not a number at all. Four sides carry the focal length and the principal point; the near and far planes move nothing a reader can see; and the bottom row, (0, 0, 1, 0), is the only place the depth divides — set it to (0, 0, 0, 1) and the same machine draws a parallel projection.
A tile is an off-centre frustum
Rendering a picture in tiles is exact, and the way to do it is one line of arithmetic: a tile's sides are the whole frustum's sides read at the tile's own pixel bounds. Aiming the camera at each tile instead is defensible at every step and is a different picture, out by about a tenth of a tile whatever the tile size.
The near plane can be any plane
Rewrite one row of a projection matrix and the near plane stops being perpendicular to the axis and becomes whatever plane is asked for. Every x and every y is untouched — it is the same projection of the same scene from the same eye — and the depth order is wrecked, which is a clean separation of the two things a projection matrix does.
One plane is nearly free
The near and far planes enter a depth buffer's precision through 1/near − 1/far, and one of those reciprocals is enormous. Pushing the far plane out by a factor of a thousand costs a tenth of a per cent; bringing the near plane in by the same factor costs a factor of a thousand — and an infinite far plane is the limit of the first rather than a separate case.
A curved screen is eight flat ones
A projection matrix is a plane and nothing else, so a curved display cannot be rendered — it has to be driven as several planes and assembled. The gap between chord and arc is the whole error, it goes as the square of the angle each piece spans, and the piece count therefore goes as the inverse root of the tolerance — three for eight pixels, eight for one, fifteen for a quarter.
One depth per sample is not enough
A depth buffer keeps a single distance at each sample, so a post-process blur can only ask how far away the thing at this pixel is. Across an occluding edge that answer is two depths and an occlusion, and the gather it produces differs from the pupil's own integral by 70 per cent of full scale over a band eleven pixels wide.
Along a line of constant depth the page is affine
Stepping a texture by a constant amount per pixel is wrong across a receding floor and exactly right along any line of it that stays at one depth — and on every plane those lines run parallel to its own vanishing line. Turn a 120 px span 1° away from that direction and it is 0.79 px out; roll the camera a hundredth of a degree and a floor drawn to 30 m is out by 0.69 px on its worst scanline.
A point at infinity is an ordinary vertex
Give a vertex a zero in its fourth slot and it stops being a point and becomes a direction — and the projection matrix draws it anyway, through the same multiply and the same divide, on that direction's vanishing point to seven trillionths of a pixel. Slide the eye ten metres and it does not move. Two of them bound a ground that reaches the horizon, where a ground drawn to ten kilometres stops a fifth of a pixel short.
A shadow map's texels land by two distances and two cosines
A renderer finds its shadows by taking a second picture from the lamp and storing a depth in every texel. Each texel reaches the screen through the surface it falls on, and how many pixels it covers there is a closed form — two focal lengths, two distances and two cosines. From a lamp beside the eye every texel lands at 0.79 px; from a lamp 40 m ahead facing back, the same map lands texels of 8.69 px on the floor 5 m out.
A tilted span walks a staircase
A span along a banked floor's constant-depth direction is exact, and a renderer visits pixels rather than the span. Snapped to the grid, a 120 px span at a 20° bank costs 0.577 px where the same span along a page row costs 13.26 — twenty-three times better — and it never rises above 1.22 px at any bank. The price is bookkeeping: a band of twenty-four such spans draws 53 of its 1,368 pixels twice.
One warped shadow map, and what it cannot reach
The lamp facing the eye needed a shadow map 9,530 texels square — 90.8 million texels — for none of them to land on the near floor larger than a pixel. Fitted to the floor and warped by one projective parameter, the same map needs 805 thousand; nothing can do better than one texel per pixel of the eye's picture, 148 thousand. What stays out of reach is not the cosine of a surface, which a warp absorbs, but two surfaces that want different densities along one ray from the lamp.
A texture reaches the horizon as a rate
A ground drawn to infinity cannot carry a texture coordinate at its far corners, because a repeating texture has no coordinate there. It can carry a rate — so many checks per metre along the direction — and a second number that is 1 at points and 0 at directions. Interpolated like every other attribute and divided once, that pair is exact half a pixel from the horizon. Give the same corner a value instead and the ground is drawn in reverse perspective.
One homography makes a shadow map the eye's picture
A shadow map for a lamp facing the eye needed 706 thousand texels however its rows were re-spaced, nearly five times the one texel per pixel no map can beat. Warp the whole map by a projective transformation, not only its rows, and it needs 154 thousand — within five per cent of the bound — and every texel, carried into the eye's picture, lands at one pixel. The reason is exact: the eye's picture of a floor and the lamp's picture of the same floor are one homography apart.
A sky is carried as a direction
A triangle of sky has all three vertices at infinity, so the weight that lets a texture reach the horizon is zero everywhere and there is nothing to divide by. The attribute that belongs to such a triangle is the direction itself: carried over w like any other, it is every pixel's own ray to 5e-14 degrees. Carry the vertices' azimuth and elevation instead and a 60° triangle is 7.7° out, a 4096-texel sky needs triangles under ten degrees wide, and a triangle across the seam where azimuth wraps is painted with the opposite sky.
Ground and sky meet at the horizon without a crack
A ground of rate triangles and a sky of direction triangles share their vertices along the horizon, and rasterised as a graphics processor does it — positions snapped to a fraction of a pixel, every centre given to one triangle by the top-left rule — they lose no pixel and claim none twice, at any roll and any snapping. What the horizon does have is a sliver: a centre that falls within half a snapping step of it goes to whichever side the snapped edge puts it, and one that falls exactly on it, rolled one way, is given to the ground at a weight of zero. A ground stopped at a far plane leaves the crack the shared vertex never does: f·h/D rows.