Jittered frames rebuild a camera's horizon, saw-tooth and all
Worth reading first: The divide is postponed, not avoided · A texture does not interpolate on the page.
A fog band softens the horizon only by repainting the ground took a walker’s hazy horizon — a distance fog on the ground meeting a sky of uniform haze — and asked what a band of blended haze, carried a few rows down into the ground, does to the step where they meet. The step is spread rather than removed, and a soft horizon needs tens of rows that repaint every piece of ground beyond the end of the street. But four rows did one useful thing. Along a horizon rolled three degrees, the fog of the ground pixel nearest the horizon runs in a saw-tooth from column to column, jumping 0.061 a column at five kilometres of visibility; four rows of band brought that to 0.012. The essay called the saw-tooth the shimmer along a moving horizon, the part of the step a camera would not record.
It ended on a tool that costs no repainting. A renderer can move each frame’s sample position by a fraction of a pixel and average successive frames, so that over a few frames every pixel is averaged over its area even though each frame samples it at a point. The question was how many frames bring the saw-tooth down to the band’s 0.012, and how much of the step the average keeps.
None do, and the reason corrects the premise. The saw-tooth the earlier essay measured was already averaged over each pixel’s area. It is what a camera records.
Three ways to take a pixel’s fog
The scene is the earlier essays’: an eye 1.62 metres up, 740 pixels of focal length, the camera rolled three degrees, ground and sky rasterised along a shared row of horizon vertices snapped to 1/256 of a pixel. The ground’s fog is distance fog, , and the sky’s is one. A pixel’s fog can be taken three ways. A renderer with no anti-aliasing takes one sample at the pixel’s centre. A camera’s pixel integrates the scene over its square, which is computed here as the average of sixty-four points spread over it. And a renderer with temporal anti-aliasing takes one sample a frame, each moved by the next offset of a sequence, and shows the average of recent frames. The offsets are a Halton sequence in bases two and three, the sequence such renderers commonly walk through.
The saw-tooth is plain in every curve. The horizon crosses a row of pixels every nineteen columns, so the ground pixel nearest it is covered by sky for a share that sweeps from almost nothing to almost half and back, and its ground part sees the steepest stretch of the fog’s rise. Each time the horizon crosses into the next row, the pixel nearest it changes, and its fog drops from about 0.85 to 0.25. That drop, nineteen columns apart, is the saw-tooth.
The camera’s curve has it, at 0.061 a column on average. The centred sample has it more sharply, 0.082: a single point either lands on the sky’s one or on the ground’s fog at one depth, and does not blend them. Eight jittered frames averaged sit almost exactly on the camera’s curve, 0.063, and within 0.009 of it pixel by pixel. The band’s curve, high and gentle, jumps 0.012.
Averaged frames converge on the camera
Jitter is an estimator of the pixel’s area average. The question is only how fast it gets there.
One jittered frame is further from the camera than a centred sample — 0.11 against 0.067 on average over the last ground pixels of three hundred columns — because its one sample may land in the sky above the horizon or deep in the ground’s rise. Averaging closes on the area as the samples visit more of the pixel. With Halton offsets two frames are within 0.033, eight within 0.009, sixty-four within 0.009. With random offsets the approach is the slow one, as one over the square root of the count: 0.24 at one frame, 0.10 at four, 0.037 at sixty-four.
The Halton curve is not monotone, which is worth knowing because it is easy to assume otherwise. At sixteen frames it is 0.019, twice as far from the camera as at eight. A prefix of the sequence covers the pixel’s square more or less evenly depending on its length, and on a pixel whose content changes as steeply as the last ground pixel’s does, an uneven cover shows. The eight-frame point is a good one; a renderer that averages a fixed history should choose its length with the sequence in mind.
The saw-tooth that stays is the camera’s
The question the earlier essay asked was how many frames bring the saw-tooth down to the band’s 0.012. The answer is in what the frames converge on.
One frame of jitter leaves the saw-tooth at 0.077. Two frames bring it to 0.066, eight to 0.063, sixty-four to 0.062. The camera’s pixel is 0.061. The averaged frames settle on the camera and stay there, and no number of them goes lower, because what they are averaging towards is the area average itself.
So the earlier essay’s reading of its own number needs correcting. The 0.061 it measured was the area-averaged fog, the camera’s, and the saw-tooth in it is no artefact of sampling. The sky’s share of the last ground pixel sweeps from nothing to nearly half of it every nineteen columns, and the ground under it rises from about a tenth to one inside that pixel; a camera pixel integrating that scene records exactly the sweep. What sampling adds is the difference between 0.082 and 0.061, about a quarter, and that is the part jitter removes. The band’s 0.012 is not a better sampling of the same scene. It is a different scene — a ground painted hazier than its distance — whose last pixels have much less of the fog’s rise in them to sweep through.
That is the honest division the earlier essay’s last section reached for, now with the sampling separated out. A renderer that wants a camera’s horizon can have it with eight frames of jitter and nothing repainted. A renderer that wants a horizon quieter than a camera’s has to change the scene, and the band is one way.
At every visibility, the same division
The saw-tooth depends on how much of the fog’s rise the last pixel holds, and that depends on the visibility. The division does not.
At a kilometre of visibility the camera’s saw-tooth is 0.027 a column, at two kilometres 0.047, at five 0.061, at ten 0.062 and at thirty 0.058. It grows with the visibility at first, because clearer air leaves the ground’s fog low everywhere except in the last fraction of a row, and the last pixel then holds more of the rise; past ten kilometres it levels off, the last pixel already holding nearly all of it. One centred sample sits above the camera at every visibility, by between a tenth and a half again. Sixty-four Halton frames sit on the camera to within a few thousandths everywhere. The band sits far below, at 0.005 to 0.014, at every visibility — always a quarter or less of the camera’s — because at every visibility it has changed what the last pixel contains.
Why a camera’s saw-tooth reads as a straight horizon
It may seem odd that a saw-tooth of 0.061 a column is “what a camera records”, since nobody looking at a photograph of a hazy plain sees a saw-tooth along its horizon. The two facts are the same fact. Along a tilted edge, the pixel nearest the edge records how much of it the edge covers, and that share is exactly what tells the eye where, inside the pixel, the edge runs. A column whose last ground pixel is three-quarters ground and a column twenty columns on whose last pixel is a quarter ground are drawing the same straight horizon a fraction of a row higher. The saw-tooth in the numbers is the anti-aliasing of the edge, read one pixel at a time; looked at as a picture, it is a smooth tilted line.
What a centred sample adds is the part that does not encode the edge. A single point either lands on the sky or on the ground, and when it lands on the ground it reports the fog at one depth rather than averaged over the depths the pixel spans; so the last pixel’s fog jumps more steeply than the edge’s position can explain, and the extra quarter of the saw-tooth is noise about where the edge is. In motion the difference is sharper. As the horizon slides past a pixel, an area average changes continuously with the share covered; a centred sample does not change at all until the horizon crosses the centre, and then changes by most of the fog’s range at once. That jump, repeated along the horizon as it moves, is the shimmer a camera never shows.
So the earlier essay’s word was right and its measurement was of something else. The shimmer is real, and it is what a renderer with one sample at the centre of each pixel produces; the saw-tooth it measured is the shimmer’s smooth relative, the trace a camera leaves of the same tilted edge. Jitter averaged over frames turns the first into the second. Nothing it does turns the second into the band’s.
This depends on the ground and the sky meeting cleanly, which ground and sky meet without a crack established for this rasteriser: every pixel along the horizon belongs to exactly one of them at every sample point, so the area average is the true coverage. A horizon with a crack — pixels that neither owns — would average the crack’s background into the last row, and jitter would converge on a horizon a camera never sees. And it depends on the sky’s value being read the way a sky is carried as a direction carries it: at every sample point that lands above the horizon, the sky’s fog is the sky’s, not a ground value extrapolated past infinity.
The step is in the scene
The earlier question had a second half: how much of the step’s own size the average keeps, since an area average is what a camera records and the band was not.
All of it. The largest change between vertically adjacent pixels at the horizon, averaged over three hundred columns, is 0.65 with one centred sample, 0.561 with sixty-four jittered frames averaged and 0.560 with the pixel’s whole area. Four rows of band bring it to 0.33. The step is the ground’s fog rising from about a tenth to one inside the last fraction of a row, and it is a property of the scene and the eye’s height: a ground and a sky share an elevation found it physical, the same in a photograph from the same height. Averaging the pixel’s area spreads a sharp edge over a pixel, which is all anti-aliasing ever does; it cannot spread a rise that already happens within one pixel over several. A photograph of a hazy plain from a walker’s height has the step, and so does a renderer that matches the photograph.
Jitter’s price is frames, and flicker
The band’s price was repainted ground. Jitter’s is paid in time.
With nothing in the scene moving, a renderer that shows each jittered frame as it is drawn makes the last ground pixels change by 0.29 from frame to frame. That is a flicker five times larger than the saw-tooth it was meant to cure: on the last ground pixel a sample that lands in the sky reads one and a sample a little below the horizon reads the ground’s fog a long way out, and those differ by most of the fog’s range. Shown as the average of the last two frames the change is 0.14; four, 0.073; sixteen, 0.018; thirty-two, 0.009. The history is what turns jitter from noise into an average, which is why a renderer that jitters always keeps one.
That makes motion the hard case, which the measurement here does not cover. A history averages frames that show the same pixel, and when the camera rolls or the head turns, the pixel the history holds is no longer the pixel being drawn. A renderer then either reprojects its history to follow the motion, or shortens it, and a shortened history brings the flicker back. The band has no such dependence: four rows of band quiet the horizon in a single frame, still or moving, by painting a scene whose last pixels hold less of the rise. A pixel is not a point is the general argument for averaging over area; jitter is the cheapest way to do it in space and the most expensive in time.
What each tool does
Put together, there are three things that could be called the horizon’s shimmer, and the tools sort them.
The first is what point sampling adds: the difference between a centred sample’s saw-tooth, 0.082, and a camera’s, 0.061, and the frame-to-frame jumps of a horizon sampled at points as it moves. Jitter with a history removes it, eight Halton frames to within a hundredth. So does any spatial supersampling, at the cost of samples rather than frames.
The second is the camera’s own saw-tooth, 0.061 a column at five kilometres, and the step of 0.56 between the rows at the horizon. Nothing that samples the same scene more faithfully removes either, because both are in the scene as a pixel-sized camera sees it. A renderer that matches photographs should keep them.
The third is the step and saw-tooth a renderer would like the scene not to have. The band removes most of both by changing the scene — painting the ground beyond a few hundred metres hazier than its distance. That is a decision about the picture’s look, not its sampling, and the earlier essay priced it in metres of falsified ground.
A texture reaches the horizon as a rate described the sparkle of a ground texture near the horizon, which is the same three-way division for a texture instead of a fog: a texture that changes faster than a pixel at the horizon sparkles under point sampling, is averaged to grey by a camera, and is quiet only if a renderer chooses to fade it. The divide is postponed began this sequence by carrying the ground’s attributes so that nothing divides by depth until the pixel; the horizon is where the pixel’s own size is the last thing left to decide, and the decision is whether the pixel shows the world or a version of it.
What the measurement assumes
The camera is still. The flicker measurement shows a still horizon with jittered samples, and the convergence measurements average frames of one still picture. A moving camera needs a history that follows it, and how well reprojection keeps the convergence along a horizon — where the ground’s fog changes fastest per pixel — was not measured.
The offsets are shared by every pixel. A renderer moves the whole frame’s samples by one offset, as here. Per-pixel offsets would decorrelate the columns, which changes how the saw-tooth looks frame by frame but not what the average converges on.
The pixel’s area is a box. A camera’s pixel has a sensitivity that is not quite uniform over its square, and a renderer’s resolve filter is usually wider than a pixel. Both widen what the average converges on, a little like a narrow band of their own, as the earlier essay noted for the resolve filter: they spread the step over about two rows and repaint nothing beyond their reach.
The fog is the earlier essays’. A uniform distance fog on a flat ground, a uniform sky, each pixel’s fog computed from the ground’s carried attributes. A fog read from a depth buffer instead would add the buffer’s own quantisation in exactly these rows, where the precision a depth buffer has left is least; that is not modelled here. A real atmosphere that thins with height, or a horizon of hills, changes where the rise falls and how steep it is, and so the size of everything above; it does not change which tool removes which part.
Still open: whether a moving horizon keeps the camera’s saw-tooth under reprojection
Everything here is a still picture. A renderer’s history earns its keep only while the scene holds still in the frame; a horizon that rolls with the camera slides its saw-tooth along the row of pixels, and a history that does not follow it averages the last ground pixel with what was the sky’s pixel a few frames ago.
The measurement that settles what jitter is worth on a moving horizon rolls the camera at a stated rate, a degree a second say, renders each frame with its next Halton offset, and keeps a history that is either reprojected to follow the roll — each history sample moved to where its ray now falls — or simply decayed with a stated weight. It asks how far the shown fog of the last ground pixels then sits from a camera’s moving picture, frame by frame; whether the reprojected history, which resamples a steep fog at shifted points, keeps the convergence the still picture had or reintroduces a flicker of its own; and at what rate of roll a renderer would do better to drop the history and add four rows of band, which quiets a moving horizon as well as a still one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A point at infinity is an ordinary vertex — both name clip space, homogeneous coordinates, horizon, point at infinity, projective map
- What happens behind the eye — both name clip space, homogeneous coordinates, point at infinity, projective map
- A point and a line are one object — both name homogeneous coordinates, point at infinity, projective map
- An angle is a cross-ratio — both name homogeneous coordinates, horizon, point at infinity
- The conic a circle becomes — both name horizon, point at infinity, projective map
- Three conics are one conic and a choice of horizon — both name horizon, point at infinity, projective map
Named objects
A flat tag is an object no other essay names yet.
Clip spaceHomogeneous coordinatesHorizonpoint at infinityProjective mapSamplingSampling grid