Counting cloud by counting pixels
Worth reading first: The third column is area.
A camera points at the sky. Something thresholds the image and counts the pixels that came out white. The number it reports is called cloud cover, and it is a fraction of the sky.
Every part of that sentence is fine except the last clause, which is true on exactly one kind of lens.
This essay is a measurement rather than an argument, and the thing that makes it a measurement is that the true answer comes from somewhere other than a count.
Making the truth knowable
The problem with testing a counting method is that the reference is usually another count. Comparing two counts establishes that they disagree and not which is wrong.
So the sky here is built from spherical caps. A cap of angular radius subtends a solid angle of exactly , and if the caps do not overlap, their solid angles add. That gives the cover as a closed form, with no sampling anywhere in it.
Two conditions have to hold for that to be true, and both are asserted rather than arranged: each cap lies wholly inside the field being counted, and no two caps overlap. If either failed, the “true” answer would be an over-count and the entire comparison would be measuring the arithmetic instead of the surfaces.
Over a field of 70° from the zenith, the four caps used here cover 10.31% of it. Exactly, to as many digits as anyone wants.
The count
The counting is done exactly as an image-processing script does it, because that is the operation under test.
A square grid is laid over the picture. Each cell’s centre is turned back into a direction through the surface’s inverse map. Cells whose direction falls in the counted field are tallied; those whose direction is in a cloud are tallied separately. The ratio is the reported cover.
The grid is the same size for every surface, and the field counted is the same 70° for every surface — the picture’s radius is taken at whatever radius that surface puts 70° at, so every camera is being asked about the same piece of sky. Without that the comparison would be between fields of view as much as between rules.
The results:
| surface | reports | error |
|---|---|---|
| equal-area fisheye | 10.31% | −0.01% |
| equidistant fisheye | 9.37% | −9.1% |
| stereographic | 7.56% | −26.6% |
| flat plane | 2.40% | −76.8% |
Reading the table
The first row is the grid’s own error and nothing else. 0.01% relative, on a 401-square grid over a round field with curved cloud boundaries — the cells that straddle a boundary are the whole of it. That number is what every other row has to be judged against, and stating it is what turns “the equal-area one is right” from a claim into a measurement with a bound.
Every error is negative, and it is worth seeing why before the numbers. The clouds here sit near the middle of the field, and every surface except the equal-area one inflates the outer part of the picture relative to the inner. So the clouds occupy a smaller share of a picture that has been stretched at its rim, and the count comes out low. Move the clouds to the horizon ring and every sign flips — which is the point about the error being systematic rather than random. It has a direction, and the direction depends on where in the sky the thing being counted is.
The flat plane’s 77% is not a bad lens. It is the correct behaviour of a rectilinear projection asked about a 70° field: at 70° is about 25, so the rim of that picture is enormously over-weighted and everything near the axis is drowned. An ordinary wide-angle photograph is not a slightly biased instrument for this job; it is the wrong instrument.
And 9% on the equidistant fisheye is the dangerous row, because it is the lens most people counting a sky actually own, and 9% is small enough to look like a real result.
The repair
The count is recoverable, and the repair is exactly what the third column of the surface table says it should be.
Solid angle per unit picture area is the reciprocal of the area scale. So weight each cell by the reciprocal of its surface’s area scale at that cell’s direction, and the sum is a solid angle rather than an area. That is ordinary quadrature with the right Jacobian in it, and it is the thing an unweighted count is silently omitting.
Weighted, all four surfaces return the cover to better than a fifth of a per cent — the flat plane included.
That is worth stating plainly because it changes the practical advice. A count made on the wrong surface is not lost, provided the rule is known. It costs one multiplication per pixel and it is exact.
What is lost is a count made on an unknown surface. “Fisheye” is not a rule; there are four in common use and they differ by up to 42% in area weight at 80°. A dataset of sky images with no lens model attached cannot be turned into cover fractions afterwards, and no amount of processing supplies what was never recorded.
Why the answer comes out low, in detail
The sign of every error is worth an argument rather than an observation, because it is the part a practitioner can reason about without recomputing anything.
Write the reported cover as a ratio of two sums: picture area inside the clouds, over picture area inside the field. Each of those sums is the corresponding solid angle multiplied by the local area scale, integrated. If the area scale were constant the two factors would cancel and the ratio would be right, which is the equal-area case.
When the area scale grows outward, the denominator is inflated more than the numerator whenever the clouds sit inside the field rather than at its rim, because the outer annulus — which contains no cloud — is where the inflation lives. So the ratio falls. Put the clouds at the rim instead and the numerator is inflated more, and the ratio rises.
The general statement is that the bias is the correlation between where the thing being counted is and where the area scale is large. That is why the error is not a fixed correction that can be tabulated once: it depends on the sky, not only on the lens.
It also explains a case that would otherwise look paradoxical. A completely uniform sky — cover spread evenly over the whole field — is counted correctly on any surface, because then the numerator and denominator inflate together. So a test made on uniform cloud would find nothing wrong with any of these lenses, and would be exactly the wrong test.
Two errors that are not this one
Two things people reach for when a sky count comes out wrong are not the problem here, and separating them is useful.
Vignetting is not this. A lens darkens toward the rim and a threshold applied to a darkened rim misclassifies pixels. That is a photometric error, it affects which pixels are called cloud, and it is corrected with a flat field. The error in this essay is geometric: every pixel is classified correctly and the pixels are the wrong size.
And distortion in the lens sense is not this either. A fisheye’s rule is not an error to be corrected; it is the design. Applying a barrel-distortion correction to an equal-area fisheye does not improve the count, it converts the surface to another one — usually rectilinear — and makes the count much worse.
The distinction is worth carrying because both mistakes produce a plausible pipeline that quietly measures the wrong thing.
The same weight, everywhere else
Once stated, the Jacobian turns up in every measurement made by counting picture, and it is worth listing them because the sky is only the most obvious.
Canopy cover, from a fisheye pointed up through trees, has exactly this structure and the same one-sided bias — worse, because the obstruction of interest is concentrated near the horizon where the weight is furthest from 1.
Sky-view factor, the fraction of the hemisphere a point in a street can see, is a solid-angle integral and is routinely computed by counting pixels in an upward fisheye.
Illuminance from a source of known shape is an integral of a cosine against solid angle, so it needs the same weight and a second one on top of it.
And coverage of a projected image — what fraction of a dome or a wall a projector’s frame reaches — is the same integral run the other way, from a flat picture surface onto a curved receiver.
In every case the operation is “sum something over the picture”, and in every case the sum is a solid-angle integral wearing a pixel count’s clothes.
How much of the sky a camera can see at all
There is a prior question that the field-of-view convention above hides, and it decides whether any of this is available.
A sky camera measuring cloud cover wants the whole hemisphere: 90° from the zenith, a solid angle of . A flat lens cannot reach it at any focal length, because the picture is unbounded as the field approaches 180° — so a rectilinear sky camera is not a badly weighted instrument for a hemisphere, it is an instrument that does not cover the domain.
That is why the comparison here is run at 70° rather than at 90°. It is the widest field every one of the four surfaces can hold, and comparing them over a field one of them cannot reach would be comparing coverage rather than weighting.
The practical consequence is a two-stage question, in order. Can this surface hold the field? and only then does it weight the field correctly? The four rules all pass the first out to 180° and beyond; the flat plane fails it well before, and no Jacobian repairs a direction that has no image.
What the grid error bounds
There is one more number worth pinning down, because a measurement with an unstated resolution is not finished.
The grid error is the whole of the equal-area row’s departure from truth, and it comes from cells straddling a cloud’s boundary. It therefore scales like the ratio of boundary length to area — inversely with the grid’s linear size — and doubling the grid roughly halves it, which is what the coarser-grid figure shows.
That gives the claim a bound with a slope. At 401 squares the error is 0.01% relative; at 241 it is larger; and there is no grid at which the other three rows come right, because their error is not a discretisation at all.
Distinguishing those two is the useful skill. A number that shrinks when the grid is refined is a numerical artefact. A number that does not is the physics, or in this case the geometry.
Where the true answer could have come from instead
The caps are a device and it is worth saying what the alternatives would have cost, because the choice is the reason this essay is a measurement.
A dense direction sample. Scatter a few million directions uniformly over the field and count how many fall in cloud. That is a Monte Carlo estimate with its own error of order one over the square root of the count, and comparing a pixel count against it would be comparing two approximations. Better than nothing and not a reference.
A finer version of the same pixel count. Circular, obviously: it assumes the thing under test.
A closed form for an arbitrary painted mask. There isn’t one, which is why the mask is not arbitrary.
Caps. A cap’s solid angle is elementary and disjoint caps add, so the cover is exact and the only conditions are geometric ones that can be checked. The cost is that the clouds are round, which is a loss of realism and no loss of generality: nothing in the counting operation knows the shape of the region it is summing over.
That is a pattern worth naming, because this site uses it constantly. When a measurement needs a reference, arrange the scene so the reference is a closed form — even if that makes the scene artificial — rather than compare two estimates and call the difference an error.
What to do
Know the rule. Not the brand, not the word fisheye: the function from angle to radius. Fit it from a photograph of known angular structure if it was not supplied.
Weight by the reciprocal of the area scale, unless the rule is already equal-area, in which case the weight is 1 and a plain count is correct.
State the field being counted and count the same field on every instrument being compared, or the comparison includes the field of view.
And report the grid error. It is the only part of the answer that improves with effort, and knowing its size is what says whether more effort is worth spending.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A mirror ball is an equal-area fisheye — both name area scale, equidistant, equisolid, fisheye, solid angle, stereographic
- Conformal is not undistorted — both name area scale, jacobian, necessary, not sufficient
- A light far enough away — both name instrument limit, single-view metrology
- Each system answers its own question — both name instrument limit, necessary, not sufficient
- How well the floor has to be known — both name instrument limit, single-view metrology
- Measuring a room off the page — both name area scale, single-view metrology
Named objects
A flat tag is an object no other essay names yet.
Area scaleEquidistantEquisolidFisheyeinstrument limitJacobiannecessary, not sufficientQuadraturesingle-view metrologySolid angleStereographic