What a null result is worth in decades
Worth reading first: An error with two terms · Recovering the camera from the picture it drew.
An error with two terms builds an instrument that separates the part of an error responding to effort from the part that does not, and runs it over seven of this collection’s own laws. It classifies all seven correctly.
Every one of those classifications rests on an assumption the instrument cannot check, which is that the sweep was long enough. This essay measures what happens when it is not, and the first answer was wrong.
The failure that was expected and does not happen
The prediction was that a short sweep would invent a floor. The reasoning is plausible: least squares has a free parameter, a short sweep does not constrain it, and an unconstrained parameter will absorb whatever is going.
It does not happen. Generate a pure power law — floor exactly zero, by construction, no noise anywhere — sample it over a third of a decade, and fit. The floor comes back at , which is correctly zero.
The reason is worth having because it corrects a piece of folklore. Least squares does not spend a free parameter on nothing; it spends it on whatever is there, and on exact data there is nothing. The fitted floor is zero because zero is the value that minimises the residual, and it minimises it exactly.
So a fit reporting no floor on clean data is telling the truth, and the first draft of this essay was about a failure that does not exist.
The failure that does
What a short sweep cannot do is rule a floor out, and that is a different claim.
Hold at a series of values, refit and around each, and ask which values leave a residual inside the measurement’s own noise. What comes back is an interval rather than a point, and the interval’s width is what “no floor found” is worth.
At one per cent noise:
- over a third of a decade, floors up to 21.5% of the first sample survive;
- over six tenths, 4.5%;
- over one decade, 1.0%;
- over one and a half, under a twentieth of a per cent.
The fit reports no floor at every one of those sweeps. It is telling the truth at all of them and saying something different each time.
How the interval is computed
The procedure is worth stating, because the obvious implementation is wrong in a way that flatters the answer.
The wrong version fixes , holds and at their best-fit values from the free fit, and asks how much the residual grew. That measures how badly the original model does with a floor bolted on, which is not the question.
The right version refits. For each candidate , subtract it from the data and fit a pure power law to what remains — which is a straight line in logs and is exactly fitPowerLaw, the same routine every exponent in this collection comes from — then compute the residual of that combination against the data.
The difference is large. Refitting lets and move to accommodate the floor, which is what a real alternative hypothesis would do, and the interval that comes out is several times wider than the one the frozen version reports.
That is the general shape of every profile-likelihood argument and it is easy to get backwards. An alternative hypothesis is allowed its own best parameters, and comparing it at the null’s parameters is comparing it with a handicap.
Why an interval rather than a residual
The natural thing to report is the fit’s residual, and it is worth saying why that does not answer the question.
A residual says how well the best-fitting model explains the data. It says nothing about how much worse a different model would be, and that is the quantity a null result is about. A floor of a fifth of the first sample, over a third of a decade, produces a fit whose residual is inside the noise — so the residual is small under the true model and small under the alternative, and it cannot distinguish them.
The interval asks the right question directly: which alternatives are still standing? That is a much older idea than any of this — it is what a confidence interval is — and the point here is only that error laws are almost never reported that way. They are reported as an exponent and a residual, and neither of those is the number a reader needs.
The arithmetic, and it is not deep
The interval’s width follows from the same inequality the instrument’s certificate uses, run backwards.
To see a floor , the falling term has to get below it: , so , so the sweep needs decades.
Turn that round. A sweep of decades starting at reaches , at which the falling term is . Any floor below that is hidden underneath it. So
and the interval’s width falls as a power of ten in the number of decades, at a rate set by the price.
That is why a steep price is an expensive purchase and a cheap diagnosis. An law’s hidden floor shrinks by four orders of magnitude over two decades where an law’s shrinks by two.
What a decade costs, and which lever is the better buy
The interval’s width is one half of a decision and the other half is what a decade costs, which for most of these sweeps is simple: reaching means evaluating up to that many lamps, views or samples, so the work goes as . Setting that against the bound,
for a total cost . Certification is polynomial in the cost, with the law’s own price as the exponent. A hundredfold budget buys a hundredfold tighter bound on an law and a ten-thousandfold one on an law.
The competing lever behaves quite differently. A floor below the measurement’s own noise is invisible however long the sweep, so the noise sets a second bound; and reducing noise by averaging costs evaluations for a factor of , giving .
So the comparison is one exponent against a half:
Extending the sweep beats reducing the noise whenever the law’s price exceeds one half — which is every law in this set.
That is a clean answer to the budget question and it is worth stating because the instinct runs the other way. Faced with a null result a reader’s first thought is usually to measure more carefully; the arithmetic says to measure further instead, and by a margin that grows with the exponent. Only for a law falling more slowly than — which would be a very shallow price indeed — does care beat range.
The two levers do have to be spent in the right order, though, and the order is set by which bound is currently binding. Sweeping past the point where the falling term drops below the noise buys nothing, because the noise’s bound has taken over and no further range improves it. So the procedure is: sweep until the curve reaches the noise floor, then average, then sweep again — and the cross-over between the two is visible in the data as the point where the measured curve stops falling and starts scattering.
Two levers, and they are not the same lever
The interval depends on two things and it is worth separating them, because a reader with a fixed budget has to choose.
The range enters as — exponentially. Doubling the number of decades squares the reduction in what can hide.
The noise enters linearly. Halving the measurement noise halves the interval and no more.
So a reader who can either measure twice as carefully or sweep one more decade should almost always sweep, and the margin is enormous rather than marginal. One more decade of an law buys a factor of a hundred; halving the noise buys two.
That ordering is the reverse of what most experimental effort goes into, and the reason is not stupidity — precision is often the only thing under control, and the range is set by what the apparatus can reach. But when both are available, the arithmetic is not close.
The control
A law that really is all floor returns its floor from the shortest sweep there is, exactly.
Sample a constant over a range of spanning nothing at all — nine points from 1.0 to 1.4 — and the fitted floor comes back at the constant, to nine decimal places. That is not a coincidence of the fit’s conditioning; a constant is exactly representable in the model with , and least squares finds it.
The control matters because without it the essay would be arguing that short sweeps are untrustworthy, and they are not. They are untrustworthy in one direction. A short sweep that finds a floor has found one. A short sweep that finds none has found nothing.
What this does to the seven laws
Every classification in the previous rung acquires a number, and four of the seven survive it comfortably.
The two pure-floor laws need zero decades, because their falling term is nothing and there is no covering to get out from under. They are certified by any sweep at all.
The hull excess and the sagitta-against-gaps have steep prices, so their sweeps of two decades hide only floors below of the first sample. Certified.
The seam ghost is the awkward one. Its price is about one, its sweep is two and a half decades, and its certificate says sixteen decades would be needed to see the floor its fit reports — which is another way of saying its fitted floor is arithmetic noise rather than a measurement. What the sweep does establish is that no floor larger than about of the first sample is hiding there, which is a real and useful bound and is not the same as zero.
Where this has already bitten this collection
The pattern is not hypothetical here and it is worth naming the two places it has already appeared, because in both the shape of the failure was recognised and the number was not computed.
A shadow reports a curvature and sometimes invents one. A fit with one parameter, handed a defect it has no parameter for, returns a number with a small residual and no referent. That essay reports the number and the residual and does not report what range of true curvatures would have produced an indistinguishable fit — which is exactly this interval, in a different subject.
A fitted radius is wrong before it is uncertain. A paraboloid fitted to a sphere’s fan over a small aperture leaves a residual below any measurement floor and a radius wrong by a stated per cent. The aperture at which the residual clears the floor is that essay’s version of the number of decades, and it is computed there — which makes it the one place in this collection where the sweep-length question was already answered properly.
There is a third case in this collection where the same thing happens and the sweep is genuinely unreachable. A single-view height has a bias — real, measured, eight microns at a pixel of click noise — and the number of measurements before it overtakes the spread is over half a million. That is a floor that has been found rather than ruled out, and it is still, for every practical purpose, not there.
So the interval cuts both ways. A floor too small to certify is a floor too small to matter, and the same arithmetic that says a null result is weak says when a positive one is irrelevant.
What it means for a design decision
The interval turns a vague question into an arithmetical one, and the vague question is the common one.
A photographer asks whether shooting more frames will remove the ghosting. A reconstruction asks whether more views will close the gap. An optician asks whether a wider aperture will settle the shape. In each case the honest answer has two parts: what the falling term costs, and how much of the error is not falling.
Where a sweep has been run, both parts are available. Where it has not, the second part is unknown and the first is a guess — and the usual practice is to quote the first as though the second were zero.
The correction is small and it is a habit rather than a technique. Before quoting a price, say what the sweep behind it could not have seen. The panorama rung does it: it establishes the vertical miss as a set with one element rather than as a slow decline, which is the strongest form of a positive floor result. The silhouette rung does it: the notch’s area is the same number at four views and at a hundred and twenty-eight. Both essays could have reported an exponent and stopped.
What to report
Three numbers, and error laws are almost always reported with one.
The exponent, with the range it was fitted over — because a law that is only asymptotically a power law has an exponent that depends on the range, which the seam ghost demonstrates at 0.918 over the whole sweep and 0.999 over the tail.
The floor, if one was found.
And, if one was not, the largest floor the sweep cannot rule out. That is the number that makes a null result a claim rather than an absence, and it costs one extra pass over data already collected.
The uncomfortable corollary
There is a version of this that applies to the instrument itself and it should be stated rather than left implied.
The two-term model can only find floors of the shape it contains. A term falling as — a genuine power law with a very shallow price — is indistinguishable from a floor over any sweep shorter than about forty decades, so the instrument would call it a floor and be wrong in a way no amount of the same data corrects.
The same reading applies to every other identifiability question in this collection, and it is worth noting that the answers differ. A dome port’s two numbers are unidentifiable exactly — a scale invariance, holding to the last bit at every angle — so no range of any experiment separates them. A pane’s thickness and index are unidentifiable over a narrow fan and separable over a wide one, so range is exactly what buys them.
Those are the two cases this instrument has to distinguish and it distinguishes them the same way it distinguishes a floor from a falling term: by asking whether the difference between the hypotheses is bounded away from zero over the range available. That is not a defect that a better fit removes; it is the same problem one level up. An instrument that separates two hypotheses can only do so where they differ, and hypotheses that differ only outside the range of any possible experiment are, for the purposes of the experiment, one hypothesis.
That last figure is the reductio and it is worth reading as one rather than as a bug. The certificate of a zero floor is infinite, which sounds like a broken calculation and is the honest answer: proving that a quantity is exactly zero requires infinite range, and every finite experiment establishes a bound instead. The instrument returns infinity rather than a large number because rounding it to a large number would suggest an achievable sweep.
Which returns the essay to the only claim it actually supports. A fit reporting no floor is reporting that no floor of a certain size is present over a certain range. Both of those qualifications are computable and neither is usually computed, and everything else in this row rests on them.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Two distances to infinity — both name asymptotics, asymptotics, power law, sampling, tolerance
- A floor is read along curves — both name demonstration, instrument limit, residual, sampling
- What a panel says about its maker — both name error term, identifiability, residual, tolerance
- A frame is an interval — both name demonstration, instrument limit, residual
- A third eye that lands on the next post — both name demonstration, instrument limit, residual
- An ambiguity is not an uncertainty — both name degeneracy, identifiability, instrument limit
Named objects
A flat tag is an object no other essay names yet.
AsymptoticsDegeneracyDemonstrationError termIdentifiabilityinstrument limitPower lawResidualSamplingTolerance