A tilted sensor is not a distortion
Worth reading first: Straight lines that are not · The principal point is not the centre.
A lens destroys the invariant put a number on what separates a lens from a pinhole: four collinear points imaged through a barrel lens come back with their cross-ratio more than a per cent out, where a pinhole returns it to the last bit of arithmetic. The cross-ratio is the test, because it is the one quantity a projection keeps and nothing else does.
That essay was about a lens whose glass departs from a pinhole. A camera can also depart from a square pinhole in a way that has nothing to do with its glass: the sensor may not be perpendicular to the lens’s axis. A sensor mounted a few degrees out of square is a manufacturing tolerance, a deliberate design in a tilt lens, or the ordinary state of a scanned film plane that did not lie flat. Its picture is visibly different from the square camera’s. The question this essay puts to it is the one the invariant exists to answer: is that difference a distortion?
A plane cut through a pencil of rays
The argument is short, and it settles the matter before any measurement does.
A pinhole camera is a centre and a plane. Every ray from the scene passes through the centre, and the picture is where those rays meet the plane. Nothing in that description requires the plane to be perpendicular to anything. A plane at any angle, cut through the same bundle of rays, records a picture — a different picture, because the rays meet it in different places, but a picture made by exactly the same kind of operation.
And a picture made by a centre and a plane is a pinhole picture, full stop. The centre is the same, so every straight world line still projects to the line where its plane through the centre meets the sensor. Four collinear points still project to four collinear points with the same cross-ratio. What changes is the camera’s description: its principal point is the foot of the perpendicular from the centre to the sensor, and its focal length is the length of that perpendicular. Tilt the sensor by τ about an axis through the point where the lens’s axis meets it, and the perpendicular’s foot moves by while its length becomes .
The two pictures — square sensor and tilted sensor — are therefore two planes cut through one pencil of rays, and two such pictures are related by a homography. A projection of a projection measured the same fact from the outside: photograph a photograph and the composite map is a homography, four marks determine it, and every other mark lands where they say to px. Here the photograph and the second camera share a centre, so the homography is of the special kind a turn of the camera produces.
So a tilted sensor is a pinhole camera turned by τ, with its principal point moved and its focal length shortened. That is a claim a direct trace can check without assuming it.
The picture, computed ray by ray
The figure computes the tilted picture with no appeal to pinholes at all: for each direction it builds the ray, builds the tilted plane, intersects them in three dimensions, and reads the intersection’s coordinates along the plane’s own axes. Every point of the picture moves — up to 7.5 px across the grid drawn, with the pattern of arrows sweeping downward and fanning at the corners — and three straight world lines drawn on the tilted sensor are straight to px.
Then it compares that picture with the turned pinhole’s. They agree to px over every direction in the frame. The principal point has moved 22.30 px, which is for the frame’s 426 px focal length, and the focal length is 0.584 px shorter, which is . The cross-ratio of four collinear points, computed before and after, agrees to .
Every point moves and no invariant breaks. That is the precise opposite of the barrel lens, which moved points by comparable amounts and broke the cross-ratio by more than a per cent. The size of the displacement says nothing about whether a picture is still a projection; the invariant does.
At 6° the displacement doubles, to 15.5 px, the principal point moves 44.53 px, and the focal length loses 2.334 px — and the lines are straight to px and the cross-ratio survives to . A larger tilt is a larger change of camera, not a larger distortion.
What four calibration models make of it
A calibration handed this picture does not know a sensor was tilted. It fits a model, and which model it fits decides what it concludes. The figure fits four, each to the tilted picture of a 24 × 14 grid of directions.
No fit. The tilted picture sits 3.65 px, root-mean-square, from the square camera’s picture of the same directions.
The principal point held at the centre of the frame, the camera’s turn and focal length free. This is what a calibration does when it judges the principal point too weakly determined to fit — a common and often sensible judgement. It leaves 2.99 px. A turned camera with its principal point in the wrong place cannot reproduce a picture whose principal point is 22 px down.
Held, with tangential distortion terms free. The Brown–Conrady model’s tangential coefficients and exist to describe distortion that is not symmetric about the principal point, and a tilted sensor’s displacement pattern is not. The fit uses them: comes back at −0.01652, and the misfit falls to 1.87 px.
Held, with tangential and radial terms free. Adding and changes the misfit in the fourth decimal place. The radial terms have nothing to describe, because a radial map moves every point along its own radius and nothing in a tilt’s displacement is symmetric about a centre.
The principal point free, no distortion terms at all. The misfit is px — the arithmetic floor — and the fitted principal point sits 22.30 px down the frame, exactly where the geometry puts the foot of the perpendicular.
The ranking is the argument. A model with four distortion coefficients reaches 1.87 px. A model with two extra numbers saying where the principal point is reaches zero. The tilt was never a distortion, and the parameters that describe distortion can only approximate a change of camera they were not built to express.
At 1° every misfit is about a third of its 3° value: 1.21 px with no fit, 0.99 px held, 0.62 px with tangential terms, and the free principal point still at the floor. That scaling matters in practice. A sub-degree tilt leaves a held-centre tangential fit well under a pixel, inside what a calibration would accept as noise, and the calibration would report a camera with a small tangential distortion and no evidence that anything else was wrong.
Why a calibration holds the principal point at all
The sequence of fits makes freeing the principal point look free, and it is not, so the reason calibrations hold it deserves a fair hearing.
The principal point is two of the four numbers that a projection matrix carries in its sides, and of the four it is the one most calibration data pins least well. A focal length is read from how fast the picture’s scale changes across the frame; a principal point is read from where that change is centred, and a small error in the centre is nearly compensated by a small turn of the camera. On a noisy calibration from few views, a freed principal point can wander by many pixels, and holding it at the middle of the frame trades that variance for a bias that is usually smaller. That is a sound trade when the bias really is small.
What the sequence of fits shows is that the trade has a hidden premise: that the principal point really is near the middle. The proportion is the assumption found how far that premise reaches — sweeping only where the picture’s centre is assumed to be takes one drawn quadrilateral from one part in fourteen to slightly wider than square, every answer a genuine rectangle. And one conic calibrates the camera is the reason the premise cannot be checked from inside a held fit: the conic that is a calibrated camera has the principal point as its centre, so holding the centre is holding part of the camera, not a nuisance beside it.
So holding the principal point is a claim about the camera, and a tilted sensor is one of the ordinary ways it becomes false. The honest procedure is to fit both ways and compare. If freeing it collapses the residual, as it does here from 1.87 px to the floor, the held fit was wrong, not merely less precise.
Used as a correction, the tangential model bends straight lines
A calibration’s coefficients are not the end product. They are used to correct pictures, and the correction is where the wrong model does its damage.
The figure takes the tilted picture — in which every straight line is straight — and applies the inverse of the fitted tangential model, about the held principal point, with the fitted focal length. Every row goes in straight. Every row comes out bent: 3.61 px at the top of the frame, rising to 4.12 px at the bottom. The columns barely move, under 0.07 px. And the cross-ratio of four collinear points, which the tilted picture preserved to , is out after the correction.
So the correction does not merely fail to undo the tilt. It takes a picture that satisfied every projective test there is and produces one that fails them — curved lines, a broken invariant — by applying a polynomial correction for a distortion that was never there. A plumb-line check run on the corrected picture would report it as worse calibrated than the uncorrected one, which is the one test that would catch the error, and it is a test people run on the raw picture rather than on the correction.
The direction of the bend follows from the displacement pattern. The tilt’s displacement is mostly vertical and grows across the frame’s height; a tangential polynomial has to reproduce a vertical displacement with terms in , and , and a term in bends a horizontal line into a parabola. The columns, which the tilt displaces nearly uniformly along their length, are left almost alone.
A self-consistent answer with a parameter missing
The held-centre tangential fit has one more property that makes it dangerous, and it has been met before in a different camera.
The pixel that is not square handed the round trip of recovering a camera and redrawing its picture a photograph from a camera with two focal lengths, and it came back with a focal length 49.6% out and three independent estimates of it agreeing with each other to . Agreement among estimates is not evidence of correctness when every estimate shares the same missing parameter; it is evidence that the model is internally consistent, which a wrong model can be perfectly.
The tangential fit is the same shape. Its coefficients are stable, its residual is small at small tilts, and the corrected picture is smooth. The only sign of the missing parameter is the residual itself — 0.62 px at 1°, which a calibration would pass — and the correction’s failure on a plumb-line test, which is rarely run on a correction. A picture’s own consistency cannot say which parameter a model lacks; only a model that has the parameter can.
Where this sits beside the shift
A displaced principal point has been met before, from other directions. The principal point is not the centre measured a shifted lens — the sensor slid across the lens’s axis rather than tilted — and found that assuming the principal point is the middle of the frame costs one and a half per cent of a recovered focal length at a fifth of a frame’s shift. The centre of the picture is not the centre of the paper found a crop doing the same by moving the rectangle rather than the lens. Straightening does not move the eye found that correcting a photograph’s converging verticals is exactly equivalent to a level camera with a shifted principal point.
A tilted sensor completes that family. A shift or a crop moves the principal point and leaves the focal length alone. A tilt moves the principal point by and shortens the focal length by , which at small angles is a shift with a second-order change of scale. A rotation of the whole camera, sensor and lens together, moves nothing about the intrinsics at all. All of them are changes of pinhole camera. None is a distortion, and all are absorbed exactly by a model that frees the principal point and fits the camera’s orientation.
That is the practical rule, stated once. Before adding distortion terms to explain an asymmetric pattern, free the principal point. If the pattern disappears, it was geometry. If it remains, it may be distortion — and the residual after freeing the principal point is the honest measure of how much.
What the tangential terms are for
It would be a misreading to conclude that tangential coefficients are a mistake. They exist because real lenses have elements that are decentred or tilted relative to each other, and a lens whose glass is misaligned does produce a distortion that is not symmetric about any point — one that a free principal point does not absorb and that does break the cross-ratio. Nothing here measures that case, and nothing here says it is small.
What the measurement establishes is narrower. A sensor tilt and an element tilt produce displacement patterns that a held-centre calibration cannot tell apart, and the model built for the second is used to describe the first. The two leave different signatures only in a test that frees the principal point. A calibration that reports tangential coefficients with its principal point held has not established that the lens has decentring distortion; it has established that something asymmetric is present, and the cheapest candidate is that the principal point is not where it was held.
A further limit on the argument: the tilt here is about the horizontal axis through the point where the lens’s axis meets the sensor. A tilt about an axis elsewhere adds a displacement along the axis as well, which is a change of focus as much as of geometry, and a real sensor is also not perfectly flat. The deliberate control on the comparison is exactly that: a sensor curved by a fiftieth of a focal length is not a pinhole, and the comparison with the turned pinhole reports it by more than half a pixel.
Still open: whether a model that inverts the other way reaches further
This essay and the one before it are about the same polynomial. A barrel model folds at a radius it sets itself found that stops being invertible at a radius its own coefficients fix; this one found that its tangential companions describe a change of camera badly. Many calibration tools use a different radial form altogether: the division model, which writes the undistorted radius as the distorted one divided by . It is invertible in closed form, which is why it is used, and it has its own limit in the opposite direction. The measurement still to be made finds where the division model stops being a lens, fits it to the same four fisheye laws the polynomial could not follow past 65°, and asks whether a model chosen for easy inversion buys a field the polynomial cannot reach or merely moves the fold to the other side of the equation.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Recovering the camera from the picture it drew — both name camera calibration, homography, picture plane, principal point
- The lens a pavement can hide — both name brown–conrady, cross-ratio, principal point, radial distortion
- The ceiling that is not a plane — both name cross-ratio, homography, projective invariant
- The plane is a choice — both name homography, picture plane, principal point
- The render is distorted on purpose — both name brown–conrady, camera calibration, radial distortion
- A floor anamorph is three numbers — both name homography, picture plane
Named objects
A flat tag is an object no other essay names yet.
Brown–ConradyCamera calibrationCross-ratioHomographyPicture planePrincipal pointProjective invariantRadial distortion