The real instrument

A tilted sensor is not a distortion

Tilt a sensor 3° out of square with its lens and every point of the picture moves — up to 7.5 px on the frame drawn here — yet every straight line stays straight to 10⁻¹³ px and the cross-ratio survives to 10⁻¹⁶. The picture is an ordinary pinhole picture whose principal point has moved 22.30 px. A calibration that frees its principal point absorbs it exactly; one that holds the principal point and reaches for tangential distortion terms leaves 1.87 px, and used as a correction it bends straight rows by 4 px.

Worth reading first: Straight lines that are not · The principal point is not the centre.

A lens destroys the invariant put a number on what separates a lens from a pinhole: four collinear points imaged through a barrel lens come back with their cross-ratio more than a per cent out, where a pinhole returns it to the last bit of arithmetic. The cross-ratio is the test, because it is the one quantity a projection keeps and nothing else does.

That essay was about a lens whose glass departs from a pinhole. A camera can also depart from a square pinhole in a way that has nothing to do with its glass: the sensor may not be perpendicular to the lens’s axis. A sensor mounted a few degrees out of square is a manufacturing tolerance, a deliberate design in a tilt lens, or the ordinary state of a scanned film plane that did not lie flat. Its picture is visibly different from the square camera’s. The question this essay puts to it is the one the invariant exists to answer: is that difference a distortion?

The invariant, against the coefficient that destroys itFour collinear points, imaged through a pinhole, return the world's cross-ratio to 2e-16. Through a lens with k₁ = -0.32 they return it 1.29% out. The curve is zero at k₁ = 0 and at no other coefficient, because a radial polynomial is not a projective map.00.50011.502-0.400-0.20000.200k₁departure of the cross-ratio from the world's value (%)pinhole1.29%the pinhole's own error, on the same four points2e-16 — the control
Fig. 1 The test a distortion fails, from the essay that introduced it: the cross-ratio of four collinear points against the barrel coefficient that bends them, zero only where the coefficient is. A tilted sensor is about to be put through the same test.

A plane cut through a pencil of rays

The argument is short, and it settles the matter before any measurement does.

A pinhole camera is a centre and a plane. Every ray from the scene passes through the centre, and the picture is where those rays meet the plane. Nothing in that description requires the plane to be perpendicular to anything. A plane at any angle, cut through the same bundle of rays, records a picture — a different picture, because the rays meet it in different places, but a picture made by exactly the same kind of operation.

And a picture made by a centre and a plane is a pinhole picture, full stop. The centre is the same, so every straight world line still projects to the line where its plane through the centre meets the sensor. Four collinear points still project to four collinear points with the same cross-ratio. What changes is the camera’s description: its principal point is the foot of the perpendicular from the centre to the sensor, and its focal length is the length of that perpendicular. Tilt the sensor by τ about an axis through the point where the lens’s axis meets it, and the perpendicular’s foot moves by fsinτf\sin\tau while its length becomes fcosτf\cos\tau.

The two pictures — square sensor and tilted sensor — are therefore two planes cut through one pencil of rays, and two such pictures are related by a homography. A projection of a projection measured the same fact from the outside: photograph a photograph and the composite map is a homography, four marks determine it, and every other mark lands where they say to 101310^{-13} px. Here the photograph and the second camera share a centre, so the homography is of the special kind a turn of the camera produces.

So a tilted sensor is a pinhole camera turned by τ, with its principal point moved and its focal length shortened. That is a claim a direct trace can check without assuming it.

The picture, computed ray by ray

A sensor tilted 3° moves every point up to 7.5 px and bends nothingThe wide frame's picture with its sensor tilted 3° about the horizontal, computed ray by ray in three dimensions. Each arrow runs from where a direction lands on the untilted sensor toward where it lands on the tilted one, drawn 4 times its length; the largest move on this grid is 7.5 px. Three straight world lines are drawn on the tilted sensor and are straight to 1e-13 px. The picture is the pinhole picture of a camera turned 3°, with its principal point moved 22.30 px and its focal length shortened by 0.584 px, to 1e-13 px; the cross-ratio of four collinear points survives to 3e-16.principal point moved 22.30 px · focal length 0.584 px shorterlines straight to 1e-13 px
Fig. 2 The wide frame’s sensor tilted 3°, computed ray by ray in three dimensions. Each arrow runs from where a direction lands on the square sensor toward where it lands on the tilted one, drawn four times its length; the largest move on this grid is 7.5 px. Three straight world lines are drawn on the tilted sensor and are straight to 1e-13 px.

The figure computes the tilted picture with no appeal to pinholes at all: for each direction it builds the ray, builds the tilted plane, intersects them in three dimensions, and reads the intersection’s coordinates along the plane’s own axes. Every point of the picture moves — up to 7.5 px across the grid drawn, with the pattern of arrows sweeping downward and fanning at the corners — and three straight world lines drawn on the tilted sensor are straight to 101310^{-13} px.

Then it compares that picture with the turned pinhole’s. They agree to 101310^{-13} px over every direction in the frame. The principal point has moved 22.30 px, which is fsin3°f\sin 3° for the frame’s 426 px focal length, and the focal length is 0.584 px shorter, which is f(1cos3°)f(1 - \cos 3°). The cross-ratio of four collinear points, computed before and after, agrees to 3×10163 \times 10^{-16}.

Every point moves and no invariant breaks. That is the precise opposite of the barrel lens, which moved points by comparable amounts and broke the cross-ratio by more than a per cent. The size of the displacement says nothing about whether a picture is still a projection; the invariant does.

A sensor tilted 6° moves every point up to 15.5 px and bends nothingThe wide frame's picture with its sensor tilted 6° about the horizontal, computed ray by ray in three dimensions. Each arrow runs from where a direction lands on the untilted sensor toward where it lands on the tilted one, drawn 4 times its length; the largest move on this grid is 15.5 px. Three straight world lines are drawn on the tilted sensor and are straight to 2e-13 px. The picture is the pinhole picture of a camera turned 6°, with its principal point moved 44.53 px and its focal length shortened by 2.334 px, to 1e-13 px; the cross-ratio of four collinear points survives to 2e-16.principal point moved 44.53 px · focal length 2.334 px shorterlines straight to 2e-13 px
Fig. 3 The same sensor tilted 6°. The displacement doubles, to 15.5 px on this grid, and the principal point moves 44.53 px; the lines are still straight to 2e-13 px and the cross-ratio still survives, because a larger tilt is the same operation.

At 6° the displacement doubles, to 15.5 px, the principal point moves 44.53 px, and the focal length loses 2.334 px — and the lines are straight to 2×10132 \times 10^{-13} px and the cross-ratio survives to 2×10162 \times 10^{-16}. A larger tilt is a larger change of camera, not a larger distortion.

What four calibration models make of it

A calibration handed this picture does not know a sensor was tilted. It fits a model, and which model it fits decides what it concludes. The figure fits four, each to the tilted picture of a 24 × 14 grid of directions.

A 3° sensor tilt, left behind by four calibration modelsThe root-mean-square misfit, over a 24 × 14 grid of directions, of four calibration models fitted to a sensor tilted 3°. With no fit the picture is 3.65 px from the untilted camera's. Holding the principal point at the centre and fitting the turn and the focal length leaves 2.99 px. Freeing tangential p₁ and p₂ lowers that to 1.87 px, with p₁ = -0.01652; freeing radial k₁ and k₂ as well leaves 1.87 px. Freeing the principal point instead, with no distortion terms at all, leaves 5e-14 px.untilted camera, no fit3.65 pxprincipal point held at the centre2.99 pxheld, with tangential p₁ and p₂1.87 pxheld, tangential and radial k₁, k₂1.87 pxprincipal point free5e-14 pxtilt 3° · root-mean-square over 336 directionsprincipal point free: 5e-14 px
Fig. 4 Four calibration models fitted to a sensor tilted 3°, by the root-mean-square misfit each leaves over 336 directions. With no fit the picture is 3.65 px from the square camera’s. Holding the principal point at the centre leaves 2.99 px; adding tangential terms leaves 1.87 px, and radial terms on top still leave 1.87 px. Freeing the principal point, with no distortion terms at all, leaves 5e-14 px.

No fit. The tilted picture sits 3.65 px, root-mean-square, from the square camera’s picture of the same directions.

The principal point held at the centre of the frame, the camera’s turn and focal length free. This is what a calibration does when it judges the principal point too weakly determined to fit — a common and often sensible judgement. It leaves 2.99 px. A turned camera with its principal point in the wrong place cannot reproduce a picture whose principal point is 22 px down.

Held, with tangential distortion terms free. The Brown–Conrady model’s tangential coefficients p1p_{1} and p2p_{2} exist to describe distortion that is not symmetric about the principal point, and a tilted sensor’s displacement pattern is not. The fit uses them: p1p_{1} comes back at −0.01652, and the misfit falls to 1.87 px.

Held, with tangential and radial terms free. Adding k1k_{1} and k2k_{2} changes the misfit in the fourth decimal place. The radial terms have nothing to describe, because a radial map moves every point along its own radius and nothing in a tilt’s displacement is symmetric about a centre.

The principal point free, no distortion terms at all. The misfit is 5×10145 \times 10^{-14} px — the arithmetic floor — and the fitted principal point sits 22.30 px down the frame, exactly where the geometry puts the foot of the perpendicular.

The ranking is the argument. A model with four distortion coefficients reaches 1.87 px. A model with two extra numbers saying where the principal point is reaches zero. The tilt was never a distortion, and the parameters that describe distortion can only approximate a change of camera they were not built to express.

A 1° sensor tilt, left behind by four calibration modelsThe root-mean-square misfit, over a 24 × 14 grid of directions, of four calibration models fitted to a sensor tilted 1°. With no fit the picture is 1.21 px from the untilted camera's. Holding the principal point at the centre and fitting the turn and the focal length leaves 0.99 px. Freeing tangential p₁ and p₂ lowers that to 0.62 px, with p₁ = -0.00552; freeing radial k₁ and k₂ as well leaves 0.62 px. Freeing the principal point instead, with no distortion terms at all, leaves 4e-14 px.untilted camera, no fit1.21 pxprincipal point held at the centre0.99 pxheld, with tangential p₁ and p₂0.62 pxheld, tangential and radial k₁, k₂0.62 pxprincipal point free4e-14 pxtilt 1° · root-mean-square over 336 directionsprincipal point free: 4e-14 px
Fig. 5 The same four models on a sensor tilted only 1°. Every misfit shrinks by about a third — 1.21 px with no fit, 0.99 px held, 0.62 px with tangential terms — and the free principal point still reaches the floor at 4e-14 px.

At 1° every misfit is about a third of its 3° value: 1.21 px with no fit, 0.99 px held, 0.62 px with tangential terms, and the free principal point still at the floor. That scaling matters in practice. A sub-degree tilt leaves a held-centre tangential fit well under a pixel, inside what a calibration would accept as noise, and the calibration would report a camera with a small tangential distortion and no evidence that anything else was wrong.

Why a calibration holds the principal point at all

The sequence of fits makes freeing the principal point look free, and it is not, so the reason calibrations hold it deserves a fair hearing.

The principal point is two of the four numbers that a projection matrix carries in its sides, and of the four it is the one most calibration data pins least well. A focal length is read from how fast the picture’s scale changes across the frame; a principal point is read from where that change is centred, and a small error in the centre is nearly compensated by a small turn of the camera. On a noisy calibration from few views, a freed principal point can wander by many pixels, and holding it at the middle of the frame trades that variance for a bias that is usually smaller. That is a sound trade when the bias really is small.

What the sequence of fits shows is that the trade has a hidden premise: that the principal point really is near the middle. The proportion is the assumption found how far that premise reaches — sweeping only where the picture’s centre is assumed to be takes one drawn quadrilateral from one part in fourteen to slightly wider than square, every answer a genuine rectangle. And one conic calibrates the camera is the reason the premise cannot be checked from inside a held fit: the conic that is a calibrated camera has the principal point as its centre, so holding the centre is holding part of the camera, not a nuisance beside it.

So holding the principal point is a claim about the camera, and a tilted sensor is one of the ordinary ways it becomes false. The honest procedure is to fit both ways and compare. If freeing it collapses the residual, as it does here from 1.87 px to the floor, the held fit was wrong, not merely less precise.

Used as a correction, the tangential model bends straight lines

A calibration’s coefficients are not the end product. They are used to correct pictures, and the correction is where the wrong model does its damage.

Corrected with tangential terms, a tilted picture's straight rows come out 4.1 px bentThe straight rows and columns of a picture from a sensor tilted 3°, passed through the correction a calibration with its principal point held at the centre would apply — the inverse of its fitted tangential model. Every row went in straight; each comes out bent, by 3.61, 3.73, 3.85, 3.98, 4.12 px from top to bottom, drawn here 10 times its bend. The columns bend by at most 0.062 px. The cross-ratio of four collinear points, exact on the tilted picture, is 2.4e-4 out after the correction.rows bent up to 4.12 px, drawn ×10cross-ratio 2.4e-4 out
Fig. 6 The straight rows and columns of the 3° tilted picture, passed through the correction a held-centre tangential calibration would apply, drawn ten times their bend. Every row went in straight and comes out bent, by 3.61 px at the top to 4.12 px at the bottom; the columns bend by under 0.07 px. The cross-ratio, exact on the tilted picture, is 2.4e-4 out afterwards.

The figure takes the tilted picture — in which every straight line is straight — and applies the inverse of the fitted tangential model, about the held principal point, with the fitted focal length. Every row goes in straight. Every row comes out bent: 3.61 px at the top of the frame, rising to 4.12 px at the bottom. The columns barely move, under 0.07 px. And the cross-ratio of four collinear points, which the tilted picture preserved to 101610^{-16}, is 2.4×1042.4 \times 10^{-4} out after the correction.

So the correction does not merely fail to undo the tilt. It takes a picture that satisfied every projective test there is and produces one that fails them — curved lines, a broken invariant — by applying a polynomial correction for a distortion that was never there. A plumb-line check run on the corrected picture would report it as worse calibrated than the uncorrected one, which is the one test that would catch the error, and it is a test people run on the raw picture rather than on the correction.

The direction of the bend follows from the displacement pattern. The tilt’s displacement is mostly vertical and grows across the frame’s height; a tangential polynomial has to reproduce a vertical displacement with terms in x2x^{2}, y2y^{2} and xyxy, and a term in x2x^{2} bends a horizontal line into a parabola. The columns, which the tilt displaces nearly uniformly along their length, are left almost alone.

A self-consistent answer with a parameter missing

The held-centre tangential fit has one more property that makes it dangerous, and it has been met before in a different camera.

The pixel that is not square handed the round trip of recovering a camera and redrawing its picture a photograph from a camera with two focal lengths, and it came back with a focal length 49.6% out and three independent estimates of it agreeing with each other to 101610^{-16}. Agreement among estimates is not evidence of correctness when every estimate shares the same missing parameter; it is evidence that the model is internally consistent, which a wrong model can be perfectly.

The tangential fit is the same shape. Its coefficients are stable, its residual is small at small tilts, and the corrected picture is smooth. The only sign of the missing parameter is the residual itself — 0.62 px at 1°, which a calibration would pass — and the correction’s failure on a plumb-line test, which is rarely run on a correction. A picture’s own consistency cannot say which parameter a model lacks; only a model that has the parameter can.

Where this sits beside the shift

A displaced principal point has been met before, from other directions. The principal point is not the centre measured a shifted lens — the sensor slid across the lens’s axis rather than tilted — and found that assuming the principal point is the middle of the frame costs one and a half per cent of a recovered focal length at a fifth of a frame’s shift. The centre of the picture is not the centre of the paper found a crop doing the same by moving the rectangle rather than the lens. Straightening does not move the eye found that correcting a photograph’s converging verticals is exactly equivalent to a level camera with a shifted principal point.

A tilted sensor completes that family. A shift or a crop moves the principal point and leaves the focal length alone. A tilt moves the principal point by fsinτf\sin\tau and shortens the focal length by f(1cosτ)f(1 - \cos\tau), which at small angles is a shift with a second-order change of scale. A rotation of the whole camera, sensor and lens together, moves nothing about the intrinsics at all. All of them are changes of pinhole camera. None is a distortion, and all are absorbed exactly by a model that frees the principal point and fits the camera’s orientation.

That is the practical rule, stated once. Before adding distortion terms to explain an asymmetric pattern, free the principal point. If the pattern disappears, it was geometry. If it remains, it may be distortion — and the residual after freeing the principal point is the honest measure of how much.

What the tangential terms are for

It would be a misreading to conclude that tangential coefficients are a mistake. They exist because real lenses have elements that are decentred or tilted relative to each other, and a lens whose glass is misaligned does produce a distortion that is not symmetric about any point — one that a free principal point does not absorb and that does break the cross-ratio. Nothing here measures that case, and nothing here says it is small.

What the measurement establishes is narrower. A sensor tilt and an element tilt produce displacement patterns that a held-centre calibration cannot tell apart, and the model built for the second is used to describe the first. The two leave different signatures only in a test that frees the principal point. A calibration that reports tangential coefficients with its principal point held has not established that the lens has decentring distortion; it has established that something asymmetric is present, and the cheapest candidate is that the principal point is not where it was held.

A further limit on the argument: the tilt here is about the horizontal axis through the point where the lens’s axis meets the sensor. A tilt about an axis elsewhere adds a displacement along the axis as well, which is a change of focus as much as of geometry, and a real sensor is also not perfectly flat. The deliberate control on the comparison is exactly that: a sensor curved by a fiftieth of a focal length is not a pinhole, and the comparison with the turned pinhole reports it by more than half a pixel.

Still open: whether a model that inverts the other way reaches further

This essay and the one before it are about the same polynomial. A barrel model folds at a radius it sets itself found that r(1+k1r2+k2r4)r(1 + k_{1}r^{2} + k_{2}r^{4}) stops being invertible at a radius its own coefficients fix; this one found that its tangential companions describe a change of camera badly. Many calibration tools use a different radial form altogether: the division model, which writes the undistorted radius as the distorted one divided by 1+λrd21 + \lambda r_{d}^{2}. It is invertible in closed form, which is why it is used, and it has its own limit in the opposite direction. The measurement still to be made finds where the division model stops being a lens, fits it to the same four fisheye laws the polynomial could not follow past 65°, and asks whether a model chosen for easy inversion buys a field the polynomial cannot reach or merely moves the fold to the other side of the equation.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Brown–ConradyCamera calibrationCross-ratioHomographyPicture planePrincipal pointProjective invariantRadial distortion