The rectangle behind the lens

A turning frame can be straightened; a travelling one cannot

Read a frame row by row while the camera turns at a radian a second and every point is 21 px from where a global shutter would put it, at every depth alike. Turn each row's rays back and every point returns to six trillionths of a pixel, with no depth known. Travel at 3 m/s instead, and the best correction that needs no depth is exact at one distance and 21 px wrong at 2 m.

Worth reading first: Every row is a different camera · A scroll is a camera that moves.

Every row is a different camera modelled a rolling shutter as a stack of projections indexed by height, one camera per row, and ended with a claim it made from the structure of the model rather than from experiment. A frame read while the camera only turns can be undone exactly, given the turn, by mapping each row back through a homography that needs no depth. A frame read while the camera travels cannot, because each row is a projection from a different place and moving between places needs to know how far away things are.

Both halves can be measured, and measuring them turns the second into something more useful than a refusal: a law for how the correction everybody actually uses goes wrong, and where.

A turning frame, as read and as undoneFifteen points 3 m away and fifteen 20 m away, in a frame read while the camera turns at 1 rad/s. Rings are where a global shutter draws each point, filled dots where the rolling frame drew it, and crosses where turning each row back puts it: 2.8e-12 px from the ring at worst for the near points and 2.8e-12 px for the far ones.undone: near 3e-12 px · far 3e-12 pxturning 1 rad/s
Fig. 1 Fifteen points 3 m away and fifteen 20 m away, in a frame read over 33.3 ms while the camera turns at 1 rad/s. Rings are where a global shutter draws each point, filled dots where the rolling frame drew it, and crosses where turning each row’s rays back to the first row’s camera puts it — 2.8e-12 px from the ring at worst, for the near points and the far ones alike.

Why a turn needs no depth

A camera that turns about its own centre does not move any ray; it only changes which rays land on which pixels. Straightening does not move the eye measured the consequence for a single picture: re-rendering a photograph as if the camera had been tilted is a homography, and it is exact because the rays themselves are unchanged. Nothing about the scene enters, because nothing about the scene decides which ray goes where.

A rolling frame read during a turn is a stack of such pictures. Row by row, the camera faced a slightly different direction when that row was read. Each row’s pixels are rays from the same eye, and mapping each row back to the orientation of the first row is one homography per row, fixed by the turn the camera had made by that row’s instant.

The first figure does exactly that for a camera turning at a radian a second, a fast pan, with a 33.3 ms readout. As read, each point is up to 20.95 px from where a global shutter would have drawn it. After each row’s rays are turned back, every point is within 2.8×10122.8 \times 10^{-12} px of the global shutter’s position — the near points at 3 m and the far ones at 20 m alike.

At every depth, the same

The depth-independence is the whole claim, so it is worth sweeping.

A frame read while the camera turns at 1 rad/s, undone at every depthA rolling frame read over 33.3 ms while the camera turns at 1 rad/s, with points at 2 to 32 m. As read, every point is up to 20.95 px from where a global shutter would draw it, at every depth alike. Turning each row's rays back to the first row's camera — a homography per row, with no depth in it — puts every point back to 6.2e-12 px, and so does mapping each row through a plane at any depth, because a turn about the centre moves no ray.234681216243210⁻¹²10⁻⁹10⁻⁶0.0011depth of the point (m, log scale)worst distance from the global-shutter picture (px, log scale)as readrows turned backthrough a plane at 8 mundone to 6.2e-12 px at every depthturning 1 rad/s · 33.3 ms readout
Fig. 2 Points from 2 m to 32 m in a frame read while the camera turns at 1 rad/s. As read, every point is up to 20.95 px from where a global shutter would draw it, at every depth alike. Turning each row back puts every point to 6.2e-12 px, and mapping each row through a plane at 8 m does exactly the same.

From 2 m to 32 m, the rolling frame’s error is 20.95 px at every depth — a turn moves every point by the same angle regardless of how far away it is. Turning rows back leaves 6.2×10126.2 \times 10^{-12} px at every depth. So does the other obvious correction, mapping each row through a plane at 8 m: for a pure turn the plane’s depth is irrelevant, because any plane’s homography between two cameras at the same centre is the same rotation homography. The two corrections are the same correction written two ways.

That coincidence is the thing to carry into the travelling case. When the camera only turns, “map through a plane” and “turn the rays back” agree for every plane. When it travels, they stop agreeing, and the plane starts to matter.

A travelling frame

Now the camera moves sideways at 3 m/s and does not turn.

A frame read while the camera travels at 3 m/s, undone only at one depthThe same frame while the camera travels sideways at 3 m/s. As read, points are 27.5 px, 18.4 px, 13.8 px, 9.2 px, 6.9 px, 4.6 px, 3.4 px, 2.3 px, 1.7 px out at 2, 3, 4, 6, 8, 12, 16, 24, 32 m — one over the depth. Turning rows back does nothing useful: 27.5 px, 18.4 px, 13.8 px, 9.2 px, 6.9 px, 4.6 px, 3.4 px, 2.3 px, 1.7 px. Mapping each row through a plane at 8 m is exact for points at 8 m, 1.1e-13 px, and wrong elsewhere: 20.66 px, 11.48 px, 6.89 px, 2.30 px, 0.00 px, 2.30 px, 3.44 px, 4.59 px, 5.17 px.234681216243210⁻¹⁴10⁻¹¹10⁻⁸10⁻⁵0.0110depth of the point (m, log scale)worst distance from the global-shutter picture (px, log scale)as readrows turned backthrough a plane at 8 mexact only at the plane's depthtravelling 3 m/s · 33.3 ms readout
Fig. 3 The same frame while the camera travels sideways at 3 m/s. As read, points are 27.5 px out at 2 m and 1.7 px out at 32 m, one over the depth. Turning rows back changes nothing. Mapping each row through a plane at 8 m is exact for points at 8 m, 1.1e-13 px, and wrong elsewhere: 20.66 px at 2 m and 5.17 px at 32 m.

As read, the error now depends on depth: 27.5 px at 2 m, 13.8 at 4, 6.9 at 8, 3.4 at 16 and 1.7 at 32 — one over the depth, the same law depth is a reciprocal found for a disparity, because a row read later sees every point from a camera shifted sideways, and a sideways shift displaces near points more than far ones.

Turning rows back does nothing. The camera did not turn, so the turn the correction removes is zero, and the error after correction is the error before it, at every depth.

Mapping each row through a plane does something, and exactly one thing. With the plane at 8 m, points at 8 m come back to 1.1×10131.1 \times 10^{-13} px. Points anywhere else do not: 20.66 px at 2 m, 11.48 at 3, 6.89 at 4, 2.30 at 6, 2.30 at 12, 3.44 at 16, 4.59 at 24 and 5.17 at 32. The correction has been bought for one depth and paid for at every other.

The law the plane follows

The residuals obey a short law, and reading it off the numbers is easier than deriving it.

A row read at time tt sees the scene from a camera shifted by vtvt. A point at depth ZZ is displaced by fvt/Zfvt/Z in that row. The plane correction undoes the displacement a point at the plane’s depth ZpZ_p would have had, fvt/Zpfvt/Z_p. What is left is

fvt1Z1Zp,f v t \left|\frac{1}{Z} - \frac{1}{Z_p}\right|,

which is the raw error times 1Z/Zp|1 - Z/Z_p|. At 2 m, 27.5 px times 12/81 - 2/8 is 20.6; at 32 m, 1.72 px times 32/8132/8 - 1 is 5.17. The residual is zero at the plane and grows on both sides of it, faster toward the near side in pixels and without limit toward the far side in proportion.

It is the same law two marks off a known plane find the other eye found for two ordinary photographs: a plane’s homography is exact for the plane and leaves every other point displaced by a parallax proportional to the difference between one over its depth and one over the plane’s. A travelling rolling frame is a pair of pictures within itself, one per row, and the same parallax appears between its first row and every later one.

A travelling frame, as read and as undone through one planeFifteen points 3 m away and fifteen 20 m away, in a frame read while the camera travels at 3 m/s. Rings are where a global shutter draws each point, filled dots where the rolling frame drew it, and crosses where mapping each row through a plane at 8 m puts it: 10.33 px from the ring at worst for the near points and 3.72 px for the far ones.undone: near 10.33 px · far 3.72 pxtravelling 3 m/s, plane at 8 m
Fig. 4 Fifteen points 3 m away and fifteen 20 m away, in a frame read while the camera travels at 3 m/s, undone through a plane at 8 m. The crosses fall 10.33 px from the global shutter’s rings at worst for the near points and 3.72 px for the far ones.

Drawn in the picture, with a near and a far set of points, the correction through an 8 m plane leaves the near points 10.33 px from where they belong and the far ones 3.72 px, pointing in opposite directions: the near ones under-corrected, the far ones over-corrected. A frame with a near object and a distant background, corrected this way, has one of the two displaced whichever plane is chosen.

Choosing the plane

The plane is a free choice, and moving it moves the zero.

Which depth a travelling frame is undone at is a choiceThe travelling frame undone through planes at 4, 8, 16 m. Each is exact at its own depth and wrong away from it, by an amount that grows with the difference between one over the point's depth and one over the plane's. Through the plane at 4 m the far points are 12.05 px out; through the plane at 16 m the near ones are 24.11 px out. No plane is right for a scene with depth in it.23468121624320.0010.010.1110depth of the point (m, log scale)worst distance after undoing through the plane (px, log scale)plane at 4 mplane at 8 mplane at 16 mzero only at the plane's own depthtravelling 3 m/s
Fig. 5 The travelling frame undone through planes at 4, 8 and 16 m. Each is exact at its own depth and wrong away from it. Through the plane at 4 m the far points are 12.05 px out; through the plane at 16 m the near ones are 24.11 px out.

Through a plane at 4 m, the near points come back and the far ones are 12.05 px out. Through a plane at 16 m, the far points come back and the near ones are 24.11 px out. No choice is right for a scene with depth in it; each is a decision about which depth to get right.

The law also says what would make the travelling case correctable after all. The residual vanishes wherever the plane’s depth equals the point’s, so a correction that maps each pixel through a plane at that pixel’s own depth is exact everywhere: given a depth for every point, a travelling frame comes back as completely as a turning one. And the law prices an imperfect depth map in the same terms. A depth wrong by δ(1/Z)\delta(1/Z) in its reciprocal leaves fvtδ(1/Z)fvt\,\delta(1/Z) px, so at 3 m/s a depth map good to a hundredth of a reciprocal metre — 4 cm at 2 m, 1.4 m at 12 — leaves at most about half a pixel, near or far. The requirement is stated in the same reciprocal units depth is a reciprocal found a stereo pair measures in, which is no accident: what the correction needs is the parallax of each point between the first row’s camera and its own.

In practice the plane is usually put at infinity, where the homography is a pure rotation and the correction is the one that suits a turning camera. That is the assumption every row is a different camera described as pretending the uncorrectable case is the correctable one, and the law above says what the pretence costs: a plane at infinity makes 1/Zp1/Z_p zero, so the residual is the whole raw error, fvt/Zfvt/Z, uncorrected at every finite depth. The distant background comes back and anything nearby keeps all of its shear.

How fast is fast

The two cases differ in kind, and they also differ in size, and the sizes decide which one a real camera needs to worry about.

For a turn, the error at a given row is the focal length times the angle turned by that row’s instant. With a 55° field on a 690 px frame the focal length is about 662 px, so a turn of a radian a second over a 33.3 ms readout moves the bottom row by about 22 px against the top — the 20.95 px the figures measure at the points sampled, which sit a little above the bottom row. A slow, deliberate pan of a tenth of a radian a second gives about 2 px; the wobble of a hand-held camera, a few hundredths of a radian a second, gives well under a pixel. The error does not depend on distance at all, so a turning camera shears near and far subjects alike.

For a travel, the error at a row is the focal length times the distance travelled by that row’s instant, divided by the depth. At 3 m/s — a fast walk, or a slow vehicle — a point 2 m away is 27.5 px out and a point 32 m away 1.7 px. At a walking 1.5 m/s the same sums give about 3 px at 10 m and about 13 px at 2 m. So a travelling camera shears nearby things and leaves the distance nearly alone, and for most scenes photographed on foot the travel’s contribution is a few pixels on the nearest subjects.

So which half dominates a real frame depends on how the camera was moving, not on which half is harder to undo. A deliberate pan at a radian a second shears the frame by 20.95 px everywhere, and a 3 m/s travel matches that only for points nearer than about 2.6 m: a pan is a turning problem, and the exact correction is the one it needs. A camera carried at a walk, turning only by a hand’s wobble, is the reverse — well under a pixel from the turn and nearly 14 px from the travel on anything 2 m away. The case that can be undone exactly dominates a pan; the case that cannot dominates a walk past something close.

What a camera can measure about its own motion

The same split appears in what a camera can know about how it moved, and it is not a coincidence.

A turn is an angular rate, and an angular rate is measured directly: the gyroscope in a phone reports it hundreds of times a second, and integrating a rate once gives the angle turned by each row’s instant with an error that grows slowly over a readout of a few tens of milliseconds. So the correction that needs the turn and nothing else can be given the turn, and a frame read during a pan can be straightened from the camera’s own sensors with no reference to what it photographed.

A travel is a velocity, and an inertial sensor measures acceleration; velocity is its integral from some known starting value, and position the integral of that. Over a readout the travel is small and its estimate poor. And even a perfect estimate of the travel is not enough, because the correction needs depth as well, which no sensor on the camera reports for every pixel.

So the geometry and the hardware agree. What can be undone without depth is also what can be measured without looking; what needs depth also needs a motion estimate the camera cannot easily make. A correction built from a gyroscope alone is a correction for the turning part of the motion, and whatever travelling part was present stays in the frame as the depth-dependent residual the figures measured.

The shear the correction leaves

It is worth being clear about what the residual of a travelling correction looks like, because it is not a blur and not a random scatter.

After a plane correction, every point at a given depth is displaced by the same fraction of its row’s travel, in the direction of the travel. That is a shear again — a row-dependent sideways shift — but one whose size depends on each object’s depth. A vertical pole at the plane’s depth stands straight; a nearer pole leans one way and a further pole the other. Every row is a different camera found that in an uncorrected frame every vertical at one depth leans by the same angle; after correcting through a plane, verticals at the plane’s depth lean by nothing and the rest lean in proportion to how far their depth’s reciprocal is from the plane’s.

That is also the reason a frame is an interval could not describe a travelling frame’s blur with one kernel: a travelling camera assigns every depth its own motion, and any single operation applied to the whole frame is right for one depth only.

Why the two cases part

The distinction is the one the scroll field drew between a panorama and a scroll, and it is exact here for the same reason.

A scroll is not a panorama separated the two by whether the eye moves. A panorama is made by an eye that turns in place, keeps a single centre and can be re-rendered onto any surface without knowing the scene. A scroll is made by an eye that travels, has no centre at all, and cannot be re-rendered from one viewpoint without depth. A rolling frame read during a turn is a very short panorama drawn one row at a time; read during a travel, it is a very short scroll.

So whether a frame can be straightened is decided before any correction is attempted, by whether the camera’s centre moved during the readout. The eye is a place, not a point measured the same question for a panorama head: turn about the entrance pupil and the frames stitch at every depth; turn about a point behind it and the eye travels a few millimetres, and near objects break the stitch. A hand-held camera turning about the photographer’s wrist is doing both at once, and its correction is exact only for the part that is a turn.

What this does not settle

The motion is known. Every correction here is given the turn or the travel exactly. A real correction estimates it — from a gyroscope, or from the frame itself — and an error in the estimated turn leaves a residual that does not depend on depth, which is distinguishable from the travelling residual by exactly that. How precisely the motion must be known for the turning correction to stay below a pixel was not measured.

Pure cases. The camera either turns or travels. A real camera does both, and the correction is exact for the turning part and follows the law above for the travelling part, but how the two combine in a single frame’s residual was not drawn.

The row’s time is taken from the row. Each point’s time is its row over the frame height times the readout, which is the model’s assumption about a sensor that reads its rows at a constant rate from the top.

Still open: whether a frame’s own shear measures its travel

The travelling residual has a depth in it, and so does the raw shear before any correction. Every row is a different camera found that a vertical line leans by an amount set by the image speed at its depth. Two vertical lines at different depths therefore lean by different amounts in a travelling frame and by the same amount in a turning one.

The question that leaves is whether that difference can be read back: given two verticals whose leans differ, whether the frame itself separates the turning part of the camera’s motion from the travelling part, and the travelling part from the depths of the two lines — and how far apart in depth two verticals must be before, at a readout of 33 ms and marks read to a pixel, the difference in their leans says anything at all.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

centre of projectionHomographyMoving viewpointParallaxPushbroomResidual