The rectangle behind the lens

A frame's shear knows travel only over depth

Read a frame row by row while the camera turns and travels, and every vertical post leans — the near ones more. The lean is the turn plus the travel over the post's depth, and that sum is all the frame holds: twice the travel past posts twice as far draws the same frame to eighteen decimal places. Two posts cannot separate turn from travel. A facade can, because a turn leans the edges of the frame more than its middle, but the two signals are 99.8 per cent alike, and reading them apart takes a pixel on every row.

Worth reading first: Every row is a different camera · A scroll is a camera that moves.

Every row is a different camera described a rolling-shutter frame as what it is: a stack of pictures, one per row, each taken from wherever the camera was when that row was read. A vertical line in the world comes out leaning, because its lower rows were read later, after the picture had moved on. A turning frame can be straightened; a travelling one cannot then found the asymmetry between the two ways a camera moves. A turn moves every point of the picture by an amount that does not depend on its depth, so each row can be turned back without knowing anything about the scene. A travel moves near points more than far ones, so no correction that ignores depth can undo it.

The essay ended on the obvious next question. If a travel leans near verticals more than far ones and a turn leans them alike, then two verticals at different depths lean differently in a travelling frame and identically in a turning one. Can the frame read that back — separate the turning part of the camera’s motion from the travelling part, and the travelling part from the depths of the lines? And how far apart in depth must two verticals be, at a readout of 33 ms with marks read to a pixel, before the difference in their leans says anything at all?

The frame throughout is 690 by 400 pixels, read top to bottom in 33.3 ms, through a 55° lens — a focal length of 663 px. Every lean is measured off the frame by solving, for each end of a post, for the row that drew it at the instant it was read.

A lean is a turn plus a travel over a depth

The hero figure takes a camera turning at 0.4 radians a second while travelling sideways at two metres a second, and five posts at depths from three metres to effectively infinity, each spanning two thirds of the frame’s height.

Read row by row while the camera turns at 0.4 rad/s and travels at 2 m/s, a near post leans 3.6° and a far one 1.5°A frame 690 by 400 px read top to bottom over 33.3 ms by a camera with a 55° lens that turns at 0.4 rad/s and travels sideways at 2 m/s. Five vertical posts at 3, 6, 12, 24, 400 m, each spanning two thirds of the frame, are drawn where the rolling frame puts them; the dashed lines are where a global shutter would. The posts lean by 3.58°, 2.37°, 1.79°, 1.58°, 1.47°: every lower row was read later, after the picture had moved on, and the picture moves faster for a near post, because a travel moves near things more. The far post leans by the turn alone. Each lean is (T/H)·(f·ω·(1 + x²/f²) + f·v/Z) to first order in the motion; measured off the frame they agree with it to 0.8 per cent.3 m3.6°6 m2.4°12 m1.8°24 m1.6°far1.5°no single viewpoint — the rays miss by the camera's travel during the readoutturning 0.4 rad/s, travelling 2 m/s
Fig. 1 A frame read over 33.3 ms while the camera turns at 0.4 rad/s and travels at 2 m/s. Five posts at 3, 6, 12, 24 and 400 m lean by 3.58°, 2.37°, 1.79°, 1.58° and 1.47°; the dashed lines are where a global shutter would draw them. The far post leans by the turn alone.

The near post leans three and a half degrees, the far one a degree and a half, and the ones between fall in order. The far post is leaning by the turn alone: a camera that travels two metres a second moves a post four hundred metres away by almost nothing in 33 ms, but a camera that turns moves every direction alike.

Each lean follows one law to first order in the motion. A post’s lower row is read later than its upper by the number of rows between them times the time per row, and in that time the picture moves sideways at the image’s speed at that place. The image’s speed has two parts: a turn at ω\omega moves the column xx at fω(1+x2/f2)f\omega(1 + x^2/f^2) pixels a second, and a travel at vv moves a post at depth ZZ at fv/Zfv/Z. So the lean, in columns per row, is

TH(fω(1+x2f2)+fvZ),\frac{T}{H}\left(f\omega\left(1 + \frac{x^2}{f^2}\right) + \frac{f v}{Z}\right),

with TT the readout time and HH the number of rows. Measured off the frame, the five posts agree with that expression to 0.8 per cent; the remainder is the second-order effect of the picture moving while the solver finds each row’s time.

Two things in that expression carry the whole of what follows. The travel appears only as v/Zv/Z. And the turn carries a factor, 1+x2/f21 + x^2/f^2, that the travel does not.

The frame cannot see a travel, only its ratio to a depth

Put the posts at the centre of the frame, where x=0x = 0 and the turn’s extra factor vanishes, and the lean is a straight line in one over the depth: the turn is where the line starts and the travel is how steeply it climbs.

A post's lean is a straight line in one over its depth: the turn is where the line starts and the travel is how steeply it climbsThe lean of a post at the centre of the frame against one over its depth, from 2 m to 200 m, measured off frames read over 33.3 ms while the camera turns at 0.4 rad/s, travels at 2 m/s, or does both. A turn leans every post alike — 1.27° near and far — so its line is flat; a travel leans a post in proportion to one over its depth, from 0.03° far off to 3.16° at 2 m; both together add. Twice the travel past posts twice as far gives every lean again to 7e-18 — the frame measures the travel only as its ratio to a depth, so the line's slope can be read as a speed only once the depths are known.02400.1000.2000.3000.4000.500one over the post's depth (1/m)lean of a post at the frame's centre (°)turning 0.4 rad/stravelling 2 m/sboth55° lens, 400 rows in 33.3 msintercept: turn · slope: travel
Fig. 2 The lean of a post at the frame’s centre against one over its depth, from 2 m to 200 m, for a camera turning at 0.4 rad/s, travelling at 2 m/s, or both. The turn is flat at 1.27°; the travel rises from 0.03° to 3.16°; both add. Twice the travel past posts twice as far gives every lean again to 7e-18.

A turn alone is the flat line: every post at the centre leans 1.27°, near or far. A travel alone runs through the origin and climbs to 3.16° for a post two metres away. Both together add, and the posts measured off the frame lie on the lines.

The same figure carries the limitation. Double the travel and double every post’s depth, and every lean comes back identical to within 7×10−187 \times 10^{-18} — the arithmetic’s last digits. That is not a coincidence of these numbers; the lean contains the travel only as v/Zv/Z, so a frame of a camera travelling at four metres a second past posts at six, twelve and twenty-four metres is the frame of one travelling at two past posts at three, six and twelve. The line’s slope can be read as a speed only once the depths are known, and the frame does not contain them.

That is the scale ambiguity of every picture taken from a moving viewpoint, arriving here in a single frame. A scroll is a camera that moves found that a handscroll’s eye moves continuously along its track; a rolling frame is a very short handscroll run down the picture, and like every picture made by a moving eye it can say how far the eye moved only in units of how far away things are.

Two posts choose nothing

Two posts at the centre of the frame, three and twelve metres out, lean by 3.372° and 1.793° in the frame drawn by a camera turning at 0.4 rad/s and travelling at 2 m/s. The question the earlier essay asked is whether those two numbers give back the motion.

Two leaning posts fix neither the turn nor the travel: every guess at their depths gives a different motion that draws the same frameTwo posts at the frame's centre, 3 m and 12 m out, lean by 3.372° and 1.793° in a frame read while the camera turns at 0.4 rad/s and travels at 2 m/s. Each dot is the motion those two leans imply for one guess at the posts' depths: 1.5 and 6 m: 0.400 rad/s and 1.00 m/s; 3 and 12 m: 0.400 rad/s and 2.00 m/s; 6 and 24 m: 0.400 rad/s and 4.00 m/s; 3 and 6 m: 0.066 rad/s and 3.00 m/s; 3 and 30 m: 0.511 rad/s and 1.67 m/s; 3 and infinitely far: 0.567 rad/s and 1.50 m/s. Doubling both depths doubles the travel and leaves the turn; changing their ratio changes both. Only a post known to be very far away pins the turn by itself, since the travel cannot lean it.012340.2000.400turn the two leans imply (rad/s)travel they imply (m/s)1.5 m & 6 m3 m & 12 m6 m & 24 m3 m & 6 m3 m & 30 m3 m & fareach dot: one guess at the two depthstrue: 0.4 rad/s, 2 m/s
Fig. 3 Two posts at the frame’s centre leaning 3.372° and 1.793°, and the motion their leans imply for six guesses at their depths. With 1.5 and 6 m, 6 and 24 m, or the true 3 and 12 m, the turn is 0.400 rad/s and the travel 1, 4 or 2 m/s; with 3 and 6 m it is 0.066 rad/s and 3 m/s; with 3 m and a far post, 0.567 rad/s and 1.5 m/s.

They give back a family. Each dot in the figure is the motion the two leans imply for one guess at the two posts’ depths, and every dot draws exactly the observed frame. Guess the true depths and the true motion comes back. Guess depths half as large and the turn comes back right while the travel halves; guess depths twice as large and the travel doubles. Guess the right near depth and the wrong far one and both change: with the far post assumed at six metres instead of twelve, the implied turn is 0.066 rad/s and the travel 3 m/s; with the far post assumed infinitely far, 0.567 rad/s and 1.5 m/s.

So two posts fix one combination and not the motion. The difference of their leans is the travel times the difference of their inverse depths, (T/H) fv (1/Z1−1/Z2)(T/H)\,fv\,(1/Z_1 - 1/Z_2), and it contains no turn at all, because a turn leans both alike. The sum mixes both. The one arrangement that pins the turn by itself is a post known to be very far away: its lean is the turn and nothing else, and every nearer post’s excess lean over it is then its v/Zv/Z. A frame with a distant tower in it can read its own turn; a frame without one cannot.

Across a facade, the turn draws a bowl

The turn’s factor 1+x2/f21 + x^2/f^2 is the one handle the frame offers. A turn moves the picture faster toward its edges than at its middle, because an off-axis direction swings through more of the picture for the same turn of the camera. A travel past a surface at one depth moves every column of it alike. So a facade — many verticals at one depth, across the frame — should show the two motions differently.

Across one facade, a turn leans the edges more than the middle and a travel leans them all alike — which is how the two separateThirteen verticals of a facade 6 m away, across the frame from -300 to 300 px, and their leans in frames read while the camera turns at 0.4 rad/s, travels at 2 m/s, or both. A turn moves the picture faster toward its edges, by the factor 1 + x²/f², so its leans rise from 1.27° at the centre to 1.53° at the edge; a travel past a facade at one depth moves every column alike, 1.05° throughout. Fitted to that shape, each set of leans returns its motion: turning only, 0.400 rad/s and 0.000 /s of travel over depth; travelling only, 0.000 rad/s and 0.333 /s of travel over depth; both, 0.400 rad/s and 0.333 /s of travel over depth. The whole separation rests on the difference between a bowl and a line.11.5022.50-2000200column of the vertical from the frame's centre (px)its lean (°)turning onlytravelling onlybotha facade 6 m away, 55° lensbowl: turn · level: travel
Fig. 4 Thirteen verticals of a facade 6 m away, across the frame, and their leans. A turn at 0.4 rad/s rises from 1.27° at the centre to 1.53° at the edge; a travel at 2 m/s leans every column 1.05°. Fitted to that shape, each returns its motion: 0.400 rad/s and 0.333 per second of travel over depth for both together.

It does. Across a facade six metres away, a turn’s leans rise from 1.27° at the centre to 1.53° at the edge — a shallow bowl — while a travel leans every column the same 1.05°. The two shapes are different, so a fit of the leans to “bowl plus level” returns each motion: the turn alone, the travel alone, and both together, each to within its own arithmetic. The facade’s depth is still unknown, and the travel still comes back only as v/Zv/Z — a third of a metre per second per metre of depth, here — but the turn comes back by itself, with no distant tower needed.

That is the separation the earlier essay was looking for, found in a different place from where it looked. Two verticals at different depths separate the travel from nothing: their difference is pure travel-over-depth, and their sum is ambiguous. Many verticals at one depth and different columns separate the turn from the travel, because the turn has a shape across the picture and the travel does not.

And the two are nearly the same signal

The bowl is shallow. At the edge of a 55° lens, x2/f2x^2/f^2 is 0.27, so the turn’s lean at the edge exceeds its lean at the centre by about a quarter — against a travel that contributes a level line the same size as the turn. To separate them, the fit has to read a quarter of one signal’s variation across the frame against the whole of both signals’ level.

Across one facade the turn and the travel are 99.8 per cent the same signal: a pixel read at every row gives the turn to ±0.075 rad/s, two ends to ±0.70The uncertainty in the turn fitted from seven verticals of a facade 6 m away spanning 84 per cent of the frame's width, each two thirds of its height, computed exactly from the fit's covariance, against the lens's field of view. With each vertical's column read to a pixel on every row it spans, the turn is known to ±0.146, ±0.107, ±0.075, ±0.056, ±0.039, ±0.033 rad/s from 30° to 100°; read to half a pixel, half that; read from each vertical's two ends only, ±1.37, ±1.01, ±0.70, ±0.52, ±0.37, ±0.31 — more than the 0.4 rad/s turn itself (the dashed line) with any lens narrower than about 85°. The fitted turn and travel are correlated at -0.9977 with a 55° lens: the bowl a turn adds is a small correction to a nearly level line, and a wider lens deepens the bowl. A readout twice as slow halves every number.0.030.10.31406080100the lens's horizontal field of view (°)uncertainty in the turn (rad/s, log scale)the turn itself, 0.4 rad/severy row read, ±0.5 pxevery row read, ±1 pxtwo ends read, ±1 pxseven verticals, 6 m away · 33.3 ms readoutcorrelation -0.998 at 55°
Fig. 5 The uncertainty in the turn fitted from seven verticals of a facade 6 m away, computed from the fit’s covariance, against the lens’s field of view. Read to a pixel on every row, ±0.146 rad/s at 30° to ±0.033 at 100°; to half a pixel, half that; from each vertical’s two ends only, ±1.37 to ±0.31 rad/s. Turn and travel are correlated at −0.9977 with a 55° lens.

The fit’s covariance says how hard that is, exactly. With a 55° lens, the fitted turn and the fitted travel-over-depth are correlated at −0.9977: the two signals across a facade are 99.8 per cent the same. The precision of the turn then depends entirely on how well each lean is read.

Read each vertical’s column to a pixel on every one of the 265 rows it spans, and fit a line to those columns, and each lean is known to about 0.0008 of a column per row. Seven such verticals give the turn to ±0.075 rad/s — a fifth of the 0.4 rad/s being measured. Read each vertical only at its two ends, to the same pixel, and each lean is known to 0.0075, ten times worse, and the turn to ±0.70 rad/s: nearly twice the turn itself. The same frame, the same facade, and the difference between a measurement and noise is whether the whole edge is read or only its ends.

The lens decides the rest. A wider lens deepens the bowl, since the edges of the frame are further off-axis, and the precision improves from ±0.146 rad/s at a 30° field of view to ±0.033 at 100°. A slower readout helps in proportion: the lean is the image speed times the time per row, so doubling the readout time doubles every lean and halves every uncertainty. The fast sensors whose frames lean least are, for exactly that reason, the ones whose shear says least about how they moved.

How far apart two posts must stand

The earlier essay asked for one number specifically: how far apart in depth two verticals must be, at 33 ms and a pixel, before the difference in their leans says anything. The difference is pure travel-over-depth, so the answer depends on the speed.

Two posts' leans differ visibly only when their depths do by a margin that grows with depth and falls with speedFor a nearer post at 2, 3, 5, 8, 12, 20 m, how much farther a second post must stand for the two leans to differ by twice the standard error of their difference, in a frame read over 33.3 ms by a camera travelling at 1, 3 or 10 m/s, with each post's column read to a pixel on all of its 264 rows (solid) or at its two ends only. At 3 m/s and every row read, a post 3 m out is told from one 4.3 per cent farther; read at two ends, from one 63 per cent farther. The difference measures travel times the difference of the inverse depths, so the margin, as a share of the depth, grows in proportion to the depth and shrinks in proportion to the speed; where a line stops, no second post however far away leans visibly less. The difference contains no turn at all: a turn leans both posts alike.23581220110100depth of the nearer post (m, log scale)how much farther the other must be to lean visibly less (%, log)1 m/s3 m/s10 m/ssolid: every row read · dashed: two ends · 1 pxa turn leans both alike
Fig. 6 For a nearer post at 2 to 20 m, how much farther a second post must stand for their leans to differ by two standard errors, at 1, 3 and 10 m/s, with each column read to a pixel on all 264 rows or at two ends. At 3 m/s a post 3 m out is told from one 4.3 per cent farther with every row read, and 63 per cent farther from two ends.

With a camera travelling three metres a second and each post’s edge read to a pixel on every row, a post three metres out leans visibly differently from one 4.3 per cent farther away. Read at two ends only, the second post must be 63 per cent farther. At ten metres a second the margins fall by a factor of three; at one metre a second they rise by three. The margin, as a share of the depth, grows in proportion to the depth — the difference measures inverse depth, and a fixed step in inverse depth is a large step in depth far away — and where a line in the figure stops, no second post, however far, leans visibly less than the first: the nearer post is itself too far away for the travel to lean it more than a reading error.

So two posts can say that a camera travelled and roughly which of them is nearer, at modest separations and ordinary speeds. They cannot say how fast, and they cannot say how far away either one is. Depth is a reciprocal found exactly the same structure for a stereo pair’s disparity: a fixed error in what is read maps to a fixed interval in inverse depth, which is a small interval near and a large one far.

With the turn known, the shear is a depth map

Most cameras that read their rows one after another now carry a gyroscope, and a gyroscope measures exactly the part of the motion the frame finds hardest: the turn. Given the turn from outside, the whole difficulty of the facade figure disappears. Every vertical’s lean, less the turn’s known contribution at its own column, is its travel-over-depth, and the frame becomes a map of inverse depth with one unknown scale — the speed.

The pair figure already says how good that map is. At three metres a second, with every row of an edge read to a pixel, a post three metres away is told from one 4.3 per cent farther; at ten metres a second, from one 1.3 per cent farther. That is a single frame, with no second picture and no baseline, ranking the verticals in it by distance. It is not a range-finder, because the speed sets the scale and the frame does not know the speed; a wheel counter or a second measurement of anything at a known distance supplies it. But the ranking itself comes free with every frame a moving camera reads row by row.

The centre a scroll does not have found that a scroll’s rays miss any common centre by the spread of the eye’s track, and that the spread is what makes the scroll a measuring instrument rather than a picture. A rolling frame’s rays miss theirs by the few centimetres the camera travels during one readout. That is a very short baseline, and the shear is its disparity.

What the frame holds, and what it lacks

A rolling frame’s shear is a record of the image’s speed at every place in the picture, integrated over one readout. The image’s speed, for a camera that turns and travels, is the turn’s contribution — the same at every depth, growing toward the edges of the frame — plus the travel’s contribution, divided by each point’s depth. Everything the shear can say follows from that sum.

It can say the turn exactly when the frame contains something very far away, or when it contains many verticals at one depth across a wide field and every row of their edges is read. It can say the travel only as a ratio to depth, whatever the frame contains, because doubling the travel and every depth draws the same frame. And it can say which of two verticals is nearer, when their inverse depths differ by more than the reading error allows at the camera’s speed.

That is also why a turning frame can be straightened; a travelling one cannot came out as it did. Straightening needs the image’s motion at every point, and the turn’s part of it is known from the frame once measured; the travel’s part is known only up to each point’s depth, which is the thing a single frame does not have. Turning and travelling blur different worlds found the same division in the smear of a long exposure, which is the same integral taken over time instead of over rows. A frame is an interval is the general statement: a frame is never an instant, and whatever the camera does during it is written into the picture.

What the measurement assumes

The camera turns only about its vertical axis and travels only sideways. A camera that also pitches, rolls or moves forward adds image motion with other shapes — a forward travel spreads the picture from a point, a roll rotates it — and each adds its own term to every lean. The bowl-and-level separation holds for the two motions measured and says nothing about the others.

The posts are truly vertical and the camera level. A post that leans in the world leans in every frame, rolling or global, and its lean would be read as motion. A reader of a real frame needs to know which lines are plumb; buildings usually are, trees usually are not. Straightening does not move the eye is the reminder that correcting a picture’s verticals changes the picture and not the place it was taken from.

Every edge is read independently to the stated pixel. Real edge detectors do better than a pixel along a long straight edge and worse on texture; the precision figure is the scaling, not a promise for any particular matcher.

Still open: what a bounce does to the rows

Everything above concerns motion across the rows — a turn or travel that moves the picture sideways while the rows are read downward. Motion along the rows is different in kind. A camera that pitches, or bounces up and down on a moving vehicle, moves the picture in the same direction the readout runs, so its effect is not a lean but a compression or stretch of the rows themselves: a feature is read earlier or later than its row would suggest, and a vertical span comes out shorter or longer than a global shutter would draw it.

The same division should apply. A pitch moves every depth alike, and more toward the top and bottom of the frame; a bounce moves near things more than far ones, and every row alike. The measurement that settles it draws a frame read while the camera pitches and bounces, measures how much each post’s height is compressed against its depth and its row, and asks whether the compression separates pitch from bounce the way the lean separated turn from travel — by a shape across the frame — and whether, because the compression acts along the readout rather than across it, it is larger or smaller than the lean for the same motion.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Conditioningfield of viewinstrument limitMoving viewpointParallaxPushbroomscale ambiguity