A frame's shear knows travel only over depth
Worth reading first: Every row is a different camera · A scroll is a camera that moves.
Every row is a different camera described a rolling-shutter frame as what it is: a stack of pictures, one per row, each taken from wherever the camera was when that row was read. A vertical line in the world comes out leaning, because its lower rows were read later, after the picture had moved on. A turning frame can be straightened; a travelling one cannot then found the asymmetry between the two ways a camera moves. A turn moves every point of the picture by an amount that does not depend on its depth, so each row can be turned back without knowing anything about the scene. A travel moves near points more than far ones, so no correction that ignores depth can undo it.
The essay ended on the obvious next question. If a travel leans near verticals more than far ones and a turn leans them alike, then two verticals at different depths lean differently in a travelling frame and identically in a turning one. Can the frame read that back — separate the turning part of the camera’s motion from the travelling part, and the travelling part from the depths of the lines? And how far apart in depth must two verticals be, at a readout of 33 ms with marks read to a pixel, before the difference in their leans says anything at all?
The frame throughout is 690 by 400 pixels, read top to bottom in 33.3 ms, through a 55° lens — a focal length of 663 px. Every lean is measured off the frame by solving, for each end of a post, for the row that drew it at the instant it was read.
A lean is a turn plus a travel over a depth
The hero figure takes a camera turning at 0.4 radians a second while travelling sideways at two metres a second, and five posts at depths from three metres to effectively infinity, each spanning two thirds of the frame’s height.
The near post leans three and a half degrees, the far one a degree and a half, and the ones between fall in order. The far post is leaning by the turn alone: a camera that travels two metres a second moves a post four hundred metres away by almost nothing in 33 ms, but a camera that turns moves every direction alike.
Each lean follows one law to first order in the motion. A post’s lower row is read later than its upper by the number of rows between them times the time per row, and in that time the picture moves sideways at the image’s speed at that place. The image’s speed has two parts: a turn at moves the column at pixels a second, and a travel at moves a post at depth at . So the lean, in columns per row, is
with the readout time and the number of rows. Measured off the frame, the five posts agree with that expression to 0.8 per cent; the remainder is the second-order effect of the picture moving while the solver finds each row’s time.
Two things in that expression carry the whole of what follows. The travel appears only as . And the turn carries a factor, , that the travel does not.
The frame cannot see a travel, only its ratio to a depth
Put the posts at the centre of the frame, where and the turn’s extra factor vanishes, and the lean is a straight line in one over the depth: the turn is where the line starts and the travel is how steeply it climbs.
A turn alone is the flat line: every post at the centre leans 1.27°, near or far. A travel alone runs through the origin and climbs to 3.16° for a post two metres away. Both together add, and the posts measured off the frame lie on the lines.
The same figure carries the limitation. Double the travel and double every post’s depth, and every lean comes back identical to within — the arithmetic’s last digits. That is not a coincidence of these numbers; the lean contains the travel only as , so a frame of a camera travelling at four metres a second past posts at six, twelve and twenty-four metres is the frame of one travelling at two past posts at three, six and twelve. The line’s slope can be read as a speed only once the depths are known, and the frame does not contain them.
That is the scale ambiguity of every picture taken from a moving viewpoint, arriving here in a single frame. A scroll is a camera that moves found that a handscroll’s eye moves continuously along its track; a rolling frame is a very short handscroll run down the picture, and like every picture made by a moving eye it can say how far the eye moved only in units of how far away things are.
Two posts choose nothing
Two posts at the centre of the frame, three and twelve metres out, lean by 3.372° and 1.793° in the frame drawn by a camera turning at 0.4 rad/s and travelling at 2 m/s. The question the earlier essay asked is whether those two numbers give back the motion.
They give back a family. Each dot in the figure is the motion the two leans imply for one guess at the two posts’ depths, and every dot draws exactly the observed frame. Guess the true depths and the true motion comes back. Guess depths half as large and the turn comes back right while the travel halves; guess depths twice as large and the travel doubles. Guess the right near depth and the wrong far one and both change: with the far post assumed at six metres instead of twelve, the implied turn is 0.066 rad/s and the travel 3 m/s; with the far post assumed infinitely far, 0.567 rad/s and 1.5 m/s.
So two posts fix one combination and not the motion. The difference of their leans is the travel times the difference of their inverse depths, , and it contains no turn at all, because a turn leans both alike. The sum mixes both. The one arrangement that pins the turn by itself is a post known to be very far away: its lean is the turn and nothing else, and every nearer post’s excess lean over it is then its . A frame with a distant tower in it can read its own turn; a frame without one cannot.
Across a facade, the turn draws a bowl
The turn’s factor is the one handle the frame offers. A turn moves the picture faster toward its edges than at its middle, because an off-axis direction swings through more of the picture for the same turn of the camera. A travel past a surface at one depth moves every column of it alike. So a facade — many verticals at one depth, across the frame — should show the two motions differently.
It does. Across a facade six metres away, a turn’s leans rise from 1.27° at the centre to 1.53° at the edge — a shallow bowl — while a travel leans every column the same 1.05°. The two shapes are different, so a fit of the leans to “bowl plus level” returns each motion: the turn alone, the travel alone, and both together, each to within its own arithmetic. The facade’s depth is still unknown, and the travel still comes back only as — a third of a metre per second per metre of depth, here — but the turn comes back by itself, with no distant tower needed.
That is the separation the earlier essay was looking for, found in a different place from where it looked. Two verticals at different depths separate the travel from nothing: their difference is pure travel-over-depth, and their sum is ambiguous. Many verticals at one depth and different columns separate the turn from the travel, because the turn has a shape across the picture and the travel does not.
And the two are nearly the same signal
The bowl is shallow. At the edge of a 55° lens, is 0.27, so the turn’s lean at the edge exceeds its lean at the centre by about a quarter — against a travel that contributes a level line the same size as the turn. To separate them, the fit has to read a quarter of one signal’s variation across the frame against the whole of both signals’ level.
The fit’s covariance says how hard that is, exactly. With a 55° lens, the fitted turn and the fitted travel-over-depth are correlated at −0.9977: the two signals across a facade are 99.8 per cent the same. The precision of the turn then depends entirely on how well each lean is read.
Read each vertical’s column to a pixel on every one of the 265 rows it spans, and fit a line to those columns, and each lean is known to about 0.0008 of a column per row. Seven such verticals give the turn to ±0.075 rad/s — a fifth of the 0.4 rad/s being measured. Read each vertical only at its two ends, to the same pixel, and each lean is known to 0.0075, ten times worse, and the turn to ±0.70 rad/s: nearly twice the turn itself. The same frame, the same facade, and the difference between a measurement and noise is whether the whole edge is read or only its ends.
The lens decides the rest. A wider lens deepens the bowl, since the edges of the frame are further off-axis, and the precision improves from ±0.146 rad/s at a 30° field of view to ±0.033 at 100°. A slower readout helps in proportion: the lean is the image speed times the time per row, so doubling the readout time doubles every lean and halves every uncertainty. The fast sensors whose frames lean least are, for exactly that reason, the ones whose shear says least about how they moved.
How far apart two posts must stand
The earlier essay asked for one number specifically: how far apart in depth two verticals must be, at 33 ms and a pixel, before the difference in their leans says anything. The difference is pure travel-over-depth, so the answer depends on the speed.
With a camera travelling three metres a second and each post’s edge read to a pixel on every row, a post three metres out leans visibly differently from one 4.3 per cent farther away. Read at two ends only, the second post must be 63 per cent farther. At ten metres a second the margins fall by a factor of three; at one metre a second they rise by three. The margin, as a share of the depth, grows in proportion to the depth — the difference measures inverse depth, and a fixed step in inverse depth is a large step in depth far away — and where a line in the figure stops, no second post, however far, leans visibly less than the first: the nearer post is itself too far away for the travel to lean it more than a reading error.
So two posts can say that a camera travelled and roughly which of them is nearer, at modest separations and ordinary speeds. They cannot say how fast, and they cannot say how far away either one is. Depth is a reciprocal found exactly the same structure for a stereo pair’s disparity: a fixed error in what is read maps to a fixed interval in inverse depth, which is a small interval near and a large one far.
With the turn known, the shear is a depth map
Most cameras that read their rows one after another now carry a gyroscope, and a gyroscope measures exactly the part of the motion the frame finds hardest: the turn. Given the turn from outside, the whole difficulty of the facade figure disappears. Every vertical’s lean, less the turn’s known contribution at its own column, is its travel-over-depth, and the frame becomes a map of inverse depth with one unknown scale — the speed.
The pair figure already says how good that map is. At three metres a second, with every row of an edge read to a pixel, a post three metres away is told from one 4.3 per cent farther; at ten metres a second, from one 1.3 per cent farther. That is a single frame, with no second picture and no baseline, ranking the verticals in it by distance. It is not a range-finder, because the speed sets the scale and the frame does not know the speed; a wheel counter or a second measurement of anything at a known distance supplies it. But the ranking itself comes free with every frame a moving camera reads row by row.
The centre a scroll does not have found that a scroll’s rays miss any common centre by the spread of the eye’s track, and that the spread is what makes the scroll a measuring instrument rather than a picture. A rolling frame’s rays miss theirs by the few centimetres the camera travels during one readout. That is a very short baseline, and the shear is its disparity.
What the frame holds, and what it lacks
A rolling frame’s shear is a record of the image’s speed at every place in the picture, integrated over one readout. The image’s speed, for a camera that turns and travels, is the turn’s contribution — the same at every depth, growing toward the edges of the frame — plus the travel’s contribution, divided by each point’s depth. Everything the shear can say follows from that sum.
It can say the turn exactly when the frame contains something very far away, or when it contains many verticals at one depth across a wide field and every row of their edges is read. It can say the travel only as a ratio to depth, whatever the frame contains, because doubling the travel and every depth draws the same frame. And it can say which of two verticals is nearer, when their inverse depths differ by more than the reading error allows at the camera’s speed.
That is also why a turning frame can be straightened; a travelling one cannot came out as it did. Straightening needs the image’s motion at every point, and the turn’s part of it is known from the frame once measured; the travel’s part is known only up to each point’s depth, which is the thing a single frame does not have. Turning and travelling blur different worlds found the same division in the smear of a long exposure, which is the same integral taken over time instead of over rows. A frame is an interval is the general statement: a frame is never an instant, and whatever the camera does during it is written into the picture.
What the measurement assumes
The camera turns only about its vertical axis and travels only sideways. A camera that also pitches, rolls or moves forward adds image motion with other shapes — a forward travel spreads the picture from a point, a roll rotates it — and each adds its own term to every lean. The bowl-and-level separation holds for the two motions measured and says nothing about the others.
The posts are truly vertical and the camera level. A post that leans in the world leans in every frame, rolling or global, and its lean would be read as motion. A reader of a real frame needs to know which lines are plumb; buildings usually are, trees usually are not. Straightening does not move the eye is the reminder that correcting a picture’s verticals changes the picture and not the place it was taken from.
Every edge is read independently to the stated pixel. Real edge detectors do better than a pixel along a long straight edge and worse on texture; the precision figure is the scaling, not a promise for any particular matcher.
Still open: what a bounce does to the rows
Everything above concerns motion across the rows — a turn or travel that moves the picture sideways while the rows are read downward. Motion along the rows is different in kind. A camera that pitches, or bounces up and down on a moving vehicle, moves the picture in the same direction the readout runs, so its effect is not a lean but a compression or stretch of the rows themselves: a feature is read earlier or later than its row would suggest, and a vertical span comes out shorter or longer than a global shutter would draw it.
The same division should apply. A pitch moves every depth alike, and more toward the top and bottom of the frame; a bounce moves near things more than far ones, and every row alike. The measurement that settles it draws a frame read while the camera pitches and bounces, measures how much each post’s height is compressed against its depth and its row, and asks whether the compression separates pitch from bounce the way the lean separated turn from travel — by a shape across the frame — and whether, because the compression acts along the readout rather than across it, it is larger or smaller than the lean for the same motion.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A mirror ball does not know its size — both name conditioning, instrument limit, scale ambiguity
- A pane gives a product before it gives two numbers — both name conditioning, instrument limit, scale ambiguity
- A plane's coefficient reaches as far as its parallax — both name conditioning, instrument limit, parallax
- A scroll can be asked its own radius — both name instrument limit, moving viewpoint, pushbroom
- A scroll of a climbing road measures its grade — both name instrument limit, moving viewpoint, pushbroom
- A scroll round a bend loses its straight-line depth — both name instrument limit, moving viewpoint, pushbroom
Named objects
A flat tag is an object no other essay names yet.
Conditioningfield of viewinstrument limitMoving viewpointParallaxPushbroomscale ambiguity