The second eye

A stereo pair and a lane wander add as two readings of one depth

A pair of cameras 0.3 metres apart across a car ranges the road twenty metres ahead to 15.8 per cent a pixel from one moment's two pictures — and to 4.6 per cent from the last second's sixty, because every picture is another reading. Add the car's ordinary drift across its lane and the second's pictures give 4.5: exactly what the pair's look and a single camera's wander would give as two independent readings. The two baselines neither waste each other nor multiply, while the pair's is the larger; only a narrow pair on a car that drifts half a metre does better than the rule, by a quarter.

Worth reading first: The image of the other eye · Depth is a reciprocal.

A drift across the lane closes the blind point, slowly looked for a sideways movement that would give a forward-driving camera depth in the one direction it cannot range: the direction it is driving towards, where the image of the other eye sits on the point itself. Steering swings a camera too little, as steering swings a camera too little to see ahead found. A car’s ordinary wander across its lane — twenty centimetres either side every four seconds — does better: in one second it moves the camera 28 centimetres sideways, and the frames of that second range the road twenty metres ahead to 23 per cent a pixel. Better, but not usable, and slow to improve with a longer look, because every older frame stands further back.

The essay ended on the instrument a car is more likely to carry: two cameras a fixed distance apart across it, a stereo pair, which ranges the road ahead in an instant. A pair thirty centimetres apart ranges a point twenty metres ahead about as well as a second of wander does, and the wander’s baseline comes and goes with the driver. The question was what the two are worth together — whether the pair’s baseline and the wander’s add, so that a modest pair on a car that keeps its lane ranges the road to a few per cent, or whether the pair’s baseline is spent across a direction the wander already covers and the combination is no better than the larger of the two.

They add, as two independent readings of one depth add — no more and no less, while the pair is the larger. But the measurement turned up something the question had overlooked, and it is larger than either.

A pair is already a moving pair

The car is the earlier essay’s: fifteen metres a second, thirty frames a second, a wander of twenty centimetres either side with a four-second period, its heading following its path. Now it carries two cameras a stated distance apart across it, square to its heading at every moment. At each moment the question is the earlier essay’s: with the pictures it has, how well is the road twenty metres straight ahead of the car known? The depth error a pixel of reading in every picture costs is computed from all the pictures at once, and the median over where in its wander the car happens to be is reported.

A pair 0.3 m apart ranges the road 20 m ahead to 15.8% a pixel from this moment and 4.6% from its last second; the car's wander takes that to 4.5%, where independent readings would give 4.5%A car at 15 m/s wandering ±0.2 m across its lane every 4 s, filmed at 30 frames a second by a pair of cameras 0.3 m apart across it; the depth error a pixel of reading in every picture leaves on the road straight ahead 20 m from the car now, median over where in its wander the car is. The pair from this moment's two pictures: 15.78% (the faint line). The pair over looks of 0.1, 0.25, 0.5, 1, 1.5, 2 s, the car holding its lane: 8.46%, 6.24%, 5.28%, 4.61%, 4.37%, 4.25%. One camera over the same looks, wandering: 244.8%, 81.7%, 44.7%, 22.8%, 16.1%, 13.5%. The pair, wandering: 8.45%, 6.23%, 5.24%, 4.53%, 4.27%, 4.05%. The pair's look and the single camera's wander combined as two independent readings of one depth: 8.45%, 6.23%, 5.24%, 4.52%, 4.22%, 4.05%. The pair's own frames do most of the work — every picture adds a reading — and the wander adds to them as one more independent reading would, no more and no less. The slider sets the pair's baseline.0.010.020.050.10.20.51200.50011.502how far back the car looks, secondsdepth error a pixel costs 20 m ahead (log)one camera, wanderingthe pair, no wanderthe pair, wanderingthe two added as independentthe pair, this moment only±0.2 m every 4 s at 15 m/s, 20 m aheadpair 0.3 m
Fig. 1 A pair 0.3 m apart: 15.8% a pixel from this moment’s two pictures (faint line); over looks of 0.1 to 2 s with the car holding its lane, 8.5% to 4.3%; wandering, 8.5% to 4.1%; one camera wandering, 245% to 13.5%. The dotted line, the pair’s look and the wander added as independent readings, lies on the solid one. The slider sets the baseline.

The faint line is the pair as a stereo rig is usually imagined: two pictures taken at one moment, a baseline of thirty centimetres, and the road twenty metres ahead ranged to 15.8 per cent a pixel. That is the square of the distance over the focal length times the baseline, the reciprocal reading depth is a reciprocal found for every pair.

But a car does not keep one moment’s pictures. Over the last second it has thirty pairs of them, each a fresh reading of the road ahead from a little further back. Holding its lane exactly, it ranges the road from a second of pictures to 4.6 per cent — three and a half times better than one moment’s pair, and short of the five and a half that thirty-one equal readings would give only because the older frames stand further from the road. Two seconds of pictures give 4.3: the frames further back add little, as the earlier essay found for the single camera, because they stand far from the road that matters.

The shortfall has a simple count behind it. A reading from a frame taken k thirtieths of a second ago is a reading of a point that frame saw from half a metre further back for every thirtieth, and a reciprocal depth read over a fixed baseline costs the square of the distance, so its variance grows as the fourth power of how far back it stood. Weighted that way, the thirty-one pictures of the last second are worth about eleven and a half readings from the present moment, and the square root of eleven and a half is 3.4 — the gain the figure shows. Far enough away, a pair is one eye is the limit of the same arithmetic: the frames from two seconds back stand fifty metres from the road twenty metres ahead, and a thirty-centimetre baseline at fifty metres reads almost nothing.

That is the first answer, and it reframes the question. A pair ranges the road ahead about as well as a second of wander only if it is allowed one moment’s pictures. Allowed the same second, it ranges it five times better. A single camera’s wander is not competing with a stereo rig’s baseline; it is competing with a stereo rig’s own history.

The wander adds as an independent reading

The dashed line in the figure is one camera wandering, the earlier essay’s reading: 23 per cent from a second, 13.5 from two. The solid line is the pair wandering, its cameras moved sideways by the drift as well as held apart by the bracket: 4.5 per cent from a second, 4.1 from two. The dotted line is the rule for two independent readings of one depth — the pair’s own look and the single camera’s wander, combined by their variances — and it lies on the solid line at every look: 4.52 per cent predicted against 4.53 measured at one second.

A pair and a wander range the road ahead exactly as two independent readings would while the pair's baseline is the larger, and up to 24% better when the wander's isThe depth error a pixel costs 20 m ahead from a second's look at 15 m/s, for pairs 0.12, 0.3, 0.6, 1 m apart and a 4 s wander of ±0.05, ±0.2, ±0.5 m. Across: the pair's own second of frames with the car holding its lane, and one camera's second of wander, combined by their variances as two independent readings of one depth. Up: the pair's frames with the wander, as measured. ±0.05 m with 0.12 m: 11.44% predicted, 11.54% measured; ±0.05 m with 0.3 m: 4.61% predicted, 4.62% measured; ±0.05 m with 0.6 m: 2.31% predicted, 2.31% measured; ±0.05 m with 1 m: 1.38% predicted, 1.38% measured; ±0.2 m with 0.12 m: 10.29% predicted, 9.42% measured; ±0.2 m with 0.3 m: 4.52% predicted, 4.53% measured; ±0.2 m with 0.6 m: 2.29% predicted, 2.31% measured; ±0.2 m with 1 m: 1.38% predicted, 1.38% measured; ±0.5 m with 0.12 m: 7.14% predicted, 5.45% measured; ±0.5 m with 0.3 m: 4.11% predicted, 3.77% measured; ±0.5 m with 0.6 m: 2.24% predicted, 2.21% measured; ±0.5 m with 1 m: 1.37% predicted, 1.38% measured. Wherever the pair's baseline exceeds the wander's, the points lie on the diagonal: the pair's baseline is not spent along a direction the wander already covered, and the two add as separate readings of one depth. Where the wander is the larger — a narrow pair on a car that drifts half a metre — they lie below it, because one camera's picture now and the other camera's picture a moment ago stand further apart than either pair, and the look holds many such pairings.0.010.020.050.10.010.020.050.1the pair's look and the wander, combined as independent readings (log)the pair, wandering, as measured (log)dots: wander ±5, 20, 50 cmdashed: equala second's look, 20 m aheadbelow the diagonal: better than independent
Fig. 2 Twelve combinations of pair (0.12 to 1 m) and wander (±5 to ±50 cm), a second’s look, 20 m ahead: predicted as independent readings across, measured up. With the pair’s baseline the larger, every point is on the diagonal — 4.52% against 4.53% for 0.3 m and ±20 cm. A 0.12 m pair wandering ±50 cm measures 5.45% against 7.14% predicted.

The figure repeats the comparison for pairs from twelve centimetres to a metre and wanders from five centimetres to fifty. Wherever the pair’s baseline is larger than the wander’s sideways movement, the measured error is the independent-readings rule’s to within a per cent: a metre-wide pair wandering twenty centimetres, 1.38 against 1.38; a sixty-centimetre pair wandering five, 2.31 against 2.31. That answers the earlier essay’s question on its own terms. The pair’s baseline is not spent across a direction the wander already covered, even though both are sideways; and the two do not combine into anything larger than the sum of what each knows. They are separate readings of one depth, and they add as separate readings add.

Where the wander’s movement is the larger, the points fall below the diagonal. A pair twelve centimetres apart on a car drifting half a metre either side is predicted at 7.14 per cent and measures 5.45, a quarter better than independent readings would be. The reason is the pairings the rule does not count: one camera’s picture now and the other camera’s picture a moment ago stand further apart than either the pair or one camera’s drift, when the drift is large and the pair narrow, and a second’s look holds many such pairings. A narrow pair on a loose car gets a little more than its parts; a wide pair on a careful car gets exactly its parts.

Why two sideways baselines do not interfere

The earlier essay’s worry, that the pair’s baseline might be spent across a direction the wander already covered, was a reasonable one: both baselines are sideways, and two readings that measure the same thing in the same way are only as good as their combined precision. That is what the independent-readings rule says, and it is what the measurement finds. What the worry missed is that this is all a reading of a single depth needs. A depth is one number; every sideways baseline constrains it the same way, as a disparity divided into a constant; and two readings of one number combine by their variances whatever produced them. The rule would fail only if one reading’s error depended on the other’s, and a pixel of error in a frame of the pair and a pixel of error in a frame of the drift are independent.

The blind point is where the two kinds of picture differ. Each of the pair’s cameras, alone, drives towards the point it is asked about, and an epipole in the picture leaves a blind disc round it, as the single camera’s did; the pair’s cross-pairings — the left camera’s picture against the right camera’s — have their epipole far off to the side, where the bracket points, and range the road ahead at every moment. The wander moves the single camera’s blind disc about a little. The bracket removes it.

The pair is the floor, the wander the bonus

A car that keeps its lane perfectly still ranges the road 20 m ahead to 4.6% from a 0.3 m pair; wandering ±0.2 m, to 4.5%: the pair is the floor, the wander the bonusA pair of cameras 0.12 or 0.3 m apart on a car at 15 m/s looking back one second, its lane wander ±0, ±0.05, ±0.1, ±0.2, ±0.3, ±0.5 m every 4 s; the depth error a pixel costs straight ahead 20 m from the car now. Pair 0.12 m: 11.53%, 11.54%, 11.05%, 9.42%, 7.83%, 5.45%. Pair 0.3 m: 4.61%, 4.62%, 4.62%, 4.53%, 4.31%, 3.77%. With no wander a second's frames add the pair's own forward travel, which ranges nothing straight ahead, so the pair's instant reading is the whole of it; every centimetre of wander then adds to it. A car that keeps its lane well loses nothing it had from the pair and gains less from its driving.051000.0500.1000.2000.3000.500how far the car wanders either side of its lane, metresdepth error a pixel costs 20 m ahead (%)pair 0.12 mpair 0.3 ma second's look at 15 m/s, 20 m aheadthe pair is the floor
Fig. 3 A second’s look, 20 m ahead, wander ±0 to ±0.5 m. A pair 0.3 m apart: 4.61% with no wander, 4.53% at ±0.2 m, 3.77% at ±0.5 m. A pair 0.12 m apart: 11.53%, 9.42%, 5.45%. A car that holds its lane loses nothing the pair gives it.

What the wander adds depends on how much of it there is against the pair’s own reading. A pair thirty centimetres apart ranges the road to 4.61 per cent from a car that holds its lane exactly; a wander of twenty centimetres takes it to 4.53, two per cent better; fifty centimetres, to 3.77. A pair twelve centimetres apart — the spacing of a compact stereo camera — gains more: from 11.5 per cent to 9.4 at twenty centimetres of wander and to 5.5 at fifty.

So the wander is a bonus the pair collects without asking and never depends on. A driver who keeps the lane well loses nothing the pair gave; a driver who drifts gives the pair more to work with, most of it when the pair is narrow. The earlier essay’s single camera had nothing else, and so depended entirely on how badly its driver kept the lane.

A car cannot choose its moment

Over the wander's cycle one camera's second of frames ranges the road ahead between 16% and 57% a pixel; with a pair 0.3 m apart, between 4.0% and 4.6%A car wandering ±0.2 m every 4 s at 15 m/s, looking back one second, the road 20 m ahead, at 24 moments through the wander's cycle. One camera: 20.1%, 15.6%, 22.8%, 56.7%, 20.1%, 15.6%, 22.8%, 56.7% at every eighth of the cycle. A pair 0.3 m apart: 4.20%, 4.04%, 4.27%, 4.60%, 4.60%, 4.52%, 4.62%, 4.56%. The wander alone is worst where the last second's movement sideways has been smallest — as the car turns at the edge of its lane and the camera has barely moved across in that second — and a car cannot choose where in its wander it is when it needs a depth. The pair's own baseline is there at every moment, so the combination is steady.0.050.10.20.500.2500.5000.7501where in its 4 s wander the car is now, as a share of the cycledepth error a pixel costs 20 m ahead (log)one cameraa pair 0.3 m aparta second's look, 20 m aheada car cannot pick its moment
Fig. 4 A second’s look through the wander’s 4 s cycle, 20 m ahead. One camera: between 16% and 57% a pixel, worst where the last second’s sideways movement was least. A pair 0.3 m apart, wandering: between 4.0% and 4.6%.

The earlier essay reported medians over the wander’s phase, and the median hides the worst moments. One camera looking back one second ranges the road twenty metres ahead to sixteen per cent at the best point of the cycle and fifty-seven at the worst — where the car has just turned at the edge of its lane and the last second’s movement across it has been small. A car needs the depth of the road ahead when the road demands it, not when its wander happens to be favourable.

The pair’s own baseline is there at every moment, so with the pair the reading is steady: between 4.0 and 4.6 per cent through the whole cycle. That steadiness is worth more than the median gain. A single camera’s wander makes a depth reading that is fair on average and bad a quarter of the time; a pair’s makes one that is good all the time and slightly better when the car drifts.

Near and far

Over a second a pair 0.3 m apart ranges the road 8 m ahead to 2.56% a pixel, 2.57% with the car's wander; 60 m ahead, 10.5% and 9.9%The depth error a pixel costs straight ahead at 8, 12, 20, 30, 45, 60 m, a car wandering ±0.2 m every 4 s at 15 m/s. A pair 0.3 m apart over a second, the car holding its lane: 2.56%, 3.30%, 4.61%, 6.13%, 8.34%, 10.50%. One camera over a second of wander: 19.0%, 20.4%, 22.8%, 26.8%, 33.4%, 40.2%. The pair, wandering: 2.57%, 3.31%, 4.53%, 5.91%, 7.90%, 9.93%. The pair's error grows faster with distance than the wander's: a reciprocal depth read over a fixed baseline costs the square of the distance, while the wander's baseline is the car's own drift, reaching the far road from viewpoints spread over fifteen metres of travel. Near, the pair carries the reading and the wander adds little; far, the wander's share grows.812203045600.0050.010.020.050.10.20.512distance straight ahead, metres (log)depth error a pixel costs there (log)the pair, no wanderone camera, wanderingthe pair, wandering±0.2 m every 4 s at 15 m/sa second's look
Fig. 5 A second’s look straight ahead at 8 to 60 m. A pair 0.3 m apart holding its lane: 2.56% to 10.5%. One camera wandering: 19% to 40%. The pair wandering: 2.57% to 9.9%. The pair’s error grows faster with distance; the wander’s share grows with it.

The pair’s error grows with distance faster than the wander’s. Eight metres ahead, a second of the pair’s pictures ranges the road to 2.56 per cent and the wander changes nothing, 2.57; sixty metres ahead, the pair gives 10.5 and the wander takes it to 9.9. A reciprocal depth read over a fixed baseline costs the square of the distance, while the wander’s baseline is the car’s own drift seen from viewpoints spread over fifteen metres of travel, which reaches the far road more evenly. The wander’s share of the reading therefore grows with distance, and on a long straight road ahead it is the part of the combination that matters most — where the pair’s own error has grown largest and the wander has least to compete with.

What the two baselines are together

Put together, the measurements answer the question the wander essay left, and replace its premise. A stereo pair and a lane wander combine as two independent readings of one depth: neither wastes the other, neither multiplies the other, while the pair’s baseline is the larger; a narrow pair on a loose car gets a quarter more than the rule. But the premise that a pair is worth about what a second of wander is worth compared the pair’s single moment with the wander’s whole second. Given the same second, a pair is a moving pair, and its own thirty frames range the road five times better than the wander does.

The general point is about what a baseline is. The image of the other eye began this sequence with two cameras at one moment, and the blind point ahead of a camera driving forwards exists because every later picture stands on the same line as the first. Any sideways separation breaks the line: a steering yaw, a lane wander, a bracket holding two cameras apart. All of them are baselines, and the information each carries about the depth ahead is set by how far apart the pictures are and how many there are. A sway gives the blind centre a depth, not a good one found a small sway giving a little; here a fixed bracket repeated thirty times a second gives a great deal, and a drift on top gives what it would have given alone.

So a car that wants the depth of the road ahead should carry a pair and keep its pictures for a second or so — the second second of pictures takes the error only from 4.6 to 4.3 per cent, and the memory spent on it buys almost nothing, as the range a pair cannot see past would predict for the pictures that stand furthest back. It need not hope for its driver’s drift, and it should not count on it: the drift is worth two per cent to a wide pair and up to half to a narrow one, and it vanishes when the driver is careful.

What the arithmetic assumes about the car

The car’s motion between frames is known. Every picture is placed where the camera was when it took it, which a car does from its own odometry and steering, or by solving for its motion from the pictures. An error in that motion is an error in every baseline but the pair’s own, which the bracket fixes; the pair’s look depends on knowing the forward travel, the wander’s on knowing the drift.

The pair’s two cameras are synchronised and calibrated. Two cameras that take their pictures a frame apart on a car at fifteen metres a second stand half a metre apart along the road as well as across it, and a bracket that flexes changes the baseline it was calibrated at.

The road ahead is a point that holds still. The depth is asked of a fixed point on the road twenty metres ahead of the car now. A vehicle ahead that moves is not a fixed point, and the older pictures see it somewhere else; for moving things the pair’s single moment is the only honest reading, and the earlier comparison with one moment’s pair is the right one there.

The wander is a sine. A real driver drifts irregularly, in steps and corrections rather than a steady swing. The combination rule does not care: it needs only how far apart the pictures stood, whatever path took them there. The single camera does care, because its worst moments are where its last second’s drift was smallest, and an irregular driver who holds a line for a few seconds and then corrects leaves it blind for those seconds. The pair has no such moments.

Every reading has a pixel of independent error. A matcher that tracks a feature from frame to frame carries correlated errors, which would make thirty frames worth fewer than thirty readings for the pair as for the wander.

Still open: what the pair gains from looking at a moving car

The pair’s advantage over its own present pictures came from treating the road ahead as still. The thing a car most needs to range is the car ahead of it, which moves, and for which older pictures are of a target that has since moved.

The measurement that settles what that costs puts a vehicle a stated distance ahead, moving at a stated speed relative to the car, and asks how many of the pair’s past pictures still help when each must be read against where the vehicle was at its own moment — whether a pair that also estimates the vehicle’s speed from its own pictures keeps most of its five-fold gain, how it falls as the relative speed grows, and whether the lane wander, which helped the still road only a little, helps the moving target more or less, since a moving target changes its own bearing and so lends the pictures a baseline of its own.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineCovarianceDegenerate configurationDepth uncertaintyEpipoleMoving viewpointTriangulation