A stereo pair and a lane wander add as two readings of one depth
Worth reading first: The image of the other eye · Depth is a reciprocal.
A drift across the lane closes the blind point, slowly looked for a sideways movement that would give a forward-driving camera depth in the one direction it cannot range: the direction it is driving towards, where the image of the other eye sits on the point itself. Steering swings a camera too little, as steering swings a camera too little to see ahead found. A car’s ordinary wander across its lane — twenty centimetres either side every four seconds — does better: in one second it moves the camera 28 centimetres sideways, and the frames of that second range the road twenty metres ahead to 23 per cent a pixel. Better, but not usable, and slow to improve with a longer look, because every older frame stands further back.
The essay ended on the instrument a car is more likely to carry: two cameras a fixed distance apart across it, a stereo pair, which ranges the road ahead in an instant. A pair thirty centimetres apart ranges a point twenty metres ahead about as well as a second of wander does, and the wander’s baseline comes and goes with the driver. The question was what the two are worth together — whether the pair’s baseline and the wander’s add, so that a modest pair on a car that keeps its lane ranges the road to a few per cent, or whether the pair’s baseline is spent across a direction the wander already covers and the combination is no better than the larger of the two.
They add, as two independent readings of one depth add — no more and no less, while the pair is the larger. But the measurement turned up something the question had overlooked, and it is larger than either.
A pair is already a moving pair
The car is the earlier essay’s: fifteen metres a second, thirty frames a second, a wander of twenty centimetres either side with a four-second period, its heading following its path. Now it carries two cameras a stated distance apart across it, square to its heading at every moment. At each moment the question is the earlier essay’s: with the pictures it has, how well is the road twenty metres straight ahead of the car known? The depth error a pixel of reading in every picture costs is computed from all the pictures at once, and the median over where in its wander the car happens to be is reported.
The faint line is the pair as a stereo rig is usually imagined: two pictures taken at one moment, a baseline of thirty centimetres, and the road twenty metres ahead ranged to 15.8 per cent a pixel. That is the square of the distance over the focal length times the baseline, the reciprocal reading depth is a reciprocal found for every pair.
But a car does not keep one moment’s pictures. Over the last second it has thirty pairs of them, each a fresh reading of the road ahead from a little further back. Holding its lane exactly, it ranges the road from a second of pictures to 4.6 per cent — three and a half times better than one moment’s pair, and short of the five and a half that thirty-one equal readings would give only because the older frames stand further from the road. Two seconds of pictures give 4.3: the frames further back add little, as the earlier essay found for the single camera, because they stand far from the road that matters.
The shortfall has a simple count behind it. A reading from a frame taken k thirtieths of a second ago is a reading of a point that frame saw from half a metre further back for every thirtieth, and a reciprocal depth read over a fixed baseline costs the square of the distance, so its variance grows as the fourth power of how far back it stood. Weighted that way, the thirty-one pictures of the last second are worth about eleven and a half readings from the present moment, and the square root of eleven and a half is 3.4 — the gain the figure shows. Far enough away, a pair is one eye is the limit of the same arithmetic: the frames from two seconds back stand fifty metres from the road twenty metres ahead, and a thirty-centimetre baseline at fifty metres reads almost nothing.
That is the first answer, and it reframes the question. A pair ranges the road ahead about as well as a second of wander only if it is allowed one moment’s pictures. Allowed the same second, it ranges it five times better. A single camera’s wander is not competing with a stereo rig’s baseline; it is competing with a stereo rig’s own history.
The wander adds as an independent reading
The dashed line in the figure is one camera wandering, the earlier essay’s reading: 23 per cent from a second, 13.5 from two. The solid line is the pair wandering, its cameras moved sideways by the drift as well as held apart by the bracket: 4.5 per cent from a second, 4.1 from two. The dotted line is the rule for two independent readings of one depth — the pair’s own look and the single camera’s wander, combined by their variances — and it lies on the solid line at every look: 4.52 per cent predicted against 4.53 measured at one second.
The figure repeats the comparison for pairs from twelve centimetres to a metre and wanders from five centimetres to fifty. Wherever the pair’s baseline is larger than the wander’s sideways movement, the measured error is the independent-readings rule’s to within a per cent: a metre-wide pair wandering twenty centimetres, 1.38 against 1.38; a sixty-centimetre pair wandering five, 2.31 against 2.31. That answers the earlier essay’s question on its own terms. The pair’s baseline is not spent across a direction the wander already covered, even though both are sideways; and the two do not combine into anything larger than the sum of what each knows. They are separate readings of one depth, and they add as separate readings add.
Where the wander’s movement is the larger, the points fall below the diagonal. A pair twelve centimetres apart on a car drifting half a metre either side is predicted at 7.14 per cent and measures 5.45, a quarter better than independent readings would be. The reason is the pairings the rule does not count: one camera’s picture now and the other camera’s picture a moment ago stand further apart than either the pair or one camera’s drift, when the drift is large and the pair narrow, and a second’s look holds many such pairings. A narrow pair on a loose car gets a little more than its parts; a wide pair on a careful car gets exactly its parts.
Why two sideways baselines do not interfere
The earlier essay’s worry, that the pair’s baseline might be spent across a direction the wander already covered, was a reasonable one: both baselines are sideways, and two readings that measure the same thing in the same way are only as good as their combined precision. That is what the independent-readings rule says, and it is what the measurement finds. What the worry missed is that this is all a reading of a single depth needs. A depth is one number; every sideways baseline constrains it the same way, as a disparity divided into a constant; and two readings of one number combine by their variances whatever produced them. The rule would fail only if one reading’s error depended on the other’s, and a pixel of error in a frame of the pair and a pixel of error in a frame of the drift are independent.
The blind point is where the two kinds of picture differ. Each of the pair’s cameras, alone, drives towards the point it is asked about, and an epipole in the picture leaves a blind disc round it, as the single camera’s did; the pair’s cross-pairings — the left camera’s picture against the right camera’s — have their epipole far off to the side, where the bracket points, and range the road ahead at every moment. The wander moves the single camera’s blind disc about a little. The bracket removes it.
The pair is the floor, the wander the bonus
What the wander adds depends on how much of it there is against the pair’s own reading. A pair thirty centimetres apart ranges the road to 4.61 per cent from a car that holds its lane exactly; a wander of twenty centimetres takes it to 4.53, two per cent better; fifty centimetres, to 3.77. A pair twelve centimetres apart — the spacing of a compact stereo camera — gains more: from 11.5 per cent to 9.4 at twenty centimetres of wander and to 5.5 at fifty.
So the wander is a bonus the pair collects without asking and never depends on. A driver who keeps the lane well loses nothing the pair gave; a driver who drifts gives the pair more to work with, most of it when the pair is narrow. The earlier essay’s single camera had nothing else, and so depended entirely on how badly its driver kept the lane.
A car cannot choose its moment
The earlier essay reported medians over the wander’s phase, and the median hides the worst moments. One camera looking back one second ranges the road twenty metres ahead to sixteen per cent at the best point of the cycle and fifty-seven at the worst — where the car has just turned at the edge of its lane and the last second’s movement across it has been small. A car needs the depth of the road ahead when the road demands it, not when its wander happens to be favourable.
The pair’s own baseline is there at every moment, so with the pair the reading is steady: between 4.0 and 4.6 per cent through the whole cycle. That steadiness is worth more than the median gain. A single camera’s wander makes a depth reading that is fair on average and bad a quarter of the time; a pair’s makes one that is good all the time and slightly better when the car drifts.
Near and far
The pair’s error grows with distance faster than the wander’s. Eight metres ahead, a second of the pair’s pictures ranges the road to 2.56 per cent and the wander changes nothing, 2.57; sixty metres ahead, the pair gives 10.5 and the wander takes it to 9.9. A reciprocal depth read over a fixed baseline costs the square of the distance, while the wander’s baseline is the car’s own drift seen from viewpoints spread over fifteen metres of travel, which reaches the far road more evenly. The wander’s share of the reading therefore grows with distance, and on a long straight road ahead it is the part of the combination that matters most — where the pair’s own error has grown largest and the wander has least to compete with.
What the two baselines are together
Put together, the measurements answer the question the wander essay left, and replace its premise. A stereo pair and a lane wander combine as two independent readings of one depth: neither wastes the other, neither multiplies the other, while the pair’s baseline is the larger; a narrow pair on a loose car gets a quarter more than the rule. But the premise that a pair is worth about what a second of wander is worth compared the pair’s single moment with the wander’s whole second. Given the same second, a pair is a moving pair, and its own thirty frames range the road five times better than the wander does.
The general point is about what a baseline is. The image of the other eye began this sequence with two cameras at one moment, and the blind point ahead of a camera driving forwards exists because every later picture stands on the same line as the first. Any sideways separation breaks the line: a steering yaw, a lane wander, a bracket holding two cameras apart. All of them are baselines, and the information each carries about the depth ahead is set by how far apart the pictures are and how many there are. A sway gives the blind centre a depth, not a good one found a small sway giving a little; here a fixed bracket repeated thirty times a second gives a great deal, and a drift on top gives what it would have given alone.
So a car that wants the depth of the road ahead should carry a pair and keep its pictures for a second or so — the second second of pictures takes the error only from 4.6 to 4.3 per cent, and the memory spent on it buys almost nothing, as the range a pair cannot see past would predict for the pictures that stand furthest back. It need not hope for its driver’s drift, and it should not count on it: the drift is worth two per cent to a wide pair and up to half to a narrow one, and it vanishes when the driver is careful.
What the arithmetic assumes about the car
The car’s motion between frames is known. Every picture is placed where the camera was when it took it, which a car does from its own odometry and steering, or by solving for its motion from the pictures. An error in that motion is an error in every baseline but the pair’s own, which the bracket fixes; the pair’s look depends on knowing the forward travel, the wander’s on knowing the drift.
The pair’s two cameras are synchronised and calibrated. Two cameras that take their pictures a frame apart on a car at fifteen metres a second stand half a metre apart along the road as well as across it, and a bracket that flexes changes the baseline it was calibrated at.
The road ahead is a point that holds still. The depth is asked of a fixed point on the road twenty metres ahead of the car now. A vehicle ahead that moves is not a fixed point, and the older pictures see it somewhere else; for moving things the pair’s single moment is the only honest reading, and the earlier comparison with one moment’s pair is the right one there.
The wander is a sine. A real driver drifts irregularly, in steps and corrections rather than a steady swing. The combination rule does not care: it needs only how far apart the pictures stood, whatever path took them there. The single camera does care, because its worst moments are where its last second’s drift was smallest, and an irregular driver who holds a line for a few seconds and then corrects leaves it blind for those seconds. The pair has no such moments.
Every reading has a pixel of independent error. A matcher that tracks a feature from frame to frame carries correlated errors, which would make thirty frames worth fewer than thirty readings for the pair as for the wander.
Still open: what the pair gains from looking at a moving car
The pair’s advantage over its own present pictures came from treating the road ahead as still. The thing a car most needs to range is the car ahead of it, which moves, and for which older pictures are of a target that has since moved.
The measurement that settles what that costs puts a vehicle a stated distance ahead, moving at a stated speed relative to the car, and asks how many of the pair’s past pictures still help when each must be read against where the vehicle was at its own moment — whether a pair that also estimates the vehicle’s speed from its own pictures keeps most of its five-fold gain, how it falls as the relative speed grows, and whether the lane wander, which helped the still road only a little, helps the moving target more or less, since a moving target changes its own bearing and so lends the pictures a baseline of its own.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A turn moves the picture, not the blind point — both name baseline, degenerate configuration, depth uncertainty, epipole, moving viewpoint, triangulation
- Square to the camera is the worst mirror — both name baseline, degenerate configuration, depth uncertainty, epipole, triangulation
- A third ray is worth what its picture is worth — both name baseline, covariance, depth uncertainty, triangulation
- The spread a point gets — both name baseline, covariance, depth uncertainty, triangulation
- A scroll of a climbing road measures its grade — both name baseline, depth uncertainty, moving viewpoint
- A scroll round a bend loses its straight-line depth — both name baseline, depth uncertainty, moving viewpoint
Named objects
A flat tag is an object no other essay names yet.
BaselineCovarianceDegenerate configurationDepth uncertaintyEpipoleMoving viewpointTriangulation