The second eye

Rectification is a family, not an operation

Turn both pictures of a pair so their epipolar lines become shared rows. A turn about the line between the eyes and a focal length are left free, and every choice puts all 44 matches on common rows to a tenth of a trillionth of a pixel and every point back where it was. What the choices disagree about is the pixels — one stretches its pictures unevenly by 1.77, another by 4.86.

Worth reading first: A point is a line over there · Depth is a reciprocal.

Two photographs taken from different places have epipolar lines at every angle, converging on points far off the frame. Matching along those lines is possible and awkward. The standard remedy is to warp both pictures first, so that every epipolar line becomes a horizontal row and a mark and its match sit on the same row of their respective pictures. Depth is a reciprocal worked in exactly that arrangement, and its whole arithmetic — disparity, then depth as the focal length times the baseline over the disparity — assumes it.

The warp is called rectification and is usually described as a step: the pictures are rectified, and then matched. The two pencils keep one number ended by asking how many ways there are to do it. There are infinitely many, in a precise and small sense, and they are not interchangeable.

The courtyard pair rectifiedBoth pictures of the standing pair turned, each about its own camera's centre, to one orientation whose horizontal axis is the line between the eyes. Every one of the 44 matches now sits on the same row of both pictures, to 1.1e-13 px, and the pair's rectified cameras triangulate every match back to its point to 3.1e-14 m. The dashed rows are the epipolar lines, which have become the rows themselves. Disparities run 198 to 613 px.left, rectifiedright, rectifiedrows agree to 1.1e-13 px · points to 3.1e-14 mturned 0° about the baseline
Fig. 1 The standing pair, 2.6 m apart, with both pictures turned about their own cameras’ centres to one orientation whose horizontal axis is the line between the eyes. All 44 matches sit on the same row of both pictures to 1.1e-13 px, and the rectified cameras triangulate every match back to its point to 3.1e-14 m. The dashed lines are rows that are also epipolar lines; disparities run from 198 to 613 px.

What rectifying is

A camera turned about its own centre does not move any ray. Straightening does not move the eye measured the consequence for one picture: a photograph of converging verticals and a photograph with the camera tilted level are the same rays drawn on different planes, related by a homography, and the correction is exact because nothing about the scene or the viewpoint changes. Rectification is that operation done to both pictures of a pair at once, with the new orientation chosen for a purpose.

The purpose fixes some of the choice. For epipolar lines to become rows, both epipoles must go to infinity in the horizontal direction, which means the line between the eyes must be parallel to both new picture planes and to their rows. So both cameras are turned to a common orientation whose horizontal axis is the baseline. That fixes two of the new orientation’s three angles.

The third is not fixed. Any orientation whose horizontal axis lies along the baseline will do, and those orientations differ by a turn about the baseline itself — a rotation that swings where the rectified pair is looking, up or down, while keeping its rows parallel to the line between the eyes. The new focal length is not fixed either: the rectified pictures can be drawn at any magnification, as long as both share it.

That is the family. Every member is a pair of homographies, one for each picture, and every member rectifies.

The courtyard pair rectified, turned 30° about the baselineBoth pictures of the standing pair turned, each about its own camera's centre, to one orientation whose horizontal axis is the line between the eyes, turned 30° about that line. Every one of the 44 matches now sits on the same row of both pictures, to 1.7e-13 px, and the pair's rectified cameras triangulate every match back to its point to 7.0e-14 m. The dashed rows are the epipolar lines, which have become the rows themselves. Disparities run 220 to 636 px.left, rectifiedright, rectifiedrows agree to 1.7e-13 px · points to 7.0e-14 mturned 30° about the baseline
Fig. 2 The same pair rectified with the common orientation turned 30° about the baseline. The rows still agree, to 1.7e-13 px, and the points still come back, to 7.0e-14 m; the pictures are drawn looking further up, and disparities now run from 220 to 636 px.

Turned 30° about the baseline, the rectified pair looks further up. Its pictures are different pictures — the courtyard sits lower in each frame and is foreshortened differently — and its disparities run from 220 to 636 px instead of 198 to 613. Its rows agree to 1.7×10131.7 \times 10^{-13} px and its cameras put every point back to 7.0×10147.0 \times 10^{-14} m. It is exactly as rectified as the first.

What every member agrees on

The rays are the same in every member, because every member only turns cameras about their centres. So the one thing no member can change is where two matched rays cross, which is to say the point.

The same three matches in seven rectified pairsSeven members of the rectification family, turned from -30° to 30° about the baseline. Three matches' disparities change with the turn — mark 1 from 297.7 to 321.3 px — but in every member the matches share a row to 5e-13 px and triangulate to their true points to 7e-14 m. A disparity is a number that belongs to one rectification; the point it gives back does not.225250275300325-30-20-100102030turn of the rectified pair about the baseline (degrees)disparity of a match (px)mark 1mark 2mark 3points agree to 7e-14 m in all seventhe disparities do not
Fig. 3 Seven members of the family, turned from −30° to 30° about the baseline, and the disparities of three matches in each. Mark 1’s disparity runs from 297.7 to 321.3 px across the family; in every member the matches share a row to 5e-13 px and triangulate to their true points to 7e-14 m.

Across seven members turned from −30° to 30°, three matches’ disparities change with the turn — one of them from 297.7 px to 321.3 px — while in every member all matches share a row to 5×10135 \times 10^{-13} px and triangulate to their true points to 7×10147 \times 10^{-14} m.

The disparity therefore is not a property of a match. It is a property of a match and a rectification. Depth is a reciprocal turned disparity into depth as Z=fB/dZ = fB/d, and that is correct within any one member: but the ZZ it gives is the distance along that member’s viewing direction, which turns with the member, and the ff is that member’s focal length. Change the member and dd, ff and the axis along which ZZ is measured all change together, in exactly the way that leaves the point where it is.

That is worth saying plainly because stereo results are often reported as disparity maps. A disparity map is a picture of depth in one rectification’s coordinates. Two disparity maps of the same scene from the same cameras, rectified differently, disagree at every pixel and describe the same scene.

What the members disagree about

Every member keeps the geometry. What distinguishes them is what they do to the pixels.

A homography stretches a picture by different amounts in different places. At a pixel (x,y)(x, y) the local area scale is detH/w3\det H / w^3, where ww is the third row of HH applied to the pixel — the same quantity that makes a picture of a receding floor show its near tiles large and its far tiles small. Where the scale is above one, the rectified picture spreads a region of the original over more pixels than it had, and the extra pixels have to be interpolated. Where it is below one, several original pixels are squeezed into one and detail is lost.

How unevenly each rectification resamples the picturesFor each member of the family, the ratio between the largest and smallest local area scale its two homographies apply anywhere in the two pictures. Turned 0° about the baseline it is 1.77: some part of a picture is spread over 1.77 times the area of another. Across the seven members it is 4.86 at -30°, 3.31 at -20°, 2.38 at -10°, 1.77 at 0°, 2.36 at 10°, 3.22 at 20°, 4.61 at 30°. Every member gives the same points; they differ in how much of the pictures they blur and how much they invent.024-30-20-100102030turn of the rectified pair about the baseline (degrees)largest over smallest local area scale, both picturesstretch ratio 1.77 at 0° → 4.86 at the endsthe standing pair, 2.6 m apart
Fig. 4 For each member of the family, the ratio between the largest and smallest local area scale its two homographies apply anywhere in the two pictures. It is 1.77 for the member turned 0° about the baseline, 4.86 at −30° and 4.61 at 30°.

For the member whose orientation is the average of the two cameras’ — turned 0° here — the largest local area scale anywhere in the two pictures is 1.77 times the smallest. Turned 30° either way, the ratio is 4.86 and 4.61. One member resamples its pictures nearly evenly; another spreads some regions nearly five times as much as others.

The cost is invisible in the geometry and very visible in a matcher. A patch taken from a region stretched by four and compared with its partner stretched by one is being compared across a change of scale that rectification introduced and the scene did not have.

What each member crops

Stretch is one cost. Where each member puts the picture is another.

Where each rectification puts the left picture's frameThe outline of the left picture after rectification, for turns of -20°, 0°, 20° about the baseline, drawn against the output frame (the grey rectangle) that a rectified picture of the same size would occupy. 28 % of the outline lies inside it at -20°, 44 % of the outline lies inside it at 0°, 28 % of the outline lies inside it at 20°. What falls outside is lost unless the output grows; what the outline does not cover has to be invented.-20°20°inside the output frame: -20° 28 % · 0° 44 % · 20° 28 %left picture's outline
Fig. 5 The outline of the left picture after rectification by members turned −20°, 0° and 20° about the baseline, drawn against a rectified output frame of the original size. At 0°, 44 % of the outline lies inside that frame; at ±20°, 28 %.

Drawn against an output frame the same size as the original picture, the rectified left picture does not fit. For the average orientation, 44 per cent of its outline lies inside the frame; for members turned 20° either way, 28 per cent. What falls outside is lost unless the output frame is enlarged, and what the outline does not cover inside the frame has to be filled with nothing.

Part of that is a property of this pair: its two cameras are verged, each turned about 11° toward the middle of the courtyard, and undoing that toe-in slides each picture sideways in the rectified frame. Part is a property of the member: turning the common orientation about the baseline swings the pictures up or down on top of that. A rectification that keeps as much of both pictures as possible has to choose both the turn and a principal point to recentre them, and the principal point is one more free number the principal point is not the centre found a picture can carry without changing any ray.

The second free number

The rectified pictures’ focal length is free too, and it trades squeezing for spreading everywhere at once.

The second free number: the rectified pair's focal lengthThe family's other free choice, the focal length the rectified pictures are drawn at, from 0.6 to 1.4 times the original. The smallest and largest local area scales run 0.29–0.52 at 0.6, 0.52–0.92 at 0.8, 0.81–1.43 at 1, 1.16–2.07 at 1.2, 1.58–2.81 at 1.4. At 0.6 and 0.8 every part of both pictures is squeezed and detail is thrown away; at 1.2 and 1.4 every part is spread and pixels are interpolated that the camera never recorded; in between, both happen in the same picture. The points recovered are the same throughout, to 4e-14 m.01230.6000.80011.201.40focal length of the rectified pair, as a multiple of the originallocal area scale applied to the pictureslargestsmallest1 = a pixel of the picture becomes a pixel of the rectified pictureturned 0° about the baseline
Fig. 6 The smallest and largest local area scales for the member turned 0°, with the rectified focal length from 0.6 to 1.4 times the original: 0.29–0.52 at 0.6, 0.52–0.92 at 0.8, 0.81–1.43 at 1, 1.16–2.07 at 1.2 and 1.58–2.81 at 1.4. The points recovered are the same throughout, to 4e-14 m.

At 0.8 times the original focal length every part of both rectified pictures is squeezed, the largest local scale being 0.92; at 1.2 every part is spread, the smallest being 1.16. At the original focal length both happen in the same picture, from 0.81 to 1.43. The points recovered do not change, to 4×10144 \times 10^{-14} m.

There is no setting at which nothing is squeezed and nothing is spread, because the two original pictures are verged and the rectified orientation is not either camera’s own. Any member resamples its pictures unevenly by at least the ratio the stretch figure found, and the focal length only chooses where that unevenness sits relative to one.

Choosing a member

The family has a natural member, and the measurement says why it is natural and what it does not guarantee.

Taking the rectified orientation as the average of the two cameras’ orientations, projected to have the baseline as its horizontal axis, gave the smallest stretch ratio of the seven turns tried — 1.77, against 2.36 and 2.38 for members turned 10° either way, 3.22 and 3.31 at 20°, and 4.61 and 4.86 at 30°. That is what an average should do: it turns each camera as little as possible from where it was. The sweep was in 10° steps, so the true minimum could lie a few degrees from zero, and it was not searched for.

It does not keep the pictures in frame. Even the least-stretching member put only 44 per cent of the left picture’s outline inside a same-size output, because rectifying a verged pair slides the pictures sideways. Keeping the pictures needs a larger output or shifted principal points, and shifting principal points changes every disparity by a constant — which, like everything else here, leaves every point where it was.

So choosing a member is an engineering decision about pixels, made with complete freedom as far as geometry is concerned. That freedom is easy to mistake for a lack of choices, because every member “rectifies”. It is the opposite: the geometry is indifferent, so the choice is made entirely on grounds the geometry cannot see.

Why turning a camera stretches its picture

The stretch has a short explanation, and it says why the ratios climb the way they do.

A homography that turns a camera about its centre, with the focal length kept, maps a pixel whose ray makes an angle θ\theta with the old viewing direction to a pixel whose ray makes an angle θ\theta' with the new one. The local area scale there is (cosθ/cosθ)3(\cos\theta / \cos\theta')^3. A ray straight down the old axis and straight down the new one is not stretched at all; a ray near the edge of the old picture that ends up further from the new axis is squeezed, and one that ends up nearer is spread, each by the cube of the ratio of cosines.

So the stretch a member applies is set by how far it turns each camera, and a turn toward one edge of a picture always squeezes that edge and spreads the other. Each of the standing pair’s cameras is verged about 11.5° toward the middle of the courtyard, so even the average orientation turns each of them 11.5° about a vertical axis, in opposite directions. The pictures’ edges are 25° either side of each old axis. After the turn, one edge is 36.5° from the new axis and the other 13.5°, and the cube of the ratio of cosines is (cos25°/cos36.5°)3=1.43(\cos 25° / \cos 36.5°)^3 = 1.43 at one and (cos25°/cos13.5°)3=0.81(\cos 25° / \cos 13.5°)^3 = 0.81 at the other — the member’s smallest and largest local scales, 0.81 and 1.43, and their ratio of 1.77, from the verging alone. Turning the common orientation up or down about the baseline adds a second turn to both cameras on top of it, and the ratio climbs — 2.36 at 10°, 3.22 at 20°, 4.61 at 30° on one side and 2.38, 3.31 and 4.86 on the other — faster than linearly, as a cube of cosines should once the edge rays are pushed far from the new axis.

The same formula explains the focal length’s effect. Changing the focal length multiplies every local area scale by the square of the change, uniformly, which is why the smallest and largest scales moved together in the last figure while their ratio stayed at 1.77 throughout.

What rectifying bought this pair

It is worth being honest about how much the geometry needed rectifying in the first place.

The search along an epipolar line that a point is a line over there describes is already a one-dimensional search, whether the line is a row or not. Rectification makes it a search along a row. For the standing pair, the epipolar lines were already close to rows: two marks off a known plane find the other eye measured the parallax lines of the same side-by-side pair — which run along its epipolar lines — and found that no two of them cross at more than 2.3°. Every epipolar line in the pair’s pictures is within a couple of degrees of every other.

Turning lines within a couple of degrees of horizontal into exact rows cost, at the best member tried, a resampling that spreads one part of a picture 1.77 times as much as another, and a crop that leaves less than half of the left picture’s outline inside a same-size frame. The geometric simplification is real, since a row is simpler to walk along than a line of slope 0.04, and small. The resampling cost is not small. For a pair this close to parallel, the right question is not which member to choose but whether to rectify at all, or to walk the epipolar lines where they are.

Where the family runs out

Every member sends the epipoles to infinity along the rows. For the standing pair the epipoles are already thousands of pixels off the frame, and moving them the rest of the way costs the stretches measured above.

For a camera that steps forward, an epipole in the picture leaves a blind disc: the epipole is inside the frame. Sending a point inside the picture to infinity requires a homography whose line at infinity passes through the picture, and along that line the local area scale is infinite. No member of this family rectifies a forward-moving pair without tearing the picture in two. Forward-moving cameras are rectified, where they are, by a different kind of resampling organised around the epipole instead of along rows, which is not measured here.

The pencil map says the same from the other side. The two pencils keep one number found that lines through one epipole map to lines through the other by a three-number projective map. Rectification makes both pencils into the same family of horizontal rows and makes that map the identity, row for row. It can do that only if both pencils can be made parallel, which is to say only if both epipoles can be moved to infinity without passing through the picture.

What this does not settle

The calibrated family only. Both cameras’ focal lengths are known here, so rectification is a rotation of each camera. Without calibration, a rectifying pair of homographies has further freedom — shears and horizontal scalings that no rotation produces — and that larger family’s stretch was not measured.

Stretch, not matching. The figures measure how unevenly each member resamples its pictures. How much that unevenness costs a particular matcher, and how interpolation interacts with sub-pixel disparity, are questions about matching, which every figure in this field hands over rather than performs.

One pair. The standing pair, verged about 11° each way, 2.6 m apart. A pair already nearly parallel would need far less turning, and every member’s stretch would be closer to one.

Still open: where a shifted rectification puts the shells

Every member of the family keeps every point, and the principal point of the rectified pictures is one more free number, which slides each picture sideways and adds a constant to every disparity. A constant added to disparity changes nothing about where points are.

It does change what a quantised disparity means. Whole pixels cut space into shells found that a pair reading disparity to whole pixels divides space into shells at depths set by the focal length and baseline over each integer disparity. If a rectification adds a disparity offset that is not a whole number of pixels, the integers fall at different disparities, and the shells sit at different depths — the same scene, the same points, and a different set of depths at which a whole-pixel reading can land. The question that leaves is how far a half-pixel shift of the principal point moves those shells at the depths a pair cares about, and whether choosing the shift is a way to put a pair’s finest shells where its scene actually is.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

BaselineDisparityEpipolar lineHomographyPrincipal pointrectified pair