Rectification is a family, not an operation
Worth reading first: A point is a line over there · Depth is a reciprocal.
Two photographs taken from different places have epipolar lines at every angle, converging on points far off the frame. Matching along those lines is possible and awkward. The standard remedy is to warp both pictures first, so that every epipolar line becomes a horizontal row and a mark and its match sit on the same row of their respective pictures. Depth is a reciprocal worked in exactly that arrangement, and its whole arithmetic — disparity, then depth as the focal length times the baseline over the disparity — assumes it.
The warp is called rectification and is usually described as a step: the pictures are rectified, and then matched. The two pencils keep one number ended by asking how many ways there are to do it. There are infinitely many, in a precise and small sense, and they are not interchangeable.
What rectifying is
A camera turned about its own centre does not move any ray. Straightening does not move the eye measured the consequence for one picture: a photograph of converging verticals and a photograph with the camera tilted level are the same rays drawn on different planes, related by a homography, and the correction is exact because nothing about the scene or the viewpoint changes. Rectification is that operation done to both pictures of a pair at once, with the new orientation chosen for a purpose.
The purpose fixes some of the choice. For epipolar lines to become rows, both epipoles must go to infinity in the horizontal direction, which means the line between the eyes must be parallel to both new picture planes and to their rows. So both cameras are turned to a common orientation whose horizontal axis is the baseline. That fixes two of the new orientation’s three angles.
The third is not fixed. Any orientation whose horizontal axis lies along the baseline will do, and those orientations differ by a turn about the baseline itself — a rotation that swings where the rectified pair is looking, up or down, while keeping its rows parallel to the line between the eyes. The new focal length is not fixed either: the rectified pictures can be drawn at any magnification, as long as both share it.
That is the family. Every member is a pair of homographies, one for each picture, and every member rectifies.
Turned 30° about the baseline, the rectified pair looks further up. Its pictures are different pictures — the courtyard sits lower in each frame and is foreshortened differently — and its disparities run from 220 to 636 px instead of 198 to 613. Its rows agree to px and its cameras put every point back to m. It is exactly as rectified as the first.
What every member agrees on
The rays are the same in every member, because every member only turns cameras about their centres. So the one thing no member can change is where two matched rays cross, which is to say the point.
Across seven members turned from −30° to 30°, three matches’ disparities change with the turn — one of them from 297.7 px to 321.3 px — while in every member all matches share a row to px and triangulate to their true points to m.
The disparity therefore is not a property of a match. It is a property of a match and a rectification. Depth is a reciprocal turned disparity into depth as , and that is correct within any one member: but the it gives is the distance along that member’s viewing direction, which turns with the member, and the is that member’s focal length. Change the member and , and the axis along which is measured all change together, in exactly the way that leaves the point where it is.
That is worth saying plainly because stereo results are often reported as disparity maps. A disparity map is a picture of depth in one rectification’s coordinates. Two disparity maps of the same scene from the same cameras, rectified differently, disagree at every pixel and describe the same scene.
What the members disagree about
Every member keeps the geometry. What distinguishes them is what they do to the pixels.
A homography stretches a picture by different amounts in different places. At a pixel the local area scale is , where is the third row of applied to the pixel — the same quantity that makes a picture of a receding floor show its near tiles large and its far tiles small. Where the scale is above one, the rectified picture spreads a region of the original over more pixels than it had, and the extra pixels have to be interpolated. Where it is below one, several original pixels are squeezed into one and detail is lost.
For the member whose orientation is the average of the two cameras’ — turned 0° here — the largest local area scale anywhere in the two pictures is 1.77 times the smallest. Turned 30° either way, the ratio is 4.86 and 4.61. One member resamples its pictures nearly evenly; another spreads some regions nearly five times as much as others.
The cost is invisible in the geometry and very visible in a matcher. A patch taken from a region stretched by four and compared with its partner stretched by one is being compared across a change of scale that rectification introduced and the scene did not have.
What each member crops
Stretch is one cost. Where each member puts the picture is another.
Drawn against an output frame the same size as the original picture, the rectified left picture does not fit. For the average orientation, 44 per cent of its outline lies inside the frame; for members turned 20° either way, 28 per cent. What falls outside is lost unless the output frame is enlarged, and what the outline does not cover inside the frame has to be filled with nothing.
Part of that is a property of this pair: its two cameras are verged, each turned about 11° toward the middle of the courtyard, and undoing that toe-in slides each picture sideways in the rectified frame. Part is a property of the member: turning the common orientation about the baseline swings the pictures up or down on top of that. A rectification that keeps as much of both pictures as possible has to choose both the turn and a principal point to recentre them, and the principal point is one more free number the principal point is not the centre found a picture can carry without changing any ray.
The second free number
The rectified pictures’ focal length is free too, and it trades squeezing for spreading everywhere at once.
At 0.8 times the original focal length every part of both rectified pictures is squeezed, the largest local scale being 0.92; at 1.2 every part is spread, the smallest being 1.16. At the original focal length both happen in the same picture, from 0.81 to 1.43. The points recovered do not change, to m.
There is no setting at which nothing is squeezed and nothing is spread, because the two original pictures are verged and the rectified orientation is not either camera’s own. Any member resamples its pictures unevenly by at least the ratio the stretch figure found, and the focal length only chooses where that unevenness sits relative to one.
Choosing a member
The family has a natural member, and the measurement says why it is natural and what it does not guarantee.
Taking the rectified orientation as the average of the two cameras’ orientations, projected to have the baseline as its horizontal axis, gave the smallest stretch ratio of the seven turns tried — 1.77, against 2.36 and 2.38 for members turned 10° either way, 3.22 and 3.31 at 20°, and 4.61 and 4.86 at 30°. That is what an average should do: it turns each camera as little as possible from where it was. The sweep was in 10° steps, so the true minimum could lie a few degrees from zero, and it was not searched for.
It does not keep the pictures in frame. Even the least-stretching member put only 44 per cent of the left picture’s outline inside a same-size output, because rectifying a verged pair slides the pictures sideways. Keeping the pictures needs a larger output or shifted principal points, and shifting principal points changes every disparity by a constant — which, like everything else here, leaves every point where it was.
So choosing a member is an engineering decision about pixels, made with complete freedom as far as geometry is concerned. That freedom is easy to mistake for a lack of choices, because every member “rectifies”. It is the opposite: the geometry is indifferent, so the choice is made entirely on grounds the geometry cannot see.
Why turning a camera stretches its picture
The stretch has a short explanation, and it says why the ratios climb the way they do.
A homography that turns a camera about its centre, with the focal length kept, maps a pixel whose ray makes an angle with the old viewing direction to a pixel whose ray makes an angle with the new one. The local area scale there is . A ray straight down the old axis and straight down the new one is not stretched at all; a ray near the edge of the old picture that ends up further from the new axis is squeezed, and one that ends up nearer is spread, each by the cube of the ratio of cosines.
So the stretch a member applies is set by how far it turns each camera, and a turn toward one edge of a picture always squeezes that edge and spreads the other. Each of the standing pair’s cameras is verged about 11.5° toward the middle of the courtyard, so even the average orientation turns each of them 11.5° about a vertical axis, in opposite directions. The pictures’ edges are 25° either side of each old axis. After the turn, one edge is 36.5° from the new axis and the other 13.5°, and the cube of the ratio of cosines is at one and at the other — the member’s smallest and largest local scales, 0.81 and 1.43, and their ratio of 1.77, from the verging alone. Turning the common orientation up or down about the baseline adds a second turn to both cameras on top of it, and the ratio climbs — 2.36 at 10°, 3.22 at 20°, 4.61 at 30° on one side and 2.38, 3.31 and 4.86 on the other — faster than linearly, as a cube of cosines should once the edge rays are pushed far from the new axis.
The same formula explains the focal length’s effect. Changing the focal length multiplies every local area scale by the square of the change, uniformly, which is why the smallest and largest scales moved together in the last figure while their ratio stayed at 1.77 throughout.
What rectifying bought this pair
It is worth being honest about how much the geometry needed rectifying in the first place.
The search along an epipolar line that a point is a line over there describes is already a one-dimensional search, whether the line is a row or not. Rectification makes it a search along a row. For the standing pair, the epipolar lines were already close to rows: two marks off a known plane find the other eye measured the parallax lines of the same side-by-side pair — which run along its epipolar lines — and found that no two of them cross at more than 2.3°. Every epipolar line in the pair’s pictures is within a couple of degrees of every other.
Turning lines within a couple of degrees of horizontal into exact rows cost, at the best member tried, a resampling that spreads one part of a picture 1.77 times as much as another, and a crop that leaves less than half of the left picture’s outline inside a same-size frame. The geometric simplification is real, since a row is simpler to walk along than a line of slope 0.04, and small. The resampling cost is not small. For a pair this close to parallel, the right question is not which member to choose but whether to rectify at all, or to walk the epipolar lines where they are.
Where the family runs out
Every member sends the epipoles to infinity along the rows. For the standing pair the epipoles are already thousands of pixels off the frame, and moving them the rest of the way costs the stretches measured above.
For a camera that steps forward, an epipole in the picture leaves a blind disc: the epipole is inside the frame. Sending a point inside the picture to infinity requires a homography whose line at infinity passes through the picture, and along that line the local area scale is infinite. No member of this family rectifies a forward-moving pair without tearing the picture in two. Forward-moving cameras are rectified, where they are, by a different kind of resampling organised around the epipole instead of along rows, which is not measured here.
The pencil map says the same from the other side. The two pencils keep one number found that lines through one epipole map to lines through the other by a three-number projective map. Rectification makes both pencils into the same family of horizontal rows and makes that map the identity, row for row. It can do that only if both pencils can be made parallel, which is to say only if both epipoles can be moved to infinity without passing through the picture.
What this does not settle
The calibrated family only. Both cameras’ focal lengths are known here, so rectification is a rotation of each camera. Without calibration, a rectifying pair of homographies has further freedom — shears and horizontal scalings that no rotation produces — and that larger family’s stretch was not measured.
Stretch, not matching. The figures measure how unevenly each member resamples its pictures. How much that unevenness costs a particular matcher, and how interpolation interacts with sub-pixel disparity, are questions about matching, which every figure in this field hands over rather than performs.
One pair. The standing pair, verged about 11° each way, 2.6 m apart. A pair already nearly parallel would need far less turning, and every member’s stretch would be closer to one.
Still open: where a shifted rectification puts the shells
Every member of the family keeps every point, and the principal point of the rectified pictures is one more free number, which slides each picture sideways and adds a constant to every disparity. A constant added to disparity changes nothing about where points are.
It does change what a quantised disparity means. Whole pixels cut space into shells found that a pair reading disparity to whole pixels divides space into shells at depths set by the focal length and baseline over each integer disparity. If a rectification adds a disparity offset that is not a whole number of pixels, the integers fall at different disparities, and the shells sit at different depths — the same scene, the same points, and a different set of depths at which a whole-pixel reading can land. The question that leaves is how far a half-pixel shift of the principal point moves those shells at the depths a pair cares about, and whether choosing the shift is a way to put a pair’s finest shells where its scene actually is.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- Turning the cameras inwards — both name baseline, disparity, principal point, rectified pair
- The range a pair cannot see past — both name baseline, disparity, rectified pair
- A rig is right on one surface — both name baseline, disparity
- A scroll round a bend loses its straight-line depth — both name baseline, disparity
- A scroll through two slits ranges in a straight line — both name baseline, disparity
- A third eye that lands on the next post — both name baseline, epipolar line
Named objects
A flat tag is an object no other essay names yet.
BaselineDisparityEpipolar lineHomographyPrincipal pointrectified pair