Straightening does not move the eye
Worth reading first: The plane is a choice · Flattening a façade out of the photograph.
Photograph a building from across the street, point the camera up to fit the top in, and the verticals converge. Every photo editor has a slider that takes the convergence out.
What that slider does is worth stating precisely, because “correcting the perspective” is a name that suggests something was wrong. Nothing was wrong. The picture was a correct projection of the building onto a leaning plane, and the correction re-casts it onto an upright one from the same point.
Built from the picture, checked against a camera it never saw
The construction takes two things, both available to a reader with a photograph and a straightedge.
Where the imaged verticals meet. Two lines that were vertical in the world, extended until they cross. That intersection is the vertical vanishing point.
And the focal length. Which enters exactly once, to turn that vanishing point into a direction: a vanishing point is where the ray through the eye parallel to a world direction pierces the picture, so the direction in the camera’s own frame is normalised.
From those, the smallest rotation carrying that direction to image-up, wrapped in the calibration, is the correction. Nothing else is needed — no known object, no measured height, no second photograph.
The check is the site’s standard one. Build a level camera at the same eye, project the same scene through it, and compare. The correction was never shown that camera, and its output agrees with it to 3e-13 px.
So “straighten the verticals” and “re-cast the picture onto a vertical plane through the same point” are the same operation, measured. The second phrasing is the one that says what has changed.
The eye has not moved, and the cross-ratio says so
The claim that matters for anyone using a corrected photograph as evidence is that the correction changes what the picture looks like and not what it measures.
A homography of the picture cannot change a cross-ratio. So take four collinear points in the scene — four marks along a ground line — and read their cross-ratio in the original picture and in the corrected one: 1.3333 both times.
That is not a numerical coincidence about these four points; it is the invariant a projectivity is defined by preserving. And it settles the question, because every measurement this site makes from a single photograph is a cross-ratio underneath: a height, a plan, a rectified façade.
Every measurement the original supported, the corrected picture supports, with the same answer. And no measurement the original could not support becomes available.
In particular the picture is still taken from the same place. A straightened photograph of a building is a photograph from the street, cast on an upright plane — not a photograph from a first-floor window, which would show the parts of the building the street view could not see.
What it costs
The correction is exact and it is not free, and the costs are the same ones a projector pays.
Resolution, unevenly. A homography stretches some parts of the picture and squashes others. Straightening a keystone stretches the top of the frame — the part that was most compressed by the tilt — so the corrected picture has less detail per drawn millimetre up there than it does at the bottom. The stretch factor at a point is the map’s local area scale, and it varies across the frame.
Frame, at the corners. The corrected picture is a quadrilateral rather than a rectangle, so a rectangular result means cropping to the largest rectangle inside it — which is exactly what a projector’s keystone correction does, and there the cost is measured at 18.7% of the panel’s pixels for 15° off square.
And nothing about the geometry. No cross-ratio, no vanishing point, no measurement. The costs are photographic rather than projective, which is the sense in which the operation is free.
Why the focal length is needed, and what a wrong one does
There is one input a reader might expect to be dispensable, and the way it fails is instructive.
The correction needs to convert the vanishing point into a direction. Guess wrong and the resulting map still sends the vanishing point to infinity — the verticals still come out parallel — and it is a different map from the correct one.
So the picture looks straightened and is not the picture a level camera would have taken. Every vertical is parallel; the horizontal scale across the frame is wrong; and a measurement made on the corrected picture returns a wrong answer with no sign of trouble.
That is the same trap the principal point essay records for a focal recovery — the assumed centre costing one and a half per cent at a fifth of a frame’s shift — and it has the same shape as the anamorph whose recovery was believed over-determined. A construction with a free parameter that does not affect the visible criterion will pass the visible criterion at every value of the parameter.
The defence is the one this site uses everywhere: check the correction against something it was not given. Here that is the level camera; on a real photograph it is a known rectangle in the scene, whose proportions after correction say whether was right.
When there is nothing to correct
The refusal is worth stating because it is the one case a slider will silently mishandle.
A picture taken with the plane vertical has no vertical vanishing point. The imaged verticals are parallel, they do not meet, and the correction has no input. The machinery refuses rather than intersecting two nearly-parallel lines and returning an enormous number.
That is the same refusal the projector’s recovery makes: a square-on projector throws a rectangle, its opposite edges are parallel on the panel, and there is no vanishing point to read a focal length from. A picture with no keystone in it contains no evidence about the tilt, and returning a value anyway would be reporting the arithmetic.
Between the two extremes the intersection is well defined but poorly conditioned: a nearly-vertical plane puts the vanishing point thousands of pixels off the frame, where a pixel of error in either line moves it a long way. The correction is then still exact in principle and increasingly uncertain in practice — the pattern the light field measures at length as a lamp walks toward infinity.
What happens to the correct viewing distance
The site’s one piece of non-negotiable furniture is the distance a picture is correct from: the focal length scaled to the width the picture is displayed at. Correcting a keystone interacts with that in a way worth following, because the answer is no change and then a change.
The correction leaves the focal length alone. It is a rotation wrapped in the same calibration at both ends, so the corrected picture has the same and the same principal point as the original. Displayed at the same width, it is correct from the same distance.
The crop changes it. The corrected picture is a quadrilateral, so a rectangular result means keeping a smaller rectangle — and the correct viewing distance is scaled to the displayed width, so a picture cropped to 80% of its width and then printed at the original size is correct from 1.25 times as far away.
That is not a subtlety about this operation; it is the general rule, and it catches every crop. A picture’s correct viewpoint is a property of the picture and the size it is shown at, and cropping changes the second while leaving the scene alone.
The practical residue: a straightened, cropped and re-enlarged architectural photograph is correct from further back than the original was, by exactly the crop factor. Which is usually an improvement — the original, taken with a wide lens close to the building, was correct from an uncomfortably short distance to begin with.
Three names for one operation
The correction has been available in three technologies and it is the same map in all three, which is worth laying out because the differences are photographic rather than geometric.
The view camera’s rising front avoids the problem instead of correcting it: keep the back — the picture plane — vertical, and slide the lens up. The plane never tilts, so the verticals never converge, and there is nothing to undo. The cost is field: the lens has to cover an image circle much larger than the frame.
The enlarger’s tilted easel corrects optically at printing time, by projecting the negative onto a plane at an angle chosen to undo the original tilt. Same homography, applied by geometry rather than by arithmetic, and it costs focus — the tilted easel is not perpendicular to the enlarger’s axis, so the whole print cannot be sharp at once without tilting the lens as well.
And the slider applies the map to a sampled image, which costs resolution unevenly and nothing else.
All three implement the same matrix. The differences are in what is paid: field, focus, or resolution — which is a useful summary of what changed when the operation moved from optics into arithmetic.
Straightening is not rectifying
Two operations that both turn a photograph into a squarer-looking photograph, and they are different maps with different requirements. The distinction is the foundations field’s census, applied.
Straightening is a plane change: same eye, new picture plane, and the map is — determined by the tilt and the calibration, with no reference to anything in the scene. It works on a whole solid scene at once, because a plane change is scene-independent.
Rectifying takes a photographed plane and produces a picture of it from square on: four correspondences, or a vanishing line. It works for one plane at a time, and applying a façade’s rectification to the rest of the scene produces nonsense, because the map is only correct for points on that plane.
So a corrected photograph of a building is still a photograph of a building — everything in it, at every depth, correctly placed. A rectified photograph of one façade is a picture of that façade and of nothing else.
The wrong-focal-length failure, and how to catch it
The free parameter above deserves its own procedure, because it is the one way a corrected photograph can be quietly wrong and the way to catch it is cheap.
The visible criterion — “are the verticals parallel?” — is satisfied by a one-parameter family of maps, one for each value of . Every member straightens; only one is the plane change. So the criterion cannot select, and a reader who applies the slider until the building looks upright has selected nothing.
The test is a known rectangle. Anything in the scene whose true proportions are known — a window, a paving slab, a door — has a shape after correction that depends on . Correct with the right focal length and its aspect ratio comes out right; correct with a wrong one and it is stretched horizontally or vertically by a factor that grows with the tilt being undone.
Or a second family of verticals at a different depth. The correct map straightens all of them; a wrong one straightens the family it was fitted to and leaves the others slightly convergent, because the map’s error is depth-independent in the picture and the two families sit at different image positions.
Both are instances of the site’s standard defence: check the construction against something it was not given. The figures here use a level camera, which a photograph does not have; a real photograph has rectangles and second families, which serve the same purpose.
What the slider is actually doing
Compressed, for anyone who has one open:
It rotates the picture plane about a horizontal axis through the eye. The scene, the eye and the rays are untouched.
It is exact. Not a fit, not an approximation, not a warp chosen to look right — a homography with three of its degrees of freedom pinned by the vanishing point and the calibration.
It cannot invent information. Not the view from anywhere else, not the parts of the building the eye could not see, not a depth the picture never had.
It is reversible. A homography has an inverse, so the original picture can be recovered from the corrected one exactly — up to whatever was cropped away and whatever resampling cost. Nothing has been destroyed except frame and detail.
And it needs the focal length. Which most editors take from the file’s metadata, and which is why a corrected picture from a scanned print — no metadata, no calibration — is straight and is not the picture a level camera would have made.
One more consequence, for anyone reading a corrected picture rather than making one. The straightening is invisible in the result — there is no mark on a corrected photograph saying it was corrected, and the geometry it now shows is a geometry the camera never recorded directly. What it is recording is a projection from the original eye onto a plane nobody chose at the time, which is a perfectly good picture and is not the picture that was taken.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A projection of a projection — both name cross ratio, homography, picture plane, projective map, rectification
- A shadow can be un-cast — both name centre of projection, cross ratio, homography, projective map, rectification
- Recovering the camera from the picture it drew — both name focal length, homography, picture plane, principal point
- The divide is postponed, not avoided — both name centre of projection, focal length, principal point, projective map
- The eye is a place, not a point — both name centre of projection, focal length, homography, principal point
- The floor that is not a plane — both name homography, picture plane, projective map, rectification
Named objects
A flat tag is an object no other essay names yet.
centre of projectionCross ratioFocal lengthHomographyKeystonePicture planePrincipal pointProjective mapRectificationvertical vanishing point