What survives

One conic calibrates the camera

This site has recovered a focal length from two perpendicular vanishing points since its first phase, by an orthocentre construction with a square root in it. There is a second derivation with no construction and no square root — two vanishing points of perpendicular directions must be conjugate with respect to one conic in the picture — and the two agree to the last bit. They are not two methods. The conic is what a calibrated camera is.

Worth reading first: The two points a picture hides · Recovering the camera from the picture it drew · An angle is a cross-ratio.

The foundation phase of this site recovers a camera from the picture it drew. Three mutually perpendicular vanishing points are read off the drawn edges of a box, the orthocentre of the triangle they form is the principal point, and the focal length falls out of a construction involving that triangle’s altitudes. It returns the focal length the camera was built with to one part in 101510^{15}.

That construction is correct and it is unmotivated. It works, the derivation can be followed, and nothing in it says why an orthocentre should have anything to do with a lens.

There is a second route with no construction in it at all, and following it explains the first.

One conic, and the focal length falls out of itThe image of the absolute conic for a camera with square pixels is a circle of radius f about the principal point. Two vanishing points of perpendicular directions must be conjugate with respect to it, and solving that for f gives 812.769 px — the same number the orthocentre construction gives, and 1.1e-13% from the focal length the camera was built with.horizonprincipal pointv_zorthocentre: 812.7691 px · vᵀωu = 0: 812.7691 pxconjugacy residual 5.9e-10 in focal-length unitscorrect from 19 cm, at 160 mm wide46° across
Fig. 1 The image of the absolute conic for a camera with square pixels: a circle of radius ff about the principal point. Two vanishing points of perpendicular directions must be conjugate with respect to it, and solving that for ff gives 812.7691 px — the same number the orthocentre construction gives, to the last digit.

The conic

The plane at infinity carries a conic on it — an imaginary one, with no real points — called the absolute conic. Its defining property is the one that makes it useful: every rigid motion of space fixes it.

That is a strong statement. Move the camera anywhere, point it anywhere, and the absolute conic is unmoved, because it is a property of space’s metric structure rather than of anything in space. So its image in a photograph depends on the camera’s internals and on nothing else — not on where the camera is, not on where it points, not on the scene.

That invariance is the whole reason calibration is possible. A quantity depending only on the internals can be measured from any picture, and pictures of different scenes from different places all measure the same thing.

Written out, the image of the absolute conic is

ω=(KKT)1\omega = (\mathbf{K}\mathbf{K}^{\mathsf T})^{-1}

where K\mathbf{K} is the calibration matrix. For the pinhole this site builds — square pixels, no skew, the principal point where the camera says it is — that comes out as a circle of radius ff centred at the principal point, which is why it can be drawn.

The condition

Two directions in the world are perpendicular. Their vanishing points are the images of their points at infinity, so both lie on the image of the plane at infinity, and their perpendicularity is a statement about the absolute conic: two perpendicular directions are conjugate with respect to it.

Conjugacy carries through the projection unchanged, so in the picture

vTωu=0\mathbf{v}^{\mathsf T}\,\omega\,\mathbf{u} = 0

for the vanishing points of any two perpendicular directions. One equation, one unknown when the principal point is known, and it solves in a line:

f2=(vp)(up)f^2 = -(\mathbf{v} - \mathbf{p})\cdot(\mathbf{u} - \mathbf{p})

No construction, no altitudes, no orthocentre. The focal length is the square root of minus the dot product of the two vanishing points about the principal point.

The minus sign is the whole content

That expression has a sign in it and the sign is not decoration.

f2f^2 is positive, so the dot product must be negative: the two vanishing points have to lie on opposite sides of the principal point in the direction joining them. If they do not, no focal length makes those two image directions perpendicular, and the equation returns an imaginary ff.

An imaginary focal length is a refusal, not a small error. A camera recovered from a pair that fails the test is not a slightly wrong camera; it is a statement that the pair cannot be the images of perpendicular directions under any pinhole with that principal point. The machinery refuses rather than taking the square root of a negative number and returning its magnitude, and the site’s gate checks both halves — a pair that straddles is accepted, a pair that does not is refused.

That refusal is the thing the orthocentre construction expresses awkwardly. The altitudes of the vanishing triangle meet inside the triangle for an acute triangle and outside it for an obtuse one, and the case distinction the construction has to make is this sign.

What conjugacy means, concretely

The word conjugate is doing the work in the condition and it is worth unpacking, because it turns the algebra into a picture.

A conic sets up a correspondence between points and lines: the polar of a point is the line of points harmonically conjugate to it with respect to the conic, and this site met that correspondence two phases ago when the pole of the horizon turned out to be the image of a circle’s centre, at 2.5e-13 px.

Two points are conjugate when each lies on the other’s polar. So the condition vTωu=0\mathbf{v}^{\mathsf T}\omega\mathbf{u} = 0 says: the vanishing point of one direction lies on the polar of the vanishing point of the other, with respect to ω\omega.

That is a drawable statement. Given ω\omega and one vanishing point, its polar is a line in the picture, and the vanishing point of every direction perpendicular to the first lies on it. The polar of v\mathbf{v} is the vanishing line of the plane perpendicular to v\mathbf{v}'s direction — which is a fact about the picture that can be constructed with a straightedge once the conic is drawn.

For a square-pixel camera and a vanishing point at distance dd from the principal point, the polar is a line perpendicular to that direction at distance f2/df^2/d on the far side. So the horizon of the ground plane and the vertical vanishing point are pole and polar of each other, which is a relation between two things this site draws in almost every figure and had never connected.

The two agree

The site’s own recovery, written four phases and one field before any of this, is compared against the conic route directly.

At the camera drawn above, the orthocentre construction returns 812.7691 px and vTωu=0\mathbf{v}^{\mathsf T}\omega\mathbf{u} = 0 returns 812.7691 px. The relative difference is under 101510^{-15}, which is the arithmetic floor, and the recovered camera’s own ω\omega puts the pair in conjugate position with a residual of 5.9e-10 in focal-length units.

Two derivations written far apart, agreeing to the last bit, is the evidence that the conic is a re-description rather than a second method. Had they been two methods there would be a difference somewhere in the tenth digit — a different linearisation, a different ordering of operations — and there is not.

A box drawn from a known camera, and the camera recovered from the drawingThree vanishing points found from the twelve drawn edges alone give back the focal length to 2e-15 relative.recovered principal pointused to drawrecoveredgapfocal length947.88947.882e-15principal x345.0345.02e-12angle40.0°40.0°correct from 22 cm, at 160 mm wide40° across
Fig. 2 The construction this replaces the motivation for. Twelve drawn edges of a box give three vanishing points, and the focal length comes back to one part in 101510^{15} — which is the same statement as the conjugacy above, arrived at with a compass.
The three vanishing points of one box, drawn to scale with the boxThe picture is the small rectangle. Two of the three vanishing points fall well outside it, which is why they are computed rather than located by eye.the pictureVP₁VP₂VP₃orthocentrefocal length from the triangle — 707.4 pxspread 0e+0% across three routes
Fig. 3 The three points the construction reads. Any two of them being conjugate with respect to ω\omega is one equation; three mutually perpendicular directions give three equations, which is what over-determines the principal point as well.

What the conic buys that the construction does not

If it were only a tidier derivation it would be worth a paragraph rather than a rung. It is not.

A calibrated camera supplies every plane’s circular points at once.

The circular points of a plane are where that plane meets the absolute conic. So their images are where the plane’s vanishing line meets ω\omega — one line-conic intersection per plane, needing no circle in that plane, no known angle in it, and no further information from the picture at all.

Calibrate once, from any three perpendicular directions anywhere in the picture, and every plane in it becomes metric. The ground, each wall, the roof, a tilted table: each has a vanishing line, and each vanishing line cut by ω\omega gives that plane’s two points.

Checked directly: taking ω\omega from the camera’s own focal length, intersecting it with the ground plane’s horizon, and running Laguerre’s formula on two ground lines returns 90.000000000° for a right angle, with a residue of zero, and no circle anywhere in the construction.

Two points, and everything metric followsThe imaged circular points are where the horizon meets the image of any circle in the plane, and they are a conjugate pair — the first coordinate here is 169.5 − 446.0i. A rectification built from them and nothing else returns the world's angles to 2.2e-13° and its length ratios to 5.6e-15, and no length at all.horizonthe horizon does not cut the circle — the pair is complexrectified from the two points aloneangles: 2.2e-13°ratios: 5.6e-15length: —circle of radius 1.05 ma dash is a quantity two points cannot buy
Fig. 4 The route this replaces. An imaged circle plus the horizon gives one plane’s pair; ω\omega plus the horizon gives the same pair, and gives every other plane’s too.
Two ground lines at 90°, and the picture says soThe two lines cross at 80.43° on the paper. Taking the cross-ratio of the pair with the two lines from their crossing point to the imaged circular points, and halving the logarithm's imaginary part, returns 90.000000° — the angle in the world, with no rectification anywhere.horizonv_zthe horizon misses the imaged circle, so the two points are a conjugate pairprotractor on the paper: 80.43° · cross-ratio: 90.000000°correct from 21 cm, at 160 mm wide42° across
Fig. 5 And what the pair is for. With the two points located, a world angle is a cross-ratio in the picture — so a calibrated camera turns every plane in a photograph into a protractor.

The counting, and why three directions is the right number

ω\omega is a symmetric 3×3 matrix up to scale: five degrees of freedom, matching a general calibration matrix’s five. Each pair of perpendicular directions gives one linear equation in its entries.

So five independent perpendicularity constraints determine a general camera, and fewer determine a constrained one. The usual constraints and their costs:

  • zero skew removes one, leaving four;
  • square pixels removes another, leaving three — ff, pxp_x, pyp_y;
  • the principal point at the centre removes two more, leaving one, which is ff alone.

Three mutually perpendicular directions give three pairs, hence three equations, which is exactly enough for the square-pixel case and is what the orthocentre construction uses. That is why a box works and a single rectangle does not: a box has three directions and a rectangle in one plane has two.

The accounting also says what a fourth direction would buy — nothing, unless it is not perpendicular to the others, in which case it buys a residual. And a residual is what turns a recovery into a measurement, which is the difference between a calibration that returns a number and one that returns a number with something to say about it.

Where the principal point comes in

The focal formula above assumes the principal point is known, and this site has a whole essay on the fact that it usually is not — the principal point is not the centre of the image, and assuming it is costs a measurable amount.

The conic form says exactly what the assumption is doing. ω\omega for a square-pixel camera has three unknowns: ff, pxp_x, pyp_y. Each pair of perpendicular directions gives one linear equation in the entries of ω\omega. So three mutually perpendicular directions give three equations and three unknowns, and the system determines all three — which is what the orthocentre construction is doing when it uses the whole triangle rather than one pair.

Assuming the principal point is the image centre spends two of those equations on nothing and leaves one over-determined. The cost of the assumption is then measurable as the residual of the equations it did not need to make, and that is a considerably more informative thing to report than a focal length with no error bar.

One conic, and the focal length falls out of itThe image of the absolute conic for a camera with square pixels is a circle of radius f about the principal point. Two vanishing points of perpendicular directions must be conjugate with respect to it, and solving that for f gives 1383.719 px — the same number the orthocentre construction gives, and 5.4e-13% from the focal length the camera was built with.horizonprincipal pointorthocentre: 1383.7194 px · vᵀωu = 0: 1383.7194 pxconjugacy residual 4.3e-10 in focal-length unitscorrect from 32 cm, at 160 mm wide28° across
Fig. 6 A longer lens. ω\omega is a larger circle and the two vanishing points move outward; the conjugacy holds at every setting, which is the invariance the whole construction rests on.
One conic, and the focal length falls out of itThe image of the absolute conic for a camera with square pixels is a circle of radius f about the principal point. Two vanishing points of perpendicular directions must be conjugate with respect to it, and solving that for f gives 474.852 px — the same number the orthocentre construction gives, and 1.2e-14% from the focal length the camera was built with.horizonprincipal pointv_zorthocentre: 474.8518 px · vᵀωu = 0: 474.8518 pxconjugacy residual 3.5e-10 in focal-length unitscorrect from 11 cm, at 160 mm wide72° across
Fig. 7 And a much wider one. The circle shrinks toward the principal point and the vanishing points close in on it — and the pair still straddles, which is what makes the recovery possible at all.

Why it is a circle here and not in general

The figure draws ω\omega as a circle, and that is a property of the camera rather than of the conic.

For a camera with square pixels and no skew, KKT\mathbf{K}\mathbf{K}^{\mathsf T} is f2f^2 times the identity in its top-left block, so ω\omega is a circle of radius ff about the principal point. Give the camera a non-square pixel aspect and ω\omega becomes an ellipse whose axis ratio is the pixel aspect; give it skew and the ellipse tilts.

So the shape of ω\omega is the camera’s internal geometry, drawn in the picture. A photograph with a known ω\omega is a photograph carrying its own calibration on it, and the number of degrees of freedom in a conic — five — is exactly the number in a general K\mathbf{K}.

That correspondence is the reason the object is worth having rather than the formula. Five numbers in a matrix nobody can picture become one conic that can be drawn on the photograph, and the constraints turn from algebra into incidences.

Squeeze the pixels and every alarm this site has stays silentThe same box, drawn by a camera whose pixels are 1.12 times wider than they are tall. The recovery is shown the drawn edges and nothing else; it returns a focal length of 390.5 px against a true 436.4, a principal point 13.0 px away — and a spread across its three independent estimates of 1.5e-16, with bundle residuals of 2.6e-13 px.pixels 1.12 : 1recovered focal length390.5 pxthe camera's actual one436.4 pxwhat the recovery returnsspread across three estimates 1.5e-16worst bundle residual 2.6e-13 pxboth are what a wrong picture would trippixel aspect 1.12 unmodelledfocal length 10.5% short, every diagnostic clean
Fig. 8 What a non-square pixel does, from the sensor field. An unmodelled pixel aspect is absorbed into a confident wrong camera — which in this language is ω\omega being an ellipse and being fitted as a circle.
The cost of assuming the principal pointThe two-vanishing-point route to a focal length needs a principal point supplied, and every textbook supplies the middle of the frame. On a shifted or cropped picture that is wrong, and f² = −(v₁ − p)·(v₂ − p) turns the error into a product of two large numbers: 1.4% at 150 px of shift. The three-point recovery does not assume it and has no such error.00.50011.50050100150how far the principal point really is from the middle of the frame (px)error in the focal length, from assuming it is not (%)unshifted: the assumption is truea 42° lens1.40% at 150 px of shift
Fig. 9 And the cost of assuming the principal point. The conic form says why it is a cost rather than a convenience: it is two of the three equations thrown away.

What it does not do

Two limits.

It needs perpendicular directions in the scene, or something equivalent. The conic is a fact about the camera and the picture only measures it where the scene provides a constraint. A photograph of a featureless plane constrains nothing.

It says nothing about lens distortion. ω\omega is the image of the absolute conic under a pinhole, and a real lens is not one. Every claim here survives with the distortion removed and none of it survives with the distortion present, which is what the lens field’s fitting essays are for.

Projective, affine, metric — what each stage buysThe photograph fixes the plane only up to a projectivity: the midpoint of a receding side lands 0.3970 of the way along. Supplying the plane's vanishing line buys the midpoint back exactly and nothing else. Supplying the image of one circle buys the last three numbers, at which point the right angle is 90.000° and two equal sides measure 1.000000. The cross-ratio is 1.333333 in all three, because it was never lost.projectiveaffinemetricmidpointtwo equal sidesa right anglecross-ratioprojective1.333333333affine0.5000001.333333333metric0.5000001.00000090.000°1.333333333— means the stage does not determine it at allcross-ratio 1.333333 throughout
Fig. 10 The ladder once more, now with the top rung’s price named twice over: two points per plane, or one conic for all of them.
Halving a receding rectangle two waysThe diagonals cross at the image of the rectangle's centre, 8e-14 px from it — the construction is exact at every camera because it uses only which lines meet where, and that is what a projection keeps. Halving the drawn side with a ruler instead lands 22.7 px from the image of the side's midpoint.the diagonals against a ruler, at 3.2 mthe diagonals — exactthe ruler — 22.7 px outcorrect from 23 cm, at 160 mm wideharmonic set -1.000000 · 8e-14 px
Fig. 11 And the projective machinery underneath, which none of this replaces. Conjugacy with respect to a conic is a pole-and-polar statement, and the harmonic relation is where this site first met one.

One object, four fields

It is worth listing where this conic has already been on this site under other names, because the list is the argument for the object.

The foundations field’s focal recovery is the conjugacy condition. The foundations field’s circular points are ω\omega intersected with a plane’s vanishing line. The metrology field’s rectification is the metric structure those points supply. The sensor field’s pixel aspect is ω\omega being an ellipse rather than a circle. And the manyviews field’s calibration is ω\omega estimated from more constraints than it needs.

Five constructions, written in five phases, each with its own derivation and its own figure, and one object underneath. That is the same shape of finding as the previous phase’s — three thin rows turning out to be one subject at three group levels — and it arrives the same way: by asking what the constructions are rather than what they compute.

The practical dividend is the one already stated and it is worth repeating in this company. Any one of the five determines ω\omega; ω\omega determines all five. So a picture with a single calibrated camera has every plane in it metric, every angle in every plane readable as a cross-ratio, and every rectification available without a further reference — and none of that is visible from any of the five constructions taken alone.

The site's camera, written as the matrix multi-view geometry needsK holds the focal length and the principal point; R's rows are the camera basis — right, down, forward — and t is −R·eye. Projecting all 44 scene points through P = K[R|t] and through the camera itself gives the same picture to 1.8e-13 px. Everything in this field rests on the two being one camera, so it is measured rather than assumed.K — focal length and principal point739.90345.00739.9200.0001.0000R — right, down, forward0.980000.1991-0.0182-0.99580.08970.1982-0.0915-0.9759t = −R·eye00.99586.6497focal 739.85 px · 50.0° acrossP projects 44 points where the camera does, to 1.8e-13 pxcorrect from 17 cm, at 160 mm wide50° across
Fig. 12 The pipeline field’s version of the same identification, from a different direction: the machine’s projection matrix projects every scene point where this site’s own camera does, to 1.8e-13 px, before anything is claimed about how the two differ.
A circle on the ground, and the two points that get called its centreThe image of the centre and the centre of the image ellipse are 21.1px apart — 5.1% of the ellipse's own width.centre of the ellipseimage of the centrecorrect from 22 cm, at 160 mm wideoffset 21.1 px
Fig. 13 And the conic this field first met, whose pole-and-polar relation is the one the conjugacy condition uses.

The short version

The absolute conic is fixed by every rigid motion, so its image depends on the camera’s internals and on nothing else. For a square-pixel camera that image is a circle of radius ff about the principal point.

Two vanishing points of perpendicular directions must be conjugate with respect to it. Solving that gives f2=(vp)(up)f^2 = -(\mathbf{v} - \mathbf{p})\cdot(\mathbf{u} - \mathbf{p}), whose minus sign is a refusal rather than a convention, and whose answer agrees with this site’s orthocentre construction to the last bit.

And it buys something the construction never offered: the circular points of every plane in the picture, as the intersection of that plane’s vanishing line with the conic. Calibrate once and every plane becomes metric, with no circle in any of them.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Absolute conicCamera calibrationCamera matrixCircular pointsConicDemonstrationFocal recoveryHomogeneous coordinatesline at infinityMetric rectificationOrthocentrepole and polarPrincipal pointVanishing pointvertical vanishing point