Measuring from one picture

Two pictures of a ball

Two outlines give two tangent cones, and where their axes cross is the ball's centre — with the radius following from either half-angle. There is no correspondence problem, because a ball's outline in one picture and its outline in another are guaranteed to be the same ball; and there is a degeneracy, which is the two eyes and the ball in a line, and the solver refuses it rather than returning something plausible.

Worth reading first: One picture of a ball · Two rays that do not meet · The one thing a single view cannot give.

One picture of a ball gives a direction and a ratio: the tangent cone’s axis is the direction of the centre, exactly, and its half-angle is the arcsine of radius over distance, exactly. Neither the radius nor the distance separately.

A second picture removes that, and the arithmetic is three lines.

Two pictures fix the ball, and the baseline is the conditioningTwo pictures of one ball give two tangent cones, and where their axes cross is the ball's centre — from which the radius follows from either half-angle. The angle the axes cross at is the whole of the conditioning, and it is what the baseline buys: 42.6° at 2.82 m here, against 2.38° at 0.15 m. At zero baseline the two axes are one line and the intersection is not a point at all — the solver refuses rather than returning something plausible, which is the behaviour a recovery should have at a degeneracy and usually does not.020406012345the baseline between the two eyes (metres)angle between the cone axes (°)40.3°two cones, one ballcentre to 9e-7°
Fig. 1 The angle the two cone axes cross at, against the baseline between the two eyes. That angle is the whole of the conditioning, and at zero baseline the two axes are one line.

Two cones, two axes, one point

Each picture fixes a cone through its own eye. Each cone’s axis is a line through that eye toward the ball’s centre. Two lines through two known points, both passing through the same unknown point: the intersection is the centre.

Then the radius comes from either half-angle — r=dsinθr = d \sin\theta, with dd now known — and the two pictures should agree about it, which is a check rather than an average.

The centre comes back at the arithmetic floor and the radius to nine decimal places. There is nothing clever in it; the cleverness was all in the previous rung’s step from a conic to a cone.

Two rays, 5.07 mm apart, in the plane that contains bothThe ray from the left eye through its mark and the ray from the right eye through its. With the marks placed exactly they meet, to 1.9e-15 m. With the same marks read to 1 px they miss by 5.07 mm at a range of 7.19 m. Triangulation is not an intersection; the reported point is a choice about what to minimise, and the gap is the part a residual alone will not tell you.midpoint — 5.07 mm gapfrom the left eyefrom the right eyegap 5.07 mm at 7.19 mexact marks: 1.9e-15 m
Fig. 2 The solver: two rays from two places and the point they come closest to. On exact data it is an intersection, and it is the same routine the refraction, scroll and panorama fields use on bundles that do not meet.

There is no correspondence problem, and that is unusual

An ordinary two-view reconstruction spends most of its effort on matching: finding the same feature in both pictures. A wrong match is not a small error — it puts a point somewhere else entirely — and every practical system is mostly machinery for not making one.

Here there is nothing to match. A ball’s outline in one picture and its outline in another are outlines of the same ball by construction; the two cones are two constraints on one object, and no point of either outline has to be paired with any point of the other.

That is a real advantage and it is worth understanding where it comes from. It is not that the object is easy to detect; it is that the object is the feature. A sphere is determined by four numbers and the recovery works on the object rather than on points of it, so the correspondence question never arises.

One correspondence moved 30 px, and where the damage wentThe clean fit is exact to 4.3e-13 px. Moving correspondence 9 by 30 px leaves every other point wrong too — the typical one by 0.49 px and the worst by 1.6 px — because a least-squares fit has nowhere to put a bad row except across all of them. Here the largest residual does fall on the culprit; it is not obliged to.00.50011.50010203040correspondenceepipolar error at every OTHER point, after one match is movedmedian 0.49 pxthe moved oneevery point wrong: median 0.49 px, worst 1.6 pxclean fit 4.3e-13 px
Fig. 3 What is avoided, from the depth field. A single wrong pairing in a point-based reconstruction produces a point with nothing wrong-looking about it.

The same property is why the wall under the paint works with one extra view: there the correspondence is given by the design, here it is dissolved by the object having no parts. Two different ways of not having a matching problem, and both of them are why those recoveries are exact where a general one is statistical.

The conditioning is one angle and it is the baseline

Two lines meet well when they cross squarely and badly when they are nearly parallel. Everything about how well this recovery behaves is in that one angle.

At a baseline of a few metres against a ball a couple of metres away, the axes cross at a comfortable angle. Bring the eyes together and the angle collapses toward nothing.

Two pictures fix the ball, and the baseline is the conditioningTwo pictures of one ball give two tangent cones, and where their axes cross is the ball's centre — from which the radius follows from either half-angle. The angle the axes cross at is the whole of the conditioning, and it is what the baseline buys: 10.3° at 0.65 m here, against 2.38° at 0.15 m. At zero baseline the two axes are one line and the intersection is not a point at all — the solver refuses rather than returning something plausible, which is the behaviour a recovery should have at a degeneracy and usually does not.020406012345the baseline between the two eyes (metres)angle between the cone axes (°)10.3°two cones, one ballcentre to 9e-7°
Fig. 4 A short baseline. The two axes are nearly the same line, and the point where they cross is poorly determined along that line — which is the ordinary depth-from-disparity behaviour with the disparity carried by a cone instead of a feature.

Which is exactly the depth field’s own result in a different costume: depth from two views is a reciprocal, so the uncertainty along the line of sight grows as the square of the distance and falls as the baseline. Nothing about the object being a ball changes that.

What does change is where the angular error comes from. In a point-based reconstruction it is the accuracy of a pixel; here it is the accuracy of a fitted conic, which uses hundreds of edge pixels at once and is correspondingly better. So a ball is a well-conditioned target for a given baseline, and badly conditioned for a short one, in the ordinary way.

Disparity to depth, and the wall the reader's own head puts inz = b·D/(b − d). Zero disparity puts the point on the screen at 4.0 m; crossed disparity brings it forward; and at d = 63 mm — the separation of the eyes — the point reaches infinity. Past that the display is asking the eyes to diverge, which they cannot do, so the depth budget is set by the width of the reader's head and by nothing about the scene.010203040-2502550disparity on the screen — millimetreswhere the point is depicted — metres from the eyesd = 63 mm — the eyes' separationon the glassscreen at 4.0 m · eyes 63 mm apartthe ceiling is the head, and it does not move when the screen does
Fig. 5 The reciprocal itself. A short baseline and a distant object give a shallow crossing, and a shallow crossing gives a long thin uncertainty along the line of sight.

The other route, and why it is not taken

There is a second way to fix the ball from two views and comparing them says what the cone route is for.

The other way is to use the two outlines’ extreme points. A ball’s outline has a leftmost and a rightmost point in each picture; take those four rays and solve for a sphere tangent to all of them. It is a determined problem and it works.

It also throws away almost everything. Each outline has hundreds of edge pixels and the conic fit uses all of them; picking four extreme points uses four, and picks the four whose position is least well determined, since an outline’s extremes are where the edge runs parallel to the direction being measured and the gradient across it is weakest.

So the cone route is not chosen for elegance. It is chosen because the conic is the sufficient statistic — every edge pixel of the outline contributes to it, and nothing about the ball that the picture knows survives outside it.

That is a small instance of a general preference this site has and it is worth naming. Fit the object the data is about, not a handful of features on it. The four-point predictor in the light field does the opposite deliberately, to separate a fit from a prediction; here there is nothing to separate, and using everything is right.

Five marks fix the conic, and the sixth is a predictionFive marks on a photographed circle determine one conic — five points and five coefficients, with no fitting left over. The sixth mark was withheld from the fit and the conic passes 5.5e-14 px from it. Nothing about the camera, the circle's size or the plane it lies in was used.horizon5withheldfive marks fitted, one withheldcorrect from 26 cm, at 160 mm widethe withheld mark is 5e-14 px off the fitted conic
Fig. 6 The conic as a determined object, from the foundations field: five marks fix it and a sixth is predicted. An outline supplies hundreds, which is why the fit is a measurement rather than a construction.

The degeneracy, and the refusal

Put the two eyes and the ball in a line. Both cones then have the same axis — the line joining everything — and two coincident lines determine no point.

The solver refuses. closestPointToRays asserts that the bundle spans enough directions to pin a point, and a pair of parallel rays fails that assertion rather than returning the nearest thing to an answer.

A point on the left is a line on the rightFive points in the left picture. Each one fixes a line in the right picture — the image of the ray it came along — and its match lies on that line to 2.8e-13 px. The five lines meet at the epipole, 2.9e-10 px from concurrent. Knowing where a point is in one picture does not say where it is in the other; it removes one of the two degrees of freedom.1234512345a point is chosen hereand must be on this linematch to its own epipolar line: 2.8e-13 px5 of 44 correspondences drawn
Fig. 7 The same degeneracy in the twoviews field: a point on the line joining the two eyes has an epipolar line that collapses, and the constraint that usually locates it says nothing.

That refusal is worth having explicitly rather than as a numerical accident. A least-squares solver handed two nearly-parallel rays returns a point — some point, far away, with an enormous covariance nobody looked at — and a solver handed exactly parallel rays either returns a NaN or divides by something small and returns a very large number. Both of those are worse than refusing, because both of them are answers.

And the degenerate arrangement is not exotic. Two photographs taken from the same spot at different times, or a stereo pair pointed at something on its own baseline, are both this. It is the arrangement a photographer produces by accident.

The degeneracy is in the reading, not in the algebraShape error against baseline, twice. Given exact correspondences the recovery is good to 4.7e-9 at a 2 mm baseline and improves from there — the algebra does not notice a short baseline at all. Read the same points to 1.18 px and the same geometry comes back 1e+9× further from the truth, and it does not refuse: it returns a confident wrong answer. What fails when a photographer swings a camera instead of stepping sideways is the ratio of baseline to reading error, and nothing in the computation reports it.-10-50-2-10baseline (m, log scale)worst deviation in the recovered shape (log scale)exact pointsread to 1.18 pxat 2 mm: 4.7e-9 exact, 6.7e+0 read to 1.18 px1e+9× apart
Fig. 8 The other arrangement that gives nothing, from the depth field: a turn of the head is not a step sideways, so a rotated pair has no baseline and no depth at all.

What two pictures still do not give

They do not give any point of the ball’s surface.

That sounds like pedantry and it is the thing that stops this generalising. Two outlines are two curves: the contour generator in one view is a different curve on the object from the contour generator in the other, meeting it at exactly two points. So the recovery above never learns where on the ball anything is.

For a sphere that costs nothing, because a sphere has a centre and a radius and no other shape. For any other surface it is the whole difficulty, and it is what the visual hull is about — a reconstruction from outlines converges to something that is not the object, and the part it never reaches does not shrink with the number of views.

Every outline this object has, from 8 directionsA cross-section with a bite taken out of it, and the region every one of 8 outlines admits. Each view contributes a pair of supporting lines — an outline is a pair of numbers per direction and nothing more — and the visual hull is what is left when all of them have cut. It has area 3.1098 against the object's 2.9193. Adding views shaves the corners between the tangents, and no number of them touches the notch: the convex hull, at 3.0749, is the floor. The bite is 5.3% of the object and it is invisible to outlines from every direction at once.hull 3.1098 · object 2.9193 · floor 3.07498 outlinesthe notch survives all of them
Fig. 9 Where this ends up for an object that is not a sphere: the region every outline admits, which contains the object and is not it.

So the two rungs together say something narrow and exact. A ball is completely recovered by two views — four numbers, four constraints, no ambiguity but the ordinary conditioning — and the reason is that a ball has no shape beyond its size and place. Everything harder than a ball needs more than outlines.

Why this arrangement is used to calibrate cameras

The recovery has a practical relative worth naming, because it explains why spheres turn up in calibration rigs at all.

Run the problem the other way. Put a ball of known radius in front of two cameras whose poses are unknown, photograph it in several places, and the outlines constrain the cameras rather than the ball. Each outline is a conic with five degrees of freedom; each ball position adds three unknowns and each view of it gives five constraints, so the count comes out in the calibration’s favour quickly.

What makes a sphere attractive for that is exactly what makes it useless as a measurement in one view: it has no orientation. A checkerboard has to be seen well enough for its corners to be located and its plane’s pose recovered, and it fails at oblique angles and near the edges of a wide frame. A ball looks the same from everywhere, so it can be moved anywhere in the field of view — including the corners, where a lens’s distortion is largest and where a plumb-line fit most needs data.

So the object’s poverty is its usefulness. A sphere carries no information about its own orientation, which means every picture of it carries the same kind of information, which means the pictures can be combined without any of the pose bookkeeping a planar target needs.

k₁ recovered from 5 bent lines and nothing elseThe fit is never shown the coefficient, the camera or the scene — only which sets of points came from straight edges. It returns 0.020000000 against a true 0.020000, off by 6e-16, and straightens its own input to 6e-14 px.fitted k₁ = 0.020000true 0.020000, off by 6e-16
Fig. 10 The calibration this feeds, from the lens field: a distortion recovered from straightness alone, which needs marks near the frame’s edges and gets them badly from a flat target.

And the scale is not free either

There is a caveat that the arithmetic hides and it is the site’s usual one.

The centre and the radius come out in whatever units the baseline is given in. Two pictures with an unknown baseline give the ball’s shape — which for a sphere is nothing at all, since every sphere is similar to every other — and its position relative to the two eyes, up to one overall factor.

Which is two views give shape and no size, arriving on an object where the “shape” half is vacuous. So a ball is the extreme case: two uncalibrated views of one give literally nothing but a direction pair, and the whole content of the recovery is bought with the baseline.

One reconstruction, drawn at its own scale and at 3.4×The courtyard recovered from two pictures, seen from above. Every pairwise distance ratio matches the world's to 8.7e-14, so the shape is exact. The size is not determined at all: the right-hand plan is the same reconstruction 3.4 times larger and fits the same two pictures equally well. One measured length in the scene — here 0.800 m — fixes it, and nothing in the pictures can.as recovered0.261 across× 3.4 — same two pictures0.887 acrossworst ratio deviation 8.7e-14size fixed only by a supplied 0.80 m
Fig. 11 The general statement from the twoviews field. A ball’s version of it is stark, because a sphere’s shape carries no information and everything is in the scale.

Why the check is that the two half-angles agree

The recovery uses one half-angle to get the radius and has a second one spare. That spare is a check and it is the right kind.

Both cones are computed from their own pictures, with no reference to each other. The centre is found from the axes alone. So the radius from camera A’s half-angle and the radius from camera B’s are two independent estimates of one quantity, and their agreement is a statement about the whole chain — the conic fits, the calibrations, the two cone decompositions, and the intersection.

A recovery that used both half-angles to fit the radius would have thrown that away. Using one and checking the other is the same discipline the four-point predictor applies to a shadow: determine with exactly as much data as the problem needs, and spend the surplus on a test.

Two pictures fix the ball, and the baseline is the conditioningTwo pictures of one ball give two tangent cones, and where their axes cross is the ball's centre — from which the radius follows from either half-angle. The angle the axes cross at is the whole of the conditioning, and it is what the baseline buys: 59.7° at 4.15 m here, against 2.38° at 0.15 m. At zero baseline the two axes are one line and the intersection is not a point at all — the solver refuses rather than returning something plausible, which is the behaviour a recovery should have at a degeneracy and usually does not.020406012345the baseline between the two eyes (metres)angle between the cone axes (°)59.7°two cones, one ballcentre to 9e-7°
Fig. 12 A wide baseline, where the axes cross squarely. The two radii agree here for the same reason they agree at a narrow one — the check is on the machinery rather than on the geometry — and what changes is how precisely the centre is placed.

The four numbers, counted

A count is worth writing out because it is what makes “two views finish it” a statement rather than an observation.

A sphere in space has four degrees of freedom: three for the centre and one for the radius.

One picture of it, from a calibrated camera at a known place, gives a tangent cone. A cone through a fixed apex has three: two for the axis’s direction and one for the half-angle. So one view leaves one degree of freedom, which is the slide along the ray with the radius scaled to match — the family the previous rung draws.

A second view gives three more constraints on a problem with one unknown left. Two of them fix it and the third is surplus, which is the radius check.

That arithmetic is the reason a ball saturates so early, and it is also why the check exists at all: a determined problem with no surplus has nothing to verify itself against, and this one has exactly one spare number. Using it as a test rather than folding it into a fit is the choice that turns the surplus into information about the machinery.

6 pictures, 168 parameters, and seven of them freeThe smallest 16 singular values of the reprojection Jacobian, as fractions of the largest. Seven sit at below 10⁻⁸ and the eighth jumps to 1.3e-2 — a factor of 4.6e+6. Those seven are where the reconstruction is, how it is turned and how big it is: three, three and one. No number of further pictures removes any of them.× 4.6e+6seven flatand the rest stiffsingular value ÷ the largest, log scale, smallest firstσ₈/σ₇ = 4.6e+6168 parameters · 528 residuals
Fig. 13 The counting done at full scale in the manyviews field: seven directions no quantity of pictures fixes, found by walking along them rather than by arguing about them.

What a third picture adds

Nothing, for a ball, and that is worth saying because it is unusual.

Four numbers determine a sphere. Two views supply four constraints — two axes, each of which is two numbers’ worth of direction — with one surplus already. A third view adds three more constraints to a problem with no unknowns left, so what it buys is noise reduction and a check, not identifiability.

That is a different situation from a point cloud, where every extra view buys accuracy on every point, and from a camera track, where more views buy the track as well as the scene. A ball saturates at two.

5 pictures, one track and one scene, recovered togetherEvery point read to 1 px. Bundle adjustment moves all 162 parameters at once and lands at 0.318 px of reprojection error in 5 iterations. The recovered camera centres, drawn over the true ones after the similarity the pictures cannot fix is removed, agree to 1.4e-3 of the track's own mean radius. Nothing in the solve was ever told where a camera or a point was.view 1view 5the scene, in plan — recovered points and cameras over the true onesreprojection 0.318 px · track 1.4e-3220 observations, 162 parameters
Fig. 14 The general many-view problem from the manyviews field, where more pictures determine more. A sphere is the case that stops needing them.

What this does not settle

It does not treat a ball resting on a known plane, which is a different and better-posed problem: the plane supplies the scale, and one view is then enough — which is the same trade a height from one photograph makes.

It does not treat two views of a ball that is moving. Then the two outlines are of the same object in two places and the recovery is over-determined in the wrong way.

It does not treat uncalibrated cameras, which give conics and no cones, so nothing above starts.

And it does not treat the ball whose outline is partly hidden. A conic needs five points and an occluded arc may or may not supply them, which is a question the foundations field asks about marks rather than about arcs.

A second view is worth exactly what the first one was missing, and no more. On a ball that is the scale, so two views finish the problem; on anything with a shape, outlines never do, however many there are.

The other eye, found in the picture — 3.1 m apartTwo views of one courtyard. In each panel the arrow points at the epipole: the image, in this picture, of the eye that took the other one. It is computed from the 44 correspondences alone, as the null vector of a fundamental matrix that has never been shown a camera, and it lands on the projection of the other eye to 2.5e-9 px on the left and 2.5e-9 px on the right.epipoleepipoleleft pictureright pictureepipole from 44 correspondences vs the projected eye: 2.5e-9 px3.07 m between the eyes
Fig. 15 The structure a two-view arrangement always has, from the twoviews field: each eye appears in the other’s picture, and everything degenerate happens on the line joining them.
One picture of a ball, and the family it cannot separateThe room from above. The outline in the picture fixes the cone of rays that graze the ball: its axis is the direction of the ball's centre, to 8.5e-7°, and its half-angle is asin(r/d) — 6.1640° here, against 6.1640° from the ball itself. Neither the radius nor the distance appears separately anywhere in that. The three balls drawn are 1, 2, 6 times as far away and 1, 2, 6 times as large, and their outlines are the same conic to 8.5e-12 px. This is the site's one-view scale ambiguity, on an object where it is usually assumed away — and beside it a trap, because the drawn ellipse's own centre is 2.25 px from the image of the ball's centre and reading one for the other is a real error.the eyecorrect from 9 cm, at 160 mm wideoutlines agree to 8e-12 px
Fig. 16 And what the second picture removes: the one-parameter family the first one leaves, walked along with the outline unchanged.

What links here

Computed from the collection, not written here: the essays that point at this one.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Angular sizeBaselineCalibrationConditioningDemonstrationDepth from disparityEpipolar geometryleast-squares intersectionQuadricReconstructionscale ambiguityTangent coneTriangulation