Two views give shape and no size
Worth reading first: Four cameras fit, and one of them can see · Recovering the camera from the picture it drew.
The metrology field ended on a sentence that took a whole essay to earn: one picture supplies every ratio and no size. A world 137 times larger, photographed from 137 times further away, gives an identical picture down to the last bit, so no measurement made on the picture can distinguish them.
The natural expectation is that a second picture fixes it. Two eyes, after all, are how animals judge distance, and the whole apparatus of stereo photography exists on the premise that a pair of images measures depth.
The expectation is wrong, and it is wrong in an unusually clean way. A second picture is a large addition — it converts ratios in the picture plane into a full three-dimensional shape, which is a genuine gain. It supplies no length whatever. The scale ambiguity of a single view survives the addition of a second view completely intact.
Where the missing length went
The reason is visible in the pose recovery and is not subtle once seen.
The essential matrix factors as a rotation and a translation. Its five degrees of freedom are three for the rotation and two for the translation — a direction on the sphere — not three. The length of the baseline is simply not a parameter of the object being recovered.
That is not an accident of the parameterisation. Consider the two cameras and the courtyard, and scale the entire arrangement — cameras, baseline, courtyard, everything — by a factor of 3.4. Every ray from every eye to every point makes the identical angle it made before, because scaling about any origin preserves angles. A camera records directions. So both pictures are unchanged, pixel for pixel.
Two photographs of a courtyard therefore cannot be distinguished from two photographs, taken from 3.4 times as far apart, of a courtyard 3.4 times as large. The pictures are the same pictures. There is nothing to measure that differs.
The right-hand panel of the figure is exactly that: the reconstruction multiplied by 3.4, drawn beside the original. It is not an approximation of the pictures, or a fit that is nearly as good. It fits them exactly.
What “shape” means, precisely
The claim that the shape is recovered needs a definition, because “shape” is doing real work.
The reconstruction is compared with the world by taking every pair of points, forming the distance between them, and dividing by the distance between one chosen reference pair. That gives a list of ratios for the reconstruction and a list of ratios for the world, and the comparison is between the two lists.
Those ratios agree to 6.3 × 10⁻¹⁴.
Ratios are used rather than positions because positions are meaningless here. The reconstruction is written in the first camera’s frame and at an arbitrary scale, so a recovered point’s coordinates are a statement about the convention. Its distance to another recovered point, divided by a third such distance, is not — it is a property of the configuration that no choice of frame or scale can change.
A comparison that skipped that step would report metres of error on a perfect reconstruction, which is the standing trap of this whole subject and the reason reconstruct in lib/stereo.js computes pairwise ratios rather than differences.
So the precise statement is: the reconstruction and the world differ by a similarity — a rotation, a translation and a uniform scale — and by nothing else, to fourteen digits.
One view, two views, and what changed
Setting the two results side by side makes the gain visible, and it is worth doing because “shape and no size” sounds like a restatement of “ratios and no size” and is not.
From one view, what is recovered is ratios along a line — the cross-ratio is the invariant, and it lives on a line. Getting anything about a plane requires knowing the plane is a plane, which is what the rectifying homography uses. Getting anything about three-dimensional structure requires assumptions the picture does not contain: that a surface is flat, that a line is vertical, that a shape is a rectangle.
From two views, what is recovered is every distance ratio in three dimensions, with no assumptions about the scene at all. The courtyard was not assumed to contain planes, right angles or verticals. Forty-four points went in as pairs of pixel coordinates and came out as a three-dimensional configuration correct to fourteen digits up to a similarity.
That is an enormous gain. The whole of the scene’s geometry is present except one number.
And it is exactly one number. Not “a scale factor and some uncertainty” — one real number, completely undetermined, with everything else exact.
Supplying it
Fixing the scale needs one length measured in the world, and one is enough.
The figure quotes 0.800 m, which is the distance between two chosen reconstructed points, and applying the ratio between the true length and the recovered one rescales everything at once. Every other distance in the courtyard then comes out right, because the ratios were already right.
The economy of that is worth sitting with. A tape measure applied to one edge of one block converts a shape into a survey. Nothing else needs measuring, and measuring more would add precision rather than information — a second length is a check on the first, not an additional constraint on the geometry.
This is the same arithmetic the single-view field ran on, and it is worth being clear that the two are not the same statement even though they share the sentence. In the single-view case, the supplied length was needed to make any three-dimensional statement, and it came with structural assumptions — a vertical, a ground plane, a known camera height. Here it is needed only for the last scaling, and it comes with no assumptions at all.
There is a practical version of this that is worth stating plainly, because it is how the ambiguity is resolved in the field. A stereo rig with a known baseline is a rig where the supplied length is the distance between the two cameras, measured once when the rig was built. That is why a stereo camera can report depth in metres and two handheld photographs cannot: the metres come from the tape measure applied to the rig, not from the pictures.
The ladder of what each assumption buys
The conditions attached to the result are not all-or-nothing, and the intermediate cases are worth laying out, because they are the standard classification of what a reconstruction can be and each rung on the ladder is bought by one specific piece of knowledge.
With nothing but correspondences, the reconstruction is projective. The fundamental matrix is recoverable, and from it a pair of camera matrices, and from those a set of three-dimensional points that reproject correctly. But the reconstruction is related to the truth by an arbitrary projective transformation of space — sixteen numbers, fifteen after scale — so straight lines are still straight and planes are still planes and essentially nothing else survives. Parallel lines are not parallel. Right angles are not right. Distance ratios along a line survive, because the cross-ratio does, and distance ratios between different lines do not.
Knowing where the plane at infinity went upgrades that to affine. Parallel lines become parallel again, midpoints become midpoints, and ratios of lengths along parallel directions become meaningful. Angles are still wrong and so is any comparison between lengths in different directions.
Knowing the calibration — the focal length and the principal point — upgrades that to metric, which is where this essay has been all along. Angles are right, ratios of all lengths are right, and the only remaining freedom is a rotation, a translation and a scale. That is the similarity.
And one measured length removes the scale, leaving a rotation and a translation, which is a survey in a frame of its own choosing.
Setting them out that way makes the shape of the subject clear. Each rung is bought by an assumption imported from outside the pictures, and the assumptions get more specific as the answers get stronger. The reason this site’s figures sit at the metric rung is that the calibration is genuinely known — the cameras were built with a stated focal length — and the reason they stop there is that no length in the courtyard was ever supplied to them.
It is worth noticing which rung the classical single-view results occupy. The rectifying homography recovers a façade’s proportions by assuming the façade is planar and that four points on it form a known shape; that is buying a rung with a scene assumption rather than a camera one. The height from a cross-ratio buys its rung with a known eye height. The multi-view results here buy theirs with a known lens, which is a different currency and is the reason the two fields’ assumptions do not overlap.
The one thing more data does not do
There is a version of this result that sounds like a claim about precision and is not, and separating them matters because they behave differently.
Precision improves with data. Read the marks more carefully and the recovered shape gets closer to the true shape; the residual falls; the fourteen digits become fifteen. That is ordinary and expected.
The scale does not improve with data. It is not badly determined, or weakly determined, or determined with a large error bar. It is not determined. Doubling the number of correspondences changes nothing about it; doubling the precision of every mark changes nothing about it; adding a third and a fourth photograph changes nothing about it.
The distinction is testable and the figure tests it. If the scale were merely poorly determined, the reconstruction at 3.4× would fit the pictures worse — by a small amount, perhaps, but by something. It does not. The reprojection is identical, not nearly identical, and that is what puts the scale in a different category from every other quantity in the recovery.
This distinction is the reason the many-view field measures its seven flat directions by walking a finite distance along them rather than by inspecting a derivative. A derivative that vanishes says the quantity is stationary; walking two and a half metres along the direction and finding the pictures unchanged to a hundredth of a billionth of a pixel says it is free.
The picture the pictures cannot tell apart
There is a version of this figure that would be more dramatic and less honest, and it is worth explaining why it was not drawn.
The dramatic version puts the two reconstructions in perspective, from a viewpoint outside both, so the smaller one appears smaller. That would show two objects of different sizes and invite the reader to think the pictures could tell them apart if only they were looked at properly.
The figure instead draws both in plan, at their own scales, with the reference distance printed under each. What that shows is two identical shapes with different numbers under them, which is the truth: the shapes are the same shape, the numbers are the arbitrary part, and no view of either would settle anything.
The single-view version of this argument made the same choice for the same reason, and drew its two worlds as two pictures rather than as two scenes. Drawing the scenes would have been drawing information the reader is being told does not exist.
What survives from here
The scale ambiguity is the first of several qualifiers this site attaches to its recoveries, and it will not be the last. It is worth listing the family, because the pattern is the subject of a thread rather than an accident of three fields.
One picture gives ratios and no size. Established in the metrology field, with a world 137 times larger giving an identical picture.
Two pictures give shape and no size. This essay, with 6.3 × 10⁻¹⁴ on the ratios and complete freedom in the scale.
Many pictures give a scene and a camera track up to a similarity — seven numbers rather than one, because a sequence does not fix where the reconstruction is or how it is turned either. That is the many-view field’s opening result, and it is measured there rather than asserted: the seven directions are counted in a Jacobian and then walked along, a metre and a half at a time, with every picture required to stay where it was.
Read in order, the qualifier gets larger as more data arrives, which is the opposite of what data is supposed to do. The resolution is that the qualifier is not uncertainty. It is a statement about which questions the pictures answer, and adding pictures adds parameters — where each new camera is — as fast as it adds constraints. The seven are the difference, and they are exactly the transformations that move a reconstruction without moving any picture of it.
A remark on where this leaves depth
Nothing in this essay says a stereo pair cannot measure distance. It says a stereo pair cannot measure distance without a supplied length, and in practice the supplied length is the baseline, which is why the whole of the depth field is written with a stated baseline in millimetres.
What that field then finds is that even with the baseline supplied, the measurement has a structure that is easy to describe badly. Depth is inversely proportional to disparity, so a fixed error in what is read maps to an interval in what is reported that is not centred, not symmetric, and past a computable range not bounded above at all.
So the honest sequence is: two views give a shape; a supplied baseline turns the shape into distances; and the distances come with an interval that has to be quoted as an interval rather than as a plus-or-minus. Each step is exact and each step has a qualifier, and the qualifiers are the content.
What was actually proved
It is worth restating the result carefully, because “two views give shape and no size” is a slogan and slogans lose their conditions.
Two views of a scene, with known internal calibration, given exact correspondences, determine the relative pose of the cameras up to the length of the baseline, and determine the scene up to a similarity. The recovered shape agrees with the world to the arithmetic’s floor.
Each condition earns its place. Known calibration is what turns a fundamental matrix into an essential one; without it the reconstruction is determined only up to a projective transformation, which is a far weaker statement — angles and parallelism go, and what survives is the cross-ratio and nothing more. Exact correspondences is what makes the fourteen digits fourteen; read the same marks to a pixel and the whole picture changes. And two views is the minimum; one gives less, and more gives the same thing more precisely.
The ambiguity is not among the conditions. It is the conclusion.
And it is worth ending on the reason that conclusion is a good one rather than a disappointing one. A method that reported a size would have to have got it from somewhere, and the only places available are the scene assumptions the classical single-view constructions smuggle in — a standard doorway, a person of average height, a brick of known dimensions. Those are guesses about the world dressed as measurements of the picture, and they fail silently when the doorway is not standard. A method that reports a shape and refuses to report a size is telling the truth about what two photographs contain, and it leaves the one guess visible, where somebody can decide whether it is a good one.
What links here
Computed from the collection, not written here: the essays that point at this one.
Shares its objects with
Essays that name at least two of the same things, and that neither author linked.
- A chain and an adjustment — both name baseline, relative pose
- The image of the other eye — both name baseline, correspondence
- The range a pair cannot see past — both name baseline, reference length
Named objects
A flat tag is an object no other essay names yet.
BaselineCorrespondencereconstruction ambiguityReference lengthRelative posescale ambiguitySimilarityTriangulation