Through water and glass

The stick a stereo pair puts back

Two eyes side by side reconstruct a submerged stick exactly as the sagittal image — kinked 14.96° — and two eyes one above the other exactly as the tangential one, kinked 9.59°. Roll the baseline between them and the two rays to a point miss each other by up to 3.13 millimetres, past the 2.85 a pixel covers at that range, and the reconstruction is a third stick that is neither — 551 millimetres of a one-metre stick, with its tip at 0.405 metres against a true 0.866.

Worth reading first: What a ray does at a surface · Two rays that do not meet.

A straight stick in water is a kink and a curve finds two sticks where the textbook has one. The image two eyes side by side see leaves the surface 14.96° off a stick leaning 30°, then keeps turning; the image a nodding head sees is kinked differently and sags differently; and a photograph from the eye shows the kink and almost none of the curve.

A stereo camera reconstructs a third, and which third depends on how the camera is held.

Two eyes are not a single eye twice

A submerged point sends a different ray to each of two eyes, and neither ray is straight: each bends at the surface, at its own angle, in its own plane of incidence.

Run the two rays back into the water and they do not meet. A single eye’s pencil has two focal lines, which is what a point under water has two depths measures; a pair of eyes takes two rays out of that pencil, and two rays out of an astigmatic pencil are skew unless they are chosen along one of its two principal directions.

So a stereo reconstruction is a choice made by the geometry of the rig, not a reading of a fact in the water. It takes the midpoint of the miss, which the midpoint is a choice of ruler establishes is a convention rather than a measurement.

Two baselines that are exact, and they are the two focal sticks

Put the stick in the eye’s own vertical plane — leaning directly away — and there are two baselines for which the rays do meet.

A level baseline, two eyes side by side, gives rays that meet exactly and reconstruct the stick kinked 14.957° at the surface. The sagittal image’s own kink is 14.956°. They are the same curve.

An upright baseline, one eye above the other, gives rays that meet exactly and reconstruct the stick kinked 9.590°. The tangential image’s kink is 9.578°. Again the same curve.

That is the content of the earlier essay’s sentence — two eyes side by side read the first, a head moving up and down reads the second — turned from a statement about what is seen into a statement about what is reconstructed, and it is exactly true rather than nearly.

A straight stick leaning across the line of sight: kinked 14.96° and turning 4.29°A straight stick enters water 2.0 m from an eye 1.6 m up and runs 1.0 m down at 30° from the vertical, leaning across the line of sight. Every point of it has its own two images. The locus two eyes side by side see leaves the surface 14.96° off the stick, then turns a further 4.29° along its length and sags 6.8 mm from its own chord; its tip is seen 0.549 m down against a true 0.866 m. The locus a nodding head sees is kinked 36.9° and sags 31.4 mm.eyethe sticktwo eyes side by sidea nodding headcorrect from 18 cm, at 160 mm widekink 14.96° · turn 4.29°
Fig. 1 The two focal sticks the pair chooses between, from the measurement that found them: one locus for eyes side by side and another for a nodding head.

Between them, a stick that is neither

Roll the baseline between those two and nothing meets.

At forty-five degrees of roll the two rays to a point miss each other by up to 3.13 millimetres, and the reconstruction is the midpoint of that miss all the way down the stick. Its kink is 13.06°, between the two focal kinks and equal to neither, and it runs smoothly from one to the other as the roll goes from level to upright: 14.96, 14.76, 14.13, 13.06, 11.63, 10.21, 9.59.

So a stereo camera with a baseline at any other angle reports a stick that no ray in the water traces and no eye anywhere sees. It is a compromise between two real curves, and it is the one a rig held at an angle actually produces.

A pair rolled 20° reconstructs a third stick, kinked 14.60°The same straight stick in water, with the two loci its focal images trace and the locus a stereo pair reconstructs. A pair of eyes side by side puts back the first exactly and a pair one above the other puts back the second; rolled 20 degrees between them, the rays to the two eyes miss each other by up to 2.09 millimetres and the reconstruction is the midpoint of that miss — a third curve, kinked 14.60 degrees against the first's 14.96 and the second's 9.58, and 674 millimetres long against a true 1000.eyethe stickcorrect from 18 cm, at 160 mm widerays miss by 2.09 mm
Fig. 2 The same three curves with the pair rolled twenty degrees from level. The reconstruction has left the sagittal stick and has not reached the tangential one.

The miss overtakes what a pixel can resolve

A miss of a few millimetres could be a shape a better matcher would sharpen or a shape nothing supports, and the way to tell them apart is to compare it with what one pixel covers.

The comparison has to be made in the right currency. The miss is across the line of sight, so it is to be compared with the lateral footprint of a pixel at that range — the range over the focal length — and not with the depth one pixel of disparity stands for, which is forty times larger and the wrong number.

At the stick’s range of 2.56 metres and nine hundred pixels of focal length, one pixel covers 2.85 millimetres. The miss is under that for most of the roll and reaches 3.13 millimetres at forty-five degrees, so there is a band around the middle of the roll in which the two rays are further apart than the reading error can explain. Inside that band the reconstruction is not a point estimate with noise on it; it is a compromise between two rays that genuinely do not meet.

The rays miss by 3.13 mm at 45°, past what one pixel resolvesThe largest distance between the two rays to any point of the stick, as the pair's baseline rolls from level to upright. It is zero at both ends — a level pair reads the sagittal image and an upright pair the tangential one, both exactly — and rises to 3.13 millimetres in between. The level line is what one pixel of a nine-hundred-pixel camera covers at the stick's range, 2.85 millimetres: over 45° the miss exceeds it, so the reconstruction is a shape no point in the water supports rather than one a better matcher could sharpen. The kink of the reconstructed stick runs from 14.96 degrees to 9.58 over the same roll.0123020406080how far the pair's baseline is rolled from level, in degreeshow far the two rays miss each other, in millimetreswhat one pixel coverszero at level and at uprightworst 3.13 mm
Fig. 3 The largest miss between the two rays, as the baseline rolls. Zero at both ends, and past what one pixel covers around the middle.

What the reconstruction gets wrong about the stick

The shape is only part of the damage, and the sizes are the part a reader would feel.

The stick is one metre long below the surface, entering two metres from the eye and leaning thirty degrees. A level pair reconstructs it as 715 millimetres; a pair rolled forty-five degrees as 551; an upright pair as 416. The true length is 1,000.

Its tip is 0.866 metres below the surface. A level pair puts it at 0.512, a rolled pair at 0.405, an upright pair at 0.317.

None of those is a small error and none is noise. They are what a correct triangulation of correctly-read marks reports, because the rays it triangulates are not rays to the stick.

Why exactly those two baselines

The two exact cases are not a coincidence of the numbers, and the reason says what would be exact in any other arrangement.

A submerged point’s rays to a nearby eye form an astigmatic pencil: it has two principal directions, and along each of them the rays do pass through a point — the sagittal focus and the tangential one. Those two directions are the plane of incidence and the plane perpendicular to it, which for a point in the eye’s own vertical plane are the vertical and the horizontal.

Two eyes separated along a principal direction therefore take two rays from the same focal line, and two rays through one point meet there. Two eyes separated along any other direction take one ray from each family and nothing meets.

So the exactness is a statement about the baseline lying in a principal plane of the pencil, and it transfers: for any submerged point, the two exact baselines are the ones along its own plane of incidence and across it. For a point in the median plane those are level and upright; for a point off to the side they are two other directions, tilted by however far that point’s plane of incidence is tilted. A rig looking at a scene rather than at a point therefore has no exact orientation at all, because different points want different ones.

The same absence, in three instruments

The shape of this failure is one this field keeps producing, and setting the three beside each other says what is general in it.

A picture through water has no viewpoint fits a common point to the rays arriving at a camera through a flat water surface and finds none: the rays pass through a region rather than a point. A ball of water has no eye either does the same for a refracting sphere. And two rays that do not meet establishes the base case with no water at all — two rays read to a pixel miss by millimetres.

The three failures are not the same size and they are not the same kind. The last is a reading error: exact marks give rays that meet exactly, and the miss is the noise made visible. The first two are structural: no reading error is involved and the rays still have no common point.

The measurement here is the second kind and it produces a number in the first kind’s units, which is why the pixel comparison had to be made. Three millimetres of miss between two rays is either noise or geometry, and at this range and focal length it is just past the boundary — which is the least convenient answer and the one to know.

What a diver’s camera would show

Everything above puts the eyes in air and the object in water, which is the pool-side arrangement. Turning it round is worth a paragraph, because the answer is not symmetric.

An eye in the water looking at an object in the water sees no refraction at all along the path between them: the whole path is in one medium and the rays are straight. A stereo pair under water reconstructs a submerged stick exactly, subject only to the reading error, and none of this applies.

What such a pair does have is the housing. A flat port between the lens and the water is a refracting slab at the front of every ray, and what survives a pane of glass measures what a slab does and does not preserve: it moves every point and no direction, so the camera recovered through it is the camera that took the picture. A dome port centred on the entrance pupil does better still.

So the arrangement measured here — eyes dry, object wet — is the one with the structural problem, and it is also the common one: a phone photographing a coin in a fountain, a survey of a river bed from a bank, a stereo rig over a tank. The arrangement a diver has is geometrically cleaner than the one a bystander has.

The pool’s own floor, and why the stick is worse

The apparent depth of a submerged point is smaller than its true depth, which is the field’s first result and the reason a pool looks shallow. That accounts for part of the numbers above and not for their spread.

What a ray does at a surface measures the ratio as a function of the line of sight: three quarters looking straight down, a half at sixty degrees, a fifth at eighty. The stick’s tip is seen at a steep angle and its entry point at a shallower one, so different parts of it are compressed by different amounts — which is why the reconstruction is shortened and kinked and curved rather than simply scaled.

What the roll of the baseline adds is a third effect on top of those two. The compression is a property of the water and the sight angle; the kink is a property of the water and the stick’s lean; the difference between the three reconstructions is a property of the rig, and it is the only one of the three a photographer controls.

Apparent depth is a curve, not a numberDividing by n is correct looking straight down and nowhere else. At 80° from the vertical the bottom appears at 19% of its true depth rather than 75%.00.2000.4000.600020406080angle of the line of sight from the vertical (degrees)apparent depth, as a fraction of the true depth1/n = 0.711n = 1.40771.1% straight down, 17.3% at 80°
Fig. 4 The compression itself, against the line of sight: the ratio that makes a pool look shallow, and which decides how much of the stick’s shortening has nothing to do with the rig.

A stick across the line of sight has no exact baseline at all

The two clean cases needed the stick to lie in the eye’s own vertical plane. Lean it across the line of sight instead and both of them go.

At every roll from level to upright the rays miss, by between 0.71 and 2.44 millimetres, with no zero anywhere. The reason is that the plane of incidence of an off-axis point is neither level nor upright, so neither baseline lies along a principal direction of the pencil.

That makes the two exact cases a property of the arrangement rather than a property of stereo. A rig photographing a pool from the side has no orientation that reads a real curve; it has a family of compromises, and the best it can do is to choose which compromise.

The practical conclusion is unusually clean for this subject. Hold the baseline level and keep the object in the vertical plane through the lens, and the reconstruction is at least the locus something actually sees. Any other arrangement reconstructs a shape that exists nowhere.

The rays miss by 2.44 mm at 30°, inside what one pixel resolvesThe largest distance between the two rays to any point of the stick, as the pair's baseline rolls from level to upright. It is zero at both ends — a level pair reads the sagittal image and an upright pair the tangential one, both exactly — and rises to 2.44 millimetres in between. The level line is what one pixel of a nine-hundred-pixel camera covers at the stick's range, 2.85 millimetres: the miss stays under it at every tilt, so a better matcher would still have somewhere to put the point. The kink of the reconstructed stick runs from 14.96 degrees to 36.91 over the same roll.0123020406080how far the pair's baseline is rolled from level, in degreeshow far the two rays miss each other, in millimetreswhat one pixel coverszero at level and at uprightworst 2.44 mm
Fig. 5 The same sweep for a stick leaning across the line of sight. The curve never reaches zero, so no roll of the baseline reads a real image.

What a pipeline would report, and why it is no warning

A stereo pipeline handed this scene produces numbers, and it is worth asking which of them would tell anybody something was wrong.

The matching is perfect: every point of the stick is found in both pictures, with a sharp, unambiguous correspondence and no outliers. The epipolar check is perfect too — the two marks of every correspondence sit on each other’s epipolar lines to the arithmetic floor — because the cameras are exactly where the pipeline thinks they are and the geometry between them is exactly a stereo geometry. Nothing about the cameras has been disturbed by the water.

What is disturbed is the assumption that a mark’s ray goes to the thing it is a mark of. The pipeline’s rays are the rays that arrive at the lens, which are straight in air and correct; they simply do not point at the stick.

So the only quantity with information in it is the gap between the two rays at the triangulated point, and that is exactly the quantity a pipeline computing a midpoint has thrown away by the time it reports anything. At forty-five degrees of roll it is 3.13 millimetres — measurable, larger than a pixel’s footprint, and not in any report.

That is worth carrying because it is the same shape as a degeneracy: everything that is checked is at the floor and the answer is wrong. The repair, as with the others, is to keep a residual that the wrong model cannot make small — here the ray gap — and to compare it against what the reading error can account for rather than against zero.

What this does not settle

Both eyes are in the air. A pair of cameras in housings under the water is a different problem, with the flat port’s own refraction between the lens and the water — which is the subject what survives a pane of glass takes up, and which adds a second surface to every ray traced here.

The surface is flat and still. A ripple moves the entry point of every ray and moves it differently for the two eyes, so the miss acquires a component that varies along the stick and in time. Nothing here is about that, and it is the dominant effect in any real water.

And the stick is a line. A real submerged object is a surface, and the two focal images of a surface are two surfaces; which of them a stereo rig reconstructs, and what a rolled baseline does to a shape rather than to a curve, is the same question with one more dimension in it.

Still open: what the compromise costs a fitted shape

The reconstructions above are curves, and nothing is fitted to them. The use a reconstruction is actually put to is fitting a shape — a straight line, a plane, a cylinder — and the question with a number in it is what the compromise does to a fit rather than to a point.

A straight stick fitted as a straight line to each of the three reconstructions gives three different lines, with three different residuals, and the residual is the quantity a reader would use to decide whether the reconstruction is trustworthy. The measurement fits a line to each roll from level to upright, records the residual, and asks whether the residual warns: whether the rolled reconstruction, which is the one belonging to no real image, is the one whose fit looks worst. If it does not — if the compromise curve fits a straight line as well as the exact ones do — then the residual a pipeline reports is blind to exactly the failure that matters, and the only warning available is the rays’ own miss, which a pipeline that takes midpoints has already thrown away.

The short version

A stereo pair reconstructs a submerged stick by taking the midpoint of two rays that do not meet, and which stick it gets depends on how the baseline is held. With the stick in the eye’s own vertical plane, a level baseline reconstructs the sagittal image exactly — kinked 14.957° against that image’s 14.956° — and an upright baseline reconstructs the tangential one, 9.590° against 9.578°.

Every roll between them reconstructs a third curve that no eye sees, with the rays missing by up to 3.13 millimetres — past the 2.85 one pixel covers at that range. The reconstructed length runs from 715 millimetres at level to 416 upright, against a true 1,000, and the tip’s depth from 0.512 metres to 0.317 against a true 0.866. Lean the stick across the line of sight instead and there is no exact baseline at all.

The kink is 14.96° whichever way the stick leans; the turn runs from 1.11° to 7.62°A stick leaning 30° from the vertical, entering water 2.0 m from the eye, swept round from leaning straight away to leaning straight toward it. The kink its image makes at the surface stays at 14.96° — the closed form, which has no azimuth in it — to within 7e-7°. How far the image turns along its length does depend on the direction: 1.11° leaning away, 4.29° across, 7.62° leaning toward the eye.051015050100150direction of the lean, from straight away to straight toward the eye (degrees)angle (degrees)kink at the surface: 14.96°turn along the lengthlean 30° · entry 2.0 m from the eyekink 14.96° throughout
Fig. 6 The kink itself, swept through the stick’s own lean — the quantity the three reconstructions disagree about, measured where it comes from.

Shares its objects with

Essays that name at least two of the same things, and that neither author linked.

Named objects

A flat tag is an object no other essay names yet.

Apparent-depthAstigmatismBaselineMidpointPlane of incidenceRefractionskew raysSnell's lawStereo pairTriangulation