oosioo
Cover for From Burr to Tongue, Part 2. Beside the title Four points are enough, a tilted sheet of paper with ArUco markers at its four corners and coffee grounds scattered across it
SeriesFrom Burr to Tongue · Ep. 3

Four Points Are Enough

A single coin can't straighten a tilted photo. Stick four ArUco markers on the paper, and a perspective transform (homography) can flatten the shot — in synthetic-photo tests, the tilt error dropped from 7.7 percentage points to 0.08. But some error still remains, especially a print-scale problem that no photo alone can reveal, and that's what the marker sheet v1 introduced here is designed to catch.

· SYSOP · 3 views

In the last piece, we worked out where the one-coin method for fixing magnification goes wrong. Tilt the phone just 5° off parallel to the paper, and particles of identical size get measured almost 9 percentage points differently depending on where they land in the frame. But at that tilt, a coin looks only 0.4% out of round, so you can't even spot the tilt by looking at the coin. That's because a coin really only gives you one number: magnification.

Unspecialty's column "Developing a Grind-Size Guide" (2023) solved this by sticking four markers on the paper. This piece retraces that method from the ground up. Why does it have to be four points, how do markers tell you where those points are, and how much does the error actually shrink once you flatten a photo using four points? We'll also look at the errors a perspective transform ultimately can't fix. At the end, you'll find a marker sheet I designed myself, ready to download and print.

What the Floor Tiles Showed First

At the Galleria Nazionale delle Marche in Urbino, Italy, hangs a painting called The Ideal City. It's estimated to date from the 1480s and was once attributed to Piero della Francesca, though the artist is now unknown.

A wide horizontal panel painting. At the center stands a two-story circular temple, flanked by rows of buildings with columns and arches. The plaza floor in front is marked with gray bands forming large squares and diagonal patterns, which flatten and narrow the farther back into the scene they go. There are no people in the painting

The Ideal City (probably 1480s, Galleria Nazionale delle Marche). The plaza floor's tile pattern is actually made of squares and diagonals, but in the painting it becomes a flattened trapezoid the farther back it goes. Painting: Unknown artist / Wikimedia Commons (public domain)

The plaza floor is a single flat plane. The painter moved that plane onto another plane — the canvas. The tiles in front are large; the ones further back are small. Lines that run parallel on the floor converge to a single point in the painting. And yet a straight line in the painting is still a straight line. It never bends.

This is exactly how the perspective transform that carries one plane onto another works. In computer vision, this transform is called a homography. A sheet of paper with scattered coffee grounds is a plane too, and so is a phone's camera sensor. A photo of paper taken with a tilted camera is in the same situation as the floor in The Ideal City. Squares become trapezoids, and particles closer to the camera get photographed larger.

Flip the logic around and the path becomes clear. Once you work out what the transform looks like, its inverse can restore the photo to how it would look from directly overhead. It's the same idea as knowing the tiles' true shape and using that to unflatten the trapezoids in the painting back into squares.

Eight Unknowns, Four Points

A homography is written as a single 3×3 matrix. Multiply this matrix by a point (x, y) on the paper, and you get the point (u, v) in the photo. The matrix has nine numbers, but multiplying the whole matrix by the same constant gives the same result. So one number can be fixed at 1, leaving only eight to determine.

A diagram showing two shapes and a matrix equation. On the left is a square seen from directly above the paper, with colored points numbered 1 through 4 at its corners. On the right is the same square as seen by a tilted camera, with the same four colored points now at the corners of a trapezoid. Between the two shapes is an arrow labeled H, and the grid inside the square is shown on the right with spacing changed but not curved, matching the perspective. Below is a 3×3 matrix with entries h1 through h8 and a 1, along with notes: 8 unknowns, 2 equations per point, 4 points give 8 equations

Locating, in the photo, one point whose position on the paper is known gives two equations: one for the photo's horizontal coordinate u, one for the vertical coordinate v. With eight unknowns, four points yield eight equations, which pins down the matrix. This is the basis for Unspecialty's claim that "at least four points are needed." The catch is that no three of the four points can lie on the same line — otherwise the equations overlap and the unknowns can't actually be solved.

A single coin can only tell you one of these eight values. Even if you measure the coin's edge carefully enough to capture the shape of its ellipse, you gain only a few more, and as we saw in the last piece, that information is nearly invisible at small tilt angles.

What if there are more than four points? When there are more equations than unknowns, there's no answer that satisfies every equation exactly. Instead, you pick the answer that minimizes the overall error (least squares). The small measurement errors scattered across each point cancel each other out, and the answer stabilizes. Here, plugging in raw coordinates directly mixes pixel values (in the thousands) with the constant 1 in the same equation, making the calculation unstable — so the standard approach is to first shift the coordinates near the origin and rescale them before solving (Hartley 1997). Functions like OpenCV's findHomography handle this process.

The Pattern That Tells You Where the Point Is

The equations themselves are simple. The hard part is finding the points within the photo. From a single photograph, you have to determine which point on the paper landed on which pixel — without human intervention, and with precision finer than a single pixel. Fiducial markers were built to do exactly this.

Fiducial markers first found their footing in augmented reality. To overlay a virtual object on a camera feed, you need to know at every instant where the camera is and which way it's facing. ARToolKit, released in 1999, used markers consisting of a picture inside a thick black border (Kato & Billinghurst 1999). After that, the approach of filling the inside of the border with black-and-white cells to encode a number spread widely. AprilTag (Olson 2011), widely used in robotics, and ArUco (Garrido-Jurado et al. 2014), used in this piece, are examples of this.

An image divided into four panels. Panel 1, ARToolKit, shows a stroke resembling the Chinese character 人 drawn in a white cell inside a black border. Panels 2, 3, and 4 — ARTag, AprilTag, and ArUco — all show patterns filling a black border with black-and-white square cells

A few fiducial markers used in computer vision. Image: Cmglee / Wikimedia Commons (CC BY-SA 4.0)

A single ArUco marker tells you two things.

An anatomical diagram of an ArUco marker. On the left, a large marker divided into a 6×6 grid is shown, with the outer ring entirely black border. Red dots mark the four corners, and a blue square outlines the inner 4×4 cells. On the right are three explanatory notes: the black border gives one square and four corner points; the inner 16 bits give the marker's ID number and which side is up; the dictionary's 50 markers differ from each other, and from their own 90°-rotated versions, by at least 4 bits. Below, the same marker is shown rotated by 0°, 90°, 180°, and 270°, side by side

*Marker 0 from OpenCV's DICT_4X4_50 dictionary. The bit distances within the dictionary were counted directly from the dictionary built into OpenCV*

First: the position of the points. A marker's black border forms a sharp square against the white paper. The detection algorithm finds the square's outline in the photo, fits straight lines to its four sides, and takes their intersections as the corners. Because the lines are fit across a boundary spanning tens to hundreds of pixels, the intersection can be located more finely than a single pixel. This is called subpixel precision.

Second: who that point belongs to. The 4×4 cells inside the border encode a 16-bit number. This number tells you which marker it is and which way the marker is rotated. If you place markers with different ID numbers at the four corners of the paper, each square found in the photo is automatically assigned to the correct corner of the paper — no need for a human to match up the points.

16 bits could encode 65,536 different patterns, but the DICT_4X4_50 dictionary contains only 50, because only patterns that won't be confused with one another were selected. Counting them directly, I found that the markers in this dictionary differ from every other marker, and from their own 90°-rotated versions, by at least 4 bits. So even if one cell is misread, it can still be corrected to the nearest marker, and the odds of it being read as the wrong number or upside down are low. The original ArUco paper covers how to generate dictionaries like this at any desired size (Garrido-Jurado et al. 2014).

How Accurately Are the Corners Found?

To know how accurately a marker's corners are found, you need a photo whose true answer is known. Since you can't know the true corner values in a real photograph, I built the photo directly using the pinhole camera model from the last piece. I simulated photographing the marker sheet introduced at the end of this piece with a 4032×3024-pixel phone main camera from 15 cm above, and added a bit of blur (1 pixel) and noise to the photo. Then I used OpenCV's ArUco detector to find the corners and compared them against the true values.

With subpixel refinement turned off, the corner error was 0.7–0.9 pixels RMS. With refinement turned on, it dropped to 0.25 pixels (model estimate). At this distance, one side of a marker (12 mm) spans about 240 pixels in the photo.

If the corners are off by 0.25 pixels, how far off is the measurement? Using the same model, I deliberately injected noise into the corner positions and re-estimated the perspective transform 200 times. At a 5° tilt with a 0.3-pixel corner error, 95% of the 200 trials fell within 0.03% magnification error inside the measurement window — using all 16 points, four corners from each of the four markers. Using only the four outer corners, it was 0.05%. Even when the corner error grew to 2 pixels, it stayed at just 0.17% with 16 points and 0.31% with 4 points.

How far apart the markers are placed also matters. Widening the spacing between marker centers from 72 mm to 144 mm shrank the same-condition error from 0.057% to 0.037%. The more spread out the points are, the more reliably the tilt gets captured.

In other words, once a marker is properly captured, the precision of finding its corners is no longer the bottleneck. Real photos will likely be worse than the synthetic ones due to JPEG compression, phone sharpening, and glare off the paper surface. Still, even if the corner error grows several times over, it doesn't change the conclusion.

Flattening the Photo

Now let's flatten the whole photo. Inside the measurement window of the same synthetic photo, I laid out 96 discs with a precise diameter of 4.000 mm — test particles of known size, standing in for coffee grounds. The camera was tilted 5°, and the sheet was rotated 10°.

Two photos side by side. On the left is a synthetic photo taken by a tilted camera, with the four corner markers outlined in red and labeled id 0 through 3. The grid of discs in the middle appears as a skewed trapezoid. On the right is the photo flattened by estimating a perspective transform from the 16 marker corner points; the four markers now sit at the corners of a proper rectangle, and the discs are evenly aligned in rows and columns

I measured the discs in the left photo using the method from the last piece — applying a single magnification value taken from the center of the photo to every disc. The right photo was measured after flattening it by estimating a perspective transform from the 16 corner points of the four markers.

A dot plot. The horizontal axis is measured diameter (3.7–4.3 mm), across four rows. At 5° tilt, the red dots measured with a single center magnification spread widely across 3.843–4.151 mm. The blue dots flattened using 16 marker points cluster tightly at 3.989–3.993 mm. At 10° tilt, the red dots measured with a single magnification spread even wider, 3.706–4.321 mm, while the blue dots — flattened with only three markers detected — still cluster at 3.990–3.994 mm

Synthetic photo test. Disc boundaries were found with OpenCV's Otsu threshold. All calculations are my own for this series

At 5° tilt, the disc diameters measured with a single magnification value spread from 3.843 to 4.151 mm — a gap of 7.7 percentage points between the smallest and largest measured disc. After flattening with the markers, the range was 3.989–3.993 mm, a gap of just 0.08 percentage points. At 10°, the difference is even larger: 15.4 percentage points with a single magnification, versus 0.10 percentage points after flattening.

In the 10° photo, the bottom-left marker's corner fell off the bottom edge of the frame and wasn't detected. The remaining three markers yielded 12 corners, which was enough. Since only four corner points are needed, measurement can continue even if one marker gets covered by grounds or falls outside the frame. Four markers, in other words, build in the slack to lose one.

One thing worth noting: the average of the flattened diameters was 3.991 mm, 0.2% smaller than the true value. Since all 96 discs came out equally small, this isn't a magnification problem. It comes from where you draw the line between particle and background at a blurry edge. A perspective transform doesn't fix this.

You can change the conditions yourself in the simulation below. On the left is the photo as captured by the camera, with disc colors showing the error from a single center magnification. On the right is the photo flattened using the marker corners, marked with red crosshairs. Try changing the tilt and corner detection error, and below that, the print scale, lens distortion, and marker lift. Click the image on the right to redraw the detection error.

A simulation showing side by side the view captured by a tilted phone camera photographing a marker sheet, and the view after flattening it by estimating a perspective transform from the marker corners. Both views contain 96 discs of identical size, each colored by the error in its measured size. Discs measured larger than actual are red; smaller are blue. You can change the camera tilt, sheet rotation angle, camera height, corner detection error, number of points used (4 outer points or 16), print scale (100%, 97.3%, 94.1%), lens distortion, and lift of the top-left marker. Below each view, the range of disc size error is shown as a number, and at the very bottom, how many mm a card would be off if laid on the outline at that print scale.

What a Perspective Transform Can't Fix

Even once tilt error disappears, the measurement isn't finished. Homography is only accurate under the assumption that "flat paper was photographed with a pinhole camera." Error remains wherever this assumption breaks down. Here's each one calculated at a baseline of 15 cm height and 5° tilt (all model estimates).

A horizontal bar chart. The top gray bar, shown for comparison, is the error range from measuring magnification with a single coin (Part 1, 5° tilt): 8.7 percentage points. Below it, blue bars show: corner detection error of 0.3 px with 16 points is nearly invisible at under 0.03%; lens distortion of 1% remaining at the photo's edge gives 0.67 percentage points; a marker lifted 1 mm gives 0.67 percentage points; a measurement window bulging 1 mm at the center gives 0.85 percentage points. The bottom two red bars show: printing a Letter document fit-to-page on A4 makes every particle +2.8%; printing an A4 document fit-to-page on Letter makes every particle +6.3%

Lens distortion. Homography sends straight lines to straight lines. A lens's radial distortion bends straight lines, so a single planar transform can't erase it. Even after a phone's software correction, if 1% barrel distortion remains at the photo's edges, the error within the flattened measurement window ranges from −0.15% to +0.52% depending on position. Distortion grows faster the further you move from the center of the photo, so raising the camera a bit to fit the sheet smaller within the frame reduces it. Under the same conditions, raising the height to 20 cm shrinks the range from 0.67 to 0.38 percentage points — at the cost of each pixel covering a larger area.

When the paper isn't flat. Homography assumes every point on the paper lies on a single plane. If the paper's edge curls and one marker lifts 1 mm, that marker's corners get photographed as if they were closer to the camera than they actually are. The transform trusts these displaced points as if they were on the plane, twisting the whole measurement window. Calculations show the error ranges from −0.48% to +0.19% depending on position. If the center of the measurement window bulges up 1 mm, particles there get measured up to 0.69% larger. This is the same principle as the coin's thickness in the last piece: it grows by the height difference divided by the camera distance, 1 mm ÷ 150 mm = 0.67%.

Print scale. This is where the largest error comes from. Homography trusts that the markers were printed exactly where they were designed to be on the paper. But if the printer setting "fit to page" is turned on, the printer shrinks the whole thing slightly. For example, if a document made for A4 (210×297 mm) is fit onto US Letter paper (215.9×279.4 mm), the height has to shrink from 297 mm to 279.4 mm, so the whole thing scales down to 94.1%. Conversely, fitting a Letter document onto A4 shrinks it to 97.3%.

Shoot a photo of a sheet shrunk this way, and every particle in the measurement window gets measured 6.3% (or 2.8%) larger. The problem is that the photo alone gives you absolutely no way to know this. A sheet printed at 94.1% is still flat and a proper rectangle. The ratios between markers are unchanged. As far as the perspective transform is concerned, there's nothing wrong — it's just a slightly smaller piece of paper. Unlike tilt, there's no way to notice this from the information in the photo itself. The only way is to bring in an object from outside the photo whose size you already know, and check against it.

It's also possible that the printer stretches or shrinks the horizontal and vertical directions differently, since the mechanism that feeds the paper and the one that moves the ink head are different parts of the machine. This piece didn't confirm how large this difference actually is in practice. That's why the marker sheet below includes separate rulers to check the horizontal and vertical directions.

Marker Sheet v1

I put together everything worked out above into a single marker sheet. Rather than borrowing the column's 100 mm square layout as-is, I redesigned it based on the calculations above.

A preview of a single-page A4 marker sheet. At the top are the words "Marker Sheet v1" and "Print at actual size (100%)". Below that, four ArUco markers sit at the four corners of a rectangle, with a 120×80 mm measurement window marked by gray corner brackets inside it. Below that is a 10 cm horizontal ruler, a rounded-corner card outline (85.60×53.98 mm) at the bottom left, a 10 cm vertical ruler in the middle, and on the right, six steps for how to take the photo along with the marker specifications

Open and print Marker Sheet v1

Here's what the design settles on:

  • 12 mm markers, 4×4 dictionary. With a 6×6 grid, each cell is 2 mm. From 15 cm up, each cell is captured at roughly 40 pixels, giving plenty of margin to read. Dictionaries with finer grids (5×5, 6×6) can hold more ID numbers, but at the same print size the cells shrink and become harder to read. This sheet only needs four IDs
  • Marker center spacing of 144×104 mm. Phone photos are 4:3, and from 15 cm up they capture roughly 200×150 mm. Spacing the markers out into a rectangle matching the photo's aspect ratio uses the frame to the fullest while still leaving margin outside the markers. The wider the spacing, the smaller the effect of corner error
  • 120×80 mm measurement window, 6 mm from the markers. There needs to be white margin around each marker for the border to register clearly. If grounds touch a marker, detection can fail
  • Card outline of 85.60×53.98 mm. Credit cards, transit cards, and ID cards all follow the international standard (ISO/IEC 7810 ID-1), so their size is fixed. It's an item nearly every household has, with a precisely known size. Lay the card on the printed sheet's outline and check whether all four sides line up with the edges to immediately verify the print scale. If printed at 94.1%, the card would overhang by 5.1 mm horizontally. A 1 mm mismatch corresponds to 1.2% horizontally and 1.9% vertically
  • A 100 mm horizontal ruler and a 100 mm vertical ruler. Checking against a steel ruler gives more precision than the card. Checking the two directions separately accounts for the horizontal/vertical scale difference mentioned above
  • Black and white, matte paper. Glossy paper can reflect light and blow out the marker cells to white

The print page was built to print directly from the browser. Saving it as a PDF from Chrome and re-measuring showed the marker center spacing on a single A4 sheet came out to 143.99×104.01 mm, matching the design values. That said, I haven't yet verified this by actually printing it on a real printer and photographing it with a phone. Results may vary by printer, so after printing, I'd strongly recommend checking with a card first.

Summary

  • Paper photographed by a tilted camera undergoes a plane-to-plane perspective transform (homography). Since this transform has 8 unknowns, four points of known position are enough to determine it
  • ArUco markers give you both the position of each corner and which corner it is. In the synthetic photos, turning on subpixel refinement brought the corner error down to about 0.25 pixels
  • Flattening the photo using 16 corner points shrank the disc size range from 7.7 percentage points to 0.08 at 5° tilt, and from 15.4 to 0.10 at 10° tilt. The result held up almost unchanged even with one marker missing
  • A perspective transform can't fix lens distortion, paper lift, or print scale. The first two stay under 1%, but "fit to page" printing inflates every particle by 3–6%, and there's no way to notice it from the photo alone
  • That's why Marker Sheet v1 includes a card outline and rulers in both directions. After printing, just check with a card first

Switching from a coin to markers changed the nature of the problem. In the last piece, the error depended on how you held the camera. Now, most of what's left depends on the paper — what got printed, and how flatly it was mounted, are what govern the measurement.

References

  • Unspecialty (2023). "Developing a Grind-Size Guide." Link — original text verified
  • Hartley R, Zisserman A (2004). Multiple View Geometry in Computer Vision, 2nd ed. Cambridge University Press. doi:10.1017/CBO9780511811685 — standard textbook on homography and DLT
  • Hartley RI (1997). In defense of the eight-point algorithm. IEEE Transactions on Pattern Analysis and Machine Intelligence 19(6), 580–593. doi:10.1109/34.601246 — coordinate normalization, bibliographic reference verified
  • Kato H, Billinghurst M (1999). Marker tracking and HMD calibration for a video-based augmented reality conferencing system. IWAR '99, 85–94. doi:10.1109/IWAR.1999.803809 — ARToolKit, bibliographic reference verified
  • Olson E (2011). AprilTag: A robust and flexible visual fiducial system. ICRA 2011, 3400–3407. doi:10.1109/ICRA.2011.5979561 — bibliographic reference verified
  • Garrido-Jurado S, Muñoz-Salinas R, Madrid-Cuevas FJ, Marín-Jiménez MJ (2014). Automatic generation and detection of highly reliable fiducial markers under occlusion. Pattern Recognition 47(6), 2280–2292. doi:10.1016/j.patcog.2014.01.005 — original ArUco paper, bibliographic reference verified
  • OpenCV. Detection of ArUco Markers. Link — detector and dictionary (DICT_4X4_50). Tests in this piece were done with OpenCV 5.0
  • ISO/IEC 7810:2019. Identification cards — Physical characteristics — ID-1 card, 85.60×53.98 mm
  • Brown DC (1966). Decentering distortion of lenses. Photogrammetric Engineering 32(3), 444–462 — radial distortion model, bibliographic reference verified

Image Credits

  • The Ideal City: Unknown artist (estimated 1480s), Galleria Nazionale delle Marche, Wikimedia Commons, public domain
  • Fiducial marker comparison: Cmglee, Wikimedia Commons, CC BY-SA 4.0 (SVG converted to PNG)
  • All other figures, synthetic photos, simulations, and marker sheets: created by the author

Read this series from the start: From Burr to Tongue.