In 1980, Intel introduced the 8087, a floating-point coprocessor that let the IBM PC do heavy math at something other than the speed of a damp candle. The improvement was not subtle: where the 8086 microprocessor needed about 13,000 microseconds of software grinding to compute a tangent, the 8087 did it in 90 microseconds — well over a hundred times faster. The natural next question is the one reverse-engineer Ken Shirriff has now answered: how? What, exactly, does a slab of silicon know about tangents?

Shirriff answers the way he usually does, by examining the circuitry and microcode of the 8087, he describes popping the lid off a chip with a chisel and photographing the silicon under a microscope, he writes on his blog. The large rectangular region in the centre of the die is the microcode ROM, holding the 1,648 micro-instructions that run the show; the bottom half is the datapath, a specialised neighbourhood of adders, shifters, constant ROMs and registers that crunches floating-point values 80 bits wide.

There are two classic ways to make hardware cough up a trigonometric function. One is CORDIC, an algorithm built from shifts and additions. The other is polynomial approximation. The 8087’s tangent instruction, called FPTAN, uses both — a relay race in which each method runs the leg it’s best at.

The bomber in the math

CORDIC is a quarter-century older than the 8087 and considerably more glamorous. It was developed in 1956 for the B-58 Hustler, the first bomber capable of flying at Mach 2. The Hustler’s navigation computer was analog, and analog components could only be so accurate, so engineer Jack Volder was given the task of replacing it with a digital computer. The snag was trigonometry: an analog machine generates sines and cosines almost for free with an electromechanical device called a resolver, while a digital machine built from that era’s slow transistors finds trig functions genuinely hard.

Volder’s answer was CORDIC — “COordinate Rotation DIgital Computer,” the name covering both the algorithm and the machine built around it. The insight is that instead of computing sine and cosine directly, you rotate a vector by your angle and read the coordinates off the end. Rotating by arbitrary angles is hard, but CORDIC uses only special precomputed angles — arctan(2⁻ⁿ) — kept in a table, because rotating by those powers-of-two angles reduces to shifting bits. Each iteration buys roughly one more bit of accuracy, and multiplication and division are never required, which is excellent news when your arithmetic hardware is made of 1950s transistors. The algorithm later turned up in scientific calculators, running a decimal version of the same trick.

The trigonometry is worth one sentence because it is lovely. For an angle θ, the point it picks out on the unit circle gives you X = cos θ, Y = sin θ, and Y/X = tan θ. Start at the point (1, 0), apply the shift-and-add rotation for each special angle that adds up to your target, and the tangent simply falls out as Y over X.

Sixteen bits, then a fraction

Here is the catch CORDIC keeps in its back pocket: accuracy is pay-as-you-go. Sixteen iterations yields about 16 bits of precision; 64-bit accuracy would demand 64 iterations and a table of 64 special angles. The 8087’s designers declined. After 16 CORDIC steps, the leftover angle — the sliver between the sum of the special angles and the real target — is tiny, around 2⁻¹⁶. And when x is that small, tangent becomes easy to fake, beautifully, with the ratio of two polynomials.

The 8087 uses a Padé approximant, and the specific one is almost insultingly simple: 3x/(3−x²). Its error is proportional to x⁴, which for x smaller than 2⁻¹⁶ shrinks below 2⁻⁶⁴ — comfortably inside the chip’s 64-bit accuracy requirement. A calculus student might reach for a Taylor series instead, but tangent has the bad habit of rocketing to infinity at π/2; a plain polynomial cannot blow up, while a ratio of polynomials cheerfully can. The ratio simply fits the shape of the beast.

The complete FPTAN routine runs in three movements. First, pseudo-division works out which stored angles to use, with the decision bits living in a 16-bit shift register. Then the rational approximation handles the leftover sliver. Finally, pseudo-multiplication applies all the rotations. The adder at the heart of the datapath does nearly everything — it is also used in a loop for multiplication, division and square roots.

The whole excavation delighted readers on Hacker News, where commenter jaygreco wrote, “It’s incredible that the rough equivalent of an entire 160lb digital computer was built and baked into silicon.” But my favourite detail is the exit strategy. FPTAN never performs the final division — not Y/X, and not the one lurking inside 3x/(3−x²). The chip just returns the numerator and the denominator as two separate numbers and considers its work done, making the division, as Shirriff puts it, “free.” That is correct: the machine that took tangents more than a hundred times faster than its host processor finishes by handing you a fraction and letting you carry the one yourself.