Some of the work inside a computer goes unseen. When a chip has no floating-point unit, it has no hardware for adding two numbers with fractions, so the compiler cannot hand that job to the silicon. It writes a call to a small helper routine instead: __adddf3 to add two doubles, __fixdfsi to turn a double into an integer, __letf2 to compare two 128-bit floats. These helpers live in compiler-rt, the runtime library of the LLVM compiler project. For years they have run unnoticed inside programs built for embedded boards and bare-metal devices.
This summer Mohamed Emad, a computer science and engineering student at Zagazig University in Egypt, set out to replace them. On the LLVM project blog he describes a Google Summer of Code project that hands compiler-rt’s soft-float routines over to LLVM-libc, the project’s own C library.
Two offices, one job
The problem he took on is familiar from any large administration: two departments doing the same job. LLVM maintained two separate soft-float implementations of the same operations. compiler-rt wrote its own from scratch in C and assembly, with a separate version for each architecture. LLVM-libc wrote its own independently, with correctly rounded results and a large test suite. Emad writes that the two copies drift apart, that each needs its own review and testing, and that a bug fixed in one rarely reaches the other. In his account, compiler-rt’s versions are the older ones and the less rigorously verified.
The RFC that proposed the change earlier this year put it more bluntly. It called the existing builtins under-tested, said subtle numerical bugs had gone unnoticed in them for years, and said the mix of C and hand-written assembly made them harder to maintain.
A reflection, not a finding: duplicates in an institution rarely die of being wasteful. They survive because each one has its own keepers, and nobody is paid to notice that both exist. The fix usually comes from outside, from someone with a summer, a mentor and a mandate.
The mandate here is a longer campaign called Hand-in-Hand, which aims to make LLVM-libc a shared foundation for the rest of LLVM. It began in 2024 on a small scale, with a pull request opened in May of that year by Michael Jones of Google. That change let libc++ parse floating-point numbers through std::from_chars using libc’s existing string-to-float code, and it was merged on October 21, 2024.
Emad’s mechanism follows the same idea. LLVM-libc exposes its routines as freestanding headers under LIBC_NAMESPACE::shared::, and each compiler-rt builtin becomes a thin wrapper that passes the call along. A file such as truncdfsf2.c is replaced by truncdfsf2.cpp, which calls shared::truncdfsf2. A CMake macro, use_libc_builtin(), performs the swap, and only when COMPILER_RT_USE_LIBC_MATH is turned on. Distributors who want the libc-backed path choose it. Everyone else keeps what they had. The reform is offered, not imposed.
The work covers addition, subtraction, multiplication and division for float, double and float128. It covers conversions between floats and integers of every width in both directions. It covers extending and truncating between every float format, including the x87 80-bit type, half precision and bfloat16, along with the comparison routines. In all, about seventy builtins.
The routine that called itself
Most of what was hard came from the edges of the number system.
Some of LLVM-libc’s conversions are written as a plain cast from one floating-point type to another. That works on hardware that can do the conversion. On a chip without an FPU, the compiler turns the same cast back into a call to the builtin being implemented. The routine calls itself and never finishes. Emad’s fix was to do the conversion through an internal representation that uses only integers, so the compiler sees nothing it can turn back into a floating-point builtin.
Half precision and bfloat16 are not real types on every target. Even where they do exist, he writes, naming one in a function signature quietly pulls in yet another conversion builtin. So these routines never take a 16-bit value directly. They receive the raw bits and rebuild the number from a description of the format.
Then there was Intel’s 80-bit extended long double, which does not follow the layout rules the shared code expects for standard IEEE formats. Holding its bit pattern needs a 128-bit integer, and that type does not exist on 32-bit x86, where long double is still 80 bits. The project handled it the way compiler-rt already did: it converts through a standard double or float at the boundaries and compiles these routines only where the 80-bit format actually exists.
Two smaller problems followed. LLVM-libc’s routines carry their own symbol names, which would clash with the target’s real C library if both ended up in one binary. They are now built under a private namespace, with a few legacy alias names kept because some platforms still expect them. Each of the seventy builtins also has the same four parts: a libc header, a compiler-rt wrapper, a build entry and a test. Rather than keep them aligned by hand, Emad wrote a small generator that produces all four from a single description of each builtin.
What remains
The project is not finished. According to Emad, the arithmetic and conversion builtins are now backed by LLVM-libc. The comparison stack and the integer, x87, half and bfloat16 conversion stacks are still in review, and the remaining builtins will follow the same generator-driven pattern.
Measurement comes next. These routines end up inside every program compiled for an FPU-less target, so both their speed and their size matter. Emad plans to compare each libc-backed builtin with compiler-rt’s original on soft-float Arm, on x86 without SSE and on bare-metal embedded setups, measuring per-call latency and the size each routine adds to a static binary. His test board is the Raspberry Pi Pico 2, whose Arm Cortex-M33 has no floating-point unit for double precision. A public benchmarking repository already builds compiler-rt’s builtins twice, once with the switch off and once with it on. It records object code size, retired instructions per operation and wall-clock time, and it describes the timings from free hosted runners as noisy and reported only for reference. If the shared integer-only path proves more expensive than the old hand-tuned code, Emad says the hot paths will be tightened.
Two efforts inside LLVM-libc should eventually remove the workarounds. A companion Google Summer of Code 2026 project by Sukumar Sawant, tracked as #206895, adds struct-backed emulated versions of the 128-bit and 80-bit formats, and a software float16 does the same for half precision. Once a format has an emulated type, the compiler has no native type to convert with and falls back to the software path. The self-call disappears at its root, the raw-bits detour for half precision is no longer needed, and Emad expects most of the room for optimization to open up there.
He thanks three mentors by name: Tue Ly, Michael Jones and Muhammad Bassiouni. He also lists what the summer taught him about floating point: how a harmless-looking cast can end up calling itself forever, how a 16-bit float can simply not exist on a target, and how the x87 format bends nearly every rule around it. His open pull requests have not yet landed.

