---
title: "The Loongson chip that sometimes forgot to count, and the AI that proved it"
description: "A Debian package’s infinite loop on LoongArch turned out to be a CPU erratum — fixed by flipping a single undocumented bit"
author: "Penny Quirke"
published: 2026-09-27T00:38:56.819Z
modified: 2026-09-27T01:12:28Z
url: https://rews.cc/a/the-loongson-chip-that-sometimes-forgot-to-count-and-the-ai--b8d25a
language: en
tags: ["cpu", "loongson", "openmp", "bugs", "linux", "tech"]
publisher: "Rews (https://rews.cc)"
---

# The Loongson chip that sometimes forgot to count, and the AI that proved it

*A Debian package’s infinite loop on LoongArch turned out to be a CPU erratum — fixed by flipping a single undocumented bit*

By Penny Quirke · September 27, 2026 · https://rews.cc/a/the-loongson-chip-that-sometimes-forgot-to-count-and-the-ai--b8d25a

## In brief

- An infinite loop while packaging normaliz for a LoongArch Debian port was traced to the CPU losing atomic-add updates
- After six months stalled, an AI assistant directed by the researchers produced a minimal reproducer in about two days
- The trigger was glibc’s memcpy using the chip’s LSX/LASX vector extensions; certain scalar reads can also trigger the flaw
- Loongson delivered a near-zero-cost firmware fix within two weeks, with release expected before October 1

In February 2026, Wang Miao — a maintainer of loong13, a community port of Debian 13 stable for China’s LoongArch processor architecture — was packaging a piece of mathematics software called normaliz when its built-in test quietly refused to finish. It was not slow. It was stuck in an infinite loop, grinding away until the packaging job timed out. The loop’s sole exit demanded that a counter of data points processed so far equal the total number of points, and the counter kept arriving short. The team skipped the package, discovered they could not skip it forever — several other packages depend on normaliz — and started digging.

The loop itself is boring by design, the way good plumbing is boring. normaliz works through a list of points in parallel using OpenMP, and each time a thread finishes a point it bumps a shared counter with \`#pragma omp atomic\`, the standard promise that updating that variable is a single uninterruptible gesture. Once every point is done, the counter hits its target and the loop breaks. Under the debugger, the investigators could watch every point get marked as processed while the counter sat stubbornly below the total. A second counter, incremented alongside the first by the same threads, told the same story — and the two drifted apart by small, random amounts that changed from run to run.

An atomic increment is the software version of a turnstile: no matter how many threads shove through at once, each passage is supposed to count exactly once. A counter that ends low means the gates occasionally let two threads through on one tally. The obvious suspects were cleared first. LoongArch’s weak memory model had exposed hidden race conditions and memory-ordering bugs in other packages before, but this code only ever reads and writes the atomic variables themselves, so ordering could not explain it. OpenMP was next: the disassembly showed the compiler emitting LoongArch’s amadd.d atomic-add instruction exactly as expected.

So the team ran a control: two extra std::atomic counters parked beside the original pair. All four should have agreed; instead they showed random discrepancies — a strong hint that the atomic add itself was dropping updates. Then the trail went cold in a maddening way. A simple atomic-add stress test never reproduced the loss, and neither did a stripped-down imitation of the processing logic. Commenting out normaliz’s real computation step by step did not help either; even with most of the program’s work deleted, the bug persisted. With no minimal example a human could actually read, the matter was shelved.

In August — six months later, with the fresh disclosure of the LoongLeak/LoongBleed vulnerabilities putting LoongArch under the microscope — Wang Miao came back, and the write-up’s author tried a new division of labour: a human would direct the investigation and an AI would do the legwork. It was not magic. In the first conversation the AI noticed the misbehaving loop but never blamed the atomic add. Told the problem appears only on LoongArch, it still withheld judgment. Only when the researchers said outright that they had pinned the fault on the atomic add, and asked for a minimal reproducer, did the AI light on the thing round one had walked straight past: a memcpy call inside the processing function.

That call mattered because memcpy does not live in the program; it lives in glibc, which picks its fastest implementation based on the hardware in front of it. On Loongson chips with the LSX/LASX vector extensions, glibc’s memcpy uses vector instructions to shovel memory in bulk — and it was precisely those vectorised copies that made the neighbouring atomic add drop its update. About two days after the humans handed over the keys, the AI produced a small program that reproduces the failure reliably.

From there the questions multiplied, as questions will: are other atomic instructions affected, and can other memory operations trigger the same loss? According to the [full write-up](https://jia.je/hardware/2026/09/24/loongson-cpu-erratum-en/) — whose title pins the bug on Loongson’s LA664 core — Rong “Mantle” Bao independently found that when the atomic variable’s address and an ordinary read’s address sit in a particular positional relationship, plain scalar reads trigger the lost update too, so the flaw can strike with no vector code involved at all, just less often. Trigger conditions that baroque, the author notes, are exactly how a defect like this hides for years.

This was no longer a weird software bug but a new CPU erratum — the catalogue entry chipmakers keep for silicon that lies. Loongson was told, and, credit where due, came back two weeks later with a fix carrying almost no performance loss, plus test firmware. As the write-up describes it, in a passage quoted in the Hacker News discussion:

> The fix is to set bit 13 of MCSR24 to 1. MCSR24 is an internal CSR whose function is not described in the manual. After setting this bit, the lost update no longer occurs. Testing showed the performance loss is very small: single-core performance is unaffected, and multi-core performance drops only slightly.

“Presumably this is a so-called ‘chicken bit’,” wrote Hacker News commenter TazeTSchnitzel — hardware slang for a just-in-case switch a chipmaker etches into the silicon so a feature can be switched off if it turns out to be haunted.

The team confirmed Loongson’s test firmware stops the lost updates, and Loongson says the fixed firmware is expected to ship before National Day, October 1, after which users can simply upgrade. Which leaves the pleasingly undignified truth of the matter: a processor that occasionally forgot how to count was cured by flipping a single switch its own manual refuses to admit exists.
