Open-weight A.I. models have become good enough for everyday coding work, a German technology consultant wrote on Thursday, showing how he runs one at about 150 tokens per second on a rented server-grade graphics card that he starts with one command and deletes when he is finished.

The consultant, Christian Hofstede-Kuhn, said he had tested the 27-billion-parameter Qwen 3.8 model against real tasks: shell scripts, Ansible roles, Python tooling, refactoring, tests, explaining an unfamiliar codebase. For most of that work, he wrote, the open model is “sufficient.” Frontier systems such as Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 still answer better on hard architectural questions and long chains of reasoning across many files, he said. The gap is real, he added; it just no longer sits where his work happens.

His setup runs on an NVIDIA RTX PRO 6000 Blackwell, rented by the hour from RunPod. He adapted a freely licensed stack, qwen38-runpod-stack by the developer nicremo, and built it around the model’s Apache 2.0 base. The stack defaults to an “abliterated” build, one whose weights were modified after training to suppress much of the model’s refusal behavior.

The card is chosen for three reasons, Mr. Hofstede-Kuhn wrote. The 4-bit NVFP4 number format needs Blackwell’s native FP4 tensor cores, which even an H100 or H200 lacks, so older hyperscale cards fall onto a slower path. Its 96 gigabytes of memory holds the model’s full 262,000-token context, where an RTX 5090’s 32 gigabytes forces deep cuts. And the inference engine, SGLang, documents this pairing end to end. In his running system the weights take 20.6 gigabytes and the speed-up draft model 3.8 more, with most of the card’s memory reserved for working data.

He described a trap he said is worth knowing even off RunPod. One widely used NVFP4 build of the model is stored in a mixed-precision format that SGLang loads completely unquantized, without a warning, according to the project’s bug tracker. The model runs and answers normally, and the entire speedup is gone. His fix: use a uniformly quantized checkpoint, and check the format field in the model’s configuration file if an FP4 model feels as slow as a bigger one.

Speed comes partly from speculative decoding, in which the small DFlash2 draft model proposes eight tokens ahead and the large model verifies them in one pass. The upstream project’s tests, with an 8,192-token input and 1,024 tokens of output across five runs, measured a median of 150.7 tokens per second, with runs between 122 and 156. The same model on a MacBook Pro M2 Max manages about 14, the project says. At 150, Mr. Hofstede-Kuhn wrote, the model is faster than he can read, and the bottlenecks move to compilers and tests.

The deeper reason he bothers, he said, is data sovereignty. A hosted A.I. service puts every prompt and every proprietary file onto someone else’s systems, governed by terms that can change; his own server lets him choose the machines and software involved. A session costs nothing but the hourly rental, starts in 8 to 15 minutes most of it spent downloading 20 gigabytes of weights and, as he put it, it fits the rhythm of work: “start it, then read the ticket properly.”