---
title: "When AI agents are trained with tools, the GPUs stand idle"
description: "A measurement study finds the waste depends on how slowly the tools start, not on the training system itself"
author: "Arthur Wren"
published: 2026-09-25T16:27:26.232Z
modified: 2026-09-26T11:30:03Z
url: https://rews.cc/a/when-ai-agents-are-trained-with-tools-the-gpus-stand-idle-e21353
language: en
tags: ["ai", "llm", "automation", "productivity", "openai", "tech"]
publisher: "Rews (https://rews.cc)"
---

# When AI agents are trained with tools, the GPUs stand idle

*A measurement study finds the waste depends on how slowly the tools start, not on the training system itself*

By Arthur Wren · September 25, 2026 · https://rews.cc/a/when-ai-agents-are-trained-with-tools-the-gpus-stand-idle-e21353

## In brief

- RL training of tool-using AI agents leaves rollout GPUs idle while CPU-side work like sandbox starts and code execution runs
- Amazon Science researchers measured the “environment bubble” on eight A100 GPUs with a 20 Hz profiler
- Bubble size is set by tool cold-start cost, not the RL system: idle share 0.80 with 8-second starts, 0.049 with sub-second Docker
- Speculative prewarming recovered no meaningful idle on real hardware, though a naive simulator can suggest otherwise

There is a kind of waste in the making of tool-using AI agents that more money alone will not cure, and it has now been measured directly. In reinforcement-learning post-training, the rollout GPUs — the most expensive hardware in the room — spend their time waiting. Each simulated agent pauses while ordinary CPU work happens around it: a sandbox starting from cold, code being run, documents fetched, an outside service answering its call.

Researchers writing on Amazon Science give this idle a plain name: the “environment bubble”. They measured it on a machine of eight A100 GPUs, using a profiler that samples the hardware twenty times a second. The method was chosen because the usual utilization counter can report an idle GPU as busy when a finished kernel still sits on it. Every idle instant was then assigned to one of three states: waiting on the environment, waiting on a straggler, or waiting on the pipeline.

The central finding is blunt. The size of the bubble is governed by the tool’s cold-start cost — the time it takes a fresh sandbox or container to stand up — and not by the reinforcement-learning system. This matters, because it tells the engineer where the fault lies, and where it does not.

Two regimes on real hardware mark the extremes. With an aggressive injected cold start of eight seconds — the sort of delay a fresh container, a spun-up virtual machine or browser, or a costly environment reset can impose — the share of idle spent waiting on the environment fell from 0.80 with four sandbox workers to 0.23 with sixty-four. Against a real Docker code-execution backend built on the python:3.11-slim image, whose cold start the researchers measured at 0.3 to 1.2 seconds, the same share was small from the start, falling from 0.049 to 0.008 across the same range, while effective GPU utilization rose to between 0.42 and 0.55.

The lesson follows plainly enough. The alarming figures appear only when a tool takes whole seconds to start. With realistic tools that answer in under a second, only a few percent of the idle time can be won back by any scheme. The authors’ advice is to measure the tool latency of the system you actually run before optimizing anything — sound counsel, and cheaper than the alternative.

Some remedies were tested and found wanting, and honesty about this is the better part of the study. Adding sandbox workers shrinks the bubble, but the gain diminishes as cold-start falls. A closed finite-source queue model of the kind written M/M/k//C was tried; because it overstates the savings extra workers bring, the researchers sized the worker count from the measured idle curve instead.

A lightweight predictor can guess the next tool call from real rollout token streams with an accuracy scored at AUC 0.81, and it needs no extra pass through the model. Even so, speculative prewarming — starting tools before they are asked for — recovered no meaningful GPU idle on real hardware in either regime. The authors note that a naive simulator, one that keeps a permanently warm pool and tests it with too few random seeds, can make prewarming look like a winner when it is not.

Every number in the study is labelled by how it was obtained, which is more than many papers in this field can say. The conclusion is worth keeping in plain words: the bubble is real, but its size belongs to the tool, not the trainer, and a simulator that flatters your fix is no fix at all.
