Apple researchers on Thursday introduced a system called SCLATE that compresses a month of testing for A.I. agents into a few hours, and they used it to compare competing agent memory systems head to head.

The paper, SCLATE: A Substrate for Continual-Learning Agent Training and Evaluation, takes aim at a practical problem. Continual-learning agents are systems of models, harnesses and memory working across many sessions, and judging them means interleaving tasks with events like session stops, scheduled jobs and memory consolidation. Existing benchmarks schedule only their own events, the authors wrote, so every benchmark-and-agent pair ends up needing a custom scheduling loop.

In SCLATE, benchmarks and unmodified agents each register events on a single open scheduler through an adapter. A hybrid simulated clock ticks in real time while the agent works and skips the idle gaps, which is what squeezes a monthlong scenario into hours. The platform also works as a rollout engine: it runs any agent’s harness and memory unmodified and records the tokens and log probabilities of every model call through a proxy inside the agent’s container.

The team ported seven benchmarks to SCLATE and compared 10 harness-and-memory configurations across 10 models. An added memory system did not reliably beat the harness’s built-in native memory, and models differed widely in how they used the same harness and memory.

The researchers then post-trained a small open model, Qwen3.5-4B, through unmodified harnesses and memory systems. The model learned to use both, according to the paper. It read 6.8 times fewer file lines while its pass rate on SWE-bench Verified rose 16.7 points, it wrote richer memory records, and its accuracy on MetaClaw, a benchmark held out of training, climbed by as much as 11.8 points.

The authors are Youngmok Jung, Sirajul Salekin, Henry Tran, Javier Movellan, Zhao Huang and Manjot Bilkhu.