The hard part of artificial-intelligence video is no longer making beautiful frames. It is making them fast enough to matter. A clip rendered overnight can afford to dawdle; a world a user is steering cannot. A frame that arrives late shatters the illusion, because, as the teams at Reactor and Amazon’s Neuron Science group put it in an account published by Amazon Science, generation must stay ahead of the playback timeline. Their solution involves abandoning the compiler and getting their hands dirty with the chip.
Sparing a thought for history first: Jürgen Schmidhuber, who in 1990 was the first person to propose world models as a machine-learning concept, published a paper in 2018 with David Ha describing “a predictive world model” that could “extract useful representations of space and time” and use them to train agents to drive. In the eight years since, world models—fed by vast datasets, variational autoencoders, and both diffusion and autoregressive transformers—have learned to convert text, images, audio and video into latent space, then maintain and update explorable environments on demand. They now underpin work in robotics, transport, climate modelling, video-game development, scientific simulation and design prototyping.
What they demand in return is hardware that never pauses. That is the brief of Reactor, a platform that lets developers, designers and researchers deploy, use and scale real-time interactive AI models. “It’s about real time, it’s about low latency, and it’s about doing that as efficiently as you can at scale,” says Bryce Schmidtchen, the firm’s co-founder and chief technology officer. Efficiency, he says, runs from “how you schedule the inference on the given chip, in our case Trainium, to how you think about maximally bin packing every forward pass of the model”. The job is made no easier by geography: “We think about GPU clusters in terms of regions — we have hundreds all over the world.” A pixel latent must reach the network without wandering through Kubernetes, whether the cluster is well supported or, in Mr Schmidtchen’s phrase, “bare metal in a closet”, and tie into media streaming that copes with different codecs, resolutions and interfaces. The global reach of AWS, Amazon’s cloud arm, already supplies much of that, he says.
Sequencing the dream
The Neuron Science team, meanwhile, hunts for bottlenecks. Its job, explains Jun Wu, a principal applied scientist, is to explore “new techniques for generative AI model enablement optimization”—a new architecture, a new algorithm, new code for running models on Trainium, Amazon’s in-house AI chip—and to prototype what can be ported into production. The team noticed video generation evolving from short, fixed-length clips towards infinite or dynamic lengths, a shift that produced autoregressive diffusion models. These marry the sequential next-token prediction of large language models with the iterative refinement of diffusion. “Autoregressive diffusion means a user keystroke can be absorbed as input to the model and, conditioned on the previously generated video frame, the model can correctly decide the next move or the next scene,” says Mr Wu. “Most of the interactive video generation models we are seeing today are using this.”
Mr Wu and his colleagues—Mason Fu and Lingfan Yu, both senior applied scientists—saw room to optimise such models on Trainium and found a partner in Reactor, which had been working on real-time interactive video for over 18 months. The teams settled on Rolling Forcing, a technique that can consistently generate high-quality 30-second videos while staying relatively small in parameter count. Small did not mean easy. Unlike offline generation, which renders a whole clip and hands it over when done, Rolling Forcing streams: each frame is generated and immediately consumed or fed back as the next input. “High-quality generation diffusion models pose the challenge of a very long sequence, which requires a lot of memory consumption,” says Mr Wu. “Rolling Forcing is relatively small, but its sequence length is very large.” Add a hard floor—the model generates 16 frames per second, the standard for video playback—and the goal becomes real-time generation comfortably above 16 fps, sustained by Trainium’s performance and memory, rather than good frames produced eventually.
Down to the metal
Three faults recur on every forward pass and collapse ordinary compilers: dynamic shapes, unusual memory-access patterns and heavy cache management. Shapes change because scenes do—a pitcher alone on the mound becomes, in a single pan, a stadium of thousands—and because frames are generated and evicted in a continuous sliding window. “This means attention lengths vary across forward passes, and there are two distinct passes per window: denoising, then cache cleanup,” says Mr Wu. “All of that makes static compilation hard.” Then there are operations whose access patterns are specific to the workload. Rotary position embeddings, which account for one position in a language model, must contend with three axes in video—height, width and time. “The 3D rotary embedding interleaves odd and even elements along the innermost dimension, producing many tiny data transfers when compiled generically,” Mr Wu explains. Finally, every layer reads and writes a rolling key-value cache at every step, a relentless drag on memory throughput.
The team’s answer is the Neuron Kernel Interface, which lets developers write compute kernels that run directly on NeuronCore hardware, with precise control over how data moves between memory and the compute engines. The point, says Mr Wu, is a self-service path: where profiling shows that ordinary compilation falls short, engineers replace the offending operations with hardware-tuned implementations rather than wait for the toolchain to wise up. Kernel by kernel, the bottleneck operations are swapped out.
The exercise is revealing beyond the frame rate. Interactive video is where generative AI’s economics turn punitive: a customer pays for inference in real time, second by second, so every wasted cycle is margin spent and every millisecond is product lost. That is presumably why a cloud firm that designs its own chips and a startup serving hundreds of regions found common cause at the level of cache cleanup. If world models are to graduate from impressive demos to games, robots and simulations, the contest will be won less in the model than in the kernel. The dream machine must first learn to keep time.

