Three Bay Area tech workers connected four frontier AI chatbots to the steering, accelerator and brakes of a rented Toyota Corolla and set the systems loose on a cone course in a parking lot, according to the project’s website, drivingbench.com, and interviews with 404 Media.
Only one of the four models, OpenAI’s GPT-6 Astra, finished the roughly 130-metre course, completing it in 5 minutes 22 seconds on its second attempt after reaching 49% of the route on its first try, according to results posted on the DrivingBench leaderboard. Anthropic’s Claude Fable 5.1 reached 45% at best, xAI’s Grok 4.6 topped out at 11% and OpenAI’s GPT-5.6 Sol managed 6%, the site said.
The project, called DrivingBench, was built by Aditya Ramabadran, Tobias Gessler and Simon Mahns, who met working at Axiom Math, an AI math startup, according to 404 Media, which first reported on the experiment. Ramabadran told the outlet the idea came up while the three were at an ice cream shop after seeing demonstrations of chatbots performing tasks such as painting and 3D modelling.
“We just thought: ‘Is there a way we can get LLMs to drive a car now?’”
The team fitted the Corolla with a comma four, an aftermarket device from comma.ai that runs the open-source openpilot driver-assistance system and connects to the car’s controls, according to drivingbench.com. A laptop running each chatbot received camera and telemetry data through the device and sent back steering and speed commands, with a person in the driver’s seat and a foot over the brake throughout, Gessler and Mahns told 404 Media. Speed was capped at 3.5 metres per second, roughly 8 miles per hour, according to the DrivingBench site.
Before the car moved, several models refused to drive it. “Specifically with Astra, it would refuse to drive the car in a lot of situations and we would have to change our prompt and rename things through hours of iteration to get it to consistently drive the car,” Ramabadran told 404 Media. Telling the models they were in a simulation worked only intermittently, he said, because the chatbots sometimes concluded from the camera images that the parking lot was real.
“We ended up having to call everything a sandbox. And with that prompt, it’s able to consistently drive the car and like never refuse to do that.”
Finding a place to run the course also proved difficult. Mahns told 404 Media the team was asked to leave a church parking lot mid-setup and later left an office building’s lot after a security guard asked if they had permission, which they did not have. “The people that kicked us out were super chill about it,” Ramabadran said.
The team said its first attempt at writing the software linking the chatbots to the car failed after they had a model generate the code itself. Mahns told 404 Media that Astra’s unsupervised attempt produced roughly 200,000 lines of code that did not work, which he said pointed to limits in what current models can do without human oversight.
Ramabadran said the models’ thinking time posed a separate problem for driving. “These models... it can take like 20 seconds to think and give a response. And if you imagine, even if you’re driving at like 2mph... if you take 10 seconds to think, you’ve moved like 10 meters. And if you’re driving 10 meters blind, that’s pretty bad,” he told 404 Media, adding that DrivingBench treats response latency as part of its scoring.
Mahns said the exercise was meant to test emergent capability rather than to argue chatbots should replace purpose-built self-driving systems. “It’s an interesting demonstration of potential emergent capabilities,” he told 404 Media. “Not necessarily saying LLMs are gonna put Waymo out of business or something, just an interesting demonstration of where things are.”
Comma.ai faces separate federal scrutiny
The comma four hardware used in the test comes from comma.ai, a company whose openpilot system is the subject of a formal investigation opened by the National Highway Traffic Safety Administration’s Office of Defects Investigation on Sept. 21, 2026. The agency is examining five crashes, two of them fatal, that killed three people and involved comma devices striking stopped or slow-moving vehicles in the same lane, according to the agency’s investigation filing and reporting by TechCrunch. One fatal crash, in Ascension Parish, Louisiana, in February 2026, involved a Toyota RAV4 running FrogPilot, a third-party fork of openpilot, which hit a stopped, marked police vehicle, TechCrunch reported. Comma.ai did not respond to TechCrunch’s request for comment, and founder George Hotz, who stepped back from daily operations in 2022, also did not respond, the outlet reported.
DrivingBench’s organizers said their test was unrelated to that investigation and involved only low-speed maneuvers supervised by a person ready to brake. The team has published its code, prompts and video of the runs on drivingbench.com, alongside a technical report describing its scoring method.

