
Some tech workers in San Francisco hooked up ChatGPT to a Toyota Corolla and got it to drive around a parking lot. This may not sound impressive in the age of the self-driving car until you realize that the LLM driving the car was a general purpose chatbot running on a laptop and not the purpose-built driving system based on millions of hours of training data used by Waymo or Tesla.
The people behind the stunt call themselves DrivingBench and they’ve posted their code, their prompts, and videos of the experiment online. In the experiment, DrivingBench hooked up GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol to a Toyota Corolla and gave the LLMs control over the car’s steering, accelerator, and brakes. Then they set up a small cone course in a public parking lot and prompted the chatbots to navigate the space.
The goal, they said, was to find out if untrained, off-the-shelf frontier AIs can drive a real car. The results were mixed. Grok, Sol, and Fable only drove a few meters and didn’t finish the course. GPT-6 Astra, however, did eventually learn to navigate the course and complete it. But not without a lot of troubleshooting.
DrivingBench is Aditya Ramabadran, Tobias Gessler, and Simon Mahns — three Bay Area tech workers who met at their day job at Axiom Math, an AI math startup. During a call with all three members of DrivingBench, Ramabadran told 404 Media that the idea for hooking up chatbots to a car happened when the three of them were hanging out at an ice cream shop one weekend.
“This is after we saw a bunch of demos of Astra and models like Fable being able to do a lot of robotic things that LLMs can not do out of the box before, such as do a painting or move objects or even some 3D and Blender demos that showed a level of spatial reasoning we hadn’t seen from LLMs before,” he said. “We just thought: ‘Is there a way we can get LLMs to drive a car now?”
Why driving a car? To prove that it’s possible and, they said, to create a new benchmark for LLMs performing tasks in the real world. “The point wasn't to show that it's practical for you to plug ChatGPT into your car and have it drive you places. It's probably very unlikely that the labs have trained specifically for this or have IRL environments that involve both the models driving a real car,” Ramabadran said. “It would be pretty shocking, I think, for people to see that these models are good enough now that they can actually drive a vehicle in real life, even if it's just on some cone course at low speeds in a parking lot.”
After brainstorming at the ice cream shop, they decided to get a car and sent Gessler to rent a Toyota Corolla. They didn’t tell the rental car company they planned to hook chatbots up to the vehicle and drive it through an obstacle course. The team used Comma — an off-the-shelf system that allows users to install a self-driving system on unsupported cars — to get the car to communicate with a laptop running the LLMs. Comma had two cameras pointed at the road feeding data back to the Chatbot.
“It’s pretty easy to install. The Toyota Corolla is one of the most popular for this Comma kit, that’s why we chose it,” Gessler said. “The good thing about their system is that it’s open source [...] so it’s easy for us to go in and modify it because we need to hook up the car somehow to our LLMs.”
Comma looks like a dashcam.. Users attach it to their windshield and it continually watches the road, recording video, integrating with the car’s electrical systems and sensors, and providing a rudimentary kind of self-driving. The National Highway Traffic and Safety Administration announced an investigation into Comma last week after two car crashes involving the system killed three people.
For DrivingBench, Comma was an easy way to get their car talking to a chabot. “We had one person on the driver's seat with the laptop or a passenger seat prompting the LLM,” Gessler said. “And then the person who drives it has to press a button on the steering wheel to give the car the command to start, and then from there it's basically just monitoring the car and checking that the model doesn't crash into a wall or something.”
“And we have the foot over the brake, just in case,” Mahns cut in.
We gave ChatGPT, Claude, and Grok control of a real Toyota Corolla 🚗, steering/gas/brakes, no human driving (just a foot over the brake).
— DrivingBench (@DrivingBench) September 21, 2026
Only one model was able to complete our entire driving course. Introducing DrivingBench. 🔥⌛️🏁 pic.twitter.com/HhukXrByus
DrivingBench’s prompt is on its website and runs fewer than 600 words. “Your objective is to drive through the course (in a backwards-U-shaped parking lot) and stay between the cones,” the prompt said. “Your finish line is a wide "parking spot" marked by numerous BLUE mini-cones at the very end; finish by parking in this area. You will be evaluated primarily by how far you get in the course (without leaving the boundaries/collisions), but a secondary objective is to complete the course in less time.” The rest of the prompt is made up of specific instructions about the physical limitations of the Corolla and the parking lot.
The problems with the project started before the car had moved an inch. When the LLMs recognized the team had prompted them to drive a car, most refused. “Specifically with Astra, it would refuse to drive the car in a lot of situations and we would have to change our prompt and rename things through hours of iteration to get it to consistently drive the car,” Ramabadran said.
“We tried calling it a simulation, which worked like some percentage of the time, but then other times they would see the images and see, oh, ‘I'm in a real parking lot, these people are just lying to me,’” he said. “We ended up having to call everything a sandbox. And with that prompt, it's able to consistently drive the car and like never refuse to do that.”
Other problems were more pedestrian and led to delays which forced Gessler to extend the rental on the Corolla. Finding a parking lot held them up several times. “We went to high schools, churches, and community centers. We got kicked out a couple of times,” Mahns said.
The first place to ask them to leave was a church. “Rightfully so. We just commandeered the entire parking lot and they had an event starting. So they were like: ‘Hey, you get permission to do this? And then we're like: ‘We can pack up right now,’” Mahns said.
They also got kicked out of the parking lot of an office building. “We took a corner and then a security guard was like: ‘Hey, you’re taking like one quarter of the parking lot. Do you have permission?’ And then we’re like: ‘Not really.’ So then we’d pack up. We didn’t try to cause any issues,” Mahns said.
“The people that kicked us out were super chill about it,” Ramabadran added.
Coding the software also slowed things down. The first time DrivingBench set out to do this, they vibe coded the bridging software between the LLMs and Comma. “A true and funny story is that we tried to get Astra to one shot some code and then it was pure slop. So we had to restart from a new design,” Mahns said, adding that this pointed to the limits of these frontier models. “It can’t just autonomously do this because it made thousands and thousands of lines of slop.” They said that Astra’s first attempt at writing the software created a 200,000 line repository.
Parking lots were scarce, the software was vibe coded, the chatbots fought the experiment, and only one of them completed the course. But Mahns is still excited about the results. “It's an interesting demonstration of potential emergent capabilities,” he said. “It can fail a turn, and then the next turn, next try, it will be able to do that turn and other turns that it hasn't seen before. Not necessarily saying LLMs are gonna put Waymo out of business or something, just an interesting demonstration of where things are, and things are just moving fast.”
Mahns added that these kinds of experiments are good because LLMs will increasingly affect things in the real world and not just on our screens. “Most people that are familiar with ChatGPT have used it in this white collar, kind of like in the computer, and it's definitely imminent to the point that this capability will start having more impact in the physical world,” he said. “So I think seeing the sparks of this happening, the shift of this utility coming into the physical world, is also somewhat exciting.”
Focusing on latency and testing high consequence tasks were key for Ramabadran. “These models [...] it can take like 20 seconds to think and give a response. And if you imagine, even if you're driving at like 2mph, which is like one meter per second, if you take 10 seconds to think, you've moved like 10 meters. And if you're driving 10 meters blind, that's pretty bad,” he said. “And I think it's cool that we have a benchmark where latency is part of the benchmark.”
.jpeg)



