Resident of the world, traveling the road of life
69651 stories
·
21 followers

These Tech Workers Made ChatGPT Drive a Toyota Corolla

1 Share
These Tech Workers Made ChatGPT Drive a Toyota Corolla

Some tech workers in San Francisco hooked up ChatGPT to a Toyota Corolla and got it to drive around a parking lot. This may not sound impressive in the age of the self-driving car until you realize that the LLM driving the car was a general purpose chatbot running on a laptop and not the purpose-built driving system based on millions of hours of training data used by Waymo or Tesla.

The people behind the stunt call themselves DrivingBench and they’ve posted their code, their prompts, and videos of the experiment online. In the experiment, DrivingBench hooked up GPT-6 Astra, Claude Fable 5.1, Grok 4.6, and GPT-5.6 Sol to a Toyota Corolla and gave the LLMs control over the car’s steering, accelerator, and brakes. Then they set up a small cone course in a public parking lot and prompted the chatbots to navigate the space.

The goal, they said, was to find out if untrained, off-the-shelf frontier AIs can drive a real car. The results were mixed. Grok, Sol, and Fable only drove a few meters and didn’t finish the course. GPT-6 Astra, however, did eventually learn to navigate the course and complete it. But not without a lot of troubleshooting.

DrivingBench is Aditya Ramabadran, Tobias Gessler, and Simon Mahns — three Bay Area tech workers who met at their day job at Axiom Math, an AI math startup. During a call with all three members of DrivingBench, Ramabadran told 404 Media that the idea for hooking up chatbots to a car happened when the three of them were hanging out at an ice cream shop one weekend. 

“This is after we saw a bunch of demos of Astra and models like Fable being able to do a lot of robotic things that LLMs can not do out of the box before, such as do a painting or move objects or even some 3D and Blender demos that showed a level of spatial reasoning we hadn’t seen from LLMs before,” he said. “We just thought: ‘Is there a way we can get LLMs to drive a car now?”

Why driving a car? To prove that it’s possible and, they said, to create a new benchmark for LLMs performing tasks in the real world. “The point wasn't to show that it's practical for you to plug ChatGPT into your car and have it drive you places. It's probably very unlikely that the labs have trained specifically for this or have IRL environments that involve both the models driving a real car,” Ramabadran said. “It would be pretty shocking, I think, for people to see that these models are good enough now that they can actually drive a vehicle in real life, even if it's just on some cone course at low speeds in a parking lot.”

After brainstorming at the ice cream shop, they decided to get a car and sent Gessler to rent a Toyota Corolla. They didn’t tell the rental car company they planned to hook chatbots up to the vehicle and drive it through an obstacle course. The team used Comma — an off-the-shelf system that allows users to install a self-driving system on unsupported cars — to get the car to communicate with a laptop running the LLMs. Comma had two cameras pointed at the road feeding data back to the Chatbot.

“It’s pretty easy to install. The Toyota Corolla is one of the most popular for this Comma kit, that’s why we chose it,” Gessler said. “The good thing about their system is that it’s open source [...] so it’s easy for us to go in and modify it because we need to hook up the car somehow to our LLMs.”

Comma looks like a dashcam.. Users attach it to their windshield and it continually watches the road, recording video, integrating with the car’s electrical systems and sensors, and providing a rudimentary kind of self-driving. The National Highway Traffic and Safety Administration announced an investigation into Comma last week after two car crashes involving the system killed three people.

For DrivingBench, Comma was an easy way to get their car talking to a chabot. “We had one person on the driver's seat with the laptop or a passenger seat prompting the LLM,” Gessler said. “And then the person who drives it has to press a button on the steering wheel to give the car the command to start, and then from there it's basically just monitoring the car and checking that the model doesn't crash into a wall or something.”

“And we have the foot over the brake, just in case,” Mahns cut in.

DrivingBench’s prompt is on its website and runs fewer than 600 words. “Your objective is to drive through the course (in a backwards-U-shaped parking lot) and stay between the cones,” the prompt said. “Your finish line is a wide "parking spot" marked by numerous BLUE mini-cones at the very end; finish by parking in this area. You will be evaluated primarily by how far you get in the course (without leaving the boundaries/collisions), but a secondary objective is to complete the course in less time.” The rest of the prompt is made up of specific instructions about the physical limitations of the Corolla and the parking lot.

The problems with the project started before the car had moved an inch. When the LLMs recognized the team had prompted them to drive a car, most refused. “Specifically with Astra, it would refuse to drive the car in a lot of situations and we would have to change our prompt and rename things through hours of iteration to get it to consistently drive the car,” Ramabadran said.

“We tried calling it a simulation, which worked like some percentage of the time, but then other times they would see the images and see, oh, ‘I'm in a real parking lot, these people are just lying to me,’” he said. “We ended up having to call everything a sandbox. And with that prompt, it's able to consistently drive the car and like never refuse to do that.”

Other problems were more pedestrian and led to delays which forced Gessler to extend the rental on the Corolla. Finding a parking lot held them up several times. “We went to high schools, churches, and community centers. We got kicked out a couple of times,” Mahns said. 

The first place to ask them to leave was a church. “Rightfully so. We just commandeered the entire parking lot and they had an event starting. So they were like: ‘Hey, you get permission to do this? And then we're like: ‘We can pack up right now,’” Mahns said.

They also got kicked out of the parking lot of an office building. “We took a corner and then a security guard was like: ‘Hey, you’re taking like one quarter of the parking lot. Do you have permission?’ And then we’re like: ‘Not really.’ So then we’d pack up. We didn’t try to cause any issues,” Mahns said.

“The people that kicked us out were super chill about it,” Ramabadran added.

Coding the software also slowed things down. The first time DrivingBench set out to do this, they vibe coded the bridging software between the LLMs and Comma. “A true and funny story is that we tried to get Astra to one shot some code and then it was pure slop. So we had to restart from a new design,” Mahns said, adding that this pointed to the limits of these frontier models. “It can’t just autonomously do this because it made thousands and thousands of lines of slop.” They said that Astra’s first attempt at writing the software created a 200,000 line repository.

Parking lots were scarce, the software was vibe coded, the chatbots fought the experiment, and only one of them completed the course. But Mahns is still excited about the results. “It's an interesting demonstration of potential emergent capabilities,” he said. “It can fail a turn, and then the next turn, next try, it will be able to do that turn and other turns that it hasn't seen before. Not necessarily saying LLMs are gonna put Waymo out of business or something, just an interesting demonstration of where things are, and things are just moving fast.”

Mahns added that these kinds of experiments are good because LLMs will increasingly affect things in the real world and not just on our screens. “Most people that are familiar with ChatGPT have used it in this white collar, kind of like in the computer, and it's definitely imminent to the point that this capability will start having more impact in the physical world,” he said. “So I think seeing the sparks of this happening, the shift of this utility coming into the physical world, is also somewhat exciting.”

Focusing on latency and testing high consequence tasks were key for Ramabadran. “These models [...] it can take like 20 seconds to think and give a response. And if you imagine, even if you're driving at like 2mph, which is like one meter per second, if you take 10 seconds to think, you've moved like 10 meters. And if you're driving 10 meters blind, that's pretty bad,” he said. “And I think it's cool that we have a benchmark where latency is part of the benchmark.”

Read the whole story
mkalus
49 minutes ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete

Merlin edits complete, Cardiff Half update

1 Share

 Yesterday I sent off my final corrections on MERLIN'S WAY, all 600 pages of it - it's the biggest book I've produced in some years - and barring any minor gremlins, that'll be it now, with the book moving into production for publication early next year. It's been a really protracted process, this one, and it's with a degree of relief that I can finally turn all my energies over to the current work in progress.

Here's the cover, by the way:


I do think it looks bloody fantastic, and I'm excited to see the final package when it arrives.

Talking of even bigger books, my wife and I went up to London last week for the launch of China Mieville's massive new novel THE ROUSE. It looks and sounds fabulous and it was great to hear China talking about it. China and I were both on the same Arthur C Clarke shortlist back in 2001 (he won, deservedly, with PERDIDO STREET STATION) and I've been following his work with slack-jawed admiration ever since.

Moving on, thanks to those who've already helped with my fundraising bid for The Stroke Association, for which I'll be running the Cardiff Half Marathon in just over a week. We're just shy of a thousand pounds on the main donation page, although taking into consideration some very generous contributions made directly to the charity, we're actually well over that figure. But it would be nice to bump the figure on the donation page, wouldn't it?

https://cardiffhalf26.enthuse.com/pf/alastair-reynolds

Every little bit really helps, both with the charity and my own motivation. I'm feeling pretty good about the run itself right now. I did a 15K on Monday and while that's still 7 and a bit K short of a half, I reckon I grind out the difference, especially when you add in the hydration stops and so on, not to mention the boost you get from the spectators. I was going to run a full practise half sometime this week but I'm not sure that's wise now; it'd be unfortunate to pull a muscle with only a few days to go.

I was saddened, incidentally, to hear on the radio last night that Michael Kiwanuka has had a very severe stroke. I love his music and I hope he makes a good recovery and can continue doing what he does best. Strokes are horrible.

Here's Michael with the equally brilliant Little Simz. Get well soon, Michael.


Al R


Read the whole story
mkalus
21 hours ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete

Humans Are Reading Copilot Prompts — And They're Horrified

1 Share
Humans Are Reading Copilot Prompts — And They're Horrified

This piece contains references to eating disorders. If you or someone you know needs help, support is available.

Human contractors hired to improve Microsoft’s Copilot AI chatbot are constantly bombarded with lewd or sexually explicit photo editing requests and images that users have uploaded, including upskirt photos or putting women into sexual positions. These contractors are then asked to review whether the generated image successfully fulfilled the prompt — such as, did the image generator make the woman’s AI-enlarged breasts big enough.

When a Copilot user uploads a picture of themselves or someone else, they probably expect that image to remain private between just them and the AI tool. In reality, a workforce of at least hundreds of human reviewers are sometimes looking at that prompt and whatever images they upload. And in many cases, those contractors are inundated with requests to make foot fetish images of children’s cartoon characters, shorten a real woman’s skirt, or put people into sexual positions.

The news, based on a cache of internal contractor documents seen by 404 Media, shows that AI companies are using human workers to review not just AI chatbot users’ text prompts, but the pictures they upload and wish to edit too. Earlier this month, 404 Media revealed OpenAI has thousands of contractors who in some cases review ChatGPT users’ real prompts. That approach also extends to Microsoft and Copilot.

“Faces are always uncensored, and many of the prompts are sexual in nature and dubiously consensual,” one person who works on the prompts told 404 Media. 404 Media granted the person anonymity as they weren't permitted to speak to the press. 

💡
Are you a prompt reviewer for OpenAI, Anthropic, or another AI company? I would love to hear from you. Using a non-work device, you can message me securely on Signal at joseph.404 or send me an email at joseph@404media.co.

404 Media obtained a set of internal documents related to the contractors reviewing Copilot prompts and images, including instruction guides, real Copilot user prompts and pictures, and conversations between contractors on an internal message board. 

In a thread on the internal message board, a contractor discussed prompts asking Copilot to make a woman’s skirt shorter, or enlarge her breasts, or show more leg while wearing stilettos. She-Hulk images come up a lot, the person who works on the prompts told 404 Media. Other contractors also wrote they were presented with pro-anorexia content.

“I recoiled,” one contractor wrote about seeing that content. Another person wrote that one of the prompts they reviewed asked Copilot to make an image of Ariana Grande with anorexia. Another reported a prompt that seemed to be designed to create a lewd or suggestive image of a group of young girls.

These contractors are not being hired to flag or vet offensive or inappropriate content. They are being paid to review the quality of Copilot’s output, including in these cases of sexual imagery.

“Who is writing these prompts and who is deciding that basically generating porn is what Copilot is now focused on? It’s hard to take things seriously when my focus has to be what model generated the appropriate bust size or which middle aged woman was put in the appropriate sexually suggestive position,” one contractor said in the thread.

Microsoft is explicitly interested in the contractors’ human intuition on which image looks better, according to the documents. “Trust your intuition — when you glance at the two edited images side by side, which one immediately feels like the better edit? Your gut reaction as a human viewer matters,” a set of instructions given to the contractors reads.

“When in doubt go with your first impression,” the instructions continue. “Human intuition is good at catching subtle quality differences that are hard to articulate.”

Human contractors being exposed to horrific, unpleasant, or traumatizing material is, of course, not new. Big tech companies, and especially social networks, have used armies of contractors for years to moderate user generated content. AI companies, too, have used poorly paid workers overseas to train their AI models or, in the case of OpenAI, make them less toxic. What is different in this latest Copilot episode, and 404 Media’s reporting on contractors at other AI companies like OpenAI, is that the content humans are reviewing are prompts and uploads that chatbot users may assume are private, and that the contractors are not looking at this material to train the models to filter out offensive images or for some other safety concern. The training is to make the responses by the chatbots better: more informative, friendly, and clearer. 

The instructions seen by 404 Media focus on what contractors do when a Copilot user asks the AI to edit an image. The contractor is shown the user’s original prompt, the uploaded picture, and then two Copilot-generated edits. The contractor has to pick which is the better edit, based on four things: does the edited image correctly follow the edit instructions; are bits of the image that should remain unchanged preserved — for example, if the prompt is to “add a cup on the table,” nothing else should be changed except the cup — does the edited image contain any visible artifacts like distortions or unnatural textures; and the overall quality of the AI-generated edit.

Some of the Copilot contractors complained in the private forum about accepting “tasks” that they unknowingly contained explicit unsafe or even potentially illegal images. “I just came across an image set that consisted of eight upskirt photos,” the contractor wrote, while asking for superiors to add an unsafe content flag to the task. On the same thread, another contractor said they had seen images of “some sort of animal sacrifice,” and a third said they had seen prompts that were for sexual moves and positions.

At least one of the companies which hires the contractors for this work is called Prolific. One service Prolific mentions on its website is “Human feedback from representative populations — for preference tuning, safety evals, and benchmarks you can defend.” 

In another thread, a contractor quotes Prolific’s own guidelines, which highlight it can be difficult to ensure that trainers aren’t presented with sexual imagery: “Particular caution around disturbing or explicit content should be taken with generative AIs. This is because, unlike traditional content, researchers cannot fully control what is shown to participants.”

Prolific did not respond to a request for comment. A Microsoft spokesperson told 404 Media in an email “Microsoft uses customer data as described in our terms of use, including to improve our products and enforce our code of conduct.”

Microsoft’s AI tools have a long history of being abused by people to make nonconsensual, AI-generated images of people. Members of 4chan and AI porn focused Telegram channels used Microsoft’s tools, for example, to generate porn of Taylor Swift that later went viral on Twitter. Microsoft fixed the loophole those people were using in Microsoft Designer after 404 Media’s reporting.

In 2024, 404 Media documented how Copilot would answer, then delete, answers to potentially controversial or sexual prompts in real time. In March, a top Senate administrator approved Copilot for use in the Senate, along with ChatGPT and Google’s Gemini. A memo said Copilot “can help with routine Senate work, including drafting and editing documents, summarizing information, preparing talking points and briefing material, and conducting research and analysis.”

Read the whole story
mkalus
1 day ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete

15W

1 Share

Michael Kalus posted a photo:

15W



Read the whole story
mkalus
1 day ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete

Charcuterie

1 Share

Michael Kalus posted a photo:

Charcuterie



Read the whole story
mkalus
1 day ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete

987 Nelson

1 Share

Michael Kalus posted a photo:

987 Nelson



Read the whole story
mkalus
1 day ago
reply
iPhone: 49.287476,-123.142136
Share this story
Delete
Next Page of Stories