Researchers Put ChatGPT, Claude, And Grok Behind The Wheel Of A Real Car… It Didn’t Go Well

Image Credit: Screenshot / DrivingBench.

Artificial intelligence is becoming increasingly involved in the automotive industry, particularly in the development of autonomous driving technology. However, there’s a significant difference between a system engineered specifically for driving and a general-purpose chatbot attempting to control a moving vehicle.

A group of researchers recently decided to explore that difference through an unusual experiment. Instead of relying on a driving simulator, they connected several leading AI models to a real car and challenged them to navigate a course marked with traffic cones.

The experiment, called DrivingBench, was designed to evaluate how effectively language models could interpret visual information and translate it into driving commands. The researchers weren’t asking the systems to handle highway traffic or complicated city intersections, either.

Their challenge was a relatively short course laid out in a parking lot, with a human safety driver ready to intervene. Despite those controlled conditions, most of the AI models struggled to make it beyond the opening section.

Four AI Models, One Toyota Corolla

Drivingbench e1791620156102
Image Credit: Screenshot / DrivingBench.

Researchers Aditya Ramabadran, Simon Mahns, and Tobias Gessler used a 2022 Toyota Corolla to evaluate four AI models: Claude Fable 5.1, GPT-5.6 Sol, Grok 4.6, and GPT-6 Astra. The objective was to navigate a 134.7-meter course using visual observations and a sequence of driving commands.

Rather than controlling the vehicle continuously like a conventional autonomous driving system, each model received camera images and issued instructions specifying steering, speed, and duration. Vehicle speed was limited to approximately 8 mph, and a human was ready to apply the brakes.

During the 11 attempts, eight runs failed to progress beyond 11% of the course. Only one model eventually completed the entire route, although even that successful attempt was far from smooth.

Grok And GPT Struggled With The First Turn

Grok 4.6 failed to complete even 12% of the course in its three attempts. During one run, it interpreted a gap between cones as an intended opening and directed the Corolla outside the designated route.

GPT-5.6 Sol performed even worse, reaching just 6% completion on all three attempts. The model incorrectly assumed the cones followed a consistent color arrangement, despite instructions explaining otherwise.

Claude Fable 5.1 showed some improvement, progressing from 9% on its first attempt to 45% on its third. Nevertheless, it couldn’t complete the course, frequently spending extended periods stationary while deciding what to do next.

One AI Finished, Although Very Slowly

DB e1791620172792
Image Credit: Screenshot / DrivingBench.

GPT-6 Astra was the only model to successfully navigate the course, reaching 49% on its first attempt before finishing on its second. Even then, the completed run took more than five minutes, with the Corolla moving cautiously through the turns.

According to the researchers, Astra’s successful run cost $7.74 in token usage, nearly four times as much as its unsuccessful first attempt. The figures illustrate the computational demands of using a general-purpose AI model for a task requiring repeated observations and physical control.

The experiment also revealed a fundamental limitation involving reaction times. While the models were processing camera images and generating their next instructions, the vehicle could continue moving under its previous command, potentially making their observations outdated.

What This Means For Autonomous Driving

The researchers emphasize that DrivingBench evaluates general-purpose AI models, not the specialized software used by established autonomous driving companies. Systems developed specifically for vehicle control use different architectures and are designed to respond continuously to changing road conditions.

DrivingBench also tested each model within a particular software environment, meaning the results reflect the combined performance of the model and its control interface. With only 11 attempts, the experiment is an early demonstration rather than a definitive comparison of AI driving capabilities.

Author: Andre Nalin

Title: Writer

Andre has worked as a writer and editor for multiple car and motorcycle publications over the last decade, but he has reverted to freelancing these days. He has accumulated a ton of seat time during his ridiculous road trips in highly unsuitable vehicles, and he’s built magazine-featured cars. He prefers it when his bikes and cars are fast and loud, but if he had to pick one, he’d go with loud.

Leave a Comment

Flipboard