Three AI engineers put a text-generating model in control of a 2024 Toyota Corolla long enough to reach an In-N-Out take-out window. A safety driver stayed ready to brake as the model used camera feeds and the car’s power steering to move through the drive-thru.
A language model meets a real car
Aditya Ramabadran, Simon Mahns, and Tobias Gessler, engineers at startup Axiom, set up the experiment as a project outside their work. They linked a chat interface to a server connected to several cameras mounted on the windshield and the vehicle’s power steering system.
The setup put a general-purpose model into a task usually handled by software specifically engineered for driving. The model, OpenAI’s GPT-6 Astra, was designed to generate text, code, and images. In this case, it had to interpret what its cameras showed and issue commands that moved a car.
The engineers asked it to navigate to the drive-thru window. It proceeded slowly and made it to the point where they could collect lunch. A person in the car remained prepared to intervene, underscoring that the demonstration was a controlled experiment, not an unattended ride.
What the test may show about physical reasoning
AI systems have become capable in digital settings, including answering complex questions and carrying out some virtual tasks. Their ability to make sense of the physical world is a different challenge: a model needs to connect what it sees with how objects and spaces behave, then choose actions that work outside a computer screen.
The Corolla experiment suggests that language-based models may have some early ability to do this. The engineers had not coached the models specifically to drive, and they were surprised when the systems began controlling vehicles after initially refusing. The team says the models’ driving ability may have emerged from broader training in spatial reasoning, rather than from explicit lessons in operating a car.
That possibility matters beyond driving. Andrew Dai, CEO of Elorian AI and a former Google DeepMind researcher, points to applications such as understanding whether restaurant diners are enjoying a meal and enabling robots to function in homes. He describes physical reasoning as essential for robotics.
Elorian AI and Scale AI developed Humanity’s Sixth Sense, a benchmark intended to measure how well models understand physical scenes. Its premise is that visual capability involves more than recognizing objects: researchers want to know whether a system can grasp a scene intuitively, in a way that supports sensible action.
Early driving results are still limited
The engineers also created DrivingBench, a benchmark that tests models on a simple course laid out in a parking lot. The results show a wide gap between reaching a drive-thru window under supervision and demonstrating dependable driving ability.
- GPT-6 Astra completed the course, but very slowly.
- Claude Fable 5.1 made it 45 percent of the way around.
- Grok made it 11 percent of the way around.
Ramabadran says the latest models appeared to adjust to mistakes and learn how to use the car’s controls during the task. That kind of adaptation could help explain how a model performs without task-specific preparation, but the benchmark results make clear that performance is uneven.
Promise comes with real-world stakes
Nothing overtly alarming happened during the lunch run, but a moving car leaves little room for error. The engineers acknowledge the high stakes of letting a general-purpose model control a fast-moving, two-ton vehicle. A successful short demonstration does not establish that a model can handle the range of situations involved in driving.
The experiment is better read as a sign of a developing capability than as proof that AI can safely take over the road. If models gain stronger physical understanding, they could become useful in robotics and other settings where software must respond to the world. The same shift could also bring new risks when a mistaken decision affects people and objects outside a screen.