Language Models Take the Wheel: Researchers Test GPT-6 Astra's Real-World Driving Abilities

Engineers at startup Axiom successfully commanded OpenAI's GPT-6 Astra model to autonomously navigate a Toyota Corolla through a drive-thru without specialized autonomous driving training. The experiment demonstrated that general-purpose language models are developing rudimentary physical understanding capabilities beyond their typical text and code generation functions. While the stunt succeeded, researchers acknowledge the significant safety risks of deploying untrained AI models in control of vehicles.
The experiment involved three engineers at Axiom who connected GPT-6 Astra's text-based interface to vehicle cameras and steering controls, allowing the model to autonomously navigate a Toyota Corolla through a drive-thru without prior training on driving tasks. This differs from conventional autonomous vehicles, which rely on algorithms specifically engineered for transportation. The researchers were testing whether general-purpose language models could transfer their understanding of complex systems into physical-world navigation.
Researchers at companies like Elorian AI and Scale AI are developing new measurement tools to assess how well AI models understand physical scenes intuitively, recognizing this as a critical gap in current AI capabilities. Physical reasoning is considered foundational for emerging applications including home robotics and environmental awareness systems, representing a frontier that major AI labs have yet to fully master despite claims about advanced AI capabilities.
This development could affect multiple sectors reliant on autonomous systems, from transportation and logistics to manufacturing and domestic robotics. While the successful demonstration suggests language models may acquire broader reasoning capabilities, the acknowledged safety risks indicate premature deployment could endanger public safety. The progress may influence how companies prioritize research investment, potentially accelerating development timelines for physical-world AI applications or prompting stricter safety standards before real-world implementations become widespread.