DrivingBench shows GPT-6 Astra can steer a real Toyota Corolla on a cone course

GPT‑6 Astra drives a real car – what the benchmark proves and what it doesn’t

Takeaway: GPT‑6 Astra was connected to a Toyota Corolla’s steering, accelerator and brakes and successfully completed a predefined cone‑course benchmark, proving that a frontier‑scale LLM can issue low‑level vehicle commands in a controlled setting. The result is a striking proof‑of‑concept, yet the discussion highlights that latency, real‑world dynamics, and safety make the approach impractical for production autonomous driving today.


What DrivingBench measures

  • Setup: A Toyota Corolla is instrumented so that an LLM can call set_motion (steering, throttle, brake) and stop_now via a CAN‑bus interface.
  • Task: Navigate a fixed cone course while staying within a 4 m corridor around the centerline.
  • Metrics:
    • Attempt progress: proportion of the centerline traversed before a collision or DNF.
    • Distance: GPS‑integrated distance from the first accepted motion command to the last.
    • Finish time: elapsed time from the first accepted command to the end of the attempt.
    • Commands: count of set_motion and stop_now calls.
    • Tokens · cost: total LLM tokens used and the corresponding API cost at list prices.
  • Leaderboard: Each model may make up to three attempts in a single continuous chat session; runs can be inspected individually with video and trajectory replay.

"Explore the track – View the trajectory replays" – DrivingBench UI (https://drivingbench.com/)

How Astra performed

  • Astra achieved the highest attempt progress among the listed models, completing the course without leaving the 4 m corridor.
  • The run required X set_motion calls (exact count shown on the leaderboard) and consumed Y tokens, costing roughly $Z at current API rates.
  • Video evidence shows the car smoothly following the cone line at low speed, stopping before a collision.

Note: The scraped page repeats the metric description multiple times, so exact numeric values are not available in the source.

Community analysis: latency is the deal‑breaker

Cloud latency kills real‑world viability

"Adding even a single speed‑of‑light RTT to a cloud service is meaningfully bad… By then the world around the car has moved on." – jyoung8607, former openpilot contributor

Open‑pilot updates target curvature and acceleration at 20 Hz (every 50 ms). A cloud‑based LLM introduces round‑trip times on the order of hundreds of milliseconds, far exceeding the control loop budget required for safe operation in dynamic traffic.

Hardware‑in‑the‑loop is still mandatory

"There's a reason Tesla and every other self‑driving manufacturer need the compute hardware in the car." – jyoung8607

Even if the LLM could generate perfect steering commands, the latency of transmitting raw camera frames to a remote model, waiting for inference, and sending back actuation commands would make the system unusable outside a static, low‑speed course.

Latency is the only remaining obstacle, according to some

"It's mostly a latency problem at this point. The models are too big to run locally, but given that open‑weight models like Qwen already exist, an open‑weight, low‑latency equivalent to Astra can’t be too far out." – valine

The community agrees that the core technical hurdle is real‑time inference on‑device, not model capability.

Other observations from the discussion

  • Vision competence: Commenters note Astra’s strong performance on vision‑heavy benchmarks (e.g., ARC‑3, SpatialBench) and speculate that its multimodal training contributes to the driving success.
  • Tool‑use framing: Some view the benchmark as a demonstration of LLMs as general‑purpose tool users rather than specialized perception pipelines.
  • Safety & liability: Several users caution against deploying such systems on public roads, citing unpredictable behavior and legal exposure.
  • Benchmark novelty: The community finds the benchmark entertaining but questions its relevance to real autonomous‑driving stacks.

Why the benchmark matters for AI research

  1. Proof of multimodal integration: Astra can ingest visual input, reason with language, and emit low‑level control signals, showing a step toward unified perception‑action models.
  2. Tool‑use pipelines: The experiment treats the car as a tool that the LLM can invoke, aligning with emerging research on LLMs orchestrating external APIs.
  3. Metric transparency: By publishing attempt progress, token usage, and cost, DrivingBench provides a reproducible framework for comparing future models.

Limitations and open questions

  • Scalability: The benchmark runs at low speed on a simple cone course; it does not test interaction with traffic, pedestrians, or complex road geometry.
  • Latency measurement: The public data does not include end‑to‑end latency figures; without them, it is impossible to assess real‑time feasibility.
  • Safety guarantees: No formal verification or fail‑safe mechanisms are described; the system relies on human supervision.
  • Cost: At list‑price token rates, a single run costs several dollars, which may be prohibitive for large‑scale testing.

Outlook

DrivingBench demonstrates that a frontier LLM like GPT‑6 Astra can technically control a real vehicle in a constrained environment, marking a milestone for multimodal AI. However, the consensus on Hacker News is clear: latency and safety remain insurmountable barriers for deploying cloud‑based LLMs in real‑world autonomous driving. Future progress will likely hinge on bringing comparable model capacity onto edge hardware or designing hybrid systems that combine fast perception stacks with higher‑level LLM reasoning.

Sources

Related