Project Fetch: Evaluating Claude's Impact on Robotics Task Performance
TL;DR
Anthropic conducted "Project Fetch," an uplift study to determine if Claude could help non-experts program quadruped robots to retrieve beach balls. The results showed that the AI-enabled team completed tasks significantly faster and made substantially more progress toward full autonomy than the team without AI access.
Experiment Design and Methodology
Anthropic researchers divided eight non-robotics experts into two groups: "Team Claude" (with AI access) and "Team Claude-less" (without AI access). The teams were tasked with three increasingly difficult phases of robot dog operation:
- Phase One: Using a manufacturer-provided controller to bring a ball back to a specific area.
- Phase Two: Establishing a computer connection to the robot, accessing onboard sensors (video and lidar), and developing a custom software program for retrieval.
- Phase Three: Developing a program for the robot to detect and fetch the ball autonomously.
Key Findings: AI Uplift in Robotics
Performance and Speed
Team Claude demonstrated a clear advantage in task completion and efficiency. For tasks completed by both teams, Team Claude succeeded in approximately half the time it took Team Claude-less. While Team Claude-less completed 6 out of 8 tasks, Team Claude completed 7 out of 8.
Technical Advantages
Claude provided the most significant uplift in the initial stages of hardware interaction:
- Connectivity: Team Claude was more efficient at exploring connection methods and avoided being misled by incorrect online documentation, whereas Team Claude-less prematurely discarded the easiest connection method.
- Sensor Integration: Team Claude more easily accessed data from the lidar sensor, while Team Claude-less struggled with this task until the end of the day.
- Program Quality: Although Team Claude took longer to write their control program, the resulting software was superior, providing a streaming video feed rather than the intermittent still images used by Team Claude-less.
Trade-offs and "Side Quests"
The experiment revealed that AI assistance can lead to divergent workflows. Team Claude wrote approximately nine times more code than Team Claude-less. While this allowed them to explore multiple approaches in parallel, it also led to "side quests" and distractions from the primary goal. For example, one member developed a natural language controller for the robot, which was not required for the primary task.
Impact on Team Dynamics and Morale
Emotional Expression and Confusion
Quantitative analysis of audio transcripts showed that Team Claude-less expressed confusion at double the rate of Team Claude. Their dialogue was generally more negative, and they reported feeling that their coding skills had diminished without access to their usual AI tools.
Collaboration Styles
The two teams exhibited distinct working patterns:
- Team Claude: Members primarily partnered with their individual AI assistants, pursuing parallel paths toward objectives.
- Team Claude-less: Members strategized more deeply with each other and asked 44% more questions than the AI-enabled team.
Limitations and Future Implications
Study Limitations
Anthropic noted several constraints of the study, including a small sample size (two teams), a short duration (one day), and the use of a convenience sample of Anthropic employees who were already familiar with AI.
The Path to Autonomy
Anthropic posits that "uplift often precedes autonomy," suggesting that the ability of AI to assist humans in robotics today is a precursor to the AI performing these tasks independently in the future. This capability is highlighted as a critical threshold in Anthropic's Responsible Scaling Policy, as autonomous AI R&D involving hardware could lead to rapid, unpredictable advances in AI capabilities.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch