Unreal Agent: Reducing LLM Tool-Management Overhead via Asynchronous Harnesses

Unreal Agent reduces the operational cost of AI agents by decoupling tool execution from the model's reasoning loop. By implementing an asynchronous harness, it eliminates the need for the underlying LLM to manage waits, polls, and heartbeats, resulting in cost savings of up to 40% compared to Codex and 20% compared to Pi in real-world workloads.

Asynchronous Tool-Calling Architecture

Unreal Agent replaces synchronous tool execution with an asynchronous event-log system. In traditional agent harnesses, the model often wastes tokens polling for the status of a long-running task or waiting for a tool to return a result before proceeding.

Unreal Agent optimizes this process as follows:

  • Immediate State Logging: When a tool is called, the harness immediately appends an event-log record stating the tool is "in-progress."
  • Background Execution: The tool continues to execute in the background without blocking the model.
  • Concurrent Work: The agent can schedule multiple heterogeneous tool calls (e.g., setting up a development environment while simultaneously searching the web) in a single model turn.
  • Event-Driven Resumption: Once a tool finishes, the result is appended to the session log, and the LLM is called to process the output.

This architecture allows users to steer the agent in real-time without waiting for background tool calls to complete, reducing the "token tax" associated with lifecycle management.

Performance and Cost Benchmarks

Tested using GPT-6 Astra (xhigh reasoning effort), Unreal Agent demonstrates higher cost efficiency and lower token usage across several coding benchmarks while maintaining competitive pass rates.

Terminal-Bench 4.0

Unreal Agent achieved a 57.9% pass rate with a total cost of $1,428, compared to Codex's 57.9% pass rate at a cost of $2,350. This represents a significant reduction in total expenditure for the same outcome.

SWE-Atlas Codebase QnA

Unreal Agent reached a 65.8% pass rate with a total cost of $936, outperforming both Codex (63.3% pass rate, $1,303) and Pi (64.0% pass rate, $1,033).

DeepSWE 1.1

Unreal Agent achieved a 72.4% pass rate with a total cost of $1,367, compared to Codex (69.0% pass rate, $1,633) and Pi (69.6% pass rate, $1,584).

Agents’ Last Exam (ALE-CLI)

Unreal Agent recorded a 30.0% full pass rate with a total cost of $217, whereas Codex and Pi both recorded 29.0% full pass rates at higher costs ($292 and $262, respectively).

Engineering Trade-offs and Design Philosophy

Unreal Labs identifies several limitations in existing CLI-oriented SDKs that motivated the creation of Unreal Agent:

  • Production Mismatch: Many SDKs assume local sessions and subprocesses, making them difficult to scale in production environments where reliable cancellation and background task management are required.
  • Dependency Risk: Heavy dependency trees in third-party SDKs introduce maintenance and supply-chain risks.
  • Security: The team argues that deterministic environment constraints (e.g., sandboxed hosts and granular access tokens) are more robust than harness-level security hooks.

To maximize efficiency, Unreal Agent avoids the use of sub-agents or complex workflows, relying instead on simple prompts and token-optimized tool results.

Community Insights and Technical Critiques

Discussion among developers highlights both the potential and the pitfalls of the asynchronous approach:

  • Comparison to Programmatic Tool Calling: Some users note that this is similar to "Programmatic Tool Calling" (PTC) or returning a program (e.g., TypeScript) to invoke tools asynchronously, a technique already being explored by AI labs.
  • The "Polling" Problem: Critics point out that the high token usage in competitors like Codex is often due to "hot looping" on polling tasks, and that patching these loops can yield similar savings to those claimed by Unreal Agent.
  • Self-Hosting Perspectives: Some developers argue that for those self-hosting models, cost optimization is an "anti-feature," suggesting that harnesses should instead optimize for the best possible results using the full context window of non-frontier models.
  • Naming Concerns: Multiple commenters noted a potential trademark conflict with Epic Games' Unreal Engine, despite the tool having no relation to game development.

"The headline graph is kind of bizarre. For some reason they're comparing their harness running on Astra xhigh to Codex with Astra max? ... A big part of why Codex uses so many tokens is that it basically hot loops on polling tasks it starts for... absolutely no good reason."

Availability

Unreal Agent is available as a Go library for direct integration, a runner executable for CLI use, and a benchmark runner compatible with the Harbor framework. The source code is hosted on GitHub at github.com/unreallabsai/unreal-agent.

Sources

Related