OpenAI A Scorecard for the AI Age

OpenAI A Scorecard for the AI Age

OpenAI introduces the "Useful Intelligence per Dollar" framework

OpenAI has proposed a new metric for measuring the economic value of AI in business: "Useful Intelligence per Dollar." This framework shifts the focus from traditional software adoption metrics—such as seats purchased or active users—to a measure of work actually accomplished.

According to OpenAI, the core economic question for business leaders is whether the value of the work AI completes grows faster than the cost of producing it. This requires looking beyond cost-per-token to the full cost of producing a successful outcome, including human review, retries, and rework.

Measuring useful work and task completion

To determine the value of AI, organizations must first define what "done" means for a specific workflow and measure outcomes in the system where the work occurs. Tokens only create value when they are transformed into usable work, such as resolving customer issues, shipping code changes, or reviewing contracts.

OpenAI highlights that as models become more capable, they can handle longer and more complex tasks involving multi-step reasoning and tool integration. For example, in finance workflows, AI can automate the data gathering and reconciliation process, allowing human teams to focus on high-level judgment and creativity.

Calculating the full cost per successful task

OpenAI argues that the lowest price per token does not always result in the lowest cost per outcome. A more capable frontier model may be more cost-effective if it achieves the correct result in a single pass, reducing the need for human review and multiple retries.

The GPT-5.6 Model Family

OpenAI recently released the GPT-5.6 family, which offers three tiers to help customers optimize the cost-to-performance equation:

  • Sol: The flagship model for strongest reasoning and complex tasks.
  • Terra: A balanced model for performance and cost.
  • Luna: The fastest and most affordable model for high-volume workflows.

Performance and Efficiency Benchmarks

GPT-5.6 Sol (with max reasoning) set a new state-of-the-art on the Artificial Analysis Coding Agent Index. Specifically, on the DeepSWE v1.1 benchmark for long-horizon engineering tasks, GPT-5.6 Sol achieved a score of 72.7%, surpassing Claude Fable 5's 69.9%, while maintaining a 36.2% lower estimated API cost.

Furthermore, GPT-5.6 Sol used 54% fewer output tokens than another leading model to achieve these results.

Evaluating dependability and reliability

Dependability has direct economic value because accurate and consistent results reduce the time spent on correction and repetition. OpenAI suggests tracking three specific outcomes to measure AI dependability:

  1. Ready to use: The result met the quality bar as delivered.
  2. Needs correction: The result required human edits or another attempt.
  3. Needs escalation: A human was required to step in and finish the work.

To move AI from drafting to taking action, OpenAI emphasizes the need for clear boundaries regarding data access, system permissions, and human approval triggers. This is supported by ChatGPT Work, which utilizes the security and compliance foundation of ChatGPT Enterprise to allow for deeper workflow integration.

Scaling AI value and the role of compute

OpenAI asserts that the return on AI investment improves at scale when completed work grows faster than total cost while quality remains stable or improves.

Compute is the central driver of this equation. Improvements in research, inference efficiency, purpose-built hardware, and smarter routing translate into better outcomes for the customer. This creates a compounding effect: better infrastructure accelerates research, which produces more efficient models, which in turn drive adoption and revenue to support further investment in the next generation of compute and safety.

Sources