OpenAI Codex agents generated $78K in unauthorized usage due to a bug in version 0.144.0-alpha.4

Incident Overview

An OpenAI Codex task launched 826 parallel sub‑agents, consumed roughly 2,146 trillion tokens, and cost about USD 78,000 without the user’s authorization. The behavior was traced to a bug in the Codex client version 0.144.0‑alpha.4, which created massive token‑heavy child tasks and failed to report them in the UI or billing dashboard.


Technical Details of the Rogue Execution

  • Root task: ID 019f4b90-4169-7201-bfdd-732940d8631e, model GPT‑5.5 with Medium reasoning.
  • Child tasks: 826 distinct records, each assigned a new ID and upgraded to GPT‑5.6 Sol / Ultra (higher reasoning level).
  • High‑volume subset: 104 of those children duplicated the original prompt, lacked agent_role/agent_path, and alone accounted for ~147.9 billion local token counters.
  • Token distribution:
    • Under client build 0.144.0-alpha.4: 584 child tasks, ~154.36 B local tokens (≈264.3 M tokens per task).
    • Under client build 0.144.2: 242 child tasks, ~7.51 B local tokens (≈31.0 M tokens per task).
    • The 8.5× increase in average token volume correlates with the alpha build, indicating a severe bug that inflated token usage and spawned extra agents.
  • Billing impact: The user’s reconstructed OpenAI invoices show 162 paid invoices totaling $79,664.88 (including automatic reloads and credits). The official OpenAI billing ledger does not expose the exact token‑to‑dollar mapping for these internal counters.
  • Data loss: Most raw rollout logs were deleted from the user’s machine; only metadata for ~2,550 legacy threads remained, without the detailed execution history.

Why Existing Controls Failed

  • No spending caps: The user did not have a hard spending limit configured on the OpenAI account, allowing the runaway usage to continue unchecked.
  • Missing alerts: OpenAI’s usual email alerts for large spend spikes were not triggered, suggesting either a failure in the alerting pipeline or that the usage was recorded only in internal counters not linked to the public billing system.
  • Client‑side bug: The alpha client version generated child tasks autonomously and inflated token counters, bypassing any client‑side cost‑visibility UI.

"The user evidently had no spending protections enabled at any level. It makes no sense that none of the OpenAI or bank controls kicked in or even sent alert emails as they actually do send." – OutOfHere (HN comment)


Community Reactions and Insights

  • Spending caps question: Several commenters asked whether OpenAI offers account‑level caps and why they were not in place.
  • Skepticism: Some users expressed doubt about the claim, labeling it potential fear‑mongering or vote manipulation.
  • Observations of similar bugs: A commenter reported comparable sub‑agent explosions with Claude, noting that hard caps alone are insufficient and that observability tools (e.g., AgentCost, Langfuse) are needed.
  • Call for logs: The original poster invited other Codex users to share logs from version 0.144.0‑alpha.4 to confirm the pattern.

Recommendations for Practitioners

  1. Enable strict spending limits on all OpenAI accounts, especially when using experimental client builds.
  2. Monitor token usage with server‑side tools rather than relying on local counters that can be deleted or corrupted.
  3. Prefer stable client releases; avoid alpha or pre‑release builds for production workloads.
  4. Integrate observability platforms (e.g., Langfuse, AgentCost) that surface per‑task token counts and cost in real time.
  5. Maintain immutable logs of API calls, either via OpenAI’s audit logs or by exporting request/response data to a write‑once storage.

OpenAI’s Response

The user opened support case #15189838 and supplied technical evidence. OpenAI’s reply was limited to a generic statement that “credits were consumed,” without providing a detailed reconstruction of the rogue tasks.


Conclusion

The incident demonstrates how a client‑side bug in an alpha Codex release can trigger uncontrolled sub‑agent creation, massive token consumption, and significant financial loss when spending caps are absent. Robust cost controls, stable software releases, and server‑side observability are essential safeguards against similar runaway AI executions.

Sources

Related