Basis Scales Accounting Automation with OpenAI GPT-5 and o3
Basis has developed an AI agent system for accounting firms that leverages OpenAI's latest models to automate repetitive, structured tasks such as reconciliations, journal entries, and financial summaries. By integrating GPT-5, GPT-4.1, and o3, Basis enables accounting firms to achieve up to 30% time savings, allowing professionals to shift their focus toward high-leverage advisory work and business growth.
Multi-Agent Architecture for Task Routing
Basis employs a multi-agent architecture that assigns specific OpenAI models to tasks based on their complexity, latency requirements, and input types. This system is managed by a supervising agent, which has migrated from OpenAI o3 to GPT-5, to coordinate the entire process and route steps to specialized sub-agents.
- GPT-5: Used as the supervising agent and for complex scenarios requiring deep reasoning, consistency, and explainability, such as interpreting unusual transaction patterns or managing month-end closes.
- GPT-4.1: Utilized for speed-critical interactions, including surfacing quick feedback or answering clarifying questions during a review.
This orchestration allows Basis to continuously expand the scope of tasks the agents can handle as OpenAI's model capabilities evolve.
Validating Output through Reasoning and Explainability
To ensure automation is reviewable, Basis agents surface the assumptions, data sources, and logic behind every decision. The system originally utilized OpenAI o3-Pro for scaling reasoning, but has since migrated to GPT-5 for its superior ability to reason through structured processes and explain outcomes.
For example, when preparing a journal entry, the supervising agent coordinates sub-agents to retrieve data and reference best practices. The accountant then receives the entry along with a detailed explanation of the data used, the mapping logic, and the system's confidence level. This visibility is enabled by GPT-5's ability to scale test-time compute, allowing the system to provide transparent explanations that maintain human control over the process.
Benchmarking and Parallel Tool Calling
Basis evaluates new model releases using a detailed benchmark suite that measures both accuracy and the reasoning clarity of the model. GPT-5 is currently the strongest model in the Basis stack, particularly in areas of depth and precision.
A key technical differentiator for GPT-5 is its performance in parallel tool calling. In a Basis benchmark testing the ability to use multiple tools in parallel with both web search and code interpreter enabled, GPT-5 achieved a 100% success rate, leading all other evaluated models in reasoning benchmarks.
Impact on Accounting Capacity
By delegating multi-step processes like reconciliations and journal entries through function calling, Basis has moved from simple task automation to full workflow delegation. Large accounting firms across the U.S. using the platform report an average of 30% time savings, increasing their capacity to serve more clients and explore new practice areas.