Blue J Scaling Tax Research with GPT-4.1

Blue J has developed a tax research engine using GPT-4.1 and Retrieval-Augmented Generation (RAG) to automate the synthesis of dense regulations into expert-grade answers. This system allows tax professionals to receive fully-cited guidance in seconds, replacing a manual process that previously took hours, days, or weeks.

RAG Architecture for Regulated Domains

Blue J utilizes a Retrieval-Augmented Generation (RAG) system that combines GPT-4.1 with a proprietary library consisting of millions of curated documents. This library includes authoritative primary sources and expert commentary from providers such as Tax Notes.

When a user submits a query, the system retrieves relevant source material and uses GPT-4.1 to synthesize the information into a clear, cited answer. According to CTO Brett Janssen, GPT-4.1 was selected because it consistently follows instructions, respects context, and handles edge cases more effectively than other models tested.

Feedback Loops and Accuracy Scaling

To maintain trust in a high-stakes environment where errors can lead to audits or financial loss, Blue J implemented a closed-loop feedback system:

  • User-Driven Feedback: Every response includes a "disagree" button. When a user flags a response, it is systematically categorized by issue type, tax topic, and root cause.
  • AI-Powered Triage: GPT-4.1 is used to analyze thousands of feedback points and cluster related issues, allowing the product and tax research teams to prioritize fixes.
  • Rapid Iteration: This flywheel approach has reduced the disagree rate to fewer than 1 in 700 responses.

This infrastructure allows Blue J to respond rapidly to legislative changes. For example, when a major U.S. tax bill passed in 2025, the team mapped the impact for six weeks prior to the bill's signing and deployed updated answers into production within hours.

Evaluation Frameworks and Model Selection

Blue J employs a rigorous evaluation process to ensure model quality serves as a gating function for deployment. The company uses a benchmark suite of over 350 prompts covering U.S., Canadian, and U.K. tax law to test for:

  1. Instruction adherence
  2. Source alignment
  3. Answer clarity

Brett Janssen stated that OpenAI models have consistently outperformed competitors against these internal benchmarks, particularly in instruction-following and real-world utility, leading Blue J to exclusively ship OpenAI models.

Business Impact and User Engagement

By automating the retrieval and synthesis of tax law, Blue J has shifted the workload of tax professionals toward higher-margin planning and advisory work. Key performance metrics include:

  • User Retention: More than 70% of users log in weekly.
  • Efficiency Gains: Users save an average of 2.7 hours per week on research and client communication.

Sources