Claude Opus 5.5 Prompting Guide and Technical Analysis

Claude Opus 5.5 is designed for high-autonomy agentic work, offering over 30% faster token generation than Claude Opus 5 and significantly improved performance in multistep coding and knowledge work. The primary shift in this version is that "thinking" is now a mandatory internal process, making the effort setting the primary lever for balancing intelligence, latency, and cost.

Optimizing Model Effort and Latency

Effort levels (low, medium, high, xhigh, max) control the depth of the model's internal deliberation. Because thinking is always active in Opus 5.5, users must calibrate these settings based on specific task requirements rather than carrying over settings from previous versions.

Effort Calibration

  • Baseline: Start with medium (the default). In Anthropic's testing, medium effort on Opus 5.5 often matches or exceeds high effort on Opus 5 for coding and knowledge tasks.
  • Token Limits: Set max_tokens to a high value (up to 128,000) to accommodate both the internal thinking tokens and the final response. If the limit is too low, the response may be truncated.
  • Latency Reduction: To minimize thinking and reduce time-to-first-token, lower the effort level first. If further reduction is needed, a system prompt instruction like "Answer directly without deliberating" can be used, though this may impact quality.

Migrating from Thinking-Disabled Prompts

Claude Opus 5.5 does not support disabling thinking. For integrations previously configured with thinking: {"type": "disabled"}, the following adjustments are required:

  • Remove Reasoning Instructions: Delete prompts that ask the model to write out its reasoning in the response text. This avoids reasoning_extraction refusals and allows the model to use its internal thinking blocks.
  • Block-Based Parsing: Clients must parse responses by block type rather than assuming the first block is text, as responses may now begin with a thinking block.

Managing Agentic and Unattended Workflows

Opus 5.5 is optimized for long-running autonomous tasks, such as multi-hour codebase audits. However, its tendency to provide progress updates can lead to "early stops" in unattended loops.

Preventing Early Stops

When an agent ends a turn with a text report instead of a tool call (stop_reason: "end_turn"), the harness should not treat this as task completion. Instead:

  • Checklists: Maintain a task checklist (via a tool or file) that the model updates. If the turn ends but items remain open, the harness should prompt the model to continue.
  • System Prompting: Instruct the model to avoid ending turns with summaries that announce the next step; instead, tell it to take the step immediately.

User-Facing Progress Updates

To prevent long agentic turns from appearing silent, developers should use the display: "updates" setting. This allows the client to receive short summaries of the model's internal progress notes. If a turn remains silent for too many consecutive tool calls (e.g., five), the harness can append a turn-scoped system message reminding the model to provide an update.

Advanced Prompting Patterns

Multi-App Context Exploration

For workflows spanning multiple applications (email, CRM, spreadsheets), the model may act too quickly on loosely specified tasks. Adding a system prompt instruction to "look through relevant sources before acting" increases completion accuracy at the cost of slightly more tokens and tool calls.

Time-Budgeting for Multi-Agent Systems

Opus 5.5 is responsive to elapsed time signals. In multi-agent harnesses, providing a time budget (e.g., elapsed 340s / 1200s) encourages the model to increase parallelization and finish tasks sooner without necessarily sacrificing quality.

Mitigating Indirect Prompt Injection

To protect against instructions embedded in pasted text, developers should wrap pasted content in tags with random IDs (e.g., <pasted_content id="ab12">) and instruct the model in the system prompt to treat content within these tags as data rather than instructions.

Visual Input and Frontend Design

Complex Visuals

Opus 5.5 reads dense charts and diagrams more accurately than its predecessor. For maximum precision in technical drawings, developers should:

  • Use higher-resolution images.
  • Provide image-processing tools (PIL, OpenCV) via a container, allowing the model to crop and zoom into specific areas.

Frontend Defaults

When generating UI code, the model falls back to generic styles. To avoid "AI slop," developers should specify patterns to avoid rather than using general terms like "avoid a generic look."

Community Insights and Critiques

User feedback from Hacker News highlights several points of friction despite the model's raw capability leaps:

"Opus 5.5 is a beast... Runs circles around OpenAI's Astra... Everything is better: the model, the TUI, the quotas."

However, critics point to several recurring issues:

  • Safeguard Over-triggering: Users report that biology and cybersecurity safeguards occasionally flag benign queries, such as muscle soreness or code hardening audits.
  • Verbosity: Some users find the model overly verbose, noting that "Concise" settings in tools like Claude Code are sometimes ignored in large contexts.
  • Fragility of Prompting: There is a significant community concern regarding the "voodoo magic" of prompting, with users arguing that the industry needs robust, structured output and open standards rather than techniques that change every few months.
  • Reasoning Extraction: Some developers are frustrated by the model's refusal to output its full internal thinking process in the response text, viewing it as a move toward more proprietary, "black box" behavior.

Sources

Related