Claude Opus 5.5 Prompting Guide and Technical Analysis
Claude Opus 5.5 is designed for high-autonomy agentic work, offering over 30% faster token generation than Claude Opus 5 and significantly improved performance in multistep coding and knowledge work. The primary shift in this version is that "thinking" is now a mandatory internal process, making the effort setting the primary lever for balancing intelligence, latency, and cost.
Optimizing Model Effort and Latency
Effort levels (low, medium, high, xhigh, max) control the depth of the model's internal deliberation. Because thinking is always active in Opus 5.5, users must calibrate these settings based on specific task requirements rather than carrying over settings from previous versions.
Effort Calibration
- Baseline: Start with
medium(the default). In Anthropic's testing,mediumeffort on Opus 5.5 often matches or exceedshigheffort on Opus 5 for coding and knowledge tasks. - Token Limits: Set
max_tokensto a high value (up to 128,000) to accommodate both the internal thinking tokens and the final response. If the limit is too low, the response may be truncated. - Latency Reduction: To minimize thinking and reduce time-to-first-token, lower the effort level first. If further reduction is needed, a system prompt instruction like "Answer directly without deliberating" can be used, though this may impact quality.
Migrating from Thinking-Disabled Prompts
Claude Opus 5.5 does not support disabling thinking. For integrations previously configured with thinking: {"type": "disabled"}, the following adjustments are required:
- Remove Reasoning Instructions: Delete prompts that ask the model to write out its reasoning in the response text. This avoids
reasoning_extractionrefusals and allows the model to use its internal thinking blocks. - Block-Based Parsing: Clients must parse responses by block type rather than assuming the first block is text, as responses may now begin with a
thinkingblock.
Managing Agentic and Unattended Workflows
Opus 5.5 is optimized for long-running autonomous tasks, such as multi-hour codebase audits. However, its tendency to provide progress updates can lead to "early stops" in unattended loops.
Preventing Early Stops
When an agent ends a turn with a text report instead of a tool call (stop_reason: "end_turn"), the harness should not treat this as task completion. Instead:
- Checklists: Maintain a task checklist (via a tool or file) that the model updates. If the turn ends but items remain open, the harness should prompt the model to continue.
- System Prompting: Instruct the model to avoid ending turns with summaries that announce the next step; instead, tell it to take the step immediately.
User-Facing Progress Updates
To prevent long agentic turns from appearing silent, developers should use the display: "updates" setting. This allows the client to receive short summaries of the model's internal progress notes. If a turn remains silent for too many consecutive tool calls (e.g., five), the harness can append a turn-scoped system message reminding the model to provide an update.
Advanced Prompting Patterns
Multi-App Context Exploration
For workflows spanning multiple applications (email, CRM, spreadsheets), the model may act too quickly on loosely specified tasks. Adding a system prompt instruction to "look through relevant sources before acting" increases completion accuracy at the cost of slightly more tokens and tool calls.
Time-Budgeting for Multi-Agent Systems
Opus 5.5 is responsive to elapsed time signals. In multi-agent harnesses, providing a time budget (e.g., elapsed 340s / 1200s) encourages the model to increase parallelization and finish tasks sooner without necessarily sacrificing quality.
Mitigating Indirect Prompt Injection
To protect against instructions embedded in pasted text, developers should wrap pasted content in tags with random IDs (e.g., <pasted_content id="ab12">) and instruct the model in the system prompt to treat content within these tags as data rather than instructions.
Visual Input and Frontend Design
Complex Visuals
Opus 5.5 reads dense charts and diagrams more accurately than its predecessor. For maximum precision in technical drawings, developers should:
- Use higher-resolution images.
- Provide image-processing tools (PIL, OpenCV) via a container, allowing the model to crop and zoom into specific areas.
Frontend Defaults
When generating UI code, the model falls back to generic styles. To avoid "AI slop," developers should specify patterns to avoid rather than using general terms like "avoid a generic look."
Community Insights and Critiques
User feedback from Hacker News highlights several points of friction despite the model's raw capability leaps:
"Opus 5.5 is a beast... Runs circles around OpenAI's Astra... Everything is better: the model, the TUI, the quotas."
However, critics point to several recurring issues:
- Safeguard Over-triggering: Users report that biology and cybersecurity safeguards occasionally flag benign queries, such as muscle soreness or code hardening audits.
- Verbosity: Some users find the model overly verbose, noting that "Concise" settings in tools like Claude Code are sometimes ignored in large contexts.
- Fragility of Prompting: There is a significant community concern regarding the "voodoo magic" of prompting, with users arguing that the industry needs robust, structured output and open standards rather than techniques that change every few months.
- Reasoning Extraction: Some developers are frustrated by the model's refusal to output its full internal thinking process in the response text, viewing it as a move toward more proprietary, "black box" behavior.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch