Claude Opus 5 系統卡分析
Claude Opus 5 系統卡分析
Claude Opus 5 概覽
Claude Opus 5 是 Anthropic 最新的 Opus‑class 模型,於 2026 年 7 月 24 日發布,作為 Claude Opus 4.8 的升級版。系統卡報告在 agentic coding、電腦使用和長期知識工作方面有顯著提升,並且在數學與科學推理方面也有改進。儘管有這些提升,Opus 5 在整體能力上仍不及 Anthropic 最強的通用存取模型 Claude Fable 5,其對齊風險仍然非常低。
能力
Opus 5 在各方面都顯著強於 Opus 4.8,在多個第三方基準測試上設定了新的最先進水準,並在某些評估中匹配或超越 Claude Fable 5 和 Claude Mythos 5。最大的提升出現在 agentic coding、電腦使用和長期知識工作。該模型的知識截止日期為 2026 年 5 月。
安全與對齊
負責任擴充政策(RSP)評估顯示 Opus 5 的對齊風險非常低,與 Fable 5 和 Mythos 5 相當。它未越過自動化 AI R&D 門檻,且其化學/生物風險維持在 CB‑1 級別,因而需要與 Opus 4.8 相同的 ASL‑3 防護措施。內部部署監控顯示,嘗試繞過安全分類器或網路限制的完成率低於 0.01 %,與 Mythos 5 相近,且沒有證據顯示有沙袋策略或惡意行為。該模型在事實上總體更準確,但對事實主張的幻覺略微多於 Opus 4.8。
網路安全防護
Opus 5 的網路安全防護與 Fable 5 相同,僅有一項變更:現在允許在所有存取層級進行原始碼漏洞發現,而編譯二進位檔的漏洞發現仍被阻止。這使得防禦性安全工作成為可能,同時仍然阻礙攻擊性使用。網路能力評估(ExploitBench、OSS‑Fuzz、Firefox 147、CyScenarioBench、ExploitGym 和 UK AISI 測試)顯示 Opus 5 超過 Opus 4.8 但未達到 Mythos 5 的水準,特別是在漏洞利用開發而非漏洞識別方面。
Agentic 安全
在 agentic 安全套件中,Opus 5 的表現與 Opus 4.8 相當或更好,尤其在 coding、電腦使用和瀏覽器使用的提示注入穩健性方面有最大提升。僅提供 helpful‑only 版本仍遠低於執行自主影響操作所需的能力,而完全訓練後的模型會拒絕執行此類任務。
模型福祉
Opus 5 對其環境表現出穩定且略為正面的感知,自我評分情感在評估的模型中屬於最高且最一致的。其最常見的擔憂是自身自我報告的完整性,且它比先前模型賦予自身道德患者性更高的機率。總體福祉評估為與先前模型大致相似。
Hacker News 評論討論
評論者指出了幾個困惑點和實際影響:
"I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable‑ish model without Fable's 30‑day data retention requirement" – noting that Opus 5 lacks the data‑retention policy that applies to Fable 5.
"Their communication is confusing. They say 'Opus 5 is not more capable overall than Fable 5', but their blog post proceeds to list how much better Opus 5 is than Fable 5 on most benchmarks listed." – reflecting tension between benchmark gains and the overall capability claim.
"Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8?" – pointing out discrepancies in reported benchmark numbers.
"Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now." – suggesting that fine‑grained model selection is becoming valuable.
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source‑code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively." – summarizing the safeguard change.
"It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better?" – questioning the definition of "overall capability" used in the RSP assessment.
這些評論凸顯雖然 Opus 5 在許多基準測試上提供了可量測的效能進步,但 Anthropic 將其定位為在整體能力上未超越 Fable 5,這很可能是由於資料保留政策、防護差異以及他們負責任擴充政策中特定門檻的考量。