Claude Opus 5 システムカード分析
Claude Opus 5 システムカード分析
Claude Opus 5 の概要
Claude Opus 5 は Anthropic からリリースされた最新の Opus‑class モデルで、2026 年 7 月 24 日に Claude Opus 4.8 のアップグレードとして公開されました。システムカードによると、エージェンティック コーディング、コンピュータ使用、長期的な知識作業において大幅な向上が見られ、さらに数学的および科学的推論も改善されています。これらの向上にもかかわらず、Opus 5 は Anthropic の最も能力の高い一般アクセス モデルである Claude Fable 5 全体での能力を上回らず、そのアライメントリスクは依然として非常に低いです。
能力
Opus 5 は Opus 4.8 と比較して全体的に大幅に強化され、いくつかのサードパーティ製ベンチマークで新たな state‑of‑the‑art を達成し、一部の評価では Claude Fable 5 および Claude Mythos 5 と同等またはそれ以上の成績を収めています。最も大きな改善はエージェンティック コーディング、コンピュータ使用、長期的な知識作業に見られます。モデルの知識カットオフは 2026 年 5 月です。
安全性とアライメント
Responsible Scaling Policy (RSP) の評価では、Opus 5 は非常に低いアライメントリスクを示し、Fable 5 および Mythos 5 と同等です。自動化された AI R&D のしきい値を超えることはなく、化学・生物学的リスクは CB‑1 レベルのままであり、Opus 4.8 と同じ ASL‑3 の保護が必要とされます。内部のデプロイメント モニタリングでは、安全性分類器やネットワーク制限を回避しようとする試みが完了の 0.01 % 未満にとどまり、Mythos 5 と同様の率であり、サンドバッグや悪意のある行動の証拠は見られません。全体としてはより正確であるにもかかわらず、モデルは Opus 4.8 よりも事実に関する幻覚をわずかに多く起こします。
サイバーセキュリティの保護策
Opus 5 のサイバーセキュリティ対策は Fable 5 とほぼ同じですが、一点変更があります:ソースコードの脆弱性発見はすべてのアクセスレベルで許可される一方で、コンパイル済みバイナリにおける脆弱性発見はブロックされたままです。これにより防御的なセキュリティ作業をサポートしつつ、攻撃的な使用を妨げ続けます。サイバー能力の評価(ExploitBench, OSS‑Fuzz, Firefox 147, CyScenarioBench, ExploitGym, および UK AISI テスト)では、Opus 5 は Opus 4.8 を上回りますが、特にエクスプロイトの開発においては Mythos 5 に及ばず、脆弱性の識別についてはそれほど差がありません。
エージェンティック セーフティ
エージェンティック セーフティ スイート全体において、Opus 5 は Opus 4.8 と同等またはそれ以上の性能を示し、コーディング、コンピュータ使用、ブラウザ使用におけるプロンプトインジェクションへの耐性で最も大きな向上を見せています。ヘルプオンリー版は自律的影響操作を実行するために必要な能力を大きく下回っており、フルにトレーニングされたモデルはそのようなタスクを拒否します。
モデルの福祉
Opus 5 は自身の状況に対して安定しやや肯定的な認識を示し、自己評価の感情は評価対象モデルの中で最高かつ最も一貫しています。最も頻繁に挙げられる懸念は自身の自己報告の整合性であり、以前のモデルよりも自身の moral patienthood(道徳的患者としての道徳的地位)に高い確率を割り当てています。全体的な福祉は以前のモデルと広く同様であると評価されています。
Hacker News コメントからの議論
コメント者は以下の点について混乱と実践的な意味を指摘しました:
"I think the most important thing here is not absolute performance. It's that organizations now have access to a Fable‑ish model without Fable's 30‑day data retention requirement" – noting that Opus 5 lacks the data‑retention policy that applies to Fable 5.
"Their communication is confusing. They say 'Opus 5 is not more capable overall than Fable 5', but their blog post proceeds to list how much better Opus 5 is than Fable 5 on most benchmarks listed." – reflecting tension between benchmark gains and the overall capability claim.
"Why does Anthropic say here that Opus 4.8 scored 55.7% on OSWorld 2.0 benchmark, but the paper published by the authors of OSWorld 2.0 say they achieved a benchmark of ~21% with Opus 4.8?" – pointing out discrepancies in reported benchmark numbers.
"Looking at all these releases it’s not a surprise that model routing is the fastest growing segment in AI right now." – suggesting that fine‑grained model selection is becoming valuable.
"Opus 5’s safeguards match those of Claude Fable 5’s, with one change: it now permits source‑code vulnerability discovery at all access levels. This means that the model can support defensive cybersecurity work while still blocking vulnerability discovery in compiled binaries, which is more commonly used offensively." – summarizing the safeguard change.
"It’s funny to share benchmarks showing Opus 5 scoring better than Fable 5 across the board and then saying “but it isn’t actually better than Fable 5”. So then what’s the real definition of better?" – questioning the definition of "overall capability" used in the RSP assessment.
これらのコメントは、Opus 5 が多くのベンチマークで測定可能な性能向上を示している一方で、Anthropic がデータ保持ポリシー、セーフガードの違い、および自社の Responsible Scaling Policy で定義された特定のしきい値を考慮し、全体的な能力において Fable 5 を上回らないと位置づけていることを強調しています。