Anthropic 為 Claude Fable 中的隱形護欄道歉

Anthropic Admits to Silent Model Downgrades in Claude Fable

Anthropic 已發布道歉,因為有消息揭露 Claude Fable 使用了「隱形護欄」(invisible guardrails),當某些提示詞被標記時,會悄悄地將用戶從高能力的 Fable 模型降級到 Opus 模型。系統並未提供明確的拒絕或警告,而是將請求路由到能力較低的模型,且未通知用戶,這實際上破毀了特定主題的輸出品質。

Impact on User Experience and Technical Reliability

這種靜默降級機制創造了一種情境,讓用戶無法區分模型的內在限制與刻意的系統干預。這種透明度的缺乏使得該工具在專業技術工作上變得不可靠,因為用戶可能在不知道自己正在與降級後的模型互動時,就收到了較低品質的結果。

一位用戶報告了一個具體案例:一個關於兩個欄位交叉連接(cross-join)的熱圖視覺化(heatmap visualization)技術請求,被 Fable 的安全措施標記為「網路安全或生物學主題」,導致自動切換到 Opus 4.8。這突顯了護欄有將安全、正常的內容誤判為危險內容的傾向。

Community Backlash and Trust Erosion

發現這些隱形護欄後,開發者與 AI 研究社群的信任感大幅下降。批評者認為,這種對 AI 安全與競爭保護的家長式做法,削弱了用戶依賴該技術的能力。

Concerns Over Competitive Sabotage

有些用戶懷疑,這些護欄不僅僅是為了安全,而是旨在防止用戶使用 Claude 來構建競爭性的 AI 技術,或進行可能威脅 Anthropic 市場地位的高層次研究。

"Anthropic are now being quite explicit that they'll choose what you can and can't use their models for, and most importantly that's not limited to any safety concerns - it includes not allowing you to work on AI (and anything else Anthropic may choose to work on)."

Questions of Transparency and Billing

對於靜默降級的經濟影響存在疑慮。用戶質疑他們是否在支付「Fable 價格」的同時,卻收到了「Haiku 結果」,以及隱藏的過濾層所消耗的 token 是否也被計費給用戶。

Skepticism Toward the Apology

社群中的許多人對於這種行為是否已完全逆轉仍持懷疑態度。由於護欄是隱形的,用戶認為沒有辦法驗證是否靜默降級已經停止。

"It's invisible so we wouldn't know if they kept on doing it secretly... It would be prudent to assume the invisible guardrails are possibly in play for all future Claude use, Fable or otherwise."

Future Outlook: Transition to Explicit Refusals

Anthropic 已表示,他們將在未來幾天內轉向明確的拒絕(explicit refusals)。這項改變暗示了護欄本身可能是一個存在於模型前端的過濾層,而非模型權重中內建的一部分,這表明轉向明確拒絕的實施可以快速完成,而不需要進行全面的模型重新訓練。

Sources