Quasar 438B 發佈:歐洲得分最高的推理模型

Quasar 438B 是 Artificial Analysis Intelligence Index 上得分最高的歐洲模型

Multiverse Computing 已發佈 Quasar 438B,這是一款專為企業級代理(agents)和編碼設計的推理模型。該模型在 Artificial Analysis Intelligence Index v4.1.1 中獲得 43 分,是該基準測試中任何歐洲模型的最高得分。它旨在處理需要規劃、工具使用、程式碼執行和大型上下文窗口的多步驟任務,並可透過 CompactifAI API 使用。

Perfomance and Intelligence Benchmarks

Quasar 438B 在 Artificial Analysis Intelligence Index(包含 GPQA Diamond 和 SciCode 等九項評估的綜合指標)中獲得 43 分,使其領先於其他幾款模型,包括:

  • Inkling: 42
  • NVIDIA Nemotron 3 Ultra: 38
  • Mistral Medium 3.5: 30

雖然它領先於這些模型,但仍落後於 Claude Opus 5 等前沿模型,後者以 63 分領先該指數。

Latency and Response Speed

Quasar 438B 在 400B+ 參數級別中針對低延遲進行了優化。它在 15.3 秒內回傳 500 個 token(包含思考時間)。與其他高分模型相比:

  • Mistral Medium 3.5: 18.8 秒 (Score: 30)
  • NVIDIA Nemotron 3 Ultra: 25.7 秒 (Score: 36)
  • Inkling: 48.3 秒 (Score: 42)

在比較中只有三款模型更快:Nemotron 3.5 Lightning (9.4s)、Gemini 3.5 Flash-Lite (10.8s) 和 Gemini 3.7 Flash (11.5s)。其中,只有 Gemini 3.7 Flash 在智能指數上優於 Quasar,且提供了更快的響應速度。

Long-Context and Agentic Capabilities

Quasar 438B 在長上下文推理和基於終端機的代理工作方面展現了強大的性能:

  • Long-Context Reasoning (AA-LCR): Quasar 獲得 75.0 分,與 Grok 4.6 (high) 持平,且與 Claude Opus 5 (75.7) 和 Qwen3.8 2.4T A95B (75.3) 的差距僅在一分以內。
  • Agentic Coding (Terminal-Bench v2.1): Quasar 獲得 69.3 分,領先 Mistral Medium 3.5 18.7 分,並領先 Nemotron 3 Ultra 15.4 分,儘管它落後於 Claude Opus 5 (89.1)。

Community Critique and Technical Skepticism

在發佈公告之後,Hacker News 上的技術討論對該模型的來源、透明度以及公司的聲明提出了幾點疑慮。

Model Provenance and Transparency

批評者指出,關於該模型是從頭開始預訓練還是現有模型的壓縮/微調版本,缺乏透明度。一些用戶建議,根據參數數量和變更日誌(changelog)條目,它可能是 GLM 5.2 或 MiniMax M3 的修改版本。

"I think this is GLM 5.2 with parameters removed. Its advertised in their changelog as 'capabilities are identical to GLM 5.2,' it has the same two effort settings 'high' and 'max'"

Skepticism of "Quantum AI" Claims

Multiverse Computing 使用 "quantum physics" 和 "tensor networks" 透過其 CompactifAI 技術進行模型壓縮,這引起了懷疑。用戶質疑量子算法在目前 LLM 開發中的實用性,並將該公司的部分營銷手段描述為 "technobabble"。

Closed Weights and Benchmarking

由於 Quasar 438B 僅透過 API 使用且權重不公開,一些開發者對提供的基準測試表示不信任,指出 "frontier" 模型通常包含隱藏的系統提示詞(system prompts)或封裝器(wrappers),這會使分數與 "naked" 開放權重模型相比時顯得較高。

"When the weights are closed I don't believe any benchmark... 'Frontier' models get tested as a model + whatever secret sauce they choose to put in front."

Deployment and Availability

Quasar 438B 支援 English 和 Spanish,並可透過 CompactifAI API 存取,讓企業團隊能夠將該模型整合到軟體開發代理(agents)、技術 Copilot 和研究系統中,而無需管理自己的基礎設施。

Sources

相關