MAI-Cyber-1-Flash 在 MDASH 内部 – Microsoft AI 公告

MAI-Cyber-1-Flash inside MDASH – Microsoft AI Announcement

Overview

Microsoft 在 MDASH 内部推出了 MAI-Cyber-1-Flash,相比其之前的最佳产品,在 CyberGym 基准上实现了顶尖性能,成本降低了 50%。

Model Architecture and Lineage

MAI-Cyber-1-Flash 是一个紧凑且代码密集的安全模型,源自 MAI-Thinking-1 谱系,内部基于高质量数据构建。

Performance on CyberGym

在 CyberGym 基准测试中,MDASH 使用 MAI-Cyber-1-Flash 与 GPT-5.4 结合,成功率达到 95.95%,比 Mythos 高 12 个百分点,超过了其他模型 83.2%–85.6% 的范围。

Cost Efficiency and Multi‑Model Strategy

该系统的设计使得 MAI-Cyber-1-Flash 可处理多达 90% 的任务,从而让更大且更昂贵的 GPT-5.4 模型仅保留用于剩余 10% 的极其困难任务;这种分割相比 MDASH 之前的最佳产品(GPT 5.4 + 5.4 mini + 5.3 codex)实现了 50% 的成本节约。

Data and Harness Advantages

Microsoft 的遥测提供了跨身份、终端、云和网络的万亿级每日信号,此外还有无与伦比的真实漏洞和修复记录;MDASH 的多代理 harness 由行业专家调校,包含 100+ 个使用多种领先模型的代理,提升了模型的有效性。

Safety, Governance, and Enterprise Controls

由于这是 Microsoft 的首个网络模型,MAI-Cyber-1-Flash 在安全首要的校准下开发,经过 Microsoft AI Red Team 的严格评估,通过自动化和专家领导的对抗性演练进行测试,并由第三方独立评估;通过 MDASH 提供的企业控制措施包括基于角色的访问、租户隔离、加密、可审计性以及无需互联网访问的沙盒执行。

Reinforcement Learning Loop and Future Agentic Systems

Microsoft 的安全运营每天从 160 万客户那里产生超过 100 万亿条安全信号,构建了一个实时强化学习循环,持续改进模型;新推出的 Perception agentic 安全系统将把 MAI‑Cyber‑1‑Flash 扩展到漏洞识别以外的其他安全工作流。

Community Questions and Limitations

评论者提出了若干公告尚未涉及的问题:

"If I’m being frivolous, does this mean Microsoft’s model is best at fixing Microsoft products because they have trillions of data points on problems with Microsoft products" – @gste

"Does it work on Linux" – @tesdinger

"Open weights, please?" – @LorenDB

"How much of this security related stuff is open source? Models, training data, evaluation benchmarks?" – @chvid

"the problem is they are using cybergym, which has been saturated for a few months now." – @danieltk76

这些言论凸显了关于平台兼容性、模型开放性、命名约定以及 CyberGym 基准当前饱和状态的担忧。

Sources