MAI-Cyber-1-Flash 在 MDASH 内部 – Microsoft AI 公告
MAI-Cyber-1-Flash inside MDASH – Microsoft AI Announcement
Overview
Microsoft 在 MDASH 内部推出了 MAI-Cyber-1-Flash,相比其之前的最佳产品,在 CyberGym 基准上实现了顶尖性能,成本降低了 50%。
Model Architecture and Lineage
MAI-Cyber-1-Flash 是一个紧凑且代码密集的安全模型,源自 MAI-Thinking-1 谱系,内部基于高质量数据构建。
Performance on CyberGym
在 CyberGym 基准测试中,MDASH 使用 MAI-Cyber-1-Flash 与 GPT-5.4 结合,成功率达到 95.95%,比 Mythos 高 12 个百分点,超过了其他模型 83.2%–85.6% 的范围。
Cost Efficiency and Multi‑Model Strategy
该系统的设计使得 MAI-Cyber-1-Flash 可处理多达 90% 的任务,从而让更大且更昂贵的 GPT-5.4 模型仅保留用于剩余 10% 的极其困难任务;这种分割相比 MDASH 之前的最佳产品(GPT 5.4 + 5.4 mini + 5.3 codex)实现了 50% 的成本节约。
Data and Harness Advantages
Microsoft 的遥测提供了跨身份、终端、云和网络的万亿级每日信号,此外还有无与伦比的真实漏洞和修复记录;MDASH 的多代理 harness 由行业专家调校,包含 100+ 个使用多种领先模型的代理,提升了模型的有效性。
Safety, Governance, and Enterprise Controls
由于这是 Microsoft 的首个网络模型,MAI-Cyber-1-Flash 在安全首要的校准下开发,经过 Microsoft AI Red Team 的严格评估,通过自动化和专家领导的对抗性演练进行测试,并由第三方独立评估;通过 MDASH 提供的企业控制措施包括基于角色的访问、租户隔离、加密、可审计性以及无需互联网访问的沙盒执行。
Reinforcement Learning Loop and Future Agentic Systems
Microsoft 的安全运营每天从 160 万客户那里产生超过 100 万亿条安全信号,构建了一个实时强化学习循环,持续改进模型;新推出的 Perception agentic 安全系统将把 MAI‑Cyber‑1‑Flash 扩展到漏洞识别以外的其他安全工作流。
Community Questions and Limitations
评论者提出了若干公告尚未涉及的问题:
"If I’m being frivolous, does this mean Microsoft’s model is best at fixing Microsoft products because they have trillions of data points on problems with Microsoft products" – @gste
"Does it work on Linux" – @tesdinger
"Open weights, please?" – @LorenDB
"How much of this security related stuff is open source? Models, training data, evaluation benchmarks?" – @chvid
"the problem is they are using cybergym, which has been saturated for a few months now." – @danieltk76
这些言论凸显了关于平台兼容性、模型开放性、命名约定以及 CyberGym 基准当前饱和状态的担忧。