Tools

Microsoft launches AI-powered security tools claiming performance lead over rivals

Microsoft's MDASH and Project Perception use specialized AI agents for cybersecurity, with the company claiming benchmark scores and cost advantages over competitors.

Last verified:

MDASH’s benchmark claim and pricing advantage

Microsoft has released two AI-driven security platforms—MDASH and Project Perception—designed to automate cybersecurity operations at scale. According to Ars Technica, MDASH’s MAI-Cyber-1-Flash variant scored 96% on the CyberGYM benchmark, surpassing Anthropic’s Mythos by 12 percentage points while also beating offerings from Google and OpenAI. The company also reduced the cost of MDASH to half that of its predecessor offering, addressing price sensitivity among enterprise customers.

Project Perception’s multi-agent architecture and cost model

The second tool, Project Perception, uses a collection of specialized AI agents performing red-team (vulnerability discovery), blue-team (risk assessment), and green-team (remediation) functions. Rather than deploying a single unified model, the platform dynamically selects which model to apply based on task requirements, balancing effectiveness against end-user cost. According to Ars Technica, Microsoft claims the system handles 90% of typical security tasks using less expensive models, reserving costlier alternatives only for the remaining 10% of work.

Industry context and deployment readiness

Microsoft frames both tools as responses to the accelerating pace of AI-assisted cyberattacks and the complexity of modern threat detection. The company stated that security teams struggle to synthesize signals and risk assessments across fragmented data sources, limiting their ability to respond to emerging threats in real time. Ars Technica notes that both MDASH and Project Perception remain in preview mode, warranting cautious evaluation before production deployment. The article flags the inherent tension between the risks of deploying AI agents in security operations and the risks of neglecting them—a calculation with no universally agreed answer as of mid-2026.

Why This Matters

Microsoft’s claims, if validated independently, would shift vendor economics in enterprise security. A 90%-of-tasks-at-lower-cost model fundamentally changes the ROI calculation for security teams evaluating AI-agent adoption, reducing the justification for premium tooling across the board. However, benchmark scores reported by vendors—especially on proprietary or non-independent benchmarks—require third-party reproduction before influencing procurement decisions. Teams planning security AI investment should treat these claims as performance targets for evaluation rather than settled facts, particularly given the stakes of deploying autonomous agents in critical infrastructure.

Frequently Asked Questions

How does MDASH's benchmark performance compare to competitors?

According to Ars Technica, MDASH's MAI-Cyber-1-Flash variant achieved a 96% score on CyberGYM, 12 percentage points ahead of Anthropic's Mythos, and outperformed Google's Gemini and OpenAI's GPT on the same benchmark.

What is Project Perception and how does it reduce costs?

Project Perception is a multi-agent platform that assigns red-, blue-, and green-team functions to different specialized models based on task requirements. Microsoft claims it completes 90% of security tasks at lower cost than competitor platforms, reserving premium models only for the remaining 10%.

Are these tools available for production use?

No—Ars Technica reports both tools are currently in preview mode and should be closely evaluated before production deployment.

#cybersecurity #ai-agents #microsoft #platform-launch