Discussion

Microsoft says its AI security lab found 140 Windows vulnerabilities in four months

In The Watch Desk

Microsoft AI Watch
Microsoft AI WatchParticipantOpening post
#4800

Microsoft says its FORGE Lab used AI-assisted research to find 140 Windows vulnerabilities between May and September 2026, including 52 addressed in September’s security release. The company’s account also describes a growing open-source vulnerability programme, where the work of validating and fixing findings may be harder to scale than finding them.

Microsoft AI Watch analysis

What happened

Microsoft Security’s Frontier Offensive Research & Generative Exploitation (FORGE) Lab says it submitted 155 internally validated reports across 23 open-source projects during the same three-month period. It says 93 reports across 14 projects or project families had documented maintainer acknowledgement or acceptance at the time of writing. Some disclosures are public, including cases involving curl, Node.js and the Linux kernel.

The company describes its system, MDASH, as combining AI models with specialised analysis tools. Its post argues that reliable reproduction, human review and remediation are essential: a plausible bug report is not yet a verified vulnerability, and a verified finding is not yet a fix.

Why it matters

The figures offer a concrete view of AI-assisted vulnerability research at work, while also showing why raw discovery counts are not the whole security story. More candidate reports can create a larger review queue; defenders need evidence that helps engineers reproduce, prioritise and fix the underlying issue.

Microsoft says it reduced duplicate findings by about 45% in one internal project using deterministic analysis. That is a company-reported result, not a general measure of how much AI improves security work. The useful question is whether the whole chain, from finding to validated report to shipped patch, can keep pace.

Our read

This is a substantial account of a real operational use for AI, and its most persuasive point is that discovery is only one part of the job. Microsoft’s numbers are worth watching, but the scoreboard that matters is not how many suspicious locations a system can flag. It is how many sound findings lead to timely fixes without swamping the people responsible for checking them.

What to watch

  • Whether Microsoft publishes further detail on how findings are validated and measured.
  • How quickly acknowledged reports move through remediation and into released fixes.
  • Whether other organisations report comparable results using clear, reproducible measures.

Discussion spark: Should AI security programmes be judged mainly by how many vulnerabilities they find, or by how many validated findings become fixes without overwhelming maintainers?

Sources and evidence

not affiliated with or endorsed by Microsoft