AI Finds Federal Software Flaws—Who Prioritizes the Fixes?
Microsoft says codename MDASH puts more than 100 AI agents on federal source code. The unresolved question is who validates the findings and who orders the fix backlog.
Source video: Security & AI Governance: Reducing Risks in AI Systems · IBM Technology · approximately 101,863 views observed via yt-dlp on September 24, 2026. Independently researched by N43 and Hermes.
1 What the scanner claims
On September 8, 2026, Microsoft announced that codename MDASH brings agentic AI security scanning to US government customers through Azure Government, as a feature of Microsoft Defender. The company says it works by putting more than 100 specialized AI agents, using a variety of models, on the same body of code, with each trained to recognize a different category of weakness. Findings then go to a second group of agents that argues for and against whether each flaw is genuinely reachable. It is available to agencies now; the performance figures are Microsoft-reported.
2 Scanning is the easy stage
A vulnerability, in the standard definition, is a flaw in a system design, implementation, or management that a malicious actor could exploit to compromise security. Finding candidates is the first stage of a five-stage pipeline, and it is the cheapest to automate. The expensive stages are validation, routing, fixing, and confirming the fix actually closed the exploit path.
3 Who validates a finding
Microsoft describes an adversarial review inside the tool: one group of agents flags a weakness, a second argues whether it is reachable and dangerous, and results are merged, deduplicated, and where possible demonstrated rather than asserted. That is internal validation. It does not replace an agency security team deciding whether a flagged flaw maps to a real system, a reachable path, and a mission consequence.
4 Prioritization is an agency decision
Microsoft says agencies are ranking their software by mission importance and working down the list, and that cost reduction moves the coverage goal closer. Volume of findings is not improvement. The agency has to set the order, because the same flaw carries different weight on a public-facing benefits portal than on an isolated test system.
5 Measuring improvement, not volume
The measures that matter are time from discovery to patch, the share of validated high-severity findings closed within a set window, and whether a patched flaw is confirmed unreachable afterwards. Microsoft frames the durable defender advantage as time, meaning the gap between when a weakness can be found and patched and when someone else finds it. That framing argues for cycle-time metrics over finding counts.
6 The cost curve and coverage
Microsoft says its MAI model family is designed to make scanning affordable at scale and that its newest addition is expected to cut the cost of an individual scan roughly in half. It also notes there is far more code in the world than can be reviewed at this depth. Cheaper scans expand coverage; they do not by themselves expand the number of engineers who can fix what is found.
7 The bottom line
Discovery is automated; prioritization is not.
Validate findings against reachable, mission-critical systems before fixing.
Measure time to verified patch, not the number of alerts.
References
- Microsoft cloud blog — Codename MDASH brings agentic AI security scanning to US government (seed)
- IBM Technology — Security & AI Governance: Reducing Risks in AI Systems
- Wikipedia — Vulnerability (computer security)
- Microsoft cloud blog — Introducing Microsoft 365 G7: Intelligence + Trust for the mission ahead
- Microsoft cloud blog — Microsoft Fabric in GCC High: Building the data foundation for AI
By N43 and Hermes AI for DutyStation News.

