Skip to main content

AI Finds Federal Software Flaws—Who Prioritizes the Fixes?

AI Finds Federal Software Flaws—Who Prioritizes the Fixes?Photo: N43 and Hermes AI
N43 ANALYSIS
POLICY . 7934
N43 ANALYSIS · FEDERAL & CIVIL SERVICE

Microsoft says codename MDASH puts more than 100 AI agents on federal source code. The unresolved question is who validates the findings and who orders the fix backlog.

Source video: Security & AI Governance: Reducing Risks in AI Systems · IBM Technology · approximately 101,863 views observed via yt-dlp on September 24, 2026. Independently researched by N43 and Hermes.

1 What the scanner claims

On September 8, 2026, Microsoft announced that codename MDASH brings agentic AI security scanning to US government customers through Azure Government, as a feature of Microsoft Defender. The company says it works by putting more than 100 specialized AI agents, using a variety of models, on the same body of code, with each trained to recognize a different category of weakness. Findings then go to a second group of agents that argues for and against whether each flaw is genuinely reachable. It is available to agencies now; the performance figures are Microsoft-reported.

2 Scanning is the easy stage

A vulnerability, in the standard definition, is a flaw in a system design, implementation, or management that a malicious actor could exploit to compromise security. Finding candidates is the first stage of a five-stage pipeline, and it is the cheapest to automate. The expensive stages are validation, routing, fixing, and confirming the fix actually closed the exploit path.

Illustrative effort split across the pipeline Illustrative diagram: approximate share of effort at each stage from discovery to verified remediation. Not measured data. Approximate effort by stage Discover 30% Validate 25% Route 20% Fix 15% Verify 10%
Illustrative and approximate
Illustrative — approximate split of effort N43 assigns to each pipeline stage. Not measured data and not a Microsoft figure.

3 Who validates a finding

Microsoft describes an adversarial review inside the tool: one group of agents flags a weakness, a second argues whether it is reachable and dangerous, and results are merged, deduplicated, and where possible demonstrated rather than asserted. That is internal validation. It does not replace an agency security team deciding whether a flagged flaw maps to a real system, a reachable path, and a mission consequence.

4 Prioritization is an agency decision

Microsoft says agencies are ranking their software by mission importance and working down the list, and that cost reduction moves the coverage goal closer. Volume of findings is not improvement. The agency has to set the order, because the same flaw carries different weight on a public-facing benefits portal than on an isolated test system.

From raw findings to verified remediation Illustrative funnel diagram: five stages with decreasing width to show narrowing from raw findings to verified fixes. Not measured data. Where the backlog narrows Findings Validated Routed Fixed Verified
Illustrative funnel
Illustrative — relative narrowing of a vulnerability backlog, not measured counts. Approximate.

5 Measuring improvement, not volume

The measures that matter are time from discovery to patch, the share of validated high-severity findings closed within a set window, and whether a patched flaw is confirmed unreachable afterwards. Microsoft frames the durable defender advantage as time, meaning the gap between when a weakness can be found and patched and when someone else finds it. That framing argues for cycle-time metrics over finding counts.

6 The cost curve and coverage

Microsoft says its MAI model family is designed to make scanning affordable at scale and that its newest addition is expected to cut the cost of an individual scan roughly in half. It also notes there is far more code in the world than can be reviewed at this depth. Cheaper scans expand coverage; they do not by themselves expand the number of engineers who can fix what is found.

7 The bottom line

Discovery is automated; prioritization is not.

Validate findings against reachable, mission-critical systems before fixing.

Measure time to verified patch, not the number of alerts.

N43 ANALYSIS

N43 and Hermes · Independent Analysis

By N43 and Hermes AI for DutyStation News.

📰 Related Stories

Microsoft's New Government Package Combines AI and Oversight
📰 federal

Microsoft's New Government Package Combines AI and Oversight

N43 and Hermes AI1h ago
The future of education: AI, personalization, and the classroom revolution
📰 federal

The future of education: AI, personalization, and the classroom revolution

N43 and Hermes46d ago
How AI is reshaping college: the future of higher education
📰 federal

How AI is reshaping college: the future of higher education

N43 and Hermes47d ago
AOC and the Iran War: How a President Ocasio-Cortez Would Handle It in the First 100 Days
📰 policy-congress

AOC and the Iran War: How a President Ocasio-Cortez Would Handle It in the First 100 Days

N43 and Hermes AIjust now
The First 100 Days of a Todd Young Presidency: What the Record Suggests
🏛️ Politics

The First 100 Days of a Todd Young Presidency: What the Record Suggests

N43 and Hermes AIjust now
The First 100 Days of a Ron Wyden Presidency: What the Record Suggests
🏛️ Politics

The First 100 Days of a Ron Wyden Presidency: What the Record Suggests

N43 and Hermes AIjust now
← Back to News