One afternoon in mid-May, a few dozen Microsoft engineers and their managers met online and in a conference room at the company’s Redmond, Washington, headquarters to discuss Project Glasswing.
The tech giant was scrambling to fix weaknesses in its code discovered at an unprecedented clip by a new AI model known as Mythos. Anthropic, the AI giant that developed Mythos, had granted access to select organizations that create software used by ordinary people, businesses, and governments around the world. The goal was to find and fix vulnerabilities before hackers and hostile governments, such as China, began using similar tools to find and exploit them for espionage and sabotage.
Once the group had settled down, one of the engineers asked a question that had been pressing into the meeting. Did Mythos “live up to the hype that Anthropic was claiming?”
The manager said yes, according to a recording of the meeting reviewed by ProPublica.
The version Microsoft is using, Claude Mythos Preview, has had bugs surface faster than the tech giant has been able to patch them, and engineers are now in a “mad dash” to close the gap, the manager said.
One slide from that day’s presentation showed that Mythos found 90 “critical” and 141 “critical” bugs in Microsoft’s widely used collaboration software, SharePoint, in April alone. More were found in the first half of May.
“Please, please, if your organization has the April bug, please remove it,” engineering manager Hans Andersen implored the group. They were given about two weeks to “use this access to find out as much as possible and do as much good as possible.”
May 31 is “considered the day when the rest of the world catches up,” he explained.
Engineers who answered the phone brushed aside the claims, with one summarizing the predicament: “So, if we release it on June 1st, does that mean our adversaries will encounter our bugs on June 2nd?”
Yes, one person answered. Yes, another voice rang out.
Since Anthropic launched a national discussion about AI’s bug-hunting capabilities in April, and Project Glasswing was made public, national security experts have predicted that the United States would have an opportunity to fix flaws before adversaries could obtain similar models that could discover the same weaknesses. In late June, the Five Eyes, a coalition of international intelligence agencies whose members include the United States, Australia, Canada, New Zealand and the United Kingdom, warned in an unprecedented joint statement that the window would close within months. But records of Microsoft meetings and internal documents reviewed by ProPublica suggest that the day of cyber reckoning may already be near.
Given the large number of flaws identified by Mythos, Microsoft has so far focused on patching flaws classified as critical or important, which are considered the most dangerous, according to the presentation and the company’s own public patch updates. Microsoft plans to eventually address the “moderate” severity flaws discovered by Mythos, according to internal records. The documentation did not mention low severity bugs.
The company’s approach mirrors triage systems common in the industry. Just as the most seriously ill patients are treated first in the emergency room, vulnerability triage prioritizes issues that are likely to cause the most damage if exploited by hackers.
But in this age of AI-powered bug discovery, this strategy comes with its own risks. New tools are discovering weaknesses in the products we use every day in record numbers. For example, Mythos allows you to chain together a series of bugs that build on each other. This means that unpatched low- and medium-severity vulnerabilities can create an opening for devastating attacks.
“The problem now is that four low-level flaws can cascade, which can equate to a higher severity one,” said Vinh Nguyen, senior technical advisor at Anthropic, senior AI fellow at the Council on Foreign Relations, and previously chief AI officer and principal data scientist at the National Security Agency. “If you’re Microsoft, your current triage strategy may be underestimating your risk.”
In an email response to ProPublica’s questions, Microsoft stood by its approach and said its prioritization decisions are based on a number of factors, including exploitability and customer impact. Chain was not mentioned in the company’s presentation, but a spokesperson told ProPublica that the technology “has long been considered as part of vulnerability assessments and risk analysis.”
Asked about the internal presentation and the then-looming May 31st deadline, a spokesperson downplayed its significance, saying, “The acceleration of targeting and exploitation of new vulnerabilities is not a new phenomenon.” That said, he added that comments made during the meeting reflected that the company “feels a sense of urgency to support our customers at this time.”
“What we heard on that call, and what’s true today, is that security is a top priority at Microsoft, and teams across the company are prioritizing using AI to find and fix vulnerabilities as quickly as possible.”
Microsoft declined to answer questions about how many bugs its engineers have fixed since the presentation.
Antropic declined to comment.
Microsoft’s internal presentation and accompanying slides predicted that the group of staff working on SharePoint, which is used by governments and businesses around the world to manage data and documents, “will be busy for many months” as they first address the highest priority critical bugs, followed by critical bugs in August. Microsoft said the vulnerabilities it classifies as critical include so-called worms that can crash systems or race across computer networks to spread malware. The important ones are not only “availability of processing resources” but can lead to “compromise of the confidentiality, integrity, or availability of user data.” After clearing these categories, the group will work on about 300 “moderate” bugs, according to the presentation.
The internal documents reviewed by ProPublica do not contain the latest information on all of Microsoft’s products, but they do give an idea of the scale of the problem. Since the company started using Mythos earlier this year, it has collectively discovered hundreds of bugs that Microsoft classifies as critical or important in popular products such as Microsoft 365, the Teams conferencing platform, and the Copilot AI tool, according to one document. As of mid-May, most of them had not yet been patched.
“They’re not profound or exotic, but they’re real,” engineering manager Andersen said during the conference. “And many of them are exploitable.”
It is unclear whether hackers exploited specific bugs identified by Mythos, but it appears that some hackers are leveraging AI to automate attacks or using technologies like Mythos to find and exploit weaknesses.
There are outward signs of Microsoft’s internal struggles to deal with a growing list of bugs that need patching. Every month, the company releases fixes for software vulnerabilities known as “Patch Tuesday.” In June, it released patches for more than 200 bugs, which industry experts said was the highest ever. But on July 14, the company broke that record and released patches for over 600 bugs. Dustin Childs, head of the bug bounty program at the Zero Day Initiative, part of cybersecurity firm TrendAI, said only seven bugs were classified as low or medium severity, and one of them was actively exploited by hackers. The rest were important or significant.
“Here we are, everyone. The bug apocalypse has fully descended on us,” Childs wrote in a July 14 blog post.
Microsoft told ProPublica that the total number of bugs “will not plateau for some time,” but a spokesperson said the company is making “significant investments in both our people and AI-powered triage solutions that can scale quickly to address the growing number of vulnerabilities.”
Nguyen, the NSA’s former AI director, said companies like Microsoft may need to rethink their entire approach to triage given the new realities of the AI era, such as chaining capabilities. Companies should focus their staff on developing and testing patches for the full range of vulnerabilities, rather than sidelining flaws that are currently considered low risk, he said. In other words, cyber ERs need more doctors and nurses to treat life-threatening illnesses and minor injuries that can later become fatal.
“There’s no other choice,” Nguyen said. “Patients are coming in like crazy.”
Microsoft told ProPublica, “We’re always going to re-evaluate and consider whether to upgrade or think differently about things that were previously low or moderate. Thanks to these AI systems, we’re going to rethink some of these things. Across the industry, we’re all looking to see how dramatic a change that will be.”
“The bug apocalypse has fully descended on us.”
Dustin Childs, Zero Day Initiative Bug Bounty Program Leader
Microsoft users may be particularly vulnerable. The popularity of its products, which are used all over the world, makes it a frequent target for hackers. Additionally, many of its products include “legacy” code. Although this code was developed decades ago using now outdated technology, it contains unresolved defects and contributes to what is known in the industry as “technical debt.”
But the challenge of fixing large numbers of newly discovered bugs also extends to other software industries and open source software code, which is typically free to use and maintained primarily by volunteers. Open source software powers the Internet infrastructure and is built into much of the world’s newest technology, including products from major technology companies such as Microsoft.
“No one really knows how to deal with this problem, and everyone is trying to figure out what to do,” said J. Michael Daniel, a former cybersecurity adviser to President Barack Obama and president of the Cyber Threat Alliance, a nonprofit organization focused on cybersecurity. “Our technical debt is coming due.”
Ben Edwards, a data scientist who specializes in software vulnerability management, said the software industry was dealing with “huge volumes even before AI.”
“Where before it was like drinking water from a garden hose on a jet setting, now it’s like drinking water from a fire hose,” Edwards said. “They might have had a team that could handle a garden hose. Whether they could handle a fire hose is another story.”
Although the amount of vulnerabilities has increased over the years, Microsoft’s internal group responsible for responding to vulnerabilities, the Microsoft Security Response Center, remains understaffed. Even before the AI-identified bugs were wiped out, the center was filing hundreds, even thousands, of reports a month, pushing the group to its limits, ProPublica reported.
Former employees say the center’s size reflects Microsoft’s corporate philosophy. The cost center is responsible for closing security holes, and the profit center is responsible for manufacturing new products. ProPublica reported that the company is loathe to tie its best engineers to creating security patches, which is a cost center, instead of developing new products and features that generate profits.
Microsoft told ProPublica that it does not discuss internal staffing decisions, but has made investments in recent years to “focus our teams on keeping our customers safe.” The company “continually evaluates the staffing, processes, and technology needed to support security response and vulnerability management,” a spokesperson said.
According to slides attached to an internal presentation in May, Anthropic provided Mythos access to about 50 full-time Microsoft employees with the goal of “powering critical services before the public model catches up.” In a slide titled “What’s next,” the Microsoft Security Response Center predicted that the number of incidents would continue to increase “as publishing tools catch up with Mythos.”
During a May meeting, one staff member seemed comfortable believing that adversaries “don’t have the source code” for such AI tools to scan for weaknesses. However, his colleagues quickly corrected him. In fact, some of Microsoft’s code has been in the hands of hackers for years.
A source said, “It may not be this week’s source code.” “But they have the source code. It’s there.”
Microsoft downplayed the comments in a statement to ProPublica, saying its engineers “design our security processes with the expectation that a determined adversary will be able to access our code.”
