AI is finding your bugs faster than you can patch them

AI is finding your bugs faster than you can patch them

Two stories landed this week that, taken together, should make every development and security team sit up.

First, the big one: Anthropic announced Claude Mythos Preview, a new AI model it says is too dangerous to release publicly. The reason? It's absurdly good at finding security vulnerabilities. We're not talking about toy examples in CTF challenges. Mythos found a 27-year-old bug in OpenBSD, one of the most security-hardened operating systems on the planet. It found a flaw in FFmpeg that automated testing tools had hit five million times without catching. It chained together multiple Linux kernel vulnerabilities to go from ordinary user access to full machine control. Thousands of zero-days across every major OS and browser, most of them unpatched until Anthropic reported them.

Anthropic isn't releasing the model to the public. Instead, they've launched Project Glasswing, giving access to about 40 organizations (including some of the largest tech companies in the world) to scan and fix their own code. They're also providing $100 million in usage credits and $4 million to open-source security foundations.

Second story, smaller but arguably more instructive: CVE-2026-39987, a pre-authentication remote code execution flaw in Marimo, an open-source Python notebook popular with data scientists. CVSS score 9.3. The vulnerability was in a WebSocket endpoint (/terminal/ws) that simply didn't check authentication, even though every other WebSocket endpoint in the codebase did. Sysdig deployed honeypots and observed exploitation within 10 hours of public disclosure. Ten hours.

What these two stories have in common

The thread connecting Mythos and the Marimo CVE is the same: the window between a vulnerability existing and it being exploited has collapsed.

With Mythos, we're looking at AI that can find bugs human researchers missed for decades. With the Marimo case, we're looking at attackers who weaponize a disclosed flaw before most teams have even read the advisory. Both point in the same direction: patching speed and attack surface awareness are no longer optional disciplines. They're survival.

And here's what makes me uneasy. Anthropic's head of frontier red teaming estimated that open-weight models will catch up to Mythos-level bug-finding capability within six to eighteen months. That means this isn't a capability that stays locked behind a gated research preview forever. It proliferates.

What this means for development teams

If you build web applications, APIs, or anything with a WebSocket layer, the Marimo bug should feel uncomfortably familiar. One endpoint that skipped authentication. Every other endpoint did the right thing. That's the kind of inconsistency that slips through code review, passes QA, and sits in production for months.

From our PHP and Docker work with enterprise clients, we see this pattern constantly. A team adds a new endpoint, copies from an existing one, strips the auth middleware "temporarily" during development, and it ships. The rest of the application is locked down perfectly. One route isn't. That's all it takes.

When we do penetration testing for gaming studios and travel platforms, we specifically look for these inconsistencies. The fully open debug endpoint behind a load balancer. The admin panel that checks authentication but not authorisation. The staging API that accidentally got DNS-pointed to production. These are the bugs that CVSS 9+ vulnerabilities are made of, and they're exactly the kind of logic-level flaws that Mythos-class AI models are going to start finding at scale.

The mobile angle matters too

If you're thinking "this is a server-side problem, my mobile app is fine," think again. In our native iOS and Android projects, we've seen apps that trust the server to handle all security validation. If the backend gets compromised through a flaw like CVE-2026-39987, everything the mobile app sends and receives is exposed. API tokens, user data, session credentials.

We've seen this in healthcare and IoT apps where the mobile client stores sensitive data locally and syncs with a backend that the team assumed was safe because "it's behind a firewall." The Marimo flaw was exploitable over a single unauthenticated WebSocket connection. Firewalls don't help if the door is already open from inside the application layer.

Apple's new App Store requirements for watchOS and iOS 26 SDK compliance are adding deadline pressure too. Teams rushing to meet those April deadlines while simultaneously keeping their backends secure is exactly the kind of operational squeeze where auth checks get skipped.

One thing you can do right now

Here's a concrete step: audit every WebSocket and real-time endpoint in your application for authentication consistency. Not just "does the app require login?" but "does every single endpoint, including debug, terminal, monitoring, and internal ones, actually enforce auth?"

The Marimo vulnerability existed because their terminal WebSocket endpoint used websocket.accept() without calling validate_auth(). Their other endpoints did call it. That gap, one function call missing in one file, was a CVSS 9.3 pre-auth RCE.

If you're on a team that manages web applications or APIs, take an hour this week and grep your codebase for WebSocket accept handlers. Check each one. If you use Starlette, FastAPI, Express with ws, or any framework that treats WebSocket connections as a separate path from HTTP middleware, you probably have endpoints that bypass your normal auth stack. Find them before someone else does.

The bigger picture

Anthropic is framing Project Glasswing as "defenders first." The idea is to give the good guys a head start before Mythos-class capabilities become widely available. That's a reasonable position, but it also rests on an uncomfortable premise: the only way to protect against AI that finds bugs is to build it first and hope you patch faster than attackers can exploit.

From our security work configuring Cloudflare and Akamai for DDoS protection on large-scale travel platforms, we already live in a world where automated attacks move faster than human response times. The difference with AI-assisted vulnerability discovery is that the attacks won't just be faster. They'll be smarter. Instead of brute-force scanning, you'll get targeted exploitation of logic flaws that no WAF rule can catch because the attack looks like a legitimate request.

That changes how you architect defences. Static rules aren't enough. You need layered security: authentication at every endpoint, proper network segmentation so a compromised notebook server can't reach your production database, and real monitoring, not just log aggregation but actual anomaly detection on what your WebSocket connections are doing.

We've talked to our clients about this shift before. This week made it feel a lot less theoretical.

Where we go from here

The security industry is about to get a lot busier. AI models that can find thousands of zero-days in a week will force software maintainers into a permanent sprint. Open-source projects with small teams, like Marimo, will be hit hardest because they don't have the resources to patch at the speed the threat now demands.

If your team is building on open-source data science tools, developer notebooks, or any internal tooling that was designed for trusted networks but now sits anywhere near the internet, this is your wake-up call. Treat every tool in your stack as part of your attack surface. Patch aggressively. And don't assume that because something is "internal," it's safe.

If any of this sounds like a conversation your team needs to have, let's talk.

aiexpert-analysissecuritysoftware-developmenttech-newsvulnerability-management