❌

Vue lecture

The cloud reduced operational complexity. But many teams need someone to own it completely.

Dark abstract digital wave with subtle color distortion, representing application lifecycle and cloud native infrastructure management.

For companies running serious production workloads with lean engineering teams, the real test of operational ownership comes at 3 a.m. Not whether the system stays up but whether anyone needs to be awake to make that happen. The page arrives. A service is degrading. The engineer who answers didn’t build this service, doesn’t know what thresholds were set at deploy time, and cannot tell whether the system is healing itself or waiting for a human decision. 

This is the moment that separates services that transferred operational burdens from platforms that merely deferred them. The system that passes the 3 a.m. test already knows what healthy looks like, what to do when healthy stops being true, and how to communicate what happened—because those decisions were made at deploy time, not incident time. The team sleeps through the night because nothing went wrong, but because the response was already determined.

Forty engineers, eight applications, zero dedicated ops

Cloud infrastructure evolved in two directions simultaneously, and neither arrived where mid-market teams actually stand. On one end: full control. Infrastructure-as-code, service meshes, custom pipelines. Powerful, flexible, and designed for organizations that have deliberately invested in operational staff who can absorb the cost of that flexibility. On the other end: single-app simplicity. Push code, get a URL. Elegant for a first deployment, but architecturally limited the moment a team manages more than one service, needs compliance controls, or inherits an application that doesn’t fit the platform’s opinions.

You know this company. Forty engineers. Eight production applications. Two generate 80% of revenue. One SRE who is actually a senior developer with an on-call rotation nobody else wants. A compliance audit due in Q3 that nobody has started preparing for.

“The right investment is shipping product. The cost of this gap is measured in what these teams do not ship.”

Every engineer is a full-stack contributor. The person who wrote the feature deploys it, monitors it, and gets the page when it breaks not because the team lacks sophistication, but because hiring dedicated infrastructure staff is not the right investment at their stage. The right investment is shipping product. The cost of this gap is measured in what these teams do not ship. Every sprint spent upgrading the deployment pipeline is a sprint without the feature a customer asked for. Every 3 a.m. page answered by a developer with a product standup at 9 a.m. is diminished output that never shows up in any dashboard. The gap is widening.

The applications nobody planned to operate

Not every application in a portfolio was built by the team now responsible for it. Many companies acquire products through M&A. They inherit internal tools built by engineers who left three years ago. They run commercial off-the-shelf applications customized beyond vendor support. They maintain line-of-business applications in languages nobody on the current team chose. These applications share a common trait: they are in production, they serve customers or meet compliance requirements, and nobody has the budget or the mandate to rewrite them. They need a home that accepts them as they are, not as a modernization roadmap says they should become.

“They need a home that accepts them as they are, not as a modernization roadmap says they should become.”

This is where the full lifecycle vision matters. An application management service that only serves net-new applications forces teams to maintain two operational models: one for the applications they are building today, and another for the applications they inherited yesterday. That split is where drift starts, where patching falls behind, and where audit findings accumulate.

The application management service that solves this problem must accept the full range: the Java application packaged as a WAR file, the .NET Framework service running on Windows, the Python application with dependencies pinned to a specific runtime, the containerized service already running elsewhere, and apply the same operational model, the same deployment interface, the same patching and scaling behavior to all of them. Migrate, manage, and modernize within the same experience, without requiring a different operational posture for each stage of the application lifecycle.

What the market data actually shows

Janakiram MSV, an analyst and advisor on cloud-native platforms and TNS contributor, puts it this way:

Cloud native standardized the infrastructure layer around containers and Kubernetes. But it never standardized the operational boundary between the application team and the infrastructure underneath it, and platform engineering largely emerged to re-establish that boundary. CNCF’s latest research with SlashData reports that 28 percent of organizations run a dedicated platform engineering team and 41% split those capabilities across multiple teams. Another 3% have no formal approach at all, and that is where most mid-market engineering organizations exist. They want the same outcome as platform engineering without first becoming a platform engineering organization.

The pattern that repeats is that maturity stalls at the third application rather than the first deployment. A forty-engineer team gets one service into production and keeps it healthy through familiarity. Then an acquisition brings in a .NET workload, an internal tool developed by an engineer who left two years ago, and the team is carrying three operational models with nobody owning the mandate to reconcile them. Hiring will not close that gap, because what is missing is a standardized operational posture, not headcount.

What hundreds of thousands of production deployments actually reveal

We have visibility into hundreds of thousands of production deployments across thousands of customers. That scale does not tell you what people say they want. It tells you what actually breaks, what gets escalated at 3 a.m., and what determines whether a team trusts their platform enough to stop thinking about it. The patterns are remarkably consistent.

1. The deployment that nobody touches again

A team spends two full days getting a Spring Boot application deployed with CI/CD and an SSL certificate. The deploy works. Then nobody touches it for three months because it is stable, but because touching it might break it again. They discover the problem when customers report the application is unreachable. Four hours of forensics follow. What changed? Why? How do we prevent this?

“A release that can partially succeed is a release that will partially fail.”

This is the most common failure mode we observe: not the deployment that fails loudly, but the deployment that succeeds quietly and degrades invisibly. The team cannot explain what happened because the platform did not decide what “healthy” meant before the incident arrived.

A release that can partially succeed is a release that will partially fail. The platforms that earn trust are the ones where a deployment either completes fully or reverses entirely no intermediate states, no manual rollback procedures discovered under pressure.

2. The observability sprint that ships three weeks late

A developer notices response times degrading. They want memory utilization metrics. The platform does not collect them by default. They spend a sprint writing configuration files to install a monitoring agent across every instance. The insight they needed three weeks ago ships three weeks late.

This is the second most common pattern: observability treated as an add-on rather than a default. Every team we have observed that adds monitoring after their first incident wishes they had it before. The teams that never experience this problem are the ones whose platforms shipped metrics, traces, and health signals at deploy time without code changes, without configuration, without a sprint spent on plumbing. The distinction matters: a platform that can be observed is not the same as a platform that is observed from the moment it goes live.

3. The portfolio tax

Each application gets its own infrastructure. Costs scale linearly with the portfolio. The team running eight applications pays eight times the overhead of the team running one, not because each application needs dedicated resources, but because the platform’s architecture assumes isolation rather than shared operational responsibility. The result: teams avoid migrating inherited applications because the cost model punishes breadth. The compliance audit does not care that the inherited application runs on a different operational model. It expects the same governance, patching cadence, and access controls.

A platform that rewards portfolio growth, shared infrastructure, and a consistent operational posture —and economics that improve with breadth rather than degrade—changes the calculus for teams managing applications they did not build.

4. The security configuration nobody made 

The forty-engineer company has no security team. They have a senior developer who reads the CIS benchmarks on weekends. The compliance audit arrives regardless. Teams without specialists will not configure security controls that require specialist configuration. This is not a criticism of those teams; it is a structural observation about how security actually gets implemented (or does not) in organizations where every engineer is a full-stack contributor with a product backlog that never shrinks.

The only security posture that works reliably for these teams is the one they inherit by default: compliance certifications, network isolation, access controls that ship with the platform rather than requiring a dedicated sprint to implement.

What these patterns demand today

Every failure mode we observed the deployment nobody touches, the observability sprint that ships late, the portfolio tax, the security configuration nobody made- shares a root cause: the platform asked the team to make an operational decision, and the team either made the wrong one or made none at all. The response to these patterns required starting from a different question. Not “what should we configure for the team?” but “what should the team never need to decide?”

A platform that answers that question correctly holds a specific set of commitments. It decides what healthy looks like before the first request arrives, not after the first incident. It ships observability at deploy time, not as a sprint the team schedules after something breaks. It treats the eighth application in a portfolio the same as the first, same operational model, same governance, same economics. It inherits security posture by default, because the teams it serves will never staff a dedicated security function.

These are not feature decisions. They are architectural decisions about where operational responsibility permanently resides. The team defines the application source code, a Dockerfile, a pre-built image, or an existing workload being migrated in. From that point forward, the platform owns everything underneath it. Not for the first deploy. For the life of the application. Patching, scaling, healing, certificate rotation, capacity planning, health evaluation. These are not capabilities the team enables. They are responsibilities the platform holds permanently.

This is what we rebuilt AWS Elastic Beanstalk to be. Not a deployment tool. Not a hosting layer. An application management service that takes operational responsibility for everything underneath the application. The architecture now starts from the question above and refuses to let the answer drift back toward the team over time. Elastic Beanstalk operates in two modes, a structural change from its previous single-environment architecture:

Standard Mode delivers full operational ownership for individual applications and Windows/.NET Framework workloads: the complete operational stack, owned outright, for a single service.

Cluster Mode extends the same ownership model across the portfolio, shared infrastructure, source-to-production deployment that transforms code into running applications, and economics that improve as the portfolio grows. The eighth application shares operational overhead with the first seven rather than duplicating it. For the forty-engineer company running eight production applications today and inheriting ten more next quarter, this is the difference between a platform that covers the portfolio and a platform that covers only the applications simple enough to fit its opinions.

The industry convergence

The distinction is real, though I would not draw it as a line between platforms that reduce complexity and platforms that own operations permanently. Every vendor in this market absorbs some operational responsibility at deploy time. The key question is how much of it returns to the team during an incident, a patch cycle, and an audit. A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.

“A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.”

At convergence, the direction is correct, but the shape is incorrect. This is not two camps meeting in the middle. Gartner’s 2026 Magic Quadrant for Cloud Native Application Platforms places AWS, Microsoft, Google, and Red Hat in the leaders quadrant, with Render, Netlify, and Upsun as niche players. Vendors specializing in developer experience showed this category is viable, and now the hyperscalers are adopting it. Since source-code-to-URL mapping is now standard across the entire quadrant, the key differentiation becomes who bears operational liability for the eighth application three years after its release.

The only aspect of framing I would challenge is the idea that a platform determines everything the team never has to decide. Routine infrastructure decisions should stay out of the developer’s path, and escape hatches should stay in place for the teams that genuinely need them. A platform that removes choice altogether will demo well and then stall when teams migrate applications that don’t fit its opinions.

The 3 a.m. test that actually matters

For teams already living this reality — serious production, lean staff, growing portfolios — nobody planned to operate the 3 a.m. test; it is not a nice-to-have. It is the evaluation criterion.

The platforms that define the next decade will not simply make deployment easier. They will decide, in advance, how production systems should behave when things inevitably go wrong. Because by 3 a.m., the time for deciding has already passed. The CNAP category was built to describe platforms that own the application lifecycle.

Elastic Beanstalk made those decisions before the incident arrived: what healthy looks like, what to do when it stops being true, how to communicate what happened. The cloud gave teams power. These teams needed someone to stay. Elastic Beanstalk stays.

Check the latest AWS release notes for Elastic Beanstalk here.

The post The cloud reduced operational complexity. But many teams need someone to own it completely. appeared first on The New Stack.

  •  

AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex

Multiple monitors on a desk.

Amazon Web Services (AWS) is lifting the lid on a new open source, general-purpose AI agent, designed to give developers a ready-made foundation they can run locally or deploy to the cloud.

Strands Harness, as it’s called, builds on Strands Agents, which AWS debuted in May 2025 as an open source Python SDK for building AI agents. Strands takes what AWS calls a “model-driven approach”: developers provide the model, tools and instructions, while the model determines how to tackle the task and when to call those tools. AWS later brought Strands to TypeScript, and in February created Strands Labs as a separate home for more experimental projects.

Marc Brooker, VP and distinguished engineer at AWS, tells The New Stack that Strands Harness essentially sits above the existing Strands SDK, giving developers a preconfigured Strands Agent. It brings together the tools and supporting machinery an agent needs to operate over longer-running tasks, with AWS supplying its own defaults for how those pieces work together.

“You still need to decide how to manage context, persist conversations, integrate tools, and guide the agent’s behavior.”

“An SDK like the Strands Harness SDK gives you the building blocks, but you still need to decide how to manage context, persist conversations, integrate tools, and guide the agent’s behavior,” Brooker explains.

Unpacking Strands Harness

Out of the box, Strands Harness gives developers a working agent with file, shell and web tools, alongside built-in handling for context, memory, persistent sessions, prompt caching and delegation to other agents.

Developers can install Strands Harness as a Python or TypeScript package, using pip install strands-harness or npm install @strands-agents/harness.

AWS also provides the Strands CLI as an interactive way to prototype and configure an agent. Developers can choose a model, add prompts, tools and other capabilities, then use /export to generate the resulting agent as Python or TypeScript code.

Configuring an agent with the Strands CLI.
Configuring an agent with the Strands CLI.

Individual agents can be tailored to different jobs, with developers able to change their instructions, choose which model they use, control which tools and capabilities are available to them, and decide whether they can hand work off to another agent.

Demo of Strands Harness running on a desktop
Demo of Strands Harness running on a desktop (Credit: AWS)

Most of Strands Harness itself doesn’t depend on AWS infrastructure. The agent loop, tools, context management, session handling and delegation are all included in the open source release, and AWS says those processes run on the machine that’s running the agent by default.

The exception is the call to the underlying model. Perhaps unsurprisingly, AWS routes model access through Amazon Bedrock, its managed service for accessing and running foundation models. However, while Brooker says that this is the only out-of-the-box default tied specifically to AWS infrastructure, it too can be switched out.

“This is easily overrided to use a different model provider with one line,” he says.

Strands Harness can instead use Anthropic, OpenAI or Google as its model provider, or use a locally running model through Ollama. Changing provider doesn’t necessarily mean changing the underlying model, but Brooker notes that choosing a different model will obviously affect how the agent behaves.

“Different models have different strengths on reasoning, tool use, and cost,” he continues. “What doesn’t change: context management, sessions, tools, delegation all work the same regardless of provider. No features require Bedrock.”

It’s worth noting that all the other defaults can be changed, too. Developers can bring their own tools and skills, connect MCP servers, alter how context is handled, and choose where session state is stored.

“Developers can focus on their application’s task and domain expertise, while customizing the components that need different behavior.”

“Developers can focus on their application’s task and domain expertise, while customizing the components that need different behavior,” Brooker says.

AWS benchmarks its agent

AWS says Strands Harness is intended as a general-purpose agent rather than a coding assistant, though it takes cues from harnesses such as Claude Code and Codex. The difference, AWS says, is that developers can deploy Strands Harness to whichever cloud provider they choose — addressing what it describes as a common wish among developers using Claude Code and Codex to be able to run the same setup in the cloud.

From its own testing, AWS suggests the way a harness manages the surrounding agent machinery can materially affect cost and performance, even when the underlying model stays the same. For each harness, AWS averaged its score across six benchmarks — ALFWorld, ContextBench, GAIA, WebShop, τ³-bench and Terminal-Bench 2.1– and compared that with the average cost per task across the same tests. Against Claude Code and Codex specifically, the company says Strands Harness came out 45% cheaper, with broadly comparable accuracy.

However, that figure drops to 28% once DeepSeek Harness — which AWS says ran around 14% cheaper than Strands Harness on matched runs — is folded into the wider comparison.

Strands Harness benchmark results.
Strands Harness benchmark results. (Credit: AWS)

AWS points specifically to its context-management defaults as a major reason for the result. Strands Harness truncates particularly large tool outputs, compacts context once the available window passes a set threshold, and attempts to recover within the agent loop if the context overflows.

On Terminal Bench 2.1 specifically, AWS says Strands Harness running Fable 5 cost 77% less than Claude Code, at $56.29 versus $248.05 across 89 trials, while scoring 69.7 versus 61.8. DeepSeek Harness was cheaper again at $40.30, though its score was lower at 59.5.

Terminal Bench 2.1 results.
Terminal Bench 2.1 results. (Credit: AWS)

For AWS, those results help make the case for packaging and tuning functions such as context management, versus requiring every developer to work out those decisions independently with the SDK.

“Getting a prototype working is one step; evaluating how those choices affect performance and cost is another.”

“Getting a prototype working is one step; evaluating how those choices affect performance and cost is another,” Brooker says. “The opportunity we saw was to package that engineering into a complete, general-purpose agent.”

What’s in it for AWS?

AWS also has an obvious place to run the resulting agent. Amazon Bedrock AgentCore is its managed service for deploying and operating agents, providing identity and access controls, observability and the infrastructure needed to host them.

There is, in fact, a close technical relationship between the open source project and that managed offering. AgentCore Harness and Strands Harness were built by the same team, although they live in separate codebases. Brooker says work on one can also feed improvements into the other, giving AWS a route for technology developed in the open source project to inform its managed service, and vice versa.

Brooker, again, stresses that Strands Harness can be deployed independently of AgentCore, outside of AWS altogether.

“AgentCore is an optional hosting layer for teams that want AWS to manage the infrastructure side,” he says. “However, all deployment paths are open for the developer to choose.”

Still, this arrangement gives AWS a clear commercial path: developers can adopt Strands Harness freely, while AgentCore gives the company a natural destination for teams that eventually want AWS to run the infrastructure around it.

The post AWS open-sources an AI agent it says is 45% cheaper than Claude Code and Codex appeared first on The New Stack.

  •  

[Launched] Generally Available: Azure Functions support for PowerShell 7.6

Azure Functions support for PowerShell 7.6 is now generally available. You can now develop apps using PowerShell 7.6 locally and deploy them to Azure Functions plans. Learn more: Updating your app to PowerShell 7.6 What's new in PowerShell 7.6? Azure
  •  

“Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules

Vercel announced this week that teams on its free Hobby plan will now have older, unprotected deployments deleted immediately if they exceed the 10GB Deployment Storage limit. 

When asked why Vercel decided to change its retention rules, Jas Garcha, head of pricing at Vercel, tells The New Stack the update is a move to keep the free tier viable amid rapidly growing deployment volumes: 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

Old deployments now deleted immediately

Before, users in the free tier could count on eligible deployments to stick around for up to 30 days. But that window is gone for teams over the limit, and the protections that used to spare older deployments have narrowed for every Hobby project.

Now, if users exceed the standard 10GB of Deployment Storage included on the free tier, old deployments not covered by Vercel’s retention-policy exceptions will be deleted immediately, per Vercel’s updated Deployment Retention Policy. 

“This change allows us to continue supporting a Hobby community that’s deploying at a much higher rate than it was a year ago.”

What gets to stay? 

Vercel says each Hobby project will keep the three most recent production deployments, along with its three most recent deployments of any type, regardless of age. That’s a cut from the previous Hobby exception, which preserved the 10 most recent production deployments, and it applies to every Hobby project — not only teams over the 10GB limit.

Plus, preview deployments lose a separate protection

In addition to cutting the 30-day holding period for over-limit teams, Vercel’s policy update also removes a separate retention exception for preview deployments. 

Preview deployments have not vanished from Vercel’s exception list entirely — the latest preview deployment on an active Git branch is still protected on every plan. What Hobby lost is the count-based exception: Pro and Enterprise teams keep their last 20 non-production deployments in a Ready state, and that protection no longer applies to Hobby.

Per Garcha, “Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

Why the change?

Garcha tells The New Stack that Vercel’s latest policy update is needed to keep the Hobby tier sustainable as deployment volumes dramatically rise:  

“Our former retention defaults were designed for teams that ship constantly and need deep rollback history. They made less sense for Hobby projects, where dormant deployments were quietly consuming storage that active projects need.” 

“Your current production deployment is never deleted, and aliased and active-branch deployments remain protected, along with each project’s most recent deployments.” 

And activity is much higher, even compared to a year ago. According to Garcha, Vercel now handles more than 10 million deployments every day— more than a 6x increase YoY. Immediately deleting older, unprotected deployments, Garcha says, is one way to free up storage for active projects as Vercel handles much higher deployment rates.

In other words, Vercel is moving out some of the old to make room for the new. It describes how the storage limit works in its post:

“Every deployment you keep uses some of it [Deployment Storage], and going over the limit can block you from deploying until you free some up.” 

What Hobby users should do

It’s important to note that deletion isn’t immediately permanent. Vercel gives successfully built deployments a 30-day recovery period, and users can restore them from a project’s Settings, under Security → Recently Deleted. Hobby users who want to stop old, unprotected deployments from being deleted in the first place can move off the free tier and onto the Pro plan. In the Pro tier, storage beyond the plan’s included allowance is billed at $0.10 per GB-month, and the retention exceptions stay far more generous: the last 10 deployments created in a project, the last 20 production deployments in a Ready state, and the last 20 non-production ones.

For users who can’t or don’t want to upgrade to Pro, Vercel offers guidance for optimizing Deployment Storage usage to help users stay under the free 10GB storage limit, like reducing unnecessary deployment output.

Still, Garcha says few Hobby users will feel the effects enough to warrant making a change. Pointing to Vercel’s list of exceptions that still protect certain deployments, he tells The New Stack, “The vast majority of Hobby users won’t notice the change.” 

Garcha also says Vercel’s stricter retention policy helps the company keep offering Hobby as a permanent free plan.

“We’re one of the few platforms where the free tier isn’t a trial or a credit that expires. It’s a permanent plan, and we’ve kept expanding it,” he says. “By ensuring its resources go to people actively building, we’re able to continue offering it.”

The post “Dormant deployments were quietly consuming storage”: Why Vercel tightened its free-tier rules appeared first on The New Stack.

  •  

[Launched] Generally Available: High-scale mesh in Azure Virtual Network Manager

High-scale mesh using connected group in Azure Virtual Network Manager is now in general availability. In available regions, customers may connect up to 3,000 virtual networks in a single mesh connectivity configuration by default and higher scale IP conn
  •  

Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between?

There were unusual goings-on this month at Automattic, a company known for its free and open-source app for building WordPress sites. 

The company issued a notice last Thursday confirming that CEO Matt Mullenweg was on leave.

Mullenweg, who also co-founded WordPress in 2003 before establishing Automattic in 2005, was temporarily replaced by company CFO Mark Davies before being reinstated less than a day and a half later.

By last Saturday, a new alert emerged stating that Mullenweg was back in his position and that “Matt was away for only 33 hours and 20 minutes.”

Automattic has not publicly explained what changed between the initial decision and Mullenweg’s return — and it has also declined to explain the circumstances behind the leave and return. 

By way of context, WordPress sits under the Automattic brand alongside the company’s other products, including the microblogging site Tumblr, the e-commerce WordPress plug-in service WooCommerce, and the instant messaging client Beeper. 

A 33-hour vanishing act, but Automatticians are supporting him 

Automattic director of communications Megan Fox is on the record saying, “Matt Mullenweg is the chairman and CEO of Automattic, with full support of the board. “And if you search online, you can see many top executives and Automatticians supporting him as well.”

A further report reproduced Slack messages written by Mullenweg where he said, “Happy to announce the board is back in agreement, and I’m in control of Automattic. A lot happened in the past 48 hours that we need to sort out, and I hope much of it was a misunderstanding, because I have huge respect and regard for those involved.”

Was this a failed boardroom coup?

Industry watchers may naturally suspect the knives were out and that this was a failed boardroom coup. 

CEO & CTO at HasData, Roman Milyushkevich, tells The New Stack that a company can survive a CEO departure, but it struggles when employees cannot tell which governance process is real.

“But, in terms of whether the real guns were out at Automattic, it certainly looks like a failed attempt to change control, but I would not call it a coup as an established fact,” Milyushkevich says. 

“The possibilities playing out here are all very different,” Milyushkevich adds. “There could have been a second board agreement brought into place inside that 33 hours; there could have been internal (or possibly even external) negotiations; directors could have reconsidered the practical consequences of removing the founder; or there could have been an internal resolution that has not been disclosed.”

He advises that the “most important thing Automattic can establish now” is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.

“The most important thing Automattic can establish now is not who won the dispute, but whether the board and CEO have a clearly understood process for handling the next serious disagreement.”

Behavioral scientist and visiting professor at São Paulo’s FIA Business School, Ricardo D’Olivar, tells The New Stack that what matters here is whether stakeholders have “enough information to distinguish a considered correction from an unresolved struggle” over authority. 

“A reversal of this kind can reflect responsible reconsideration,” D’Olivar says. “Reinstatement answers who is in charge today. It does not, by itself, explain how a future disagreement would be resolved. This is where the potential consequences for employees, executive recruitment and investors arise. If uncertainty persists, employees may become more cautious about committing to decisions whose backing appears unstable.”

Suggesting that, responsible corporate mechanics or not, this kind of development undoubtedly throws the cat among the pigeons, D’Olivar says that developers or executives considering a career at Automattic may now question whether they would be held responsible for decisions they were authorized to make, but could no longer count on the organization to support when challenged. 

“Investors may seek clearer evidence that oversight and succession arrangements can operate under pressure. These are possible responses, not verified effects at Automattic. The information shared internally may also be more complete than the public account,” adds D’Olivar.

This is not the first boomerang CEO bounce

Mullenweg might be the fastest CEO yo-yo switcharound in history, but he’s certainly not the first. OpenAI CEO Sam Altman famously left the company’s board after a communication dispute. An employee uprising (nearly all the company’s 700+ staff threatened to resign) and added pressure from Microsoft led to Altman returning to his position five days later.

Perhaps even more famously, Steve Jobs was ousted in 1985, only to return 12 years later to realign a then-struggling Apple and take it into its golden years. Twitter (now X) founder Jack Dorsey was moved out in 2008 before a boomerang return in 2015. Michael Dell stepped down as CEO to become chairman of the board in 2004, but by the start of 2007 he was back.

At their origins, the shenanigans at Automattic may be redolent of the technology industry’s other boomerang CEO realignments, or this may be a boardroom tussle that we’ll never know the full reason for until Mullenweg writes his memoirs. Either way, the WordPress industry just got its first movie-script idea.

The post Automattic says CEO Mullenweg was gone and back inside 33 hours. What happened between? appeared first on The New Stack.

  •  

AWS agents will suggest your new flights. Code decides what gets booked.

Close-up of an airport departure board showing flight numbers, gates, departure times, and “On Time” and “Boarding” statuses.

AWS published a new Step Functions pattern this week that gives AI agents a role in airline rebooking while keeping reservation changes and payments under code’s control. Amazon Bedrock AgentCore agents suggest new itineraries and draft compensation messages after a flight disruption. Deterministic steps in the workflow then validate those proposals before changing any reservation or issuing payment.

As AWS puts it: “The principle is that agents propose, and deterministic code validates.”

Also yesterday, AWS threw more weight behind its case for supporting model reasoning with code execution in a new Abnormal AI case study, arguing that agents need a compute environment where they can perform calculations, process data, and programmatically verify work before returning results. 

A Step Functions pattern to keep agents away from the money

Airline rebooking is a good candidate for agentic workflows because agents can help operations teams offload the tedious process of finding route alternatives, comparing constraints, and coordinating next steps — just think of the manual clicking it takes to rebook hundreds of passengers to new itineraries after a flight cancellation.

Such is the scenario put forth by AWS in its new Step Functions pattern. 

“The principle is that agents propose, and deterministic code validates.”

By orchestrating specialized Amazon Bedrock AgentCore agents with AWS Step Functions, AWS claims the pattern enables developers to get “the reasoning power of generative AI with the guardrails of deterministic validation.” 

Rather than putting orchestration, fan-out, validation, routing, and retries inside an agent’s reasoning, the pattern implements this work in Step Functions, where deterministic steps wrap each agent’s non-deterministic behavior. This way, no agent can take a direct action, like writing a reservation or issuing a payment. Instead, each agent proposal is only applied after deterministic validation passes, with Step Functions maintaining an execution history for audit and review. 

Compared to multi-agent collaboration, where a supervisor agent orchestrates sub-agent runs and tool calls, AWS’s pattern pushes those decisions out of the agent layer and into the Step Functions workflow. Per AWS, this separation could help developers use AI agents more safely — that is, using agent reasoning to generate proposals but restricting any action until deterministic code gives the go-ahead. 

“The reasoning power of generative AI with the guardrails of deterministic validation.” 

While AWS uses airline rebooking as its example, the same pattern could apply to other high-stakes financial and regulatory workflows where deterministic code should stand between agent proposals and actions. 

A compute scratch pad to bring in computation when semantic reasoning isn’t enough 

On the same day it published the Step Functions pattern to validate multi-agent decisions, AWS made a related case for supporting agentic reasoning with code execution in a new case study of Abnormal AI, a behavioral AI security platform.

Per AWS, Abnormal AI uses Amazon Bedrock AgentCore Code Interpreter, a capability of Amazon Bedrock AgentCore that provides a fully managed, serverless runtime for agents to execute code dynamically, to support its real-time inline email threat detection. In this study, AWS argues that Code Interpreter “is not merely a coding tool. It’s fundamental infrastructure that agents use to reason computationally.”

Specifically, it describes pairing the managed, secured sandbox with a large language model (LLM) to combine two different strengths: the reasoning and semantic coherence of an LLM with the calculation, data processing, and verification that come from executing code. For real-world operations, like converting data into structured reports or counting, that don’t map cleanly to semantic reasoning, a compute scratchpad lets agents work through tasks computationally and verify answers instead of relying on reasoning alone. 

Why semantic reasoning alone isn’t enough 

AWS isn’t the only one working to separate model reasoning from downstream actions. 

Last month, Perplexity shipped Portable Computer, the local-first version of its Computer agent running on an Nvidia DGX Spark workstation that puts deterministic software in charge of model actions. Rather than stacking reasoning and execution in the same layer, Perplexity separates them, using probabilistic reasoning to propose what should happen next and deterministic software to decide whether or not to execute it. 

As agents take on more consequential work, like tasks across accounts payable, procurement, and the monthly close, semantic reasoning alone can’t guarantee that agent proposals are safe enough to act on. But a layer of deterministic validation may at least add a verifiable check between proposal and execution.

The post AWS agents will suggest your new flights. Code decides what gets booked. appeared first on The New Stack.

  •  

How to attach an owner to every cloud resource you find

Dark abstract 3D rendering of intertwined rings symbolizing complex cloud resource governance and technical debt.

The engineer who knew why that cloud instance existed has left the company. The instance is still running, the bill keeps growing, and the team must now decide whether it’s safe to shut down. This is a bad time to discover that its ownership history was someone’s memory.

“Good resource governance has three pillars: continuously synced inventory, policy that blocks resources without tagged owners, and an audit trail that survives every reorg.”

A cost review flags an EC2 instance nobody remembers provisioning. Someone scours Slack for the resource ID, finds nothing, burns half a day chasing dead ends, and eventually stumbles across an exhausted engineer who mumbles the mantra that ends most of these investigations: “I think that’s from the project Priya was running before she left.” 

Nobody follows up, because nobody knows how to reach Priya anymore. The instance stays up, because tearing down a fence when you don’t know what it’s protecting you from is a good way to find out the hard way. It’s the platform team equivalent of emotional baggage; they all seem to accumulate some.

Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous. 

“Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous.”

Solving the problem requires three things that already exist: queries, policies, and logs; and maybe just a little bit of therapy.

1. The query that tells you what’s missing an owner

CloudQuery’s asset inventory syncs continuously across every provider a team runs on, into tables you can query directly: aws_ec2_instances, gcp_compute_instances, azure_compute_virtual_machines, and so on. Finding every resource without an assigned owner is as easy as:

SQL
SELECT resource_id, 'aws' AS provider, 'ec2_instance' AS resource_type
FROM aws_ec2_instances
WHERE tags ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'gcp', 'compute_instance'
FROM gcp_compute_instances
WHERE labels ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'azure', 'virtual_machine'
FROM azure_compute_virtual_machines
WHERE tags ->> 'owner' IS NULL
ORDER BY provider;

Run it regularly, and you have a (hopefully short) boring list to refer to when the CFO asks who deployed an expensive instance; instead of being thrown into a frenzied goose chase at 4:59 p.m. on a Friday.

2. The policy that prevents it recurring

A query tells you what’s already missing an owner; but how do you prevent the next ownerless resource from being deployed? The answer is a policy, and env zero evaluates Open Policy Agent rules against every plan before it applies. A rule that requires an owner tag on every new resource looks something like this:

package env0

# METADATA
# title: require owner tag
# description: A resource can't be created without a declared owner.
deny[format(rego.metadata.rule())] {
	resource := input.resource_changes[_]
	resource.change.actions[_] == "create"
	not resource.change.after.tags.owner
}

format(meta) := meta.description

Add that to the project’s policy set, and a plan that creates a resource without an owner tag doesn’t just get a warning; it doesn’t get created.

3. The record that outlives its creator

A tag tells you who owns something today. It doesn’t tell you anything about who asked for it, why, or who signed off. By the time it matters, the person who could’ve answered from memory may no longer be reachable. An audit entry records that at the moment of creation, instead of reconstructing it afterward from the scraps of recollection scattered around the rest of the team. Here’s an example:

{
  "event": "resource.created",
  "resource_id": "i-0a1b2c3d4e5f",
  "requested_by": "j.chen@company.com",
  "approved_by": "platform-lead@company.com",
  "approval_ref": "ENV-4471",
  "stated_purpose": "load test environment, Q3 capacity planning",
  "timestamp": "2026-08-14T09:12:03Z"
}

This entry answers the question this whole piece opened with, without needing Priya or Slack. env zero keeps this record attached to the resource for as long as the resource exists, specifically so it outlasts any individual’s tenure.

Why this keeps happening

Employee attrition is an age-old challenge that’s only accelerating in the modern era. US private-sector voluntary turnover runs 22 to 25% a year, so a hundred-person org loses twenty-odd people every year, each one taking a small, specific piece of “why this exists” with them. Replacing a mid-level employee costs six to nine months of salary, more than double that for senior specialists. 

“Tags were supposed to survive this. In practice they rot the way everything else does.”

That figure doesn’t touch what the departure does to everyone else’s mental model of what’s actually running. Tags were supposed to survive this. In practice they rot the way everything else does: two teams merge and bring incompatible schemas, provisioning that runs on tribal knowledge accumulates configuration drift for the same reason it accumulates ambiguous ownership, and a resource tagged under a policy that’s been replaced twice since its inception isn’t much better documented than one that isn’t tagged at all.

We wrote in April about the hour it once took our own team to answer, “what are we actually running across both clouds?” That solved a point-in-time problem. The query, the policy, and the audit entry above stop the same story from playing out again next June.

The org chart will keep changing. The record doesn’t have to.

None of this stops people from leaving or teams from reorganizing; pretending otherwise is how platform teams end up rebuilding the same spreadsheet every eighteen months. What changes is whether the next “what are we actually running, and who owns it” conversation takes hours of archaeology across three teams, or is a query that already has the answer attached. Institutional knowledge decays at a fairly predictable rate. A system of record shouldn’t.

The post How to attach an owner to every cloud resource you find appeared first on The New Stack.

  •  

OpenAI’s researchers burned $7,000 a day on AI agents — now it’s opening the floodgates

speed abstract

OpenAI rolled out its Agents API in public beta Thursday, opening the backend behind Codex to developers looking to run agents unattended for days.

Now, developers don’t have to build their own system to keep an agent going because the API tracks the job as it progresses and gives the agent somewhere to execute its work, even when a task stretches well beyond a single context window.

That makes long-running agents easier to try, but it also gives developers more ways to burn through compute. Interestingly enough, on the same day Agents API launched, OpenAI paused new sign-ups for its $200-a-month Pro plan after demand for GPT-6 Astra strained capacity.

Thibault Sottiaux, engineering lead for Codex, writes on X that Pro subscriptions “put the most strain on our systems,” adding that OpenAI was working to add capacity “as fast as we can.”

To make sure our current users have an incredible experience and continued access to Astra, we are going to pause subscriptions to our $200 Pro plan. These put the most strain on our systems and we wanted to take the smallest step that allows us to continue giving the broadest… https://t.co/WhLEm3HBL7

— Tibo (@thsottiaux) September 10, 2026

The Agents API and ChatGPT Pro are separate products, so there’s no reason to assume one is taking capacity from the other. Still, the timing stands out: the company is making it easier for developers to run agents for hours or days while pulling back access to its heaviest-use consumer plan and working to add more capacity.

Agent inference adds up fast

As a task gets longer, the API can compress earlier context, so the agent doesn’t just stop when it reaches the model’s context limit. It can also bring in tools only when they’re needed or send parts of a larger job to subagents working in parallel. The actual work can run in OpenAI’s sandbox or on infrastructure the developer controls.

The actual work can run in OpenAI’s sandbox or on infrastructure the developer controls.

As agents make progress, they go back to the model for the next step, and a job that takes hours can rack up far more inference than a typical API call. The usage climbs even faster when agents work in parallel.

OpenAI has already seen this inside its own shop. In a research report published September 6, OpenAI said its research organization was logging 3.1 agent-workdays for every human workday by mid-August, measured in standard eight-hour equivalents. The median researcher, ranked by agent usage, was spending more than $600 per day on inference at API prices, while the 90th percentile exceeded $7,000.

Before June, OpenAI’s researchers were still putting in more hours than their agent, but by mid-August, the agents were doing three times as much work.

Arguably, OpenAI’s researchers are an extreme case, but the numbers show what happens when agent use starts to scale. One person can suddenly generate far more inference than their headcount would suggest.

One person can suddenly generate far more inference than their headcount would suggest.

Friction limited compute demand

The Agents API lowers the cost of that experimentation by leaving the orchestration layer out of the bill. Developers pay for the models, tools, and hosted compute their agents actually use.

The flip side is that it’s now easier to consume more inference. Context compaction is a good example. A full context window used to force developers to decide what to discard or how to summarize the work so far. Now the API handles that automatically and the agent keeps going. That’s useful for developers, but it also means the workload doesn’t stop when the context window fills up.

Astra demand hit the ceiling

The Astra rollout offers a preview of what that could look like. OpenAI stopped accepting new Pro subscribers less than two weeks after the model launched on September 3, saying those accounts put the most strain on its systems. The Agents API has its own rate limits and usage tiers, so the Pro pause doesn’t directly affect developers using it. Still, the company is already having to manage capacity around its newest model.

Infrastructure outweighs benchmarks now

The more agents developers run, and the longer they run them, the faster that usage adds up. One developer might have several agents working at once, each going back to the model throughout the task. So headcount alone doesn’t tell you much about how much compute you’re using.

For long-running agents, the challenge is keeping the work moving without wasting tokens or losing track of the task. Cloudflare made a similar bet this summer, arguing that the infrastructure around AI workloads would eventually matter as much as the models themselves.

For long-running agents, the challenge is keeping the work moving without wasting tokens or losing track of the task.

The post OpenAI’s researchers burned $7,000 a day on AI agents — now it’s opening the floodgates appeared first on The New Stack.

  •  

The critical vulnerability was a test database. That’s the whole triage problem.

A security researcher testing a 300-person B2B company with a global footprint discovered an internet-exposed database with weak authentication during a routine scan. The database appeared to be a prime target for attackers; it had critical severity and was an obvious first-fix candidate. Upon further inspection, however, researchers learned it was a resettable test database used to test job candidates, rather than a system containing client data.

The episode represents a critical problem: Scanners and security researchers can’t infer the real cost of a compromise on their own. Increasingly, companies rely on an external partner to handle cloud vulnerability triage and keep teams from becoming overwhelmed. Scanners produce excessive noise, requiring focus and attention from engineers to understand alerts and complex attack techniques, tune out false positives, and work with teams on remediation, reducing the time they have to build new features, work with customers, or scale up their tech environment.

Tech employees have more on their to-do lists than ever: Product teams face an infinite stream of feature requests and engineering teams carry more technical debt than they can clear. When you factor in a reorg, many employees may have larger project scopes and leaner teams, all while constantly monitoring, correcting, and mentoring AI agents. Some reports cite 90-hour work weeks.

Beyond today’s structural challenges, security teams are drowning in a different type of data. Programs absorb identity events, firewall logs, endpoint alerts, vendor feeds, and threat intelligence, producing far more analysis data than even a few years ago. Every item can look urgent, high-risk, and worthy of immediate attention. With finite hours, it’s hard to know where to act first.

Jon Rose, founder of the information security and risk management advisory firm IOmergent,  has seen firsthand how AI and near-universal tooling are producing more findings than teams can triage. 

“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions… But you have to be ruthless about prioritizing and investing your time.”

“Within the span of security work, there’s an unending list of things you could tackle, and you’re pulled in so many different directions,” Rose tells The New Stack. “But you have to be ruthless about prioritizing and investing your time.”

The challenge, then, isn’t remediation. It’s allocation: Deciding where limited engineering and security capacity will make the greatest impact is the most important question security leaders answer every day. A technical severity score, used without threat and environmental context, cannot answer the business question: What should we fix first? Effective vulnerability prioritization turns raw findings into business-aware priorities, and that judgment is the real work.

CVSS limitations: a starting point, not a decision

For security teams, a Common Vulnerability Scoring System (CVSS) provides a useful baseline. Its base metrics classify a vulnerability’s severity — attack vector, complexity, required privileges, and potential impact on confidentiality, integrity, and availability. But a severity score is designed to be stable across environments, so it can’t tell a team whether an asset is exposed to the internet, shielded by compensating controls, or central to the business.

Essentially, treating a base severity score as an automatic fix-first ticket is problematic, rather than using CVSS on its own.

CVSS can include Threat and Environmental metrics that account for evolving exploit conditions and organization-specific context. However, risk-based vulnerability management still depends on accurate knowledge of the environment — and on someone applying that context consistently. 

“The piece that’s missing from any of these tools is the grounding in the business, the understanding of what actually matters,” Rose says.

Escalate the internet-facing medium

A smart decision framework deprioritizes the urgency of a score and probes reachability: whether an attacker could actually reach the vulnerable component. 

“Is the affected service exposed to the public internet or not? Or is it isolated behind network controls and accessible only to a limited set of internal users?” Rose says. “Often, issues will get flagged, but it’s not in a position where it could be triggered.”

Next is consequence. An exposed flaw on a disposable test system might be a genuine security concern. Still, it doesn’t carry the same weight as a weakness on the application that processes customer transactions, stores sensitive information, or underpins a company’s main revenue stream. So while the severity might be alarming, the business outcome might not.

Teams should also ask whether the weakness is attracting active attacker interest and where it could lead. A vulnerability listed in CISA’s Known Exploited Vulnerabilities catalog deserves urgent scrutiny, because there’s evidence of exploitation in the wild.

The Exploit Prediction Scoring System provides a forward-looking signal: An estimate of the likelihood that exploitation of a particular vulnerability will be observed over the next 30 days. Neither replaces a business decision, but both help distinguish a theoretical risk from one that demands attention now. However, as AI-driven exploitation accelerates, the window between what is known to be vulnerable and what is actively exploited is shrinking because the cost of building and weaponizing exploits is dropping.

The relevant attack path might also extend beyond the affected machine. A comparatively modest flaw can become an immediate concern if it provides a route into privileged accounts, a production system, or customer data. Conversely, a high-severity finding can be deprioritized — temporarily — if it is not reachable, has limited impact, and is protected by reliable controls.

“The speed and the depth of research and investigation into those security issues are going faster,” Rose says. “So it can change really quickly.” The worst state is unacknowledged risk sitting in the backlog. Even the best detection degrades when no one owns the trend line. Teams should track exceptions, set review dates, and name an owner. Accepted risk is still risk; the difference is that it’s explicit, time-bound, and revisited. 

An operator capability, not a weekend project

Security teams are leaning more on AI to prioritize and triage threats, and they’re uncovering vulnerabilities at an unprecedented pace. Yet the sheer volume of findings is overwhelming, even for the most efficient of humans.

It’s clear, too, that AI can make it easier to discover less obvious paths to exploitation, and to introduce new ones. An academic study of 20,000+ issues fixed by AI found that LLMs introduce nearly 9x as many new vulnerabilities as developers, exhibiting unique patterns not found in developers’ code. The answer is to use AI at the outcome level. For every alert, the goal is to have the SOC analyst’s standard questions answered quickly: new or known, what’s exposed, what data is at risk, prod or dev, and how it has evolved.

“Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”

“That’s how teams get thousands of alerts down to 10 to 20 prioritized tickets,” Rose tells The New Stack. “Effective programs start by aligning with executive teams to understand the business — where the company is going — so allocation and adjustments track the actual risk, not just the score.”

As to-do lists grow ever longer, business-context judgment that’s applied every day by someone who owns it is something worth dedicating more resources.

If you think you’d benefit from managed cloud security, book a 30-minute cloud security scoping call with IOmergent.

The post The critical vulnerability was a test database. That’s the whole triage problem. appeared first on The New Stack.

  •  

[Launched] Generally Available: Azure Virtual Network Manager IPAM in additional Azure regions

Azure Virtual Network Manager IP address management is now generally available in additional regions: US Gov Virginia, US Gov Texas and US Gov Arizona, and China North 3 and China East 3. In complex network environments, managing IP addresses effectively
  •  

MCP was supposed to solve the agent tooling problem. It missed a step.

Abstract nodes

Connecting an AI agent to a tool is relatively straightforward. Things get more complicated once an organization has hundreds or thousands of resources spread across different clouds and platforms. Agentic Resource Discovery, or ARD, is designed to help agents navigate all of that.

AWS highlighted the open specification in its August 31 Weekly Roundup after taking a deeper technical look at it a week earlier, describing the idea as “DNS, but for agents.” Instead of telling an agent where to find everything in advance, ARD lets it search across different registries for what it needs.

And despite AWS highlighting the project, ARD isn’t an AWS technology. It was authored by Junjie Bu of Google, R.V. Guha of Microsoft, and Shaun Smith of Hugging Face, and released under the Apache 2.0 license. Engineers from several other companies have helped shape the project, including Cisco, Databricks, GitHub, GoDaddy, Nvidia, Salesforce, ServiceNow, and Snowflake.

AWS’s role, at least so far, has been to provide feedback on the specification and explore how it could work with its own Agent Registry. The goal is to make the existing registries work together.

Instead of telling an agent where to find everything ahead of time, ARD lets it search across different registries for what it needs.

MCP skips the discovery step

The Model Context Protocol has become a common way for AI applications to connect to external tools and data, but it assumes the client already knows which server it wants to use. That becomes a problem as companies spread their infrastructure across clouds, SaaS platforms, and internal systems.

ARD helps an agent find a resource before it tries to use it. The specification uses the term “agentic resource” to refer to anything an AI client can connect to, from an MCP server to other external capabilities. An ARD-compatible service keeps track of what’s available, rather than requiring developers to set up every connection in advance.

The Model Context Protocol has become a common way for AI applications to connect to external tools and data, but it assumes the client already knows which server it wants to use.

Federation without forced migration

Companies can keep their own catalogs and policies while routing searches to other ARD-compatible services. An enterprise, for example, could keep internal resources private while searching approved external catalogs when needed. AWS calls this “describe once, discover everywhere.”

The current v0.91 proposal, dated August 26, uses JSON-LD and a REST interface. Its required POST /search endpoint searches by task, while optional endpoints allow clients to browse available resources.

Each discovery service can set its own rules for what it returns and which sources it trusts. This is also where AWS’s DNS comparison falls short. A domain name points to a specific location, while an ARD search could turn up several options that all appear capable of doing the job.

Route 53 engineers shaped ARD

The DNS comparison has some history behind it. Two of the three authors of AWS’s August 24 ARD post work closely with Route 53. Principal software engineer Jeffrey Damick focuses on DNS and networking technologies. At the same time, Bhargav Talluri leads product management for Route 53 and for agent identity and discovery in AWS Agent Registry.The

Agent Registry already provides AWS customers with a central view of their resources. Adding ARD could bring resources running elsewhere into that view without requiring companies to register everything with AWS.

Adding ARD could bring resources running elsewhere into that view without requiring companies to register everything with AWS.

ARD’s governance is still being worked out, with board terms and membership among the details yet to be settled. The group has also discussed eventually moving the project to a neutral organization such as the W3C or an AI foundation.

Finding tools before using them

AWS is already exploring how it could connect with Agent Registry and find resources outside its own catalog.

There may not be much time to settle on a common approach, since connecting all these directories will only get harder once companies have built their own discovery systems.

The post MCP was supposed to solve the agent tooling problem. It missed a step. appeared first on The New Stack.

  •  

JetBrains told everyone to patch. It didn’t patch itself.

Abstract black and blue

JetBrains is urging users of its Cadence cloud development service to rotate credentials and treat previous executions and their outputs as untrusted after attackers exploited a critical TeamCity vulnerability on a server the company failed to patch.

The irony is hard to miss. JetBrains disclosed CVE-2026-63077, a critical vulnerability in TeamCity On-Premises, on July 27. The flaw allows an unauthenticated attacker with HTTP or HTTPS access to a vulnerable TeamCity server to execute arbitrary operating system commands with the privileges of the TeamCity server process. 

By August 7, the company announced that attackers were already exploiting unpatched TeamCity servers.

But one vulnerable server was still exposed: JetBrains’ own.

“The server should have been patched as part of our response to the vulnerability, but it was not,” JetBrains acknowledged in its disclosure of the Cadence incident.

“The server should have been patched as part of our response to the vulnerability, but it was not,”

Cadence’s unpatched TeamCity server

Attackers targeted api.cadence.jetbrains.com, which is the server behind Cadence, JetBrains’ cloud compute service for PyCharm. JetBrains found out about the attack on August 23 and took the server offline the next day. Their investigation shows that malicious activity started on August 8, so the affected period is from August 8 to August 24.

Cadence integrates with PyCharm via an optional plugin, giving developers the ability to run projects on cloud compute resources. TeamCity sat behind the service, orchestrating those workloads.

That put the compromised server in a particularly sensitive part of the development environment.

Exposed credentials and source code

JetBrains says the attackers got hold of a complete Cadence server backup from 2024, potentially exposing everything stored in it, including credentials, configuration files, artifacts, and logs. 

The company also confirmed that multiple AWS IAM users and their associated credentials were compromised, including IAM users belonging to JetBrains employees who had used Cadence.

Attackers accessed files in S3 buckets within JetBrains AWS accounts used by the service. JetBrains is still determining the full scope and says it does not yet know whether customers’ storage buckets were accessed.

Developers using the PyCharm plugin could sync project files to Cadence before running them, so source code and any credentials or configuration files included with those projects may also have been exposed.

The breach also exposed usernames, real names, email addresses, last-login timestamps, and last-accessed IP addresses. Credentials used during Cadence executions may have provided access to other connected services as well.

Supply chain risk compounds quickly

JetBrains says users should consider any credentials or secrets stored in Cadence, included in the compromised backup or used during an execution to be compromised. 

That could mean rotating AWS, Azure, and Google Cloud credentials, as well as tokens for GitHub, GitLab, and Bitbucket. JetBrains also warns about credentials for npm, Maven, NuGet, PyPI and container registries such as Docker Hub, ECR, GCR, and ACR.

Access to those registries creates another problem. An attacker with publishing credentials could push a malicious package that gets pulled into other projects, similar to a recent npm attack that used provenance attestations to spread through the software supply chain.

The warning also covers Slack tokens, webhooks, API tokens, SSH and deployment keys, service account credentials and signing keys or certificates.

Anything run through Cadence during the affected period, including the resulting output, should also be treated as untrusted, according to JetBrains. The concern isn’t limited to exposed data and credentials; anything Cadence ran during that time could potentially have been altered.

CI/CD and remote execution systems often have access to private repositories, dependencies, cloud storage, package registries and deployment systems. Their central role in software delivery is part of what has made CI/CD infrastructure an acquisition target and also gives attackers more places to go after gaining access. Once the environment itself has been compromised, developers can’t assume the credentials that passed through it or the artifacts it produced are safe. 

Once the environment itself has been compromised, developers can’t assume the credentials that passed through it or the artifacts it produced are safe. 

Audit logs reveal lateral movement

Changing potentially exposed credentials is only part of JetBrains’ advice. Users also need to check if those credentials were used elsewhere and if anything changed in systems connected to Cadence.

JetBrains recommends checking source control audit logs for unexpected repository clones or downloads, unauthorized commits, and changes to repository secrets or webhooks. Users should also look for new or changed personal access tokens, API tokens, and SSH keys.

JetBrains recommends checking source control audit logs for unexpected repository clones or downloads, unauthorized commits, and changes to repository secrets or webhooks.

Cloud environments need the same careful review. JetBrains says users should look for unexpected IAM changes, new users or service accounts, and unusual access to storage like S3 buckets. Authentication logs can also show if credentials used in Cadence were later used from unknown places.

Package repositories and release histories should be checked for unexpected publications or changes, especially where Cadence had credentials that could publish packages or artifacts.

Since JetBrains treats previous Cadence executions and their outputs as untrusted, the investigation goes beyond just checking account logs alone. Developers may also need to review artifacts made through the service during the affected period and make sure they match trusted source code and expected build results.

The company published six IP addresses associated with detected exploitation: 150.109.230.104, 43.153.227.206, 62.210.127.48, 210.247.242.190, 15.235.225.205 and 152.233.30.18 with a warning that these indicators are not complete, so not seeing them does not mean an account or system was not compromised.

The post JetBrains told everyone to patch. It didn’t patch itself. appeared first on The New Stack.

  •  

[Launched] Generally Available: Azure VM Image Builder in sovereign and air-gapped clouds

Overview:Azure VM Image Builder is now generally available in Azure Government, China North 3, Azure Government Secret, and Azure Government Top Secret.You can now use the same managed image-building service across your sovereign and air-gapped environmen
  •  

[Launched] Generally Available: eBPF host routing in Advanced Container Networking Services for AKS

eBPF Host Routing in Advanced Container Networking Services for Azure Kubernetes Service (AKS) is now generally available. eBPF Host Routing improves Kubernetes networking performance by moving packet forwarding and routing decisions directly into the Lin
  •