❌

Vue lecture

How a forgotten node can put Oracle Java back in production

Enterprise Java company Azul announced its Azul Intelligence Cloud AI Assistant on Wednesday. The technology arrives in response to industry-wide alerts that AI has become a force multiplier for threat actors today.

The Azul service is a natural-language query interface that allows software engineering teams to see where security risk and licensing infringements are hiding in their live production Java estate. The assistant provides answers that are “grounded in live runtime data”, so it replaces static reports that grow less accurate after the day they’re generated.

How fast do code-scanning reports go stale?

Azul said that “most IT and engineering teams” still manage Java risk with static IT and software asset management (ITAM/SAM) reports and code-scanning tools that describe a moment in time. The company has insisted that these reports are “accurate on the day they’re generated, and increasingly wrong after that”, typically because Java Virtual Machines (JVMs) are spun up, patched, drifted and retired underneath the report’s scope.

“For years, enterprises have built dashboards and reports to understand what’s actually running in their Java estate, but by the time a report gets properly summarized and reviewed, the risk it describes has often already changed,” said Scott Sellers, co-founder and CEO of Azul. “That used to be a productivity problem. Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk – for security, for compliance and for the licensing exposure that shows up in an audit.”

“Now that AI can find and weaponize a vulnerability in hours instead of weeks, it’s a business risk…”

The Azul Intelligence Cloud AI Assistant lets software teams ask a direct question, in plain language, and get an answer grounded in what’s actually running in production right at that moment in time, as well as query historical information for further analysis.

The weaponization gap is closing

Where enterprises run business-critical workloads on Java alongside AI services, Azul said the practical effect is that the gap between a Common Vulnerabilities and Exposures (CVE) entry being disclosed and it being weaponized is getting shorter. 

Citing AI models such as Anthropic’s Mythos and OpenAI’s Aardvark, which have autonomously discovered real-world vulnerabilities, the company pointed to an April 2026 Cloud Security Alliance white paper (listed as unofficial AI-assisted research), which suggested that, “While organizations historically took a median of 32 days to apply patches to known vulnerabilities – a window that once roughly corresponded to the time available before exploitation began – that window has collapsed to approximately 5 days for median time-to-exploit in 2025.”

Azul highlighted its Intelligence Cloud service, which gives engineers two “continuously updated” records of their Java estate: JVM Inventory, a live catalog of every JVM instance running anywhere (on-premises, cloud or container) and Code Inventory, a runtime record of which code actually executes in production versus what is merely provisioned. The new AI Assistant puts in a conversational layer using LLM models on top of both.

Software engineers can ask questions such as: 

  • “Which JVMs are running Java versions which are not the latest updates?”
  • “Where is Oracle Java running in production right now?”
  • “What code hasn’t run in the past four quarters and is safe to remove?”

In the FAQ section of its announcement, Azul suggested that post-migration JVM “drift is common”, often due to a rollback, a forgotten node, a shadow deployment or various scripts and processes that haven’t been updated, which can reintroduce an Oracle Java runtime, exposing compliance and licensing risk, if not security risk also.

The shape of the Java runtime security market

In terms of which other vendors operate in the Java runtime analytics and security market, there are more than a handful of usual suspects. Contrast Security is known for its JVM agent and in-app bytecode instrumentation. Dynatrace offers Runtime Vulnerability Analytics as an extension of its core observability platform, which ships with OneAgent monitoring for Java vulnerable functions and JVM-level bytecode instrumentation agents.

Through its acquisitions by HP, Micro Focus, and now OpenText, Fortify remains known for its static and dynamic application testing services, including Fortify Application Defender, a runtime application self-protection (RASP) agent built to monitor Java workloads during execution. Part of Thales, Imperva’s runtime security for Java and .NET applications spans simple access control up to complex anomaly detection algorithms, though Imperva has reportedly put its standalone RASP product on an end-of-sale path. Then there’s Datadog, with its Application Performance Monitoring (APM), built to power code-level distributed tracing from browser and mobile applications to backend services and databases.

“Runtime context is absolutely critical for understanding the real risk in the production environment.”

A busy market for sure, so just how much of a problem are now-anachronistic static reports?

Head of security advocacy at Datadog, Andrew Krug, tells The New Stack that the downside of most point-in-time inventory scans is that they are “not always representative” of the runtime environment. 

“Runtime context is absolutely critical for understanding the real risk in the production environment,” Krug says. “Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable [i.e. circumventable or subvertible] to get a feature deployed. Moreover, traditional vulnerability management flows outside of SDLC can bump versions outside of CI/CD process, accidentally compounding the problem and introducing added risk by bypassing known good guardrails like dependency cooldowns.”

“Even in the most mature software development lifecycle (SDLC) flows, tooling that generates static software bill of materials (SBOMs) may be bypassable… to get a feature deployed.”

Krug further advises that Datadog now sees “an increasing rise” in automated drive-by attacks on known vulnerabilities, “particularly Java” in many cases.

“Attacks used to be added to scanners, either specifically or generically, and scans are indiscriminately against targets,” Krug clarifies. “LLMs make it cheaper to add support for new vulnerabilities. However, it also makes it easier for individual researchers/hackers to have their own custom rulesets.”

“LLMs make it cheaper to add support for new vulnerabilities.”

He advises that this fact makes trends much harder to read than “oh, someone added support to CVE-2026-whatever in FFUF”, so today the goal of many attacks is unchanged.  Attackers are looking to move laterally, establish persistence, and often automations will look for credentials to leverage to do just that.

NOTE: (Fuzz Faster U Fool) is an extremely fast web application fuzzer (written in the Go language), which is used by security testers to discover hidden files, directories and endpoints.

Dead code; it’s really a ‘thing’

The above-noted list of competitors that work in Azul’s marketplace is (arguably) substantial evidence of the real commercial licensing and security risks that exist where Java code, redundant JVMs and chunks of unsubstantiated (or more likely just untracked) Java components have been left to roam free. Azul itself noted that there’s a real maintenance overhead here that needs to be addressed because “unused and dead code that still gets tuned, tested and carried through every migration”, usually because no one can assess it’s safe to remove.

The larger the Java estate, the larger the exposures get, obviously. But that also means that the less a point-in-time report can be trusted to catch these exposures before they become an incident, an audit finding or a breach.

Dedicated compliance officers will likely enjoy wider deployment of these tools, although that role itself may now reside within a DevSecOps or platform engineering team, or both. 

Azul Intelligence Cloud AI Assistant works regardless of which JVMs are deployed, from which vendor, or how old or large the applications running on them are. JVM Inventory and Code Inventory retain component and code-use history over time, so the AI Assistant can reason over which code, JVMs and applications have actually run in production, now and in the past. 

The post How a forgotten node can put Oracle Java back in production appeared first on The New Stack.

  •  

Developers and platform teams both want Kubernetes self-service. They disagree on who owns it.

Abstract view up through bold yellow angled beams to a white skylight grid and curved ceiling panels.

What do developers want? Kubernetes environments when they need them.

What do they not want? Those environments after a week or more of tickets. 

Platform teams, meanwhile, own what those environments cost, who can access them, and whether they meet company policy.

That tension is the core of Kubernetes self-service: What can safely be handed to developers, and what still belongs to the platform team?

In a recent interview with enterprise cloud specialists — Marius Bogoevici, Senior Principal Product Manager at Hewlett Packard Enterprise (HPE), and Karthik Subramanian, Principal Product Manager for HPE Morpheus Software — The New Stack explored the core friction points of Kubernetes self-service. The conversation focused less on whether self-service is desirable than on where to draw the line.

HKS, HPE’s CNCF-certified Kubernetes distribution, is integrated with HPE Morpheus Software to help platform teams deliver and lifecycle-manage Kubernetes environments as part of a broader operating model spanning Kubernetes, VMs, infrastructure, and clouds. HPE Morpheus Advanced Software supports the on-premises private-cloud use case with HKS, while HPE Morpheus Enterprise Software extends Kubernetes and application operations across hybrid and public-cloud environments.

Together, HKS and HPE Morpheus Software extend that operating model beyond infrastructure provisioning. Through service and application catalogs, platform teams can connect approved Kubernetes environments with the CI/CD pipelines, container registries, automation tools, and other services developers already use. Developers receive a governed, ready-to-use path from code to deployment instead of manually assembling the toolchain for each project.

The self-service paradox

Open-source Kubernetes provides orchestration and declarative APIs, but not a complete operating model.

Subramanian says teams building their own self-service layer usually run into two recurring problems:

  1. Tool and package sprawl: To make upstream Kubernetes production-ready, platform teams must curate and maintain an ever-evolving ecosystem of third-party CNCF tooling for networking (CNI), storage (CSI), ingress, identity, and policy enforcement. Navigating and supporting this fragmented stack creates immense maintenance overhead for internal platform teams.
  2. Day-2 lifecycle and hybrid footprint complexity: Spinning up a Kubernetes cluster is the easy part, but keeping it current — across development, QA, staging, and production — is where the work piles up. That is why HPE says every Kubernetes upgrade must be checked against the networking, storage, ingress, identity, and policy components around it. The problem gets harder when clusters span bare metal, private clouds, edge sites, and public clouds, because one-off scripts and environment-specific configurations can quickly create drift. That maintenance burden belongs with the platform team, not with developers trying to ship applications.

Giving developers direct access to raw Kubernetes APIs just shifts the operational work — it’s far from gone for good. In fact, developers will wind up debugging manifests and storage drivers instead of writing code. 

Meanwhile, operations teams have to deal with overprovisioning, idle clusters, and configurations that reach production without review.

What developers control — and what the platform supplies

The practical answer is not unrestricted access. It is a paved path: approved Kubernetes services that developers can request themselves, with access, configuration, placement, approvals, and lifecycle controls defined by the platform team.

“The best candidates for self-service are requests that are repeatable, low-risk, and well-understood,” Bogoevici tells The New Stack. “For example, a developer should be able to request a development cluster, deploy an approved application, create a namespace, or select resources from pre-approved configurations without opening a ticket. The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

“The platform team decides what a safe configuration looks like, and the developer chooses from a supporting menu.”

Rather than asking developers to write YAML for ingress, storage classes, and RBAC, HPE Morpheus exposes those choices through service catalogs, reusable layouts and blueprints, workflows, role-based access control, approvals, APIs, and automation. Developers do not lose Kubernetes. They retain direct access through standard Kubernetes interfaces and tools where permitted, while the platform team standardizes the request, governance, and lifecycle processes around them.

Those catalog items can package more than infrastructure settings. They can also integrate the approved services and application components that support the development workflow – including CI/CD tooling, source and artifact repositories, container registries, and runtime dependencies – while the platform team controls how those components are configured and governed.

Developers choose the parameters that matter to the application:

  • Approved Kubernetes versions and cluster sizes: Select from pre-tested Kubernetes runtime releases and node count templates.
  • Resource quotas: Specify required CPU, RAM, and persistent storage capacity tailored to the workload.
  • Integrated toolsets and IDE environments: Select required developer toolchains, container registries, and runtime dependencies. 
  • Lease and duration limits: Define explicit operational lifetimes for temporary development or sandbox clusters to prevent abandoned infrastructure sprawl.

Network isolation, identity-provider integration, security policy, and cost allocation stay with the platform team and are applied automatically through the approved service configuration.

The same division of responsibility applies to the delivery toolchain: Developers choose from approved services, while the platform team manages the integrations, credentials, policies, and automation behind them. This gives developers a consistent experience without shifting toolchain maintenance and governance onto individual application teams.

The division of responsibility looks like this:

Service areaDeveloper chooses or requestsPlatform team defines and suppliesReview or exception path
Development cluster provisioningApproved Kubernetes service, version, size, target environment, and duration.Reusable layout or blueprint, access controls, placement rules, storage and network defaults, and lifecycle policy.Nonstandard versions, placements, configurations, or requests outside quota.
Production deploymentApplication artifacts, target namespace, and deployment request through the approved path.RBAC, tenancy, policy, audit, backup, and release controls appropriate to the environment.Formal review for production changes and exceptions.
Resource allocation and quotasCPU, memory, storage, and other approved capacity parameters within project limits.Project quotas, upper bounds, placement constraints, and supported resource profiles.Requests above quota or for specialized resources.
Networking and securityApplication endpoints and permitted connectivity within approved patterns.Identity integration, RBAC, tenant isolation, network policy, secrets, and audit controls.Cross-tenant access, elevated privileges, or changes to baseline security policy.
Lifecycle and cost governanceService lifetime and approved operational actions.Visibility, policy, approvals, retirement workflows, and applicable cost controls for the licensed variant.Long-running exceptions, nonstandard lifecycle actions, or budget exceptions.

From ticket queues to a repeatable paved path

In conventional IT environments, provisioning a dedicated Kubernetes environment for a new project often involves cross-departmental ticket handoffs spanning infrastructure, networking, security, and storage teams. This friction frequently stretches provisioning timelines from days to weeks.

By unifying infrastructure orchestration, role-based access controls, and multi-tenancy into a single operational experience, HPE Morpheus Software can compress these provisioning workflows down to minutes or hours, according to HPE. “Developers get a usable environment that complies with the organization’s defined controls and policies, without needing to understand all the complex infrastructure steps sitting underneath,” Bogoevici says. “When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

“When you reduce provisioning time from weeks to hours, that is super meaningful and tangible.”

The result is not only faster cluster provisioning. HPE Morpheus Software can also automate the handoff into the developer’s established delivery process by making approved CI/CD and application services available with the environment. Instead of waiting for separate teams to connect pipelines, registries, credentials, and runtime dependencies, developers receive a ready-to-use path from development through deployment.

Faster provisioning can create a different problem, too: The speed can and will cause teams to lose track of what was provisioned and why. HPE Morpheus Software gives administrators visibility into utilization and cost, while lease controls can shut down temporary development clusters when their time expires.

Bogoevici says ticket volume is a poor measure of success, particularly early on, when more developers may be trying the catalog. He recommends watching deployment success, exception rates, resource utilization, and the day-to-day effort required to keep the service running.

Security belongs in the service design

Security is another boundary that must be designed into the self-service path. If identity, access, tenancy, and policy are added only after a cluster is created, every request produces more work and more room for inconsistency.

“Security must be a core design consideration built directly into the service, not an afterthought during deployment,” Bogoevici says. “HPE Morpheus Software brings identity integration, role-based access, tenant isolation, approvals, and policy into the operational workflow.”

A newly provisioned environment should arrive through an approved configuration with the applicable identity, RBAC, tenant, policy, and audit controls attached. Platform teams can validate the paved path by testing an allowed request, a request that should be rejected, and the resulting audit record.

The operating-model test

The strongest Kubernetes self-service model does not hide Kubernetes or make it the control plane for every workload. It gives developers useful, approved choices and direct access to the Kubernetes workflows they need, while the platform team standardizes the enterprise processes around those workflows.

That matters because the enterprise still runs VMs, clouds, and existing infrastructure alongside Kubernetes. HPE Morpheus Software helps platform teams use common request, governance, automation, and lifecycle processes across these environments without forcing every workload onto one runtime or creating another operational silo.

In practice, that means self-service should deliver more than a Kubernetes cluster. With HPE Morpheus Software, a catalog request can bring together the approved environment, application services, and DevOps toolchain integrations developers need, while preserving the governance and lifecycle controls the platform team requires. Developers spend less time assembling and troubleshooting delivery infrastructure – and more time building and releasing applications.

Looking toward 2027, the goal is not unrestricted control. It is faster access, predictable results, transparent guardrails, and a clear exception path when the standard service does not fit.

The post Developers and platform teams both want Kubernetes self-service. They disagree on who owns it. appeared first on The New Stack.

  •  

Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production.

Inspecting changes on a laptop screen document

We all know that producing code is easier than ever thanks to the abundance of AI coding tools and agents. The harder part undoubtedly comes after that code is written: making sure changes are safe to ship, spotting regressions in production, and figuring out what went wrong.

And that’s why Cursor is introducing Rollouts, a new agent that follows code changes into production and monitors whether they behave as intended.

The Firetiger effect

The announcement comes a little over a month after SpaceX closed its bumper $60 billion acquisition of Cursor, giving the AI coding company access to SpaceX’s vast GPU infrastructure as it develops its own models.

The day before that deal closed, however, Cursor quietly announced an acquisition of its own: it snapped up the team behind Firetiger, a three-year-old startup building AI agents that monitor software changes from pull request through deployment.

At the time, Firetiger co-founder and CEO Rustam Lalkaka argued that coding agents had dramatically reduced the effort involved in creating software changes, while doing little to reduce the risks involved in actually deploying them.

“Over the last two years, agentic coding has changed software dramatically,” Lalkaka wrote in a LinkedIn post following the deal’s announcement. “The cost of creating changes has dropped to near zero. The cost and risk of deploying them has stayed largely the same.”

“Writing code is no longer the slow part. What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rustam Lalkaka, Cursor

Fast forward to today, and Lalkaka, now at Cursor, has unveiled the first fruits from that acquisition — including Rollouts. In a blog post published on Wednesday, Lalkaka notes that the new agent, or “bot” as the company calls it, is all about helping developers “get safe, reliable code into production faster.”

“Writing code is no longer the slow part,” Lalkaka writes. “What hasn’t sped up is everything after the PR goes up: making sure code is secure, watching the deploy, deciding whether a latency bump is real, figuring out which of eleven changes broke checkout.”

Rollouts is effectively Firetiger’s Change Monitors reborn inside Cursor, rebuilt using a tool dubbed Bot Development Kit. This kit, too, appears to be new from Cursor: an early-stage framework for building and serving Cursor bots and agents, published as the @cursor/bdk package on npm. Its documentation says developers can define agents using Markdown and TypeScript, with support for tools, skills, subagents, webhooks and scheduled runs.

Like Change Monitors before it, Rollouts starts working when a pull request opens. It examines the proposed code change, works out which systems could be affected, and produces a monitoring plan covering what the change is supposed to do, the risks it sees, the signals it intends to watch, and any holes in the available instrumentation. Developers can review and edit that plan before the code reaches production.

Rollouts in action (1)
Rollouts generates a monitoring plan for a change

Once the change is deployed, Rollouts checks the resulting telemetry — including logs, metrics and traces — against that plan. Staging and production are assessed independently, with each deployment ultimately receiving one of three verdicts: verified healthy, regression detected or inconclusive.

That means a change could, for example, pass its checks in staging before Rollouts subsequently spots a problem when the same code reaches production.

Rollouts in action (2)
Rollouts reports deployment status as changes ship

If Rollouts does detect a regression, it can identify the change it suspects, alert the developer responsible and, depending on how it’s been configured, either open a revert pull request for review or hand the problem to a Cursor cloud agent to attempt a fix. There is still a human in the consequential part of that loop for now: Rollouts doesn’t merge fixes or roll back deployments by itself, though it can pause a progressive rollout.

Lalkaka notes that Rollouts is already capable of picking up problems limited to a particular endpoint or region before they trigger a broader alert, while it can also distinguish expected changes in behavior from genuine regressions.

Also “coming soon” to Rollouts, according to Cursor, is an integration with feature flags so it can directly adapt the traffic reaching a change, while support for release trains and deployment freezes is also in the works.

Enter Security Reviewer

Alongside Rollouts, Cursor is also introducing an upgraded Security Reviewer bot, which first appeared in beta back in April.

At launch, the bot could automatically inspect pull requests for security vulnerabilities, authentication regressions, privacy and data-handling risks, agent tool auto-approvals, and prompt-injection attacks, leaving findings alongside the relevant code.

As with Rollouts, the idea is that developers don’t have to remember to invoke it manually: Security Reviewer can be set to run whenever a new pull request is opened.

Security Reviewer in action
Security Reviewer runs automatically on new pull requests

In its current guise, Security Reviewer analyzes pull requests in the context of the wider codebase, with a focus on exploitable issues such as injection flaws and broken authentication, and returns a severity rating, attack path and proposed fix.

“Security Review reads code the way a security engineer does,” Lalkaka writes. “Where does user input enter, where does it end up, what does it pass through on the way.”

“Security Review reads code the way a security engineer does.”

He says that things have sped up considerably, too: average review time has fallen 21%, from 4.8 minutes to 3.8, while developer acceptance of its comments has risen from roughly 45–50% to 60–70%.

Both Rollouts and Security Reviewer are available through Cursor’s Automations tab for customers on its Teams and Enterprise plans.

The Origin story

Digging into the nuts and bolts of Rollouts reveals how it might serve as a boon for Cursor as it builds out Origin, the fledgling Git-compatible code hosting platform it launched back in August.

Origin is essentially an effort to build an alternative to GitHub for an agent-heavy software development world. It remains early, with limited functionality, but Cursor has been clear that tighter integration with its own agents is supposed to become one of the main reasons to use it.

When Cursor announced the Firetiger acquisition last month, Maxime Prades on the Cursor product team noted in a blog post that the deal was part of a “broader investment in long-running, autonomous, context-aware agents for teams.”

And he pointed to Origin and Change Monitors as two examples of that investment.

“Agents that write code should also be able to tell whether it works in production,” Prades wrote. “Today, those systems are mostly separate. Cursor and Firetiger bring them closer together so an agent can ship a change, see how it behaves, and respond when something goes wrong.”

Rollouts offers an early glimpse of that. It can connect to either Origin or GitHub for source control, pull deployment events from continuous delivery systems, and use signals from Datadog and other telemetry providers. If it spots a regression, it can then pass the problem back to a Cursor cloud agent to investigate or attempt a fix.

Origin potentially gives Cursor a native home for more of that loop: its cloud agents can already create branches, commit and push code, and open pull requests against Origin repositories. Rollouts then adds information about what happened after.

That could become increasingly important as more companies take aim at GitHub’s central role in software development. Zed, for example, put Delta into public beta last week, with its own ideas about how source control should change for teams working heavily with agents.

Cursor also faces competition further downstream. Datadog’s Bits Release, launched in preview in June, similarly follows changes from pull request into production and checks telemetry for regressions. Harness has long offered automated deployment verification and rollback based on logs and metrics, while LaunchDarkly’s Guarded Rollouts can monitor feature releases for regressions and automatically reverse them.

What Cursor can potentially bring to the table is proximity: the coding agent, repository, pull request, security checks, and production feedback can all sit much closer together. Rollouts doesn’t require Origin — GitHub remains supported — but owning the forge gives Cursor more room to integrate those pieces over time. And that may prove more compelling than simply recreating GitHub’s existing feature set.

The post Cursor acquired Firetiger. A month later, it launched a bot that tracks code changes from PR to production. appeared first on The New Stack.

  •  

Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge.

Abstract image of thin vertical ribs against a bright orange background. In the center, a blurred rectangular glow shifts from green on the left, through dark red, to pink on the right, as if seen through ribbed glass.

With often hundreds of thousands of alerts a day, many tech organizations are buried in vulnerabilities and worn down by alert fatigue. The rise of AI has only made it harder to cut through the noise and to find actionable alerts. Manual security and site reliability engineering is not an option. 

The engineering team behind WHOOP‘s health and fitness tracker felt this pain, relying on multi-day, all-hands triage sessions to stay on top of the alert deluge. But, as a high-growth consumer health company handling sensitive user data, it couldn’t afford to miss anything. Which is why the team at WHOOP built an automated vulnerability-response workflow based on the company’s specific technical, operational, and trust considerations.

Join The New Stack on Wednesday, October 7 to learn from WHOOP staff engineer Vinay Raghu and Datadog senior product engineer Amber Tunnell how WHOOP built and implemented this workflow using Datadog Bits AI and Workflow Automation for faster, at-scale response. 

Join us on October 7, 2026, for a live Datadog x TNS event

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

DevSecOps, security, and cloud/platform engineers should bring their questions to this live demo-slash-case study to learn how to reduce friction between developer velocity and security requirements without increasing headcount. 

What you’ll take away from our live webinar

Raghu and Tunnell engineers will share how they were able to:

  1. Focus on real exposure vs. scanner noise. Not everything is critical. You’ll learn how WHOOP used Datadog’s Software Composition Analysis (SCA) to analyze runtime code execution and prioritize active threats.
  2. Route the right vulnerability to the right engineer. WHOOP automated vulnerability mapping to microservice owners, so developers received tickets with full context attached.
  3. Build automated guardrails for devs to self-resolve. This let security engineers pivot from frustrating gatekeeping and ticket-pushing to more proactive, systemic work that adds value. 
  4. Maintain a human in the loop. With such sensitive data and a demand to be always-on, WHOOP isn’t ready to automate the engineer out. Learn how they decided their team’s response had to change.

And then, of course, we will end the live discussion with how to measure it all. Don’t miss out and register to attend on October 7.

The post Vulnerability alert fatigue nearly swamped WHOOP. But its fix still keeps a human in charge. appeared first on The New Stack.

  •  

Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent 

Illustration of colorful doughnut, bar, radar, and line charts alongside slider controls on a black background.

What if engineers could spend their time building and optimizing systems rather than maintaining them?

It’s 3 a.m., and the pager goes off. Tabbing between multiple dashboards and diagnostics, the SRE struggles to determine whether what woke them is a real incident, whether they’re the right person to handle it, or whether they need to wake someone else. Digging through monitoring tools, deployment history, incident systems, and team runbooks — and chasing what might be the wrong theory about the root cause — they can’t respond fast enough to stop more customers from being affected.

Or imagine that, by the time the SRE joins the incident bridge, Azure SRE Agent has already analyzed the monitoring data, identified the root cause, and prepared a fix for approval and deployment.

Sanchit Mehta, one of the head engineers for Azure SRE Agent, tells The New Stack that “[Azure SRE Agent] starts analyzing telemetry and correlates things like blast radius, deployment changes, recent changes, any recent rollouts, to try to tell the engineers, ‘OK, this is what is causing it.'” Increasingly, it will even create the PR for that fix.

The support is just as useful during normal working hours. At InEight, correlating telemetry across tens of thousands of Azure resources can take days, if not weeks. When a support ticket reports slow performance without identifying the product, engineers must determine which of the company’s 14 products is affected, then check multiple observability and reliability tools.

InEight shared that, during its first incident using Azure SRE Agent, the agent quickly identified the affected product, traced the performance issue to its root cause, and recommended scaling Redis. The DevOps team had been considering scaling the app service as a temporary fix.

Proactive and in production

This kind of help is becoming the new normal at Microsoft, where more than 3,000 service teams already use Azure SRE Agent to investigate issues, perform root cause analysis, respond to incidents, fix code, enable automatic mitigation, support proactive detection, analyze data, and report at scale. Azure SRE Agent has already handled more than 1.8 million incidents inside Microsoft, many mitigated in minutes.

The team also uses Azure SRE Agent to develop and improve the service itself, with custom agents for code review, deployment, evaluation, and monitoring. This “agent-powered engineering” approach, as Mehta calls it, lets the team take advantage of ongoing advances in AI models. That includes proactively spotting problems, like quota issues that affected deployments, and automatically raising support tickets to resolve them. The agent recently identified the root cause of a change that broke synthetic tests as soon as the change reached the first region, he says.

 “It said, ‘OK, this was an upstream PyPI package that broke your dependency; you need to add tests for it; you should roll back immediately; here’s how you should go fix this.'”

Mehta says that kind of proactive monitoring is hard to handle with deterministic queries. “You need a level of intelligence to see when a large production payload is being deployed and if it has the potential to cause degradations.” 

For some internal teams, more than half of incidents are autonomously managed by the SRE agent and don’t need any human intervention, adds Shamir Abdul Aziz, lead program manager for Azure SRE Agent, because they’re what he calls “safe” operations and mitigations: a restart, scale-out or rollback of a service, or change order requests escalated by customers.

“The humans did the governance, set up the guidelines, gave some coaching to the agent, and then it went into auto mode to complete the entire workflow,” Abdul Aziz says.

Agents are ready to help

SREs are already drowning in repetitive toil. SREs are already drowning in repetitive toil, and coding agents add to that workload. Agentic operations are now powerful enough to help, Vyom Nagrani, one of the head PMs for Azure SRE Agent, tells The New Stack.

“As code gets written more and more by agents, it’s going to take another agent to operate it,” Nagrani says. “But why wait? If the agent can manage code which other agents write, why can’t it manage code written by humans?”

“As code gets written more and more by agents, it’s going to take another agent to operate it.”

“The reasoning loop has become mature enough that now agents can automatically start figuring out a lot of these complex problems, especially when it comes to correlating across multiple data sources, which has always been the hardest thing for humans to do,” Nagrani says.


Powerful models aren’t enough, though, and homegrown automation won’t have the production-grade governance, verification, evaluation, telemetry, and control a platform can offer.

The state of the art has progressed from prompt engineering to context engineering—which grounds AI in your infrastructure, code, and institutional knowledge — and now to harness engineering. “That is what allows you to run agents at scale, control them, and govern them,” says Abdul Aziz.

“When you combine all these things with being able to verify, audit, evaluate, and get real telemetry and metrics out of the system, where the agent claims it has done something, you can validate that agent’s claim,” Abdul Aziz says.

Instead of a non-deterministic black box that can’t explain its decisions, you can trace and learn from the agent’s reasoning so that you can correct mistakes once, not over and over again. “That’s why companies are willing to adopt it now,” Abdul Aziz says. “Because when you try the same thing ten times, you’re going to get the same output.”

“You don’t just turn on the agent, give it full access, and ask it to solve everything.”

After two years of building enterprise-grade systems that can be trusted, audited, and validated, the next step for cloud-native SRE can be agentic ops with autonomous capabilities — but you still need to know how to adopt it, Abdul Aziz warns. “You don’t just turn on the agent, give it full access, and ask it to solve everything.”

Context and connections

Azure SRE Agent is built for Azure but not limited to the Azure platform. The agent provides native access to Azure services such as Azure Monitor, Application Insights, Log Analytics, and Azure Resource Graph. Connecting the agent to your subscriptions, telemetry data, and source code gives it the operational context and institutional knowledge needed to understand how you work.

Beyond Azure, Azure SRE Agent integrates with engineering and operational tools through managed connectors for Azure DevOps and GitHub, plus MCP connectors that enable access to external knowledge sources such as Google Drive, Confluence, Cursor, Claude Code, and other third-party systems. 

Put all that knowledge into Markdown files in a repo, along with the skills and tools agents need to act on your systems (including third-party and on-premises services). That gives you artifacts that agents can version, review, test, reuse, and update.

When you want to dictate how to handle an incident — what to check and in what order, what to post, and even how to format a report — you can create a custom agent, either by using an existing runbook or by working through an incident with an agent and saving that skill. Using agents to improve agents is the shortcut to making Azure SRE Agent more useful the more you use it. Essentially, saving what agents learn during incidents helps improve their future responses.

Guidelines and guardrails

Governance covers identity, role-based access control (RBAC), and tool-access policies. These controls determine which actions are allowed, blocked, or subject to step-by-step approval, and whether an agent operates autonomously or with human review.

What makes governance both flexible and powerful are hooks, based on prompts or deterministic commands, that fire at different stages of a workflow and catch edge cases, such as allowing an agent to drop the index in a SQL database but never drop a table.

Metrics show you whether governance is working. The new live reports show time to mitigation, tool reliability, how often agents act autonomously, and cost per outcome at a glance. InEight’s metrics are typical: an 80% reduction in both incident investigation time and build failure triage time, a 67% reduction in the effort needed to investigate bugs, and an 84% reduction in cost.

To get those results, you need triggers that automatically launch agents instead of waiting for a human to open a chat window.

Bind skills and custom agents to specific alert classes so they can respond to incidents first. Start agents through pipelines, webhooks, or work items to automate delivery workflows. Schedule regular checks, reviews, and audits, and have agents automatically update their artifacts.

Agents operate; you stay in control 

By reducing repetitive tasks and technical toil, Azure SRE Agent frees engineers to focus on more interesting and innovative projects. Just as there’s a familiar maturity model for adopting site reliability engineering in the first place, you don’t jump straight into having agents rather than humans handle operations. When you give agents the context about your infrastructure, you can start using them for investigations.

“If you give agents read access to your source code, your telemetry, your resources, the time to get to the root cause is reduced to minutes rather than hours or days,” Abdul Aziz points out. “Every customer starts there.”

Once you’re happy with the answers you’re getting, you can give the agent more permissions while still approving individual steps, he says. “The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

“The fixing is easy once you understand the problem. It’s usually changing your configuration, writing a piece of code, or restarting a service.”

As you expand into other operational tasks, refine the agents’ artifacts, metrics, and governance before granting more autonomy: “Things like rolling back a release when we know there was a regression in that release, restarting a service, dropping a corrupt index on a SQL table, or scaling out a service,” Abdul Aziz suggests.

For more complicated issues, agents can deliver the entire fix, ready for approval. The Azure SRE Agent that manages the Azure SRE Agent product looks at exceptions, errors, incidents, Teams conversations, emails, and GitHub issues every night and spits out PRs. 

Avoid code review bottlenecks by having agents deploy, test, measure, and include outcomes in the PRs. Use continuous evaluation to build a self-learning system that accurately follows your existing workflows.

“The agent can self-improve because the agent learns constantly,” Abdul Aziz says. “You can configure scheduled tasks to identify which evaluation scores were low and automatically improve the custom agent, custom skills, and even your knowledge documents – because knowledge management is also a toil. The agent can automate all of that.”

The right way to start

Azure SRE Agent now offers a 30-day trial experience with no always-on charges. Make the most of that by learning from some common mistakes:

  • It’s not magic! Turning on the agent doesn’t mean you don’t have to do DevOps anymore. Don’t treat it as a chatbot or connect it to just your observability system. You need to give the agent the context it needs, the tools to do the job, and intentional triggers that tell it when to act. Otherwise, it may spend effort on low-value work or generate outputs that aren’t grounded in your environment. 
  • Don’t limit yourself to what the agent does out of the box: Customize agent skills, tools, connections, and logic to fit how your organization works, and build custom agents for specific tasks.
  • Don’t use agents for jobs a single line of code can do: Using them to explore deterministic, structured data for anomalies is an expensive waste of tokens that will only flood the context window when the agent can write that line of code itself. “Orchestrate, don’t calculate,” as Nagrani puts it. If you’re drowning in alerts, use automation to filter the noise and only send alerts that need intelligent analysis to agents.
  • Don’t stick with what you’ve always done or copy your org chart: The most effective agents have a complete picture of the system, so they need all the context, even if it crosses two teams. That might mean crossing boundaries, coordinating who has expertise and who needs to grant access, or rethinking how the organization works.

“If agents have the right context, they minimize the toil and truly make operations less costly,” lead product manager Deepthi Chelupati points out. That way you can move faster, be proactive, and give engineers more time to innovate and less maintenance work to dread.

Get started today: sre.azure.com

The post Agents operate, humans govern: Scale your operations and reduce toil with Azure SRE Agent  appeared first on The New Stack.

  •  

How to attach an owner to every cloud resource you find

Dark abstract 3D rendering of intertwined rings symbolizing complex cloud resource governance and technical debt.

The engineer who knew why that cloud instance existed has left the company. The instance is still running, the bill keeps growing, and the team must now decide whether it’s safe to shut down. This is a bad time to discover that its ownership history was someone’s memory.

“Good resource governance has three pillars: continuously synced inventory, policy that blocks resources without tagged owners, and an audit trail that survives every reorg.”

A cost review flags an EC2 instance nobody remembers provisioning. Someone scours Slack for the resource ID, finds nothing, burns half a day chasing dead ends, and eventually stumbles across an exhausted engineer who mumbles the mantra that ends most of these investigations: “I think that’s from the project Priya was running before she left.” 

Nobody follows up, because nobody knows how to reach Priya anymore. The instance stays up, because tearing down a fence when you don’t know what it’s protecting you from is a good way to find out the hard way. It’s the platform team equivalent of emotional baggage; they all seem to accumulate some.

Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous. 

“Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous.”

Solving the problem requires three things that already exist: queries, policies, and logs; and maybe just a little bit of therapy.

1. The query that tells you what’s missing an owner

CloudQuery’s asset inventory syncs continuously across every provider a team runs on, into tables you can query directly: aws_ec2_instances, gcp_compute_instances, azure_compute_virtual_machines, and so on. Finding every resource without an assigned owner is as easy as:

SQL
SELECT resource_id, 'aws' AS provider, 'ec2_instance' AS resource_type
FROM aws_ec2_instances
WHERE tags ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'gcp', 'compute_instance'
FROM gcp_compute_instances
WHERE labels ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'azure', 'virtual_machine'
FROM azure_compute_virtual_machines
WHERE tags ->> 'owner' IS NULL
ORDER BY provider;

Run it regularly, and you have a (hopefully short) boring list to refer to when the CFO asks who deployed an expensive instance; instead of being thrown into a frenzied goose chase at 4:59 p.m. on a Friday.

2. The policy that prevents it recurring

A query tells you what’s already missing an owner; but how do you prevent the next ownerless resource from being deployed? The answer is a policy, and env zero evaluates Open Policy Agent rules against every plan before it applies. A rule that requires an owner tag on every new resource looks something like this:

package env0

# METADATA
# title: require owner tag
# description: A resource can't be created without a declared owner.
deny[format(rego.metadata.rule())] {
	resource := input.resource_changes[_]
	resource.change.actions[_] == "create"
	not resource.change.after.tags.owner
}

format(meta) := meta.description

Add that to the project’s policy set, and a plan that creates a resource without an owner tag doesn’t just get a warning; it doesn’t get created.

3. The record that outlives its creator

A tag tells you who owns something today. It doesn’t tell you anything about who asked for it, why, or who signed off. By the time it matters, the person who could’ve answered from memory may no longer be reachable. An audit entry records that at the moment of creation, instead of reconstructing it afterward from the scraps of recollection scattered around the rest of the team. Here’s an example:

{
  "event": "resource.created",
  "resource_id": "i-0a1b2c3d4e5f",
  "requested_by": "j.chen@company.com",
  "approved_by": "platform-lead@company.com",
  "approval_ref": "ENV-4471",
  "stated_purpose": "load test environment, Q3 capacity planning",
  "timestamp": "2026-08-14T09:12:03Z"
}

This entry answers the question this whole piece opened with, without needing Priya or Slack. env zero keeps this record attached to the resource for as long as the resource exists, specifically so it outlasts any individual’s tenure.

Why this keeps happening

Employee attrition is an age-old challenge that’s only accelerating in the modern era. US private-sector voluntary turnover runs 22 to 25% a year, so a hundred-person org loses twenty-odd people every year, each one taking a small, specific piece of “why this exists” with them. Replacing a mid-level employee costs six to nine months of salary, more than double that for senior specialists. 

“Tags were supposed to survive this. In practice they rot the way everything else does.”

That figure doesn’t touch what the departure does to everyone else’s mental model of what’s actually running. Tags were supposed to survive this. In practice they rot the way everything else does: two teams merge and bring incompatible schemas, provisioning that runs on tribal knowledge accumulates configuration drift for the same reason it accumulates ambiguous ownership, and a resource tagged under a policy that’s been replaced twice since its inception isn’t much better documented than one that isn’t tagged at all.

We wrote in April about the hour it once took our own team to answer, “what are we actually running across both clouds?” That solved a point-in-time problem. The query, the policy, and the audit entry above stop the same story from playing out again next June.

The org chart will keep changing. The record doesn’t have to.

None of this stops people from leaving or teams from reorganizing; pretending otherwise is how platform teams end up rebuilding the same spreadsheet every eighteen months. What changes is whether the next “what are we actually running, and who owns it” conversation takes hours of archaeology across three teams, or is a query that already has the answer attached. Institutional knowledge decays at a fairly predictable rate. A system of record shouldn’t.

The post How to attach an owner to every cloud resource you find appeared first on The New Stack.

  •  

The AI-native SDLC won’t be one process 

Five fuzzy pom-poms — white, pink, magenta, purple, and teal — arranged in a horizontal row across the upper portion of a dark navy background, each casting a small shadow, with colored light washing the backdrop in magenta and teal.

Anthropic recently published its AI-Native SDLC Playbook. Its central claim is that “code is no longer the bottleneck.” When agents can produce an implementation in minutes, the constraint moves to everything around the build phase: planning, review, verification, deployment, and governance.

The risk, if organizations get this wrong, is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones. The traditional answer is that a person looks at each one, and that is exactly what stops working at this volume.

The risk… is producing ten times the changes at the same quality per change or worse, with no way to identify which changes are the bad ones.

The playbook gets the foundations right. What it doesn’t capture is that an organization’s process has nuance: it is really a family of processes that vary with the change at hand, not a single flow every change travels through.

The spec-driven wave

The playbook is part of a broader wave of spec-driven development tooling, including Amazon’s Kiro and GitHub’s Spec Kit. The tools share a common shape. Written artifacts drive the work: an intent document becomes a spec, a plan, a diff, and review findings, all committed to version control. Policy is enforced by deterministic mechanisms such as hooks, rather than by instructions in a prompt. Agents check their own work before a human sees it. Humans own the approvals.

But each of these tools also prescribes a particular process: a fixed sequence of stages that produce fixed artifacts, and that every change travels through. Adopting the tool means adopting its process.

One organization runs many processes

No real organization runs a single process. The right process for a change depends on the risk it carries and the accountability it requires. A documentation fix, a dependency upgrade, and a schema migration in a payments service should not travel the same path. They need different levels of verification, different approvers, and different records. In regulated domains, the process itself is part of the compliance obligation: auditors expect a record of who approved each change and based on what evidence. What must be recorded differs by the type of change.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible.

When a tool prescribes one process, teams route around it for changes that don’t fit, which is the worst outcome because the real process becomes invisible. Or the vendor keeps adding configuration until the tool becomes a workflow engine that nobody fully understands.

The tool should not prescribe a process. It should give the organization a way to define its own.

Processes as state machines

A better model is to define each process as a state machine. The states are facts about a change: reviewed, validated against its dependencies, approved for production. Those facts live in systems no single tool owns: the repository, CI, the cluster, the tracker. So a process cannot be a program that executes steps. It is a set of rules that react to observations about those systems. Each rule specifies:

  1. The facts it requires before it can fire.
  2. Its gate: fire automatically, or wait for a person’s approval.
  3. The permission it grants when it fires, such as merging or deploying.

The definition is this set of rules, stored as data and reviewed like code. An organization runs many small machines, one per risk class.

At runtime, this behaves nothing like a workflow engine. No component tracks “we are on step four”: the process advances when a fact appears in the system that owns it, and rules react. Events that arrive late, twice, or after a restart are handled like any other, because rules only react to current state. The gate is one of a rule’s conditions, so you can hold firing during an incident or a release freeze without editing any definition.

Gates need enforcement. A gate implemented as a prompt instruction depends on the model following it. The agent harness can provide the determinism required to run the state machine and enforce its gates, stopping the agent between actions until a gate is answered, while the infrastructure enforces the rest.

The process adapts to the change

One fixed process definition per repository is not enough: every change in that repository would still travel the same path, regardless of its risk. The path a change takes should depend on what the change is, and this routing comes from classifying the change, not from the author choosing a path. The organization defines classification using signals it already has: the paths a change touches, the repository it lives in, a label on its tracking issue, etc.

The definitions themselves also need to change over time, and that has to be safe. Because a process definition is data, editing it is itself a change, and it goes through its own gated process. Loosening an approval gate on the release process gets reviewed the way a schema migration does, not edited the way a config file does.

Click to enlarge graphic.

Concretely, consider three changes to the same service:

  • A documentation fix is classified by the paths it touches. Its process has two states: the build passes, and it merges. No person is involved.
  • A dependency upgrade skips design review, but its process requires compatibility evidence: the upgraded service runs its integration tests against real dependencies. A major version bump adds an approval that a patch bump does not.
  • A schema migration in the payments service is classified by the component it touches, no matter what kind of change it claims to be. Its process adds states the others never see: review by a payments owner, validation against production-shaped data, and a release approval from someone accountable for that domain.

Each transition, in each path, is logged with who approved it and on what evidence.

Tenets

The tenets these processes should follow:

  • Autonomy is granted per action, and grows over time. Each transition is set to fire automatically, require approval, or hold. As agents prove reliable on a class of change, that setting is relaxed, so the process absorbs agent improvements without redesign.
  • Human attention is spent only where judgment is needed. Agent effort keeps getting cheaper; supervision hours do not. A person is brought into the loop only when the decision requires human judgment, and is given the context to decide quickly.
  • Evidence comes from outside the agent. An agent’s own report never moves a change forward. Transitions fire on facts from systems the agent cannot write to, such as test results and validation in a realistic environment.
  • The process record is the audit trail. The definition is the written policy, and the transition log shows who approved each step, on what evidence, under which version of the policy.

Quality at scale

The playbook and its peers get the foundations right. What is missing is the ability for an organization to define its own processes, vary them by the risk of each change, and evolve them safely. The goal is not fewer humans in the loop. It is spending human judgment only where it is needed, backed by evidence agents cannot produce about themselves, so that quality holds while throughput multiplies.

We are building these ideas at Signadot and acting as our own guinea pigs, running our own development through this process. If you’re experimenting with these ideas too, we’d love to talk!

The post The AI-native SDLC won’t be one process  appeared first on The New Stack.

  •  

Harness rebuilt its Git repository for nonstop AI agent traffic

As one of Harness’s field CTOs, Martin Reynolds spends much of his time asking engineering leaders one question with no easy answer.

It’s how they’re keeping up with all the pull requests their coding agents now produce. One of them recently answered with two words, “we’re not,” and explained that his team’s threshold for pushing code into production had dropped, Reynolds tells The New Stack.

TNS spoke with Reynolds a few days after Harness launched a rebuilt Code Repository and a new AI Code Review product — and a few weeks after GitHub’s nearly eight-hour platform outage on August 17.

In our conversation, we discussed how the review bottleneck arose, why Harness rebuilt its Git repository for agent traffic, and which parts of the pipeline should remain deterministic.

Drowning in pull requests

Reynolds says he first ran into the bottleneck during Harness’ early trials of GitHub Copilot and Amazon CodeWhisperer.

“We were getting more PRs, but all the PRs were getting stuck,” he says, “and the test team was shouting, saying, we can’t keep up with all of this.”

“Imagine what that test team feels like right now.”
—Martin Reynolds, Harness Field CTO.

The 1.5x to 2x increase in new code pushed the testing teams to the breaking point, he says, and now, he sometimes sees teams at 10x, with some claiming 50x. “Imagine what that test team feels like right now.”

Reynolds notes that during hallway conversations with engineering leaders at the conference, drowning in pull requests was a recurring theme. And what he sees when talking to customers tends to split three ways: Some have raised their risk tolerance, some have a backlog they can’t manage, and most sit in the middle.

“The somewhere in the middle, I think, is the most common,” Reynolds says. “We’re using some kind of another AI tool to help us in that space, but it doesn’t necessarily solve the problem.”

Review what’s changing, not the scaffolding

Ideally, a reviewer opening a pull request should see the most important changes first, Reynolds says, and he suggests reviewers should come from whoever has worked on that part of the codebase before, “not the person who did the prompt or wrote the code.”

“This other stuff is like 30 files because they updated a dependency. That’s less important in terms of getting eyes on,” he says. Reviewers should “actually review what’s changing rather than a bunch of stuff that’s scaffolding around it.”

“It’s not just the model on its own,” he also notes. Harness spent “a good chunk of the last 12 months” building what it calls a software delivery knowledge graph, a map of a customer’s pipelines, deployments, incidents, and policies, so the reviewer can pull context “at speed and not burn lots of tokens.”

The company’s own example is a migration flagged because an earlier incident review found an unindexed CREATE INDEX statement had locked a production table for 14 minutes.

By the company’s own count, its engineers saved more than 10,000 hours of manual review time a month. The day Harness launched, GitHub’s Copilot code review began reviewing pull requests opened by bots, including its own coding agent.

Agents don’t work nine to five

Recently, Harness customers on GitHub “would quite often genuinely send us screenshots of GitHub being down,” Reynolds says.

The reason for GitHub’s struggles, he believes, is that GitHub “was ultimately built for people, teams of maybe up to 10, 15, who are changing code, creating pull requests. Those pull requests will be there for a few hours to maybe a couple of days.” But agents “don’t work nine to five.”

Harness has been selling a repository service since 2023, when it launched Harness Code on top of its open-source Git project, and Reynolds says the company rebuilt it as “a ground-up AI-first repository that works for humans and AI.”

He describes it as Kubernetes-based, running across multiple clouds and regions, tested at thousands of commits per second, and used by about 20 enterprise customers during beta, none of which Harness has published.

GitHub CTO Vlad Fedorov’s postmortem on the August 17 outage said “a critical infrastructure component in our Central US data center failed to scale” as traffic hit a new peak. GitHub now handles 2.9 billion commits a month, a little more than 1,000 a second on average.

That’s not the scale Harness operates at, of course, but for its enterprise users, that may just be an advantage.

Harness is also starting to look beyond the traditional process. The capabilities for an autonomous delivery lifecycle exist today, Reynolds argues, but “are organizations and companies ready for that? I’m not entirely sure.”

Either way, he says deterministic tooling needs to stay, and test results still come from the test runner. “There’s no need to rip those out and replace them. It’s like, where can you enhance them?”

The reviewer is the part of this launch most teams will touch first. It works on pull requests that already live on GitHub, and moving a repository is a long project at most enterprises.

The engineering leader who told Reynolds “we’re not” doesn’t need a new Git host to change that answer. He needs something that tells his reviewers which files in a pull request still need a human and which 30 came with a dependency bump.

That’s a much smaller promise than an autonomous delivery lifecycle, but for now, it’s also likely the more useful one.

The post Harness rebuilt its Git repository for nonstop AI agent traffic appeared first on The New Stack.

  •  

[Launched] Generally Available: Playwright Workspaces in Australia East, Japan East, and Switzerland North

Playwright Workspaces in Azure App Testing is now generally available in Switzerland North, Japan East, and Australia East.Playwright Workspaces provides fully managed cloud-hosted browsers for running end-to-end Playwright tests at scale, with parallel e
  •  

AI broke code review. Two experts disagree on what replaces it.

Detective holding up a magnifying lens in front of eye

Ask two experienced engineers about how to handle the flood of AI-generated code in their review queues, and you’ll get two different answers.

The debate remains very much unsettled. And on Tuesday, September 29, two industry leaders will join a live event to hash out what to do.

John Bristowe, Principal Developer Advocate at Octopus Deploy, will join Viktor Farcic, the platform engineering voice behind DevOps Toolkit, for the live conversation we’re calling “Human Review vs. Verified Pipelines: What Catches Bugs in the Age of AI Code.”

REGISTER NOW FOR THIS WEBINAR
By registering, you consent to The New Stack’s Privacy Policy, Terms of Use and to receiving email communication from The New Stack and our event partner. You may opt out at any time.

Here are the facts: Developers have adopted AI en masse. According to the 2026 DORA report, 90% of developers now use AI at work. The result? Developers are merging 98% more pull requests than they managed in the pre-AI era. 

But all that AI-generated code is leaving a mess. Bugs per developer are up 54%, and one analysis of 10,000 developers found that incidents per pull request have climbed a staggering 243%. Octopus Deploy’s own AI Pulse report found that while AI usage enables faster code creation, it can “degrade overall performance” because coding agents write large code updates that humans struggle to fully understand.

Part of the problem is that developers have adopted automated code generation faster than they have adopted automated code review, effectively moving the human bottleneck further down the software creation chain without removing it entirely. And AI code review may have the same shortcomings as the coding agents.

Bristowe argues that code review has quietly become little more than theater. No human reviewer can quickly audit a 40,000-line, agent-created pull request, since they were not part of the reasoning that produced it and cannot realistically understand everything it may change. 

What does Bristowe recommend? Moving the quality gate off the humans’ desks and into the delivery pipeline itself. Does that mean more AI? Not necessarily, with the developer advocate arguing that building robust “policy-as-code” rules into deployment standards can flag only what goes against those policies. Humans can handle those exceptions, without pretending they are “reviewing” the entire package.

Expect Farcic to press Bristowe on how well a policy-as-code setup can truly absorb judgment, and whether we’re simply creating another accountability sink in software development. The conversation will also explore the plight of the junior engineer, who can no longer expect to join a team of humans writing code that other humans review and discuss.

The debate kicks off at 2:30 p.m. Eastern/11:30 a.m. Pacific on Tuesday, September 29. It’s free to attend, and attendees will receive a companion resource built from Octopus Deploy’s AI Pulse data, available immediately for participants who show up live. Register today.

What you’ll take away:

  • Why AI-generated code broke the assumptions code review was built on, and why more review isn’t the fix
  • Why using AI to review AI’s own code doesn’t close the gap (same training data, same blind spots)
  • How to build a pipeline that verifies every deployment against a defined set of rules, no matter who or what wrote the code
  • Where code review still earns its keep, and where it needs to step aside for the pipeline

The post AI broke code review. Two experts disagree on what replaces it. appeared first on The New Stack.

  •  

[Launched] Generally Available: Azure Copilot Observability Agent supports Basic and Auxiliary table plans

Azure Copilot Observability Agent, an AI-powered operational companion in Azure Monitor, is now covering Log Analytics data in Basic and Auxiliary table plans during interactive analysis and deep investigations. Teams can move high-volume telemetry - incl
  •  

[Launched] Generally Available: Azure Monitor Auxiliary Logs Plan in Azure Government and China regions

The Auxiliary table plan in Azure Monitor Logs is a cost-effective option for ingesting and retaining high-volume, verbose logs used for compliance and auditing. Auxiliary Logs is now generally available in the sovereign clouds: Azure Government (Fairfax)
  •  

[Launched] Generally Available: Azure Monitor Auxiliary Logs Plan support for Azure tables and plan switching

The Auxiliary table plan in Azure Monitor Logs gives you a cost-effective way to ingest and retain high-volume, verbose logs that you keep for compliance and auditing but rarely query. Two highly requested capabilities are now generally available.Auxiliar
  •  

Your container runs. Everything around it shouldn’t be your problem.

Abstract digital binary data wave depicting cloud container orchestration and automated infrastructure.

The promise with containers was simple: if it runs on your local machine, it will run in production. And this promise holds – your container runs. But there’s a tax: setting up everything around it. To reach your container, you’ll need a load balancer to route traffic to it, and scaling policies to handle variable traffic. Then come the networking components and the minimally scoped access roles. And somewhere in that checklist, don’t forget the security configuration; unless you’d rather hear about it from your compliance team. At some point, you wonder why you can’t just go back to building.

“You give it a container image. You get a production service. And when your workload outgrows a single container, you don’t outgrow Amazon ECS Express Mode.”

Sound familiar? You’re not alone. Most teams spend cycles before they feel confident deploying containers in production. That’s time from your roadmap spent on decisions that don’t differentiate your business. Somewhere between the third Terraform module and the second IAM policy review, you’ve lost the speed containers were supposed to give you.

Amazon Elastic Container Service (ECS) is the container orchestration engine behind some of the largest production workloads on AWS. But until now, getting started with it meant understanding load balancers, networking, IAM roles, and scaling policies before you shipped anything.

We tried to change that and make it easier for you to get started and stay focused on building when we launched Amazon ECS Express Mode. Express Mode is a new interface into that same engine. You’re not trading power for simplicity. You’re getting a faster door into infrastructure that’s already battle-tested. The vision was clear: keep it simple but extensible. You give it a container image and two IAM roles; you get an HTTPS service running on Fargate with a load balancer, a TLS certificate, autoscaling, and canary deployments. Oh, and all these resources run in your account, where you have full control.

 // Sample Terraform Code
 resource "aws_ecs_express_gateway_service" "frontend_service" {
   execution_role_arn = aws_iam_role.execution.arn
   infrastructure_role_arn = aws_iam_role.infrastructure.arn
 
   primary_container {
     image = "111122223333.dkr.ecr.us-east-1.amazonaws.com/my-service:1.4.2"
   }
 }

What’s in Amazon ECS Express Mode?

Underneath the one-step deploy, this is what you’re getting:

  • No sprawl – Every service sits behind an Application Load Balancer, but doesn’t get its own. Up to 25 Express Mode services within a VPC share a single ALB. Express Mode adds load balancers only when needed and removes them when services are deleted. No dangling resources for you to worry about.
  • Safer rollouts – Out of the box, each deployment is a canary release where 5% of your traffic is routed to the new revision and bakes for 3 minutes, following which the remaining traffic is shifted. If your 4xx/5xx error rate exceeds 1%, an alarm we create for you triggers an automatic rollback.
  • Scaling – Each service ships with an auto-scaling policy that targets 60% CPU utilization, scaling from 1 task to a maximum of 20 by default. If your workloads need to scale on memory or request count, you can configure that too.
  • IaC ready – Create and manage Express services through CloudFormation, CDK, Terraform or GitHub Actions.
  • Complete ownership – The cluster, load balancer, target groups, log groups and other resources are all in your AWS account. You can inspect them, audit events, and modify them directly if needed. Nothing is a black box.

But what about my sidecars? My workloads aren’t that simple.

Sooner or later, as your application evolves, the workload becomes more than just one container. An observability agent needs to run beside the app, publishing traces to your monitoring solution. The base image gets swapped for the hardened one security maintains. The credentials move out of environment variables and into Secrets Manager. This is where extensibility comes into play. 

For this, Express Mode now supports providing a standard ECS task definition – the spec that describes your containers, their resource limits, and how they connect. Your sidecar, image, and credentials all fit right in. If your team runs ECS today, this is the spec you’re already writing. If not, Express Mode generates a task definition in your account, and when requirements arrive, you take that working spec, add what you need, and hand back the ARN. Once you associate a task definition with an Express Mode service, you can continue managing your application either through task definition updates or directly through Express Mode, whichever you prefer.

// Sample CDK snippet
const taskDef = new ecs.FargateTaskDefinition(this, 'TaskDef', {
  cpu: 1024,
  memoryLimitMiB: 2048,
  executionRole,
  taskRole,
});
 
taskDef.addContainer('Main', {
image:
ecs.ContainerImage.fromRegistry('111122223333.dkr.ecr.us-east-1.amaz
onaws.com/my-service:1.4.2'),
  essential: true,
  portMappings: [{ containerPort: 8080, name: 'main' }],
});
 
taskDef.addContainer('otel-collector', {
image:
ecs.ContainerImage.fromRegistry('public. ecr.aws/aws-observability/aw
s-otel-collector:latest'),
  essential: false,
  memoryReservationMiB: 256,
  command: ['--config=/etc/ecs/ecs-default-config.yaml'],
});
 
new ecs.CfnExpressGatewayService(this, 'FrontendService', {
  infrastructureRoleArn: infrastructureRole.roleArn,
  taskDefinitionArn: taskDef.taskDefinitionArn,
});

Wait, am I locked into this?

No. Express Mode is an interface into Amazon ECS, not a walled garden. Every resource it creates is a standard AWS resource in your account, addressable by an ARN. If required, you can modify them directly through their respective AWS APIs in the Console, CLI or SDK.

“Express Mode is an interface into Amazon ECS, not a walled garden.”

If you change the scaling policy, update a security group rule, or swap the task definition, Express honors those modifications on the next update. It does not overwrite changes you make. This means you can start with Express defaults today and customize individual resources as your requirements evolve, without migrating off Express or recreating your service.

Express also exposes the ARN of everything it manages – the load balancer, target groups, security groups, and alarms – through the Describe APIs and as CloudFormation/CDK outputs. You can reference them in your own stacks or hand them to existing constructs.

Why did we design it as a single operation?

We deliberately made Express Mode a single API call, not a multi-step workflow. One input (your image or task definition ARN), one operation, one outcome.

That design choice compounds. Your IaC is under ten lines. Your CI/CD pipeline doesn’t need custom steps. An AI coding agent can deploy and iterate on your behalf because the entire surface is one well-defined spec it already knows how to read and modify.

But there’s an engineering reason too. Because Express Mode owns the full lifecycle of what it creates, it can clean up what it sets up. Delete a service, and the target groups, scaling policies, and alarms go with it. No orphaned infrastructure. And because you didn’t wire these resources together manually, one service’s deploy can’t accidentally touch another’s.

Wrapping it up

As we continue to build on ECS Express Mode, our priority is simple – taking away the undifferentiated heavy lifting from our customers. Amazon ECS Express Mode is our attempt to make good on that: point it at an image, get a production service: load balancer, TLS, autoscaling, canary deployments, all running in your account, where you can see it and change it.

“The configuration grows with your requirements; the operational burden doesn’t.”

And when your workload outgrows a single container, you don’t outgrow Express Mode. Bring your own task definition: the sidecars, the hardened images, the secrets, and keep handing the infrastructure to us. The configuration grows with your requirements; the operational burden doesn’t.

Try it from the AWS Console, Terraform, CloudFormation, CDK, or GitHub Actions – or just ask your agent.

The post Your container runs. Everything around it shouldn’t be your problem. appeared first on The New Stack.

  •  

[Launched] Generally Available: Control plane metrics collection for AKS with Managed Prometheus

Control plane metrics collection for Azure Kubernetes Service (AKS), powered by Azure Monitor Managed Service for Prometheus, is now generally available.This capability gives AKS customers native observability into key managed control plane components, in
  •  

[In preview] Public Preview: Protect sensitive generative AI telemetry in Application Insights and Microsoft Foundry

Azure Monitor Application Insights now stores generative AI content in a dedicated GenAIContent table (AppGenAIContent in Log Analytics), making it possible to apply newly available access controls to sensitive AI telemetry. Because Application Insights p
  •  

[In preview] Public Preview: Advanced platform metrics in Azure Monitor

Starting July 15, 2026, advanced platform metrics are available in Public Preview. This update provides enhanced visibility into platform performance, resource health, and operational trends across supported Azure services. With advanced platform metrics,
  •  

[In preview] Public Preview: Export historical data from Log Analytics workspace with Export jobs

The Log Analytics Export job enables you to export historical data from a Log Analytics workspace to an Azure Storage account based on a specified query and time range. This allows you to extract only the data you need and move it to external systems for
  •  

AWS DevOps Agent adds release management capabilities to assess code changes before production (preview)

Today, we’re announcing a new release management capability in AWS DevOps Agent that is now available in preview. AWS DevOps Agent is your always-available teammate that spans software changes and operations across AWS, multicloud, and on-premises environments. The practice of DevOps aims to make software change and operations smooth and increasingly autonomous, and AWS DevOps Agent delivers on both by leveraging its deep understanding of your environment, your services, their dependencies, and how they behave in production. Already generally available for post-deployment operations, it autonomously investigates incidents, provides root cause analysis and mitigation steps, and delivers targeted recommendations to prevent recurring issues. With today’s preview, AWS DevOps Agent adds release readiness review of code changes and autonomous release testing. These new features verify every change against the natural language standards you give to the DevOps Agent and run change-specific tests in production-like environments. AWS DevOps Agent now supports teams from code creation to production, helping reviewers and testers keep pace with the volume of AI-generated code.

As development teams adopt AI coding tools, the volume of pull requests moving through delivery pipelines has increased faster than review and testing processes can handle. When teams are under pressure to keep up, reviews are approved without thorough examination, and test environments drift from production. The value that coding agents generate sits waiting in review queues instead of reaching end users. At the same time, AI models are increasingly capable of catching functional and security issues that human reviewers might miss under time pressure, making speedy and safe delivery a requirement rather than a tradeoff.

The release readiness review feature evaluates every code change against production requirements, dependency safety, and the standards and best practices you provide to the DevOps Agent. The agent checks cross-repository dependency risks that could affect other services, access control changes against AWS Well-Architected Framework best practices, and compliance with any standards you have defined. When no standards are provided, the agent applies general best practices. As part of the review, the agent also runs your software in an AWS-managed isolated environment, executing lightweight user journey tests to verify the software builds, runs, and passes basic functional checks before the change enters the pipeline. Findings appear in the AWS DevOps Agent console and as comments on pull requests in GitHub or GitLab. You can also invoke reviews directly from your IDE through the Kiro power or Claude Code plugin, so developers can identify and fix dependency risks, standards violations, and access control issues before the change is committed to version control.

The autonomous release testing feature goes further, generating and running change-specific test plans for web and API-based applications in customer-provisioned, production-like environments before the change merges. Rather than running a static test suite, the agent reasons about what the change does and constructs tests tailored to it, covering functional correctness, behavioral regressions, and integration scenarios that a manually maintained test plan might not anticipate. Every test run produces structured artifacts including metrics, logs, traces, and an execution summary, giving reviewers a consistent record of what was tested and what the results were.

Getting started with AWS DevOps Agent release management
This walkthrough shows how to run an on-demand release readiness review using the AWS DevOps Agent web app. Before you begin, confirm that you have at least one GitHub or GitLab repository connected to your Agent Space. Once your repositories are connected, AWS DevOps Agent will index your code and build a knowledge graph of cross-repository and cloud dependencies.

To open the web app, navigate to the AWS DevOps Agent console, select your Agent Space, and choose the Web app tab. Choose Operator access to open the web app.

Without standards configured, the agent applies general best practices. To tailor reviews to your internal standards, navigate to Knowledge, then choose the Instructions tab. You will see a list of instruction sets, each scoped to a specific agent or task. Choose View next to Release readiness review to edit the instructions for production-readiness change review. Write your internal standards in plain English. For example, you can define infrastructure and data standards on encryption or network access rules, best practices that warn without blocking such as logging and observability requirements, and sensitive data classification best practices that identify applications or resources requiring higher security measures. To apply instructions across all agents in your space, choose View next to All agents.

You can trigger a release readiness review in two ways: by submitting a pull request to a connected repository, or by entering an on-demand query in the chat interface. To run an on-demand review from chat, choose New chat and enter a request such as:

Perform a production risk analysis on my repository branch

The agent will ask for the repository and branch you want to analyze. You can provide a branch name, a pull request number, or a commit SHA. Once you confirm your selection, the agent queues the review and analyzes the change for production risks, including infrastructure impacts, configuration changes, and potential issues.

After the review completes, you can ask follow-up questions directly in the chat to explore the findings in more detail. For example, you can ask which downstream consumers a change affects, and the agent will return a structured breakdown of in-repository and cross-repository consumers that will break, the specific files and line numbers affected, and the recommended steps to resolve the issue before deployment.

After submitting a review request, navigate to Changes in the left navigation pane. The Proposed changes table shows each review that has run, including the proposed change description, its source, category, status, and when it was created. You can filter by category or status to find specific reviews, or search by name using the search bar. Choose any entry to open the full execution detail.

The Timeline tab shows the agent’s step-by-step reasoning process, including the tools it called, the dependencies it consulted, and the observations it made at each step. Each entry is timestamped, giving you a complete record of how the agent built its understanding of the change and reached its conclusion.

Choose the Report tab to see the final recommendation. The report opens with a summary header showing the recommended action, the number of critical issues found, the commit revision, and the number of files changed. The recommended action is either BLOCK, Proceed with Caution, or Safe to Release.

Below the summary header, the Analysis section explains why the recommendation was made, citing specific risks and the evidence the agent found to support its conclusion. The Issues section lists each finding by severity, giving you a prioritized view of what needs to be addressed before the change can proceed. The Recommendations section provides specific, actionable steps the developer can take to resolve each issue. Finally, the Changes section lists each file that was modified, with the type of change, the category it falls under, and a description of what was changed, so reviewers have a complete picture of what the change does before it merges.

You can also invoke the autonomous release testing feature directly from the chat interface. To run an autonomous release test on a web or API-based application, choose New chat and enter a query such as:

Run a release test on my application deployed at [application URL]

The agent generates a change-specific test plan and executes it in your provisioned environment. Results appear in Changes, where you can review the execution steps and a structured summary of what was tested.

Get started today
The release readiness review and autonomous release testing features for AWS DevOps Agent are available in preview. These features are available at no additional cost during preview in the US East (N. Virginia) Region. For pricing information on other AWS DevOps Agent features, visit the AWS DevOps Agent pricing page.

For configuration details, visit the AWS DevOps Agent user guide.

— Esra
  •