❌

Vue normale

Reçu avant avant-hier

How to attach an owner to every cloud resource you find

15 septembre 2026 à 16:00
Dark abstract 3D rendering of intertwined rings symbolizing complex cloud resource governance and technical debt.

The engineer who knew why that cloud instance existed has left the company. The instance is still running, the bill keeps growing, and the team must now decide whether it’s safe to shut down. This is a bad time to discover that its ownership history was someone’s memory.

“Good resource governance has three pillars: continuously synced inventory, policy that blocks resources without tagged owners, and an audit trail that survives every reorg.”

A cost review flags an EC2 instance nobody remembers provisioning. Someone scours Slack for the resource ID, finds nothing, burns half a day chasing dead ends, and eventually stumbles across an exhausted engineer who mumbles the mantra that ends most of these investigations: “I think that’s from the project Priya was running before she left.” 

Nobody follows up, because nobody knows how to reach Priya anymore. The instance stays up, because tearing down a fence when you don’t know what it’s protecting you from is a good way to find out the hard way. It’s the platform team equivalent of emotional baggage; they all seem to accumulate some.

Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous. 

“Processing this baggage doesn’t require knowing Priya’s replacement, writing better documentation, or hoping the next reorg is more rigorous.”

Solving the problem requires three things that already exist: queries, policies, and logs; and maybe just a little bit of therapy.

1. The query that tells you what’s missing an owner

CloudQuery’s asset inventory syncs continuously across every provider a team runs on, into tables you can query directly: aws_ec2_instances, gcp_compute_instances, azure_compute_virtual_machines, and so on. Finding every resource without an assigned owner is as easy as:

SQL
SELECT resource_id, 'aws' AS provider, 'ec2_instance' AS resource_type
FROM aws_ec2_instances
WHERE tags ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'gcp', 'compute_instance'
FROM gcp_compute_instances
WHERE labels ->> 'owner' IS NULL
UNION ALL
SELECT resource_id, 'azure', 'virtual_machine'
FROM azure_compute_virtual_machines
WHERE tags ->> 'owner' IS NULL
ORDER BY provider;

Run it regularly, and you have a (hopefully short) boring list to refer to when the CFO asks who deployed an expensive instance; instead of being thrown into a frenzied goose chase at 4:59 p.m. on a Friday.

2. The policy that prevents it recurring

A query tells you what’s already missing an owner; but how do you prevent the next ownerless resource from being deployed? The answer is a policy, and env zero evaluates Open Policy Agent rules against every plan before it applies. A rule that requires an owner tag on every new resource looks something like this:

package env0

# METADATA
# title: require owner tag
# description: A resource can't be created without a declared owner.
deny[format(rego.metadata.rule())] {
	resource := input.resource_changes[_]
	resource.change.actions[_] == "create"
	not resource.change.after.tags.owner
}

format(meta) := meta.description

Add that to the project’s policy set, and a plan that creates a resource without an owner tag doesn’t just get a warning; it doesn’t get created.

3. The record that outlives its creator

A tag tells you who owns something today. It doesn’t tell you anything about who asked for it, why, or who signed off. By the time it matters, the person who could’ve answered from memory may no longer be reachable. An audit entry records that at the moment of creation, instead of reconstructing it afterward from the scraps of recollection scattered around the rest of the team. Here’s an example:

{
  "event": "resource.created",
  "resource_id": "i-0a1b2c3d4e5f",
  "requested_by": "j.chen@company.com",
  "approved_by": "platform-lead@company.com",
  "approval_ref": "ENV-4471",
  "stated_purpose": "load test environment, Q3 capacity planning",
  "timestamp": "2026-08-14T09:12:03Z"
}

This entry answers the question this whole piece opened with, without needing Priya or Slack. env zero keeps this record attached to the resource for as long as the resource exists, specifically so it outlasts any individual’s tenure.

Why this keeps happening

Employee attrition is an age-old challenge that’s only accelerating in the modern era. US private-sector voluntary turnover runs 22 to 25% a year, so a hundred-person org loses twenty-odd people every year, each one taking a small, specific piece of “why this exists” with them. Replacing a mid-level employee costs six to nine months of salary, more than double that for senior specialists. 

“Tags were supposed to survive this. In practice they rot the way everything else does.”

That figure doesn’t touch what the departure does to everyone else’s mental model of what’s actually running. Tags were supposed to survive this. In practice they rot the way everything else does: two teams merge and bring incompatible schemas, provisioning that runs on tribal knowledge accumulates configuration drift for the same reason it accumulates ambiguous ownership, and a resource tagged under a policy that’s been replaced twice since its inception isn’t much better documented than one that isn’t tagged at all.

We wrote in April about the hour it once took our own team to answer, “what are we actually running across both clouds?” That solved a point-in-time problem. The query, the policy, and the audit entry above stop the same story from playing out again next June.

The org chart will keep changing. The record doesn’t have to.

None of this stops people from leaving or teams from reorganizing; pretending otherwise is how platform teams end up rebuilding the same spreadsheet every eighteen months. What changes is whether the next “what are we actually running, and who owns it” conversation takes hours of archaeology across three teams, or is a query that already has the answer attached. Institutional knowledge decays at a fairly predictable rate. A system of record shouldn’t.

The post How to attach an owner to every cloud resource you find appeared first on The New Stack.

Shai-Hulud: Whoever controls your package registry controls your pipeline

31 août 2026 à 18:00
Abstract digital artwork of a constrained data stream on a dark background, representing software pipeline security.

On September 15, 2025, npm’s registry did something unprecedented: Packages began updating themselves. 

No maintainer ran npm publish. No pull request got merged. New versions just materialized, each carrying a hidden passenger that would go on to publish even more versions of more packages, on more machines, with no human involvement whatsoever. Between September 14th and 18th, more than 500 package versions were altered. The worm’s authors had their creation leave a calling card with an intriguing literary sobriquet: every stolen credential was uploaded to a new public GitHub repo named Shai-Hulud, after the untamable apex keystone species of Frank Herbert’s Dune book series.

“What should we learn about trusting infrastructure from a worm that writes and republishes its own malware?”

That wasn’t the end of the story. Two months later, on November 24, a larger variant christened Shai-Hulud 2.0 was able to backdoor 796 packages, move its execution earlier in the install process to render developer triggers irrelevant, and salt the wound on its way out by deleting the user’s home directory if it couldn’t find credentials to steal or a way to spread. By spring of 2026, its offspring, Mini Shai-Hulud, had evolved from hunting generic developer secrets to specifically targeting credentials belonging to Claude, Codex, Cursor, and Gemini, and had come to the logical conclusion that AI coding tools are involved in all the most interesting projects, making them a ripe hunting ground. 

The latest variant, named ChainDrop, appeared a few weeks ago, on August 4, 2026. In less than four hours, it compromised more than 400 packages by riding a legitimate, cryptographically signed release pipeline, which granted each poisoned version a valid SLSA provenance attestation. This allowed them to circumvent the very mechanism built to verify that a package hadn’t been tampered with. Even its command infrastructure was parked inside an Ethereum smart contract, rendering any domain blocklists moot.

Who can you trust?

Every package manager and infrastructure registry runs on the same precarious assumption: installing something means running whatever’s in the package, sight unseen, with whatever permissions the install process has. When publishing required a person to sit down and do it, that assumption was safe enough.

“If publishing can now be automated by non-human agents already in the environment, it is no longer safe at all.”

If publishing can now be automated by non-human agents already in the environment, it is no longer safe at all. In such an environment, the worm doesn’t need to convince anyone of anything. It simply needs one compromised credential and an install script. After that, the registry confers trust and credibility, and every subsequent dependency reinforces it.

What happened

The strategy is brutally simple. A compromised npm package runs a postinstall script (in Shai-Hulud’s case, a single file called bundle.js) that searches the infected machine for anything resembling credentials, including npm and GitHub personal access tokens, AWS or GCP secrets, and whatever it can extract from a cloud instance’s metadata service. It downloads Trufflehog, a legitimate open-source secret-scanning tool, and uses it to confirm the validity of any credentials it finds. So, the silver lining is that you at least get a free security audit (of sorts) out of the infection.

Upon locating a GitHub token, it exfiltrates everything to a new public repo and, for extra credit, makes any private repos it can reach public as well, republished under the original name with a “-migration” suffix appended, perhaps to make it appear as if the victim had requested the move themselves. If it finds an npm token, it calls the registry’s API to list every package the compromised developer maintains, downloads them, injects itself into the postinstall step, bumps the version number, and republishes them. That’s it. No further input required. Palo Alto Networks’ Unit 42 is moderately confident that the malicious script was partly written by an LLM, based on stylistic tells such as code comments and emoji embedded in the payload.

We must not fear. Fear is the mind-killer.

This vulnerability isn’t specific to npm. Trade npm publish for terraform apply, and the mechanics hardly change. A Terraform provider pulled from a public registry, like an npm package pulled from its registry, is code that a publisher’s account was trusted to ship, running with whatever access the machine that requested it holds. 

A CI runner with a standing cloud credential and an open path to the internet is the same target as a maintainer’s laptop with a valid npm token: one a worm can compromise once and then use to spread itself to everything downstream of it, at whatever speed automation allows.

“Attestation tells us where a package came from; it doesn’t say anything about whether the commit should exist at all.”

Up until now, the industry’s answer to this problem was provenance: sign the package, attest to the build, and prove cryptographically that what shipped and what’s reviewed match. ChainDrop is the answer to that answer. It didn’t forge a signature. It compromised a maintainer account with legitimate write access and let that account’s legitimate, signed pipeline build and publish the malware on its behalf. Attestation tells us where a package came from; it doesn’t say anything about whether the commit should exist at all.

The fix

Every dependency should be pinned to an immutable reference. Terraform’s dependency lock file, .terraform.lock.hcl, has a recorded cryptographic checksum for every provider version since Terraform 0.14, operating on a trust-on-first-use model. Once a checksum is recorded, any future terraform init that doesn’t match it fails rather than accepting a swapped-out binary. Modules deserve the same discipline. Pin a module’s source to a full commit SHA, not a branch or even a tag, since tags can be moved or deleted at the origin in a way a commit hash cannot. Checkov’s CKV_TF_1 rule and TFLint’s module-pinned-source check both exist because reviewers keep forgetting this.

Source providers and modules through a registry your organization curates. The public Terraform registry, like the public npm registry, will resolve whatever its maintainers choose to publish. If a maintainer account is compromised, everyone downstream inherits the problem on their next init. Routing provider and module resolution through a private, versioned, allow-listed registry means the potential sources a deployment can pull from are controlled by somebody on your team, not the entire public internet’s worth of Terraform code.

Credentials should be scoped to the deployment, not the developer or runner. Every stage of Shai-Hulud’s lineage depends on a long-lived secret that sits still, whether a token in a .npmrc file or a key in a CI environment variable. Short-lived, per-deployment credentials issued through OIDC, an identity token traded for temporary cloud access at the moment a deployment runs and discarded the moment it finishes, mean there’s no standing secret for a Trufflehog pass to find. A worm can still steal a credential that lives in memory for the ninety seconds an apply takes, but it can’t steal one that was never written down in the first place.

Control what a runner is allowed to talk to on its way out. Every generation of this worm relied on an open outbound path: a webhook endpoint for exfiltration, the GitHub API for persistence, npm’s own registry API for propagation, and, in ChainDrop’s case, an Ethereum RPC endpoint chosen specifically because no domain blocklist could touch it. Restricting a runner’s egress to the handful of domains a deployment actually needs (the registry, the state backend, the cloud API), and nothing else, means any attempt the worm makes to phone home runs straight into a firewall before it ever leaves the building.

Those that can destroy a thing, they control it

A build pipeline isn’t a convenience layer sitting tasteful and demure outside your security boundary. It’s production infrastructure, running unattended, typically wielding more standing privileges than the systems it deploys.

“A build pipeline isn’t a convenience layer sitting tasteful and demure outside your security boundary. It’s production infrastructure.”

Pipelines weren’t built this way intentionally. They ended up like this because until these vulnerabilities were exposed, everyone treated them like plumbing, as if history weren’t replete with examples of plumbing used to evade defenses. Turns out, it’s also a perfect avenue for worms to get to the keys that protect your secrets.

Shai-Hulud figured this out within its first 24 hours. The industry is scrambling to catch up, in progressively more painful installments, paid every couple of months. 

Software supply chains used to get compromised. Now, they get infected; and infections don’t wait; they spread. The humble worm that started as a credential thief has now learned to forge the very cryptographic proof meant to catch it. Its latest prey is the AI tooling teams use to speed up production. An unpinned reference and open registry pull are all it would take for this to spread to Terraform providers and modules. The only remedies are: pin what can be pinned; allow-list what can’t; scope credentials for what’s left. Close all paths out. 

Otherwise, by the time you notice the next worm, it’ll have already spread to everything downstream of the one reference you forgot to pin.

The post Shai-Hulud: Whoever controls your package registry controls your pipeline appeared first on The New Stack.

❌