❌

Vue normale

Reçu avant avant-hierInfra

The cloud reduced operational complexity. But many teams need someone to own it completely.

22 septembre 2026 à 16:00
Dark abstract digital wave with subtle color distortion, representing application lifecycle and cloud native infrastructure management.

For companies running serious production workloads with lean engineering teams, the real test of operational ownership comes at 3 a.m. Not whether the system stays up but whether anyone needs to be awake to make that happen. The page arrives. A service is degrading. The engineer who answers didn’t build this service, doesn’t know what thresholds were set at deploy time, and cannot tell whether the system is healing itself or waiting for a human decision. 

This is the moment that separates services that transferred operational burdens from platforms that merely deferred them. The system that passes the 3 a.m. test already knows what healthy looks like, what to do when healthy stops being true, and how to communicate what happened—because those decisions were made at deploy time, not incident time. The team sleeps through the night because nothing went wrong, but because the response was already determined.

Forty engineers, eight applications, zero dedicated ops

Cloud infrastructure evolved in two directions simultaneously, and neither arrived where mid-market teams actually stand. On one end: full control. Infrastructure-as-code, service meshes, custom pipelines. Powerful, flexible, and designed for organizations that have deliberately invested in operational staff who can absorb the cost of that flexibility. On the other end: single-app simplicity. Push code, get a URL. Elegant for a first deployment, but architecturally limited the moment a team manages more than one service, needs compliance controls, or inherits an application that doesn’t fit the platform’s opinions.

You know this company. Forty engineers. Eight production applications. Two generate 80% of revenue. One SRE who is actually a senior developer with an on-call rotation nobody else wants. A compliance audit due in Q3 that nobody has started preparing for.

“The right investment is shipping product. The cost of this gap is measured in what these teams do not ship.”

Every engineer is a full-stack contributor. The person who wrote the feature deploys it, monitors it, and gets the page when it breaks not because the team lacks sophistication, but because hiring dedicated infrastructure staff is not the right investment at their stage. The right investment is shipping product. The cost of this gap is measured in what these teams do not ship. Every sprint spent upgrading the deployment pipeline is a sprint without the feature a customer asked for. Every 3 a.m. page answered by a developer with a product standup at 9 a.m. is diminished output that never shows up in any dashboard. The gap is widening.

The applications nobody planned to operate

Not every application in a portfolio was built by the team now responsible for it. Many companies acquire products through M&A. They inherit internal tools built by engineers who left three years ago. They run commercial off-the-shelf applications customized beyond vendor support. They maintain line-of-business applications in languages nobody on the current team chose. These applications share a common trait: they are in production, they serve customers or meet compliance requirements, and nobody has the budget or the mandate to rewrite them. They need a home that accepts them as they are, not as a modernization roadmap says they should become.

“They need a home that accepts them as they are, not as a modernization roadmap says they should become.”

This is where the full lifecycle vision matters. An application management service that only serves net-new applications forces teams to maintain two operational models: one for the applications they are building today, and another for the applications they inherited yesterday. That split is where drift starts, where patching falls behind, and where audit findings accumulate.

The application management service that solves this problem must accept the full range: the Java application packaged as a WAR file, the .NET Framework service running on Windows, the Python application with dependencies pinned to a specific runtime, the containerized service already running elsewhere, and apply the same operational model, the same deployment interface, the same patching and scaling behavior to all of them. Migrate, manage, and modernize within the same experience, without requiring a different operational posture for each stage of the application lifecycle.

What the market data actually shows

Janakiram MSV, an analyst and advisor on cloud-native platforms and TNS contributor, puts it this way:

Cloud native standardized the infrastructure layer around containers and Kubernetes. But it never standardized the operational boundary between the application team and the infrastructure underneath it, and platform engineering largely emerged to re-establish that boundary. CNCF’s latest research with SlashData reports that 28 percent of organizations run a dedicated platform engineering team and 41% split those capabilities across multiple teams. Another 3% have no formal approach at all, and that is where most mid-market engineering organizations exist. They want the same outcome as platform engineering without first becoming a platform engineering organization.

The pattern that repeats is that maturity stalls at the third application rather than the first deployment. A forty-engineer team gets one service into production and keeps it healthy through familiarity. Then an acquisition brings in a .NET workload, an internal tool developed by an engineer who left two years ago, and the team is carrying three operational models with nobody owning the mandate to reconcile them. Hiring will not close that gap, because what is missing is a standardized operational posture, not headcount.

What hundreds of thousands of production deployments actually reveal

We have visibility into hundreds of thousands of production deployments across thousands of customers. That scale does not tell you what people say they want. It tells you what actually breaks, what gets escalated at 3 a.m., and what determines whether a team trusts their platform enough to stop thinking about it. The patterns are remarkably consistent.

1. The deployment that nobody touches again

A team spends two full days getting a Spring Boot application deployed with CI/CD and an SSL certificate. The deploy works. Then nobody touches it for three months because it is stable, but because touching it might break it again. They discover the problem when customers report the application is unreachable. Four hours of forensics follow. What changed? Why? How do we prevent this?

“A release that can partially succeed is a release that will partially fail.”

This is the most common failure mode we observe: not the deployment that fails loudly, but the deployment that succeeds quietly and degrades invisibly. The team cannot explain what happened because the platform did not decide what “healthy” meant before the incident arrived.

A release that can partially succeed is a release that will partially fail. The platforms that earn trust are the ones where a deployment either completes fully or reverses entirely no intermediate states, no manual rollback procedures discovered under pressure.

2. The observability sprint that ships three weeks late

A developer notices response times degrading. They want memory utilization metrics. The platform does not collect them by default. They spend a sprint writing configuration files to install a monitoring agent across every instance. The insight they needed three weeks ago ships three weeks late.

This is the second most common pattern: observability treated as an add-on rather than a default. Every team we have observed that adds monitoring after their first incident wishes they had it before. The teams that never experience this problem are the ones whose platforms shipped metrics, traces, and health signals at deploy time without code changes, without configuration, without a sprint spent on plumbing. The distinction matters: a platform that can be observed is not the same as a platform that is observed from the moment it goes live.

3. The portfolio tax

Each application gets its own infrastructure. Costs scale linearly with the portfolio. The team running eight applications pays eight times the overhead of the team running one, not because each application needs dedicated resources, but because the platform’s architecture assumes isolation rather than shared operational responsibility. The result: teams avoid migrating inherited applications because the cost model punishes breadth. The compliance audit does not care that the inherited application runs on a different operational model. It expects the same governance, patching cadence, and access controls.

A platform that rewards portfolio growth, shared infrastructure, and a consistent operational posture —and economics that improve with breadth rather than degrade—changes the calculus for teams managing applications they did not build.

4. The security configuration nobody made 

The forty-engineer company has no security team. They have a senior developer who reads the CIS benchmarks on weekends. The compliance audit arrives regardless. Teams without specialists will not configure security controls that require specialist configuration. This is not a criticism of those teams; it is a structural observation about how security actually gets implemented (or does not) in organizations where every engineer is a full-stack contributor with a product backlog that never shrinks.

The only security posture that works reliably for these teams is the one they inherit by default: compliance certifications, network isolation, access controls that ship with the platform rather than requiring a dedicated sprint to implement.

What these patterns demand today

Every failure mode we observed the deployment nobody touches, the observability sprint that ships late, the portfolio tax, the security configuration nobody made- shares a root cause: the platform asked the team to make an operational decision, and the team either made the wrong one or made none at all. The response to these patterns required starting from a different question. Not “what should we configure for the team?” but “what should the team never need to decide?”

A platform that answers that question correctly holds a specific set of commitments. It decides what healthy looks like before the first request arrives, not after the first incident. It ships observability at deploy time, not as a sprint the team schedules after something breaks. It treats the eighth application in a portfolio the same as the first, same operational model, same governance, same economics. It inherits security posture by default, because the teams it serves will never staff a dedicated security function.

These are not feature decisions. They are architectural decisions about where operational responsibility permanently resides. The team defines the application source code, a Dockerfile, a pre-built image, or an existing workload being migrated in. From that point forward, the platform owns everything underneath it. Not for the first deploy. For the life of the application. Patching, scaling, healing, certificate rotation, capacity planning, health evaluation. These are not capabilities the team enables. They are responsibilities the platform holds permanently.

This is what we rebuilt AWS Elastic Beanstalk to be. Not a deployment tool. Not a hosting layer. An application management service that takes operational responsibility for everything underneath the application. The architecture now starts from the question above and refuses to let the answer drift back toward the team over time. Elastic Beanstalk operates in two modes, a structural change from its previous single-environment architecture:

Standard Mode delivers full operational ownership for individual applications and Windows/.NET Framework workloads: the complete operational stack, owned outright, for a single service.

Cluster Mode extends the same ownership model across the portfolio, shared infrastructure, source-to-production deployment that transforms code into running applications, and economics that improve as the portfolio grows. The eighth application shares operational overhead with the first seven rather than duplicating it. For the forty-engineer company running eight production applications today and inheriting ten more next quarter, this is the difference between a platform that covers the portfolio and a platform that covers only the applications simple enough to fit its opinions.

The industry convergence

The distinction is real, though I would not draw it as a line between platforms that reduce complexity and platforms that own operations permanently. Every vendor in this market absorbs some operational responsibility at deploy time. The key question is how much of it returns to the team during an incident, a patch cycle, and an audit. A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.

“A platform that removes infrastructure management from developers during the workweek and reintroduces it at 3 a.m. Sunday addresses only half the challenge.”

At convergence, the direction is correct, but the shape is incorrect. This is not two camps meeting in the middle. Gartner’s 2026 Magic Quadrant for Cloud Native Application Platforms places AWS, Microsoft, Google, and Red Hat in the leaders quadrant, with Render, Netlify, and Upsun as niche players. Vendors specializing in developer experience showed this category is viable, and now the hyperscalers are adopting it. Since source-code-to-URL mapping is now standard across the entire quadrant, the key differentiation becomes who bears operational liability for the eighth application three years after its release.

The only aspect of framing I would challenge is the idea that a platform determines everything the team never has to decide. Routine infrastructure decisions should stay out of the developer’s path, and escape hatches should stay in place for the teams that genuinely need them. A platform that removes choice altogether will demo well and then stall when teams migrate applications that don’t fit its opinions.

The 3 a.m. test that actually matters

For teams already living this reality — serious production, lean staff, growing portfolios — nobody planned to operate the 3 a.m. test; it is not a nice-to-have. It is the evaluation criterion.

The platforms that define the next decade will not simply make deployment easier. They will decide, in advance, how production systems should behave when things inevitably go wrong. Because by 3 a.m., the time for deciding has already passed. The CNAP category was built to describe platforms that own the application lifecycle.

Elastic Beanstalk made those decisions before the incident arrived: what healthy looks like, what to do when it stops being true, how to communicate what happened. The cloud gave teams power. These teams needed someone to stay. Elastic Beanstalk stays.

Check the latest AWS release notes for Elastic Beanstalk here.

The post The cloud reduced operational complexity. But many teams need someone to own it completely. appeared first on The New Stack.

How AWS Lambda logs every flow across thousands of microVMs per host with eBPF and Rust

11 septembre 2026 à 14:00
Abstract dark geometric corridor with glowing purple and blue neon lines, representing high-density network flow architecture in AWS Lambda.

On any compute platform, when a security alert fires, the question is always the same. Which workload talked to that endpoint, when, and how much? So, you dig through logs and hope they hold up. The hard part is knowing what to look for and where to find it. On a single server, thousands of microVMs run for a few hundred milliseconds, then shut down. 

Each one belongs to a different customer running a different function. Within milliseconds, the workload is done, and the logs captured during those few milliseconds are the only witness left.

“Within milliseconds, the workload is done, and the logs captured during those few milliseconds are the only witness left.”

That was our situation at AWS Lambda. We needed a complete network ledger for every tenant workload, no matter how short its life or how much traffic it generates. This is the story of how we replaced an aging capture system with a purpose-built pipeline written in eBPF and Rust: why the old architecture ran out of road, and the decisions that let the new one hold up at Lambda’s scale.

The job: a record you’re not allowed to touch

Every Lambda worker is a bare-metal EC2 instance packed with microVMs, each an isolated Firecracker guest. They all talk over the network to S3, other AWS services, the public internet, and back into a customer’s VPC. Something has to keep an honest, complete account of it all.

That is where a network flow log comes in. A network flow log helps with investigation, incident response, audit, and reconstructing what happened with a workload. Imagine it as the system of record for what happened to a network packet flowing across the system. The same records also feed into network usage and metering services that demand accuracy above all else. Lastly, all records must be persisted for audit and compliance purposes.

Two properties matter more than the rest. The record must be complete and correctly attributed, and capturing it must add almost no overhead to both the network flow and the platform.

Correct attribution means every packet/flow is associated with the microVM where it landed and the tenant that produced it. Completeness means no missed packets. A missing or misattributed record can cause billing, observability, and monitoring problems at scale. Imagine this happening for every microVM when Lambda serves millions of requests per second. Overhead matters for a reason that isn’t readily visible to customers but plays a major role in how you run a service at Lambda’s scale. Every extra megabyte of RAM and every microsecond of CPU consumed adds up at Lambda’s density, lowering utilization, operating margin, and the ability to serve requests under load.

Why the old way ran out of road

While building the new multi-tenant, Firecracker microVM-powered Lambda, we started with a system borrowed from the older, less-dense single-tenant EC2-era design. It had two parts: a kernel-side extension that counted packets and matched them to tenants and sandboxes per flow, and a userspace daemon that read those counters, batched them into records, serialized them, and frequently uploaded the files. This solution worked well for a small number of VMs. However, the solution broke at Lambda’s density for two mutually exclusive reasons.

First, rule explosion. iptables walks its rules more or less linearly for every packet, and each new microVM piled more rules onto the chain. One worker running a couple of thousand microVMs needed well over a hundred thousand iptables rules just to keep the record. Every packet paid a tax proportional to how crowded the host was and what slot it received. So a packet’s bookkeeping slowed as the host got busier, which is exactly the wrong direction, since the whole plan was to pack more microVMs onto a host, not fewer.

“A record that can’t see half the address space isn’t one you can trust, and the moment dual-stack IPv6 support was proposed for Lambda, the old approach was finished.”

The second reason was harder: the borrowed kernel module did not speak IPv6. A record that can’t see half the address space isn’t one you can trust, and the moment dual-stack IPv6 support was proposed for Lambda, the old approach was finished, no matter the performance improvements.

Design considerations

As we embarked on the rewrite, we had a few non-negotiable design considerations: correct attribution, meaningful overhead reduction, and support for IPv6 as a first-class citizen. Dozens of internal systems already knew how to read the Amazon Ion records the old daemon produced. If we could emit byte-for-byte identical records, we could rip out the whole capture engine and no consumer, flow-log, or metering would notice the swap.

To solve for correct attribution, we implemented isolation at the microVM level. Each microVM already had its own network: a namespace with its own virtual devices. That network is the unit everything else organizes around.

The shape of the new system

The replacement is three cooperating pieces. We’ll give each a plain name based on their role.

 ONE WORKER HOST

   control plane
       |
       |  gRPC over a Unix domain socket
       v
   orchestrator ....... one privileged process per host
       |                (loads eBPF, configures TC,
       |                 spawns one tagger per network)
       v
   --- per network (one microVM) ------------------------------

   eBPF capture  -->  ring buffer  -->  tagger  -->  Ion records
   TC hooks on        one per           Rust,
   the network's      network           userspace
   devices

   ------------------------------------------------------------
       |
       v
   same billing + flow-log pipeline as before

At the bottom sits the kernel capture layer. It’s a set of small eBPF programs attached to the traffic-control(tc) hook on each network’s relevant virtual Linux devices. They intercept packets and emit one compact event per packet into a ring buffer. They just watch. Nothing in these programs can copy, block, drop, or rewrite a packet; there’s literally no code path for it.

In the middle is the tagger: an unprivileged userspace process written in Rust, one per network. It drains its own dedicated ring buffer, rolls the raw per-packet events up into per-flow records, and writes them to disk in the legacy Amazon Ion format. On top is the orchestrator, one privileged process per host. It owns everything that needs elevated permissions: loading the eBPF programs, wiring up traffic control, spawning and supervising the fleet of taggers. It exposes a small lifecycle API over a Unix domain socket so the control plane can create, assign, recycle, and tear down tagging as microVMs come and go.

The decoupled design lets capture happen in the kernel, aggregation in userspace, and one process per host orchestrates it all. The captured records go to the existing downstream consumers unchanged.

Capturing in the kernel, without getting in the way

The capture programs attach to the clsact qdisc in traffic control(tc), on both the ingress and egress side of each of a network’s devices. A network spans two devices, so that’s four attach points per network. Traffic control is a good place to stand. It sees every packet early, before anything downstream has touched it. Every eBPF program reads the packet and returns the “keep going” action. We never drop or modify a customer’s packet, and nothing we do adds meaningful latency.

For each packet, the program walks the headers- Ethernet, then IPv4 or IPv6, then TCP, UDP, or ICMP and writes one fixed-size event into that network’s BPF ring buffer. The event is small on purpose, about two dozen bytes for IPv4. A simplified version looks like this:

/* one event per packet, ~24 bytes for IPv4 */
 struct flow_event {
     u8  ip_version;       /* 4 or 6 */
     u8  protocol;         /* TCP / UDP / ICMP */
     u8  direction;        /* ingress or egress */
     u8  device_id;        /* which of the network's devices */
     u16 local_port;       /* "local" is always the sandbox side */
     u16 remote_port;
     u32 flags_and_bytes;  /* TCP flags in bits [31:24], byte count in [23:0] */
     u32 local_addr;       /* 16 bytes for IPv6 */
     u32 remote_addr;
     u32 received_time_ms;
 };

One packet, one event, one byte count. We don’t count packets in the kernel. The flags and the byte count share a single 32-bit word: eight bits of TCP flags on top, a 24-bit byte count underneath. Doing less work per packet in the kernel is the entire point, so aggregation is somebody else’s job.

The local and remote fields get normalized by direction before the event ever leaves the kernel. Arriving or departing, local always means the sandbox side and remote means the outside world. That one small normalization means userspace never has to reason about direction when it groups flows, and the record reads the same way regardless of which way the packet was headed.

“Doing less work per packet in the kernel is the entire point, so aggregation is somebody else’s job.”

Let’s look at how eBPF plays a key role in simplifying the architecture. The eBPF program here works in a reserve-then-commit order. Each eBPF program maintains a dedicated ring buffer to record packet metadata. You ask the ring buffer for space, add the event in place, and submit. This avoids copying data through a syscall, consumes minimal CPU, and needs no per-CPU bookkeeping. Just a single consumer that drains it in order.

However, all this complex logic for parsing and event submission past the eBPF verifier took work. The eBPF verifier has a critical job: proving a program is safe before the kernel can load it, and that requires strict memory bounds checks and instruction limits. We had to make sure it could verify and load programs quickly, so we did the following: First, we made the header parser a shared subroutine, so the verifier proves it once instead of re-proving it inline at every attach point in the eBPF program. Then, we bounded the IPv6 extension-header walk to a fixed number of hops so the verifier can ensure it terminates. We also had to make sure packet fragments past the first IPv4 fragment report zero ports and flags rather than vending garbage into the ring-buffer events. We also had to make sure any coalesced super-packets from segmentation and receive offload (GSO/GRO) have their byte counts handled correctly. A garbage port or an over-counted byte would create a false log entry, so we keep the record honest at the source and byte-identical with the old system.

However, we don’t trust the verifier as the sole source of correctness, because this code produces a record critical to billing, compliance, and auditing systems; each eBPF program is written in C and runs through a formal model checker (CBMC) during every build. Its harnesses assert, among other things, that the event struct’s byte layout stays compatible with what every attached program expects. A struct that silently shifts by a byte is the kind of bug that quietly corrupts every record it touches, and nobody notices until the day they need the log.

Sizing the ring buffer from first principles

The ring buffer is the one thing the kernel producer and the userspace consumer share, and its size is a real tradeoff. Too small and you drop events under a burst, which means missing records, which is a hole in the log right when traffic matters. Too large and you waste memory, and you pay for that waste on every ring buffer on the host.

So we didn’t guess. We derived the floor from each microVM’s packet rate; for example, if the ceiling is 100,000 packets per second per direction. We drain the ring roughly every 100 milliseconds. Multiply the peak rate by the drain interval, by the event size, and by two directions:

ring bytes =~ 62,500 pps x 0.1 s x ~24 bytes x 2 directions
            =~ 300 KB

The ring buffer API requires a power of two, so the design defaults to 512 KiB. That’s the smallest buffer that can’t overflow between drains at the guest’s own maximum packet rate. Put another way, the floor ensures a guest can’t outrun the recorder, even when it’s trying to. The size is configurable per network. In the running deployment, we currently provision it more generously than that floor, on the order of a couple of megabytes, while we finish tuning the right per-workload value. The number to defend is the floor, and the floor comes from a hard system limit, not a guess.

The drain cadence has another nice property: waking a userspace process isn’t free, and a fleet of thousands of processes all waking constantly would thrash the CPU. So the kernel decides when to bother. It checks how full the ring is and only forces a wakeup once the ring crosses about one percent full. Below that, it stays quiet and lets events pile up. Userspace, on its end, won’t come back to read more than once every 100 milliseconds. A quiet flow just sits there until the next drain, basically free. A busy one trips that one-percent threshold and gets read almost right away. Nothing’s on a fixed timer, so neither case gets the timing wrong.

Draining and aggregating in Rust

The tagger turns raw per-packet events into the per-flow records the pipeline stores. One tagger per network, running unprivileged.

We picked Rust for boring, practical reasons. At this density, thousands of these processes run on a host, each holding a small amount of state that has to be correct. A garbage-collected runtime would give us pause times and memory that balloon under load, and a pause in the wrong place could create a gap in the record. Rust gives us predictable memory and no collector, plus a compiler that flat-out refuses to build whole categories of bugs that turn into misattribution. Each tagger runs in a few hundred kilobytes of RAM against a budget of about a megabyte. That’s what makes thousands of them per host affordable.

Inside, it’s a small set of cooperating tasks on a single-threaded async runtime. One task reads the ring, another owns the flow state, and a third writes parcels. The only work we fence off onto a blocking pool is the couple of operations that genuinely block: receiving the ring descriptor and serializing Ion, since the Ion writer isn’t async. It reads the ring through epoll, so it sleeps when there’s nothing to do and wakes when there is.

As events arrive, the tagger drops them into a flow map keyed by device, the five-tuple, and a tenant attribution handle. That handle comes from the metadata the control plane handed over when the flow was activated. Matching events accumulate bytes, packet counts, and OR’d TCP flags. Grouping happens as the events are read, so the hot path stays a lookup and an add.

Where does attribution actually come from? That’s because mapping is the most important property for the record. The kernel event carries no identity, and it doesn’t need to. Every network has its own dedicated ring and devices, so packets are separated long before the tagger sees them. The tagger isn’t pulling one tenant’s packets out of some shared firehose. The stream it reads was only ever that one tenant’s, because the ring and the devices feeding it belong to that tenant alone.

Once a minute, on a fixed interval lined up to the top of the second to match the old system it replaced, the tagger serializes the completed flows into Amazon Ion records in exactly the schema the old daemon produced. Each file gets written to a temporary name, flushed to disk, and renamed into place. A reader sees a complete record or nothing, never a torn one. A separate flush loop, with a little random jitter at startup so thousands of processes don’t all write at the same instant, drains completed flows even after a microVM has gone quiet. A workload that goes silent still leaves a finished record behind it.

Because the records are byte-compatible with the old format, the entire downstream world kept working untouched. And when we ran the two systems side by side, we could compare their output record for record and confirm they agreed. That’s about as direct a completeness check as you can get.

Least privilege, enforced by a file descriptor

This is the part of the design I’m most fond of, because it uses an old Unix trick to get a strong security property for almost nothing.

The processes that do the actual packet work, the thousands of taggers, hold no elevated privileges. They can’t load eBPF or touch traffic control. They can’t even open the ring buffer map on their own. All of that power lives in one place: the per-host orchestrator, and even it runs with just the two capabilities it needs rather than as root.

“Passing a file descriptor over a socket is a decades-old Unix feature, and it lets us keep thousands of processes powerless while concentrating privilege in one small place.”

So how does an unprivileged tagger read a ring buffer it isn’t allowed to open? The orchestrator opens it and hands the open file descriptor to the tagger over a Unix domain socket, using the kernel’s SCM_RIGHTS mechanism to pass descriptors between processes. The tagger gets a ready-to-use handle to the ring and nothing else. It never had, and never needs, permission to create one. The privileged surface of the whole system is one small process per host. The thousands of processes touching customer traffic are about as powerless as we can make them.

A lifecycle API, and why it has two doors

MicroVMs come and go constantly, so the control plane needs a way to tell the orchestrator when to start and stop recording a network. It does that through gRPC APIs over a Unix domain socket, with a handful of methods: create a set of flows, activate a flow, recycle one, tear one down, plus a health check.

Starting to record is split into two calls, a heavy one and a light one, and the split is deliberate. Create is a heavy call, the expensive path: it loads and attaches the eBPF programs, configures traffic control, and spawns the tagger. Attaching to network devices takes a kernel lock that every such operation on the host contends for, so when a host is standing up many networks at once, we batch these to keep everyone from serializing behind that one lock. Activate is the lighter call. By the time it runs, the machinery already exists, so it just hands over the customer metadata and flips the flow into steady-state recording. Its latency budget is tight: under 2 milliseconds at p90, under 10 milliseconds at p99.9, matching the baseline of the system it replaced.

An honest tradeoff: kill it, or reuse it

Not every decision came out clean. These are lessons for anyone building something similar.

The original design had a strict rule for recycling a network: always destroy the tagger and spawn a fresh one. From a correctness standpoint, the reasoning was airtight. A brand-new process can’t carry stale metadata from a previous tenant, so a flow from one tenant landing in another’s record across a recycle becomes structurally impossible. Kill it, don’t try to clean it. We were sure that was the right call.

Reality under Production workloads showed us that forking and exec’ing a new process thousands of times as networks churned turned into a real source of CPU spikes at scale. The safest choice showed up as a flame graph. So the shipped system needed a new knob. A workload that reuses its networks can reuse the tagger after a recycle, trading off a little of that structural guarantee for a lot less CPU churn. A workload that wants the strict, cross-tenant-proof behavior leaves the knob off. I still think the strict version is the more correct design, but the fleet’s CPU budget just didn’t allow it.

What it bought us

Start with the number that killed the old design. One host needed more than a hundred thousand firewall rules to keep the record for two thousand microVMs, and each additional microVM piled on more, taxing every packet a little further. The eBPF version swaps that linear rule walk for constant-time map lookups whose cost doesn’t climb as the host fills up. The linear tax is gone. That puts the density target, roughly double the microVMs per host, within reach, and without the burst-time gap a per-packet tax invites.

The rest of the payoff falls out of the constraints we started with:

  1. IPv6 flows, invisible to the old tool, get recorded like anything else, so the log covers the whole address space instead of half.
  2. Each tagger lives in a few hundred kilobytes of RAM against a roughly one-megabyte budget, small enough that thousands per host is practical.
  3. Activating a flow into steady-state recording stays under 2 milliseconds at p90 and under 10 milliseconds at p99.9.
  4. The capture layer is observe-only and formally checked, and the processes touching customer traffic hold no privileges. We added significant visibility while shrinking the trusted, privileged surface that could corrupt the record. 
  5. The output records are byte-for-byte identical to the old format, so every downstream flow-log and metering consumer kept working with no change

Lessons worth carrying to other systems

A few of these travel well beyond Lambda. Kubernetes pods, edge runtimes, the sandboxes people are spinning up now to run AI agents. Anywhere you’ve got many tenants sharing a host and need a trustworthy record of their traffic, the same shapes hold.

First: observe from outside the hot path. The moment your recording logic sits inline in packet forwarding, its cost becomes a tax on every packet, and that tax is heaviest right when the record matters most. That same pressure pushes teams to drop or sample data, and an audit trail can’t survive sampling. eBPF lets you watch from the side and emit a compact event, while everything expensive happens elsewhere.

Second: size buffers from something real. A buffer sized by an actual rate limit times an actual drain interval is a number you can defend in a review, and it’s what lets you promise no dropped events under a burst. We could’ve picked 512 KB because it felt about right, and it probably would’ve held most of the time, right up until some burst it wasn’t sized for.

Third: keeping tenants apart at the point of capture is the part I’d argue hardest for. Give each tenant its own ring and its own devices, and the streams never touch, so you’re labeling clean traffic instead of guessing after the fact.

Fourth: old primitives are underrated. Passing a file descriptor over a socket is a decades-old Unix feature, and it lets us keep thousands of processes powerless while concentrating privilege in one small place.

Last one, and it’s the cheapest to get wrong: when you swap out an engine, keep the bolt pattern. Byte-for-byte identical output let us replace the entire capture path with zero downstream migration, and it handed us a record-for-record way to prove the new system saw everything the old one did.

The post How AWS Lambda logs every flow across thousands of microVMs per host with eBPF and Rust appeared first on The New Stack.

Cut GPU inference cold start from 8 minutes to less than a minute

3 septembre 2026 à 20:30

We instrumented the full path from pod creation to first inference response on a GPU node running a 70B-class model. Eight minutes. Six sequential phases. We expected one bottleneck. We found six, and which one dominates depends on model size.

For a 64 GB model, 65% of the startup time is spent recompiling CUDA kernels that produce identical output every time. For a 203 GB model, 92% of the time is spent downloading weights from S3 through a calling pattern that leaves 98% of available bandwidth idle. Both are fixable with configuration changes. Neither is fixed by default.

“Eight minutes. Six sequential phases. We expected one bottleneck. We found six.”

We define time to first token served (TTFTS) as the wall-clock duration from pod creation to the first inference response leaving the GPU. Not time to first token (TTFT), which measures per-request latency once the model is warm. TTFTS is the one-time startup tax. TTFT begins where TTFTS ends.

Here’s what we achieved:

ScenarioDescriptionBeforeAfterReduction
Pod restart on warm nodeWeights loading + compilation on existing node1.5-8 minunder 30s80-93%
New node from scratchFresh node provisioned, nothing cached8-15 min~5 min40-65%

The warm-node row is what you pay on every pod restart: scale-up events, rolling updates, OOM recoveries. That’s the 80-93% win, and it requires only configuration changes. The cold-node row includes ~2 minutes of fixed infrastructure cost (node provisioning and framework initialization) that no application-layer optimization can remove. The rest is avoidable waste that we eliminated through platform and configuration fixes. The warm-node optimizations are environment variables and a volume mount that work on any Kubernetes cluster. The cold-node optimizations require EKS Auto Mode, which comes pre-configured with pre-compiled NVIDIA drivers, SOCI (Seekable OCI) parallel image pull, and NVMe instance store mounting.

All model startup measurements were taken on p5.48xlarge instances running Amazon EKS Auto Mode, with S3 traffic routed directly (bypassing the NAT Gateway) and container images in a private Amazon ECR repository (same region as compute). Model startup improvement ratios (80-93%) hold consistently across instance types (validated on P-family and G-family). Cold-node times vary with network bandwidth and CPU count. For the weights loading and compilation cache configuration, see Accelerate model loading on Amazon EKS.

The Kubernetes ecosystem has made real progress on the inference stack in 2026. OCI image volumes are now stable for model delivery. Dynamic Resource Allocation (DRA) gives GPUs structured attributes instead of opaque integer counts and provides flexibility in allocating GPUs to workloads. Gateway API has inference-aware routing extensions. But none of these primitives address the full cold-start stack: the six layers between “pod pending” and “first token served,” each with its own bottleneck and its own fix.

The six layers of cold start

When a new inference pod starts on a freshly provisioned GPU node, it passes through six distinct phases before serving its first request:

  1. Node provisioning. Karpenter launches an EC2 instance, boots it, and registers it with the Kubernetes API server (~60-90s).
  2. GPU driver initialization. The driver kernel module must load and expose accelerator devices.
  3. Container image pull. The inference engine image (8-12 GB compressed) must be transferred to the node and extracted.
  4. Model weights download. The model files must stream from object storage into GPU memory.
  5. GPU kernel compilation. torch.compile traces the model graph and generates optimized CUDA kernels.
  6. Engine initialization. CUDA graph capture, KV cache profiling, and HTTP server startup (30-120s depending on whether compilation is cached).

Each layer has a different bottleneck, a different fix, and a different owner.

Which layer dominates depends on model size

Before diving into each layer, one finding shaped every decision we made: the bottleneck is not fixed.

We instrumented the model startup path (layers 4 and 5) and measured each phase independently for two model sizes:

64 GB model (Qwen3.6-35B-A3B):

  • Weights loading: ~29s (35% of model startup)
  • torch.compile: ~53s (65% of model startup)

203 GB model (Llama-4-Scout, TP=4 where TP is tensor parallelism, splitting the model across GPUs):

  • Weights loading: ~423s (92% of model startup)
  • torch.compile: ~34s (8% of model startup)

For models under ~100 GB, compilation dominates. For larger models, network transfer dominates. torch.compile time stays roughly constant (it depends on graph complexity, not parameter count). Weights loading scales linearly with file size.

“For models under ~100 GB, compilation dominates. For larger models, network transfer dominates.”

This means any single-layer optimization has a ceiling.

Layer 1: Node provisioning

On EKS Auto Mode and Karpenter-managed clusters, node provisioning takes approximately 60-90 seconds for accelerated instances from pod pending to node Ready. Karpenter calls the EC2 Fleet API directly and reacts to pending pods within seconds, keeping provisioning at the EC2 launch floor.

Layer 2: GPU driver initialization

The NVIDIA GPU Operator in its default configuration adds 2-3 minutes to node boot while it compiles the driver kernel module from source. This cost repeats on every new node.

When the platform controls the full stack (OS image, kernel version, driver version, boot sequence) it can pre-compile driver kernel modules at image build time. The node boots, runs modprobe to load an already-compiled .ko file, and the GPU is ready in seconds.

This matters more now than it used to. Blackwell-architecture GPUs (G7, G7e instances) require NVIDIA’s open-source kernel modules exclusively. Older Maxwell/Pascal/Volta GPUs can only run proprietary modules. A cluster with both legacy and next-gen GPU nodes needs different drivers, different AMIs, different upgrade cycles. A managed platform that pre-compiles the correct module per instance family eliminates this complexity.

On EKS Auto Mode, the GPU driver loads in seconds (pre-compiled at image build time), compared to the 2-3 minutes a runtime-compilation approach requires.

Layer 3: Container image pull

A production vLLM or SGLang inference image is typically 8-12 GB compressed. Standard containerd pulls layers sequentially, decompresses them one by one in memory, and writes them to disk. At this size, sequential pull takes 2-4 minutes on a cold node depending on instance type and available CPU cores. For larger custom images (30-50 GB compressed), containerd can run out of memory entirely during decompression.

EKS Auto Mode uses SOCI’s parallel pull mode, which replaces containerd’s default snapshotter. The SOCI snapshotter downloads layer chunks concurrently via HTTP range requests and writes each chunk directly to its target byte position on disk (no in-memory ordering buffer). Decompression runs in parallel across all available CPU cores.

Pull time is bottlenecked by CPU-bound decompression, not network bandwidth. We confirmed this directly: a p4d.24xlarge with 400 Gbps networking achieved only ~1 Gbps effective pull throughput because CPU decompression was the constraint. On instances with capable, current-generation CPUs, SOCI parallel pull reduces image pull time from 2-4 minutes to 30-60 seconds. The dominant factor is per-core decompression throughput, which depends on CPU generation and instruction-set support, more than raw core count. A newer CPU with fewer cores can outperform an older one with more.

For a deeper look at how bounded-memory parallel pull handles images exceeding 30 GB without OOM, see Bounded-Memory Parallel Image Pulling for Large Container Images.

Layer 4: Model weights download

The obvious optimization for weights loading: more parallel connections. Split the model files into small chunks, download them concurrently, saturate the network pipe.

We tested it on p5.48xlarge with the 64 GB model streaming from same-region S3. The results were counterintuitive:

Chunk sizeConnections neededWeights load time
256 MB25613.98s
512 MB12814.20s
2 GB3413.62s
4 GB1713.35s
8 GB921.80s (+56%)

256 parallel connections provided no benefit over 17. The only failure mode was 8 GB chunks (exceeding shard file size), which caused a 56% regression.

Why? Because the open-source Run:ai Model Streamer (integrated into vLLM and SGLang) processes S3 range requests sequentially within each worker thread. A worker assigned to a 3.9 GB shard file downloads its byte-range requests one after another on a single connection. The parallelism comes from running multiple workers on different files, not from splitting one file into more pieces.

We settled on 4 GB chunks matching typical SafeTensors shard size (3-5 GB per file) with an aggressive timeout-and-retry for slow requests. S3 GET latency has a measurable long tail: in our testing, a meaningful fraction of requests took 2-3x longer than median, and a single stalled connection holds up the entire model load. Rather than wait, we kill stalled connections after a few seconds below a speed threshold and retry on a fresh connection. This follows S3’s own performance guidance.

For the 203 GB model, these config-only changes reduced weights loading from 423 seconds to 25 seconds (94% improvement). For the 64 GB model, from 29 seconds to 12 seconds. No code modifications, just environment variables. The tuning consists of three settings: chunk size aligned to shard file boundaries (eliminating the serial sub-request problem), a minimum-speed threshold that kills and retries stalled S3 connections, and explicit concurrency matching the number of shard files per tensor-parallel rank.

Layer 5: GPU kernel compilation

Every time a vLLM or SGLang pod starts, PyTorch traces the model’s computation graph and compiles it to optimized CUDA kernels. This takes 34-53 seconds depending on model architecture. The output is identical every time for the same model, GPU type, and tensor-parallel configuration.

And Kubernetes throws it away on every pod restart. Pods use ephemeral storage by default. When a pod terminates, its local filesystem is destroyed. The next pod recompiles from scratch.

“The output is identical every time for the same model, GPU type, and tensor-parallel configuration. And Kubernetes throws it away on every pod restart.”

Point the torch.compile cache directory at local NVMe instance store. GPU instances ship with NVMe that EKS Auto Mode mounts automatically. First pod compiles and writes ~15-30 MB of cached kernels. The second pod on the same node loads pre-compiled binaries in 4-6 seconds. One volume mount and environment variables.

The cache is safe because the compiled artifacts are deterministic: same model architecture + GPU architecture + tensor-parallel degree + PyTorch version equals valid cache. An image update or hardware change triggers exactly one recompilation.

torch.compile time is hardware independent. The same model compiles in ~52 seconds whether running on H100 or A100. The cache hit (4-6 seconds) is equally consistent across GPU types. This means the optimization works identically regardless of instance type.

Layer 6: Engine initialization

After weights are loaded and kernels compiled, the inference engine must capture CUDA execution graphs and profile KV cache memory. With compiled kernels cached, this completes in 30-45 seconds. Without cache, graph capture triggers additional JIT compilation and takes 60-120 seconds.

This is why the torch.compile cache has an outsized impact: it accelerates not just layer 5 but also layer 6. Cached compilation reduces a 2-3-minute combined phase to a 35-50-second combined phase.

Framework initialization (Python interpreter startup and PyTorch import) adds tens of seconds of fixed overhead that cannot be reduced through configuration.

The compounding effect

The six layers compound. Platform fixes (layers 1-3) eliminate 4-8 minutes of overhead: pre-compiled drivers replace 2-3 minutes of runtime compilation, parallel pull reduces image transfer time from 2-4 minutes to 30-60 seconds, and Karpenter keeps node provisioning to its hardware minimum. Configuration changes (layers 4-5) cut the remaining model startup by 80-93%. Engine initialization (layer 6) drops from 60-120 seconds to 30-45 seconds once the compile cache is warm. Together, cold-node TTFTS drops from 8-15 minutes to approximately 5 minutes.

64 GB model (Qwen3.6-35B-A3B), TP=2:

ConfigurationFirst podSubsequent pod (warm node)
Baseline (no tuning)82s82s
+ S3 chunk tuning65s65s
+ torch.compile cache65s16s
Improvement-21%-80%

203 GB model (Llama-4-Scout), TP=4:

ConfigurationFirst podSubsequent pod (warm node)
Baseline (no tuning)457s457s
+ S3 chunk tuning59s59s
+ torch.compile cache59s32s
Improvement-87%-93%

The warm-node subsequent pod number is what matters most for production. It’s what you pay on every pod restart. The 80-93% reduction is consistent across instance types because the optimizations target software bottlenecks (calling patterns, redundant compilation), not hardware limits.

The cost of cold starts at scale

Why does any of this matter? Because GPU nodes are expensive and inference traffic is bursty.

A single p5.48xlarge costs $55/hour on-demand. Even G-family instances commonly used for inference cost $10-20/hour. Every minute of cold start is GPU time you’re paying for but not using. If your autoscaler needs 8+ minutes to bring up new capacity, you must over-provision (burn money on idle GPUs) or accept latency spikes during traffic surges.

“Every minute of cold start is GPU time you’re paying for but not using.”

When model startup drops to 16-32 seconds on warm nodes, the calculus changes. You can scale more aggressively, keep fewer buffer nodes, and respond to traffic spikes without multi-minute startup delays.

What we learned

  1. Decompose before optimizing. For 64 GB models, torch.compile dominates (65%). For 203 GB models, S3 loading dominates (92%). Without measuring each phase independently, we would have optimized the wrong layer.
  2. The bottleneck flips with model size. torch.compile time is roughly constant across model sizes. Weights loading scales linearly. Every team running inference should know which regime they’re in.
  3. “More parallelism” requires understanding the execution model. 256 connections performing sequential work inside each thread is no faster than 17. The bottleneck was the calling pattern, not the concurrency limit.
  4. 15-30 MB can save 53 seconds. The most impactful optimization for smaller models was persisting a tiny cache file. Always check whether an expensive computation produces deterministic output before trying to make it faster.
  5. Platform-level control enables optimizations that configuration alone cannot achieve. Pre-compiled drivers, default-on parallel image pull, and NVMe auto-mounting are infrastructure-layer decisions that compound upward. Together with the config-only changes at the application layer, these changes reduce cold start time from minutes to seconds.
  6. The ecosystem is building the right primitives, but cold start lives between them. OCI image volumes, DRA, inference-aware routing, and local model caches are all real progress. But the compilation bottleneck and S3 tuning gaps sit in spaces that no upstream Kubernetes primitive addresses. Sometimes the highest-impact optimization is a volume mount and two environment variables, not a new API.

For the complete configuration guide, including environment variables, YAML manifests, and instance-specific recommendations, see “Accelerate model loading on Amazon EKS” in the Amazon EKS User Guide.

The post Cut GPU inference cold start from 8 minutes to less than a minute appeared first on The New Stack.

Your container runs. Everything around it shouldn’t be your problem.

29 août 2026 à 17:00
Abstract digital binary data wave depicting cloud container orchestration and automated infrastructure.

The promise with containers was simple: if it runs on your local machine, it will run in production. And this promise holds – your container runs. But there’s a tax: setting up everything around it. To reach your container, you’ll need a load balancer to route traffic to it, and scaling policies to handle variable traffic. Then come the networking components and the minimally scoped access roles. And somewhere in that checklist, don’t forget the security configuration; unless you’d rather hear about it from your compliance team. At some point, you wonder why you can’t just go back to building.

“You give it a container image. You get a production service. And when your workload outgrows a single container, you don’t outgrow Amazon ECS Express Mode.”

Sound familiar? You’re not alone. Most teams spend cycles before they feel confident deploying containers in production. That’s time from your roadmap spent on decisions that don’t differentiate your business. Somewhere between the third Terraform module and the second IAM policy review, you’ve lost the speed containers were supposed to give you.

Amazon Elastic Container Service (ECS) is the container orchestration engine behind some of the largest production workloads on AWS. But until now, getting started with it meant understanding load balancers, networking, IAM roles, and scaling policies before you shipped anything.

We tried to change that and make it easier for you to get started and stay focused on building when we launched Amazon ECS Express Mode. Express Mode is a new interface into that same engine. You’re not trading power for simplicity. You’re getting a faster door into infrastructure that’s already battle-tested. The vision was clear: keep it simple but extensible. You give it a container image and two IAM roles; you get an HTTPS service running on Fargate with a load balancer, a TLS certificate, autoscaling, and canary deployments. Oh, and all these resources run in your account, where you have full control.

 // Sample Terraform Code
 resource "aws_ecs_express_gateway_service" "frontend_service" {
   execution_role_arn = aws_iam_role.execution.arn
   infrastructure_role_arn = aws_iam_role.infrastructure.arn
 
   primary_container {
     image = "111122223333.dkr.ecr.us-east-1.amazonaws.com/my-service:1.4.2"
   }
 }

What’s in Amazon ECS Express Mode?

Underneath the one-step deploy, this is what you’re getting:

  • No sprawl – Every service sits behind an Application Load Balancer, but doesn’t get its own. Up to 25 Express Mode services within a VPC share a single ALB. Express Mode adds load balancers only when needed and removes them when services are deleted. No dangling resources for you to worry about.
  • Safer rollouts – Out of the box, each deployment is a canary release where 5% of your traffic is routed to the new revision and bakes for 3 minutes, following which the remaining traffic is shifted. If your 4xx/5xx error rate exceeds 1%, an alarm we create for you triggers an automatic rollback.
  • Scaling – Each service ships with an auto-scaling policy that targets 60% CPU utilization, scaling from 1 task to a maximum of 20 by default. If your workloads need to scale on memory or request count, you can configure that too.
  • IaC ready – Create and manage Express services through CloudFormation, CDK, Terraform or GitHub Actions.
  • Complete ownership – The cluster, load balancer, target groups, log groups and other resources are all in your AWS account. You can inspect them, audit events, and modify them directly if needed. Nothing is a black box.

But what about my sidecars? My workloads aren’t that simple.

Sooner or later, as your application evolves, the workload becomes more than just one container. An observability agent needs to run beside the app, publishing traces to your monitoring solution. The base image gets swapped for the hardened one security maintains. The credentials move out of environment variables and into Secrets Manager. This is where extensibility comes into play. 

For this, Express Mode now supports providing a standard ECS task definition – the spec that describes your containers, their resource limits, and how they connect. Your sidecar, image, and credentials all fit right in. If your team runs ECS today, this is the spec you’re already writing. If not, Express Mode generates a task definition in your account, and when requirements arrive, you take that working spec, add what you need, and hand back the ARN. Once you associate a task definition with an Express Mode service, you can continue managing your application either through task definition updates or directly through Express Mode, whichever you prefer.

// Sample CDK snippet
const taskDef = new ecs.FargateTaskDefinition(this, 'TaskDef', {
  cpu: 1024,
  memoryLimitMiB: 2048,
  executionRole,
  taskRole,
});
 
taskDef.addContainer('Main', {
image:
ecs.ContainerImage.fromRegistry('111122223333.dkr.ecr.us-east-1.amaz
onaws.com/my-service:1.4.2'),
  essential: true,
  portMappings: [{ containerPort: 8080, name: 'main' }],
});
 
taskDef.addContainer('otel-collector', {
image:
ecs.ContainerImage.fromRegistry('public. ecr.aws/aws-observability/aw
s-otel-collector:latest'),
  essential: false,
  memoryReservationMiB: 256,
  command: ['--config=/etc/ecs/ecs-default-config.yaml'],
});
 
new ecs.CfnExpressGatewayService(this, 'FrontendService', {
  infrastructureRoleArn: infrastructureRole.roleArn,
  taskDefinitionArn: taskDef.taskDefinitionArn,
});

Wait, am I locked into this?

No. Express Mode is an interface into Amazon ECS, not a walled garden. Every resource it creates is a standard AWS resource in your account, addressable by an ARN. If required, you can modify them directly through their respective AWS APIs in the Console, CLI or SDK.

“Express Mode is an interface into Amazon ECS, not a walled garden.”

If you change the scaling policy, update a security group rule, or swap the task definition, Express honors those modifications on the next update. It does not overwrite changes you make. This means you can start with Express defaults today and customize individual resources as your requirements evolve, without migrating off Express or recreating your service.

Express also exposes the ARN of everything it manages – the load balancer, target groups, security groups, and alarms – through the Describe APIs and as CloudFormation/CDK outputs. You can reference them in your own stacks or hand them to existing constructs.

Why did we design it as a single operation?

We deliberately made Express Mode a single API call, not a multi-step workflow. One input (your image or task definition ARN), one operation, one outcome.

That design choice compounds. Your IaC is under ten lines. Your CI/CD pipeline doesn’t need custom steps. An AI coding agent can deploy and iterate on your behalf because the entire surface is one well-defined spec it already knows how to read and modify.

But there’s an engineering reason too. Because Express Mode owns the full lifecycle of what it creates, it can clean up what it sets up. Delete a service, and the target groups, scaling policies, and alarms go with it. No orphaned infrastructure. And because you didn’t wire these resources together manually, one service’s deploy can’t accidentally touch another’s.

Wrapping it up

As we continue to build on ECS Express Mode, our priority is simple – taking away the undifferentiated heavy lifting from our customers. Amazon ECS Express Mode is our attempt to make good on that: point it at an image, get a production service: load balancer, TLS, autoscaling, canary deployments, all running in your account, where you can see it and change it.

“The configuration grows with your requirements; the operational burden doesn’t.”

And when your workload outgrows a single container, you don’t outgrow Express Mode. Bring your own task definition: the sidecars, the hardened images, the secrets, and keep handing the infrastructure to us. The configuration grows with your requirements; the operational burden doesn’t.

Try it from the AWS Console, Terraform, CloudFormation, CDK, or GitHub Actions – or just ask your agent.

The post Your container runs. Everything around it shouldn’t be your problem. appeared first on The New Stack.

❌