The Weekend Engineering Digest
September 5, 2026 · 5 min read

Everyone is doing efficiency work again

Stripe writes its own data plane, Uber constrains shard placement to shrink blast radius, and Cloudflare reclaims 100 TB by rearranging structs. Six items from the week in industry engineering.

Five of the six items below are cost or capacity work wearing engineering clothes: memory layout, cache compression, shard placement, CI economics, finding-prioritization. After several years of capability-first spending — especially on AI infrastructure — the public engineering conversation has swung hard toward unit economics.


Stripe built its own data plane to replace Envoy

Stripe targets 99.9995% reliability on a payments flow that now exceeds 1.6% of global GDP. At that volume the Envoy-based service mesh became the constraint, so the infrastructure team wrote a purpose-built distributed proxy rather than continuing to tune the industry-default sidecar.

The interesting part isn’t that they built a proxy. It’s the argument for leaving a mature, well-supported open-source component: at what point does a general-purpose data plane’s flexibility become overhead you pay for on every single request, and what do you surrender — ecosystem, community, xDS compatibility, the ability to hire people who already know it — to reclaim that margin? The post is worth reading as a template for how to make a build-versus-adopt case that survives scrutiny: a measured constraint, a bounded scope, and a migration path, rather than “the old thing felt slow.”

stripe.dev →


Uber constrained shard placement to shrink the blast radius

Under M3DB’s original placement algorithm, any node could hold any shard, subject only to keeping replicas in separate isolation groups. That freedom has a cost: a single node failure pulls bootstrap traffic from up to (n−1)/n of the cluster, and because nodes in different isolation groups still share shards, maintenance operations interfere with each other and have to be serialized.

The fix is a structural constraint. Nodes are partitioned into fixed-size subclusters, each owning a distinct, non-overlapping slice of the shard space, so shard sharing between arbitrary node pairs drops from O(cluster) to O(subcluster). Failures stay local, and automation can work several subclusters in parallel. The price is real: homogeneous hardware only, scaling in fixed increments, and no replica-factor changes. There’s also a nice algorithmic detail — when a new subcluster needs shards, the donor picks them by simulating which removal leaves it most balanced, which avoids moving shards twice to correct skew afterward.

The general lesson is that the strongest reliability wins often come from constraining the topology rather than from handling failures more cleverly.

uber.com →


Cloudflare reclaimed 100 TB by rearranging data structures

Five Rust-level layout optimizations to the DNS cache inside 1.1.1.1’s resolver cut per-entry memory by 56%, freeing roughly 100 TB across the fleet. No new service, no architectural change — just how the bytes are arranged.

At fleet scale, per-object overhead is a capacity decision. Struct packing and pointer elimination are unglamorous, but a 56% per-entry win compounds into hardware nobody has to buy. Useful counterweight to the reflex that efficiency work has to mean a redesign: the first question is usually “what does one unit of this cost, times how many units do we have?”

blog.cloudflare.com →


Trading CPU for cache storage, with receipts

A companion piece prototypes Zstandard compression inside Cloudflare’s edge cache — spending CPU cycles to get more effective cache capacity out of the same disks.

Two things make it worth reading. The framing: cache capacity, CPU, and hit rate are one budget, not three independent ones, and “we need more hardware” is often really a resource-substitution question in disguise. And the form: it’s published as a prototype with measurements and an honest verdict rather than as a shipped win, which is a healthier model for engineering writing than only publishing successes.

blog.cloudflare.com →


The CI bill is the next AI-adoption problem

Uber’s write-up treats the internal build, test, and deploy pipeline as a cost-managed production system in its own right — with AI-assisted development now part of the load it has to absorb.

This is the second-order effect of coding agents that most organizations have not budgeted for. If agents multiply the pull-request rate, then selective test execution, build caching, and merge-queue design stop being developer-experience niceties and become capacity planning. Plenty of teams have enthusiastically adopted agent tooling without asking what happens to CI compute when PR volume triples.

uber.com →


Grounding an agent in production telemetry, not just source code

Rather than running a model over code in isolation, Cloudflare’s new vulnerability discovery system joins WAF and production traffic signals with model-driven analysis to rank findings by what is actually reachable and exercised, then proposes edge mitigations and, where appropriate, code patches.

The generalizable pattern here is context selection, not model choice. Static analysis drowns teams in findings precisely because it lacks reachability information; production telemetry is the missing prior that makes prioritization possible. The other design decision worth stealing is the graduated action: an edge rule is reversible and cheap, a merged patch is neither, so the system reaches for the reversible one first. Both questions — what grounds the decision, and what is the reversible action — transfer directly to any agent you might build over infrastructure.

blog.cloudflare.com →


Sources

  1. Building a data plane from scratch: Stripe’s own high-performance distributed proxy — Stripe, Aug 26, 2026
  2. From Chaos to Control: Addressing Shard Distribution Challenges in M3DB with Subclusters — Uber, Sep 1, 2026
  3. How we saved 100 terabytes of memory by optimizing 1.1.1.1’s DNS cache — Cloudflare, Aug 27, 2026
  4. How we could save petabytes of cache storage with Zstandard and Pingora — Cloudflare, Sep 1, 2026
  5. Running a Software Factory Efficiently at Uber Scale — Uber, Aug 27, 2026
  6. Introducing context-aware vulnerability discovery and remediation — Cloudflare, Sep 3, 2026

New issue every Saturday. Subscribe via RSS, orbrowse the archive.