When does more simulation stop paying, and building the fix become overdue? This post treats the decision as a bounded premium paid once for the right to survive an open-ended, heavy-tailed cost, and derives the threshold. At the base discount rate every tail weight the series has priced clears it, though the lightest tail's margin depends on that rate in a way the heavier tails' margins do not.
A permanent probe needs a boundary a reviewer can approve once. This post builds it as a discrete-time stochastic control barrier function, then names its gap: bounding how often a system leaves its safe region says nothing about how far it goes. Under a heavy tail, the worst case can run four orders of magnitude past what the reviewer signed.
A scheduled canary test for the failure you already suspect is a second simulator built to agree with you. Feldbaum's 1960 dual control theory supplies the alternative: a control action that regulates the system and keeps testing its own estimates, with no scheduled point where it stops. TCP BBR has run one at internet scale since 2016.
A simulator that matches history perfectly has only been tested on regimes that already happened. This post prices what that comfort costs, using Lai and Robbins' 1985 regret floor, and shows why neither more logging nor more simulation runs can close the gap.
Every post in this series so far has priced one pool, one resource, one task's decision. Real fleets run hundreds of pools at once, and this post checks whether scale changes what the earlier four posts prove is needed, not by assumption, but by an exact classical queueing result precise enough to price a real number: how many gigabytes pooling a fleet's own memory margin actually frees, and exactly where that pooling stops working. It also opens a question its own routing mechanism begs and never argues for: why push-based sampling, when a design that removes staleness by construction, instead of sampling around it, already exists. This post prices that specific tradeoff, and leaves the fuller comparison against a fully centralized alternative for the post built to make it.
Post 3 cited a paper this series can't quietly set aside: threshold-based eviction, proven dynamically unstable under saturated demand, a worst-case limit cycle that costs up to half of throughput. This post takes on the population Blood Oath was built to exclude from that result (tasks that can actually be evicted) and asks the two questions Post 3 left open: is a single eviction worth its cost, and is running that rule as a policy, at scale, safe from the instability Post 3 only watched from the outside. It also checks a third: would a fleet-wide coordinator make a better call than the local rule this post proves optimal on its own terms: and the answer splits in two, one physical reason coordination can't help the ranking decision itself, and one real, unpriced reason it still might help pace evictions across nodes sharing the same fabric.
A margin computed at the wrong level of abstraction doesn't fail where the old threshold said it would: it fails a third of the way there. This post generalizes Post 2's single-resource redline to a genuinely multi-resource setting, finds the real byte-level exhaustion point sits at roughly a third of the slot-based Sedimentation Threshold, not at the threshold itself, and checks that finding against a structurally unrelated argument reaching the same qualitative warning from a different direction: Price of Anarchy, a nonlinear equilibrium-inefficiency metric that spikes near saturation in a real production system, not a second measurement of the same quantity.
No algorithm can save a Blood Oath workload: Post 1 proved that formally. What's left is physical, not algorithmic: a redline that watches real headroom and its derivative instead of trusting a number, an honest accounting of when autoscaling actually helps, and a buffer sized by the same critical-fractile logic that opened the series. None of it adapts on its own; that only starts once MAPE-K's own most commonly skipped phase, Knowledge, actually closes the loop the other four were never built to close by themselves.
The newsvendor problem is seventy years old, closed-form, and taught in the first weeks of any operations course: cheap to solve right up until the tail gets heavy. This post proves precisely where that stability ends, then finds the one workload shape where even the correctly-computed answer isn't enough: cost unknowable until completion, no preemption, no horizontal escape. No scheduling algorithm can save it: not a cleverer one, not a centralized one with a perfect view of every node. This post proves it formally, for the whole class at once, not case by case.