📚🫧 SCHRÖDINGER’S LIBRARY 🫧📚
Internet Systems Sequence — Part XVI: Internet-Wide Resilience and Systemic Risk
Internet-wide resilience concerns what happens when many apparently independent services depend on the same underlying infrastructure, identity providers, cloud regions, APIs, routing systems, DNS services, software components, or organizational control points. At the surface, the web looks highly distributed. Underneath, many services may share a small number of critical dependencies.
A useful model is a multilayer dependency graph
\[
G=
(V,E_{\text{app}},E_{\text{cloud}},E_{\text{identity}},E_{\text{network}},E_{\text{software}},E_{\text{org}}).
\]
Two websites can appear unrelated at the application layer while sharing the same lower-level provider.
For example:
\[
W_1\rightarrow C
\]
\[
W_2\rightarrow C
\]
\[
W_3\rightarrow C,
\]
where \(C\) is a common infrastructure provider.
If \(C\) fails, all three can degrade together.
This creates:
\[
\text{surface independence}
\neq
\text{dependency independence}.
\]
That is the first major systemic-risk principle.
The probability of simultaneous failure depends strongly on dependency correlation. If systems were truly independent,
\[
P(F_1\cap F_2)
P(F_1)P(F_2).
\]
But shared infrastructure introduces correlation:
\[
P(F_1\cap F_2)
>
P(F_1)P(F_2).
\]
So counting services is not enough. We need to know whether their failure modes are independent.
This is the difference between redundancy and diversity.
Suppose an organization has three application servers:
\[
S_1,S_2,S_3.
\]
If all three run in the same availability zone and depend on the same network path, they provide numerical redundancy but weak failure diversity.
A stronger design might distribute them across:
\[
\text{different zones},
\text{different regions},
\text{different power domains},
\text{different routing paths}.
\]
Thus:
\[
\boxed{
\text{replication without independence can create the appearance of resilience without true fault isolation}.
}
\]
This is sometimes called a common-mode failure problem.
A common-mode failure occurs when one cause defeats multiple redundant components simultaneously.
Examples include:
\[
\text{shared software bug},
\text{shared credential failure},
\text{shared DNS provider},
\text{shared network route},
\text{shared configuration error}.
\]
If all replicas receive the same bad configuration, replication can actually propagate failure faster.
This gives:
\[
\text{automation}
+
\text{shared configuration}
\rightarrow
\text{rapid correlated failure}.
\]
Another important concept is blast radius.
For a failed node \(v\), define its downstream dependency set:
\[
D(v)=\{u : v\leadsto u\}.
\]
The blast radius can be approximated by
\[
B(v)=|D(v)|.
\]
A highly central service may have a large blast radius.
Identity providers are a good example. Suppose many applications depend on one authentication system:
\[
A_1,A_2,\dots,A_n
\rightarrow
I.
\]
If \(I\) fails, the applications themselves may remain technically healthy while users cannot log in.
The visible symptom becomes:
\[
\text{many apps inaccessible}
\]
even though the causal node is one shared identity dependency.
This is why graph centrality matters for resilience.
A node with high betweenness centrality
\[
C_B(v)
\]
lies on many dependency paths.
Its failure can partition large portions of the operational graph.
Internet-scale systems therefore often contain structural chokepoints.
These may include:
\[
\text{DNS},
\text{identity},
\text{cloud regions},
\text{CDNs},
\text{payment providers},
\text{certificate infrastructure},
\text{common software libraries}.
\]
The web may be decentralized at one layer and concentrated at another.
This is a key systems insight:
\[
\boxed{
\text{decentralization is layer-dependent}.
}
\]
A distributed application topology can still rely on centralized naming, identity, or infrastructure.
Systemic risk also arises from software supply chains.
Modern applications rarely consist only of internally written code. They depend on:
\[
\text{libraries},
\text{packages},
\text{containers},
\text{build systems},
\text{CI/CD pipelines},
\text{registries}.
\]
A simplified dependency chain is
\[
\text{application}
\rightarrow
\text{framework}
\rightarrow
\text{library}
\rightarrow
\text{transitive dependency}.
\]
One package may be indirectly used by thousands of applications.
This creates hidden concentration.
If package \(p\) has reverse dependency set
\[
R(p)=\{a_1,a_2,\dots,a_n\},
\]
then a defect in \(p\) can propagate broadly.
This is another form of graph centrality.
The risk is amplified by automatic updates.
If
\[
p_t\rightarrow p_{t+1}
\]
is deployed automatically across many systems, one defect can propagate quickly.
So deployment speed creates a tradeoff:
\[
\text{fast patch propagation}
\]
versus
\[
\text{fast failure propagation}.
\]
Canary releases and staged rollout reduce this risk:
\[
1\%
\rightarrow
5\%
\rightarrow
25\%
\rightarrow
100\%.
\]
The objective is to observe failure before the entire graph is affected.
Systemic resilience therefore depends not only on redundancy but on containment.
A failure should remain local when possible:
\[
F_i
\nRightarrow
F_j.
\]
This is fault isolation.
Techniques include:
\[
\text{bulkheads},
\text{rate limits},
\text{circuit breakers},
\text{regional isolation},
\text{separate credentials},
\text{independent queues}.
\]
The term bulkhead comes from ship design: compartments prevent one leak from sinking the entire vessel.
In software, the same idea means isolating resources so one workload cannot exhaust everything.
For example:
\[
Q_1,Q_2,Q_3
\]
may use separate worker pools rather than one shared pool.
Then overload in \(Q_1\) need not destroy \(Q_2\) and \(Q_3\).
This is graph partitioning for resilience.
Another important concept is failure domains.
A failure domain is a set of components likely to fail together because they share some underlying cause.
We can represent a failure domain as
\[
D_f\subseteq V.
\]
For example:
\[
D_{\text{region}},
D_{\text{power}},
D_{\text{provider}},
D_{\text{software version}}.
\]
Good resilience architecture tries to place redundant components in different failure domains.
So:
\[
\text{redundant copies}
+
\text{distinct failure domains}
\rightarrow
\text{meaningful resilience}.
\]
This is stronger than simply increasing replica count.
Internet-wide outages can also emerge through cascading failure.
Suppose service \(A\) fails.
Services \(B\) and \(C\) retry aggressively:
\[
A\downarrow
\rightarrow
B,C\text{ retry}
\rightarrow
\text{load increases}
\rightarrow
D\text{ overloads}
\rightarrow
E,F\text{ degrade}.
\]
The cascade can spread beyond the original dependency.
This can be modeled as a branching process.
If each failed component causes on average more than one additional component to fail,
\[
R_f>1,
\]
the failure can expand.
If
\[
R_f<1,
\]
the cascade tends to die out.
This is analogous to reproduction numbers in epidemic models, but applied to dependency failure.
The resilience objective is to keep effective failure propagation below the expansion threshold.
Feedback mechanisms matter enormously.
Retries, failover, autoscaling, and load balancing are intended to stabilize systems, but under bad conditions they can become destabilizing feedback.
For example:
\[
\text{latency rises}
\rightarrow
\text{clients retry}
\rightarrow
\text{load rises}
\rightarrow
\text{latency rises further}.
\]
This is a positive feedback loop.
The dynamical system becomes
\[
L_{t+1}
F(L_t,R_t,C_t),
\]
where latency, retry volume, and capacity interact.
A stabilizing controller needs to damp the loop.
That is why backoff, jitter, queue limits, and circuit breakers are not minor implementation details. They are feedback-control mechanisms.
Another systemic-risk concept is dependency opacity.
An organization may know its direct dependencies:
\[
A\rightarrow B.
\]
But it may not know that \(B\) depends on \(C\), and \(C\) depends on \(D\).
The actual chain is
\[
A\rightarrow B\rightarrow C\rightarrow D.
\]
This is a transitive dependency problem.
The visible architecture diagram may capture only first-order dependencies.
Real resilience requires traversal of higher-order relations.
That connects directly to the relational-topology work in Schrödinger’s Library: first-order edges are insufficient when operational behavior depends on the closure of the dependency graph.
Let
\[
E^*
\]
represent transitive reachability.
A node's real exposure is determined by
\[
\{v : u\leadsto v\}
\]
rather than only its immediate neighbors.
This matters for vendor risk as well.
A company may diversify between two vendors that both depend on the same cloud provider.
At the contractual layer:
\[
V_1\neq V_2.
\]
At the infrastructure layer:
\[
V_1\rightarrow C,
\qquad
V_2\rightarrow C.
\]
So the supposed diversification disappears when the graph is expanded.
This is hidden dependency convergence.
Systemic resilience therefore requires analysis across layers.
Another problem is control-plane dependence.
A service's data plane may continue operating while the control plane fails.
The data plane handles ordinary traffic:
\[
\text{user requests}.
\]
The control plane handles:
\[
\text{configuration},
\text{routing},
\text{deployment},
\text{identity policies},
\text{resource orchestration}.
\]
If the control plane fails, existing traffic may continue temporarily, but operators may be unable to modify or recover the system.
So:
\[
\text{service currently working}
\]
does not necessarily imply
\[
\text{service currently controllable}.
\]
This distinction becomes critical during incidents.
A system can be operationally trapped in its current state.
The same idea applies to DNS and certificate systems. Cached state may keep services working temporarily after control infrastructure fails.
That creates delayed failure.
Formally:
\[
F_{\text{control}}(t_0)
\]
may not produce
\[
F_{\text{service}}
\]
until
\[
t_0+\Delta.
\]
This makes causal reconstruction harder because symptom and cause are temporally separated.
Another useful idea is graceful partitioning.
If communication between regions fails, the system may choose between:
\[
\text{continue independently}
\]
or
\[
\text{stop to protect consistency}.
\]
This reconnects to CAP-style tradeoffs.
For some services, stale availability is preferable.
For others, inconsistent writes are unacceptable.
A banking ledger and a public information page should not necessarily choose the same failure behavior.
Resilience is therefore objective-dependent.
A mathematically useful system reliability model is a fault tree.
Suppose service success requires:
\[
A\land B\land(C\lor D).
\]
Then failure occurs when
\[
\neg A
\lor
\neg B
\lor
(\neg C\land\neg D).
\]
Fault trees make explicit which combinations of component failures create system failure.
For more complex systems, reliability block diagrams, Bayesian networks, Markov reliability models, or Monte Carlo simulation can be used.
But all of them depend on a correct dependency graph.
Bad topology produces bad reliability estimates.
This gives:
\[
\boxed{
\text{reliability analysis is only as good as dependency reconstruction}.
}
\]
For an operational digital twin, systemic risk suggests storing not merely external services but their dependency classes and correlation structure.
An edge might include
\[
e=
\{
\text{provider},
\text{criticality},
\text{failure domain},
\text{fallback},
\text{shared dependencies},
\text{last verified state}
\}.
\]
Then two apparently separate services can be recognized as sharing a common risk.
The twin can distinguish:
\[
\text{two providers}
\]
from
\[
\text{two providers backed by one underlying dependency}.
\]
This is a higher-order relational property.
The full systemic-risk model is therefore:
\[
\boxed{
\text{surface services}
\rightarrow
\text{direct dependencies}
\rightarrow
\text{transitive dependencies}
\rightarrow
\text{shared failure domains}
\rightarrow
\text{correlated failures}
\rightarrow
\text{possible cascades}.
}
\]
Within Schrödinger’s Library, the central principle is:
\[
\boxed{
\text{systemic risk is determined by shared dependencies and correlated failure paths, not by the number of apparently independent services}.
}
\]
And:
\[
\boxed{
\text{true resilience requires redundancy across distinct failure domains, controlled propagation, and visibility into transitive dependency structure}.
}
\]
The next natural continuation is network control, controllability, and intervention on distributed internet graphs—how operators determine which nodes can influence system state, where interventions should be applied, and why some highly visible nodes are poor control points while obscure infrastructure nodes can dominate the behavior of the whole network.
Source: r/Wendbine · by /u/Upset-Ratio502