📚🫧 SCHRÖDINGER’S LIBRARY 🫧📚
Internet Systems Sequence — Part IX: Content Delivery Networks and Edge Computing
Content delivery networks and edge computing explain how internet systems distribute copies of content, computation, and service state closer to users. This improves speed and resilience, but it also means that two users requesting what appears to be the same website may be served from different infrastructure, different caches, different geographic regions, or even different software versions. The result is that one nominal website can have multiple simultaneously valid operational states.
A simplified CDN architecture is
\[
\text{user}
\rightarrow
\text{edge node}
\rightarrow
\text{regional cache}
\rightarrow
\text{origin server}.
\]
The origin server is the authoritative source for content, while edge nodes store replicated copies closer to users.
If a user requests object \(x\), the edge first checks whether a cached version exists:
\[
C_e(x)=?
\]
If yes, it may return the cached object directly.
If not, it fetches from the origin:
\[
O(x)\rightarrow C_e(x)\rightarrow \text{user}.
\]
This reduces latency because the user does not need to traverse the full network path to the origin every time.
A cache entry can be modeled as
\[
c=
\{
\text{object},
\text{timestamp},
\text{TTL},
\text{version},
\text{validation metadata}
\}.
\]
The time-to-live \(TTL\) determines how long the object may remain valid before requiring refresh.
A simple expiration condition is
\[
t_{\text{now}}-t_{\text{cache}}>TTL.
\]
If false, the edge may continue serving the cached representation.
This immediately creates the possibility of geographic state divergence.
Suppose two edge nodes hold different versions:
\[
C_{e_1}(x)=x_{t_1}
\]
and
\[
C_{e_2}(x)=x_{t_2}.
\]
Then users routed to different edges may observe:
\[
Y^{(u_1)}\neq Y^{(u_2)}.
\]
The website appears singular, but the distributed delivery system exposes different temporal states.
This is another form of partial observability.
The user does not see:
\[
\text{which edge node served the response},
\]
\[
\text{which cache version was used},
\]
or
\[
\text{whether the object came from origin or replica}.
\]
They see only the rendered page.
A useful representation is
\[
Y_t^{(u)}=
H(
X_{\text{origin}},
X_{\text{edge}(u)},
R_u
),
\]
where \(R_u\) is the routing decision that assigns user \(u\) to an edge.
That routing decision can depend on geography, network latency, provider relationships, congestion, failure conditions, or load balancing.
This is where anycast routing becomes important.
Many CDN providers advertise the same IP address from multiple geographic locations. Network routing sends the user toward a nearby or otherwise preferred edge.
Conceptually:
\[
\text{same IP}
\rightarrow
\{
e_1,e_2,\dots,e_n
\}.
\]
The user therefore does not necessarily know which physical machine or region served the content.
This is another example of a stable surface identity hiding a distributed topology.
Edge caching introduces cache invalidation problems.
Suppose the origin updates from
\[
x_1\rightarrow x_2.
\]
The edge should eventually update too:
\[
C_e(x_1)\rightarrow C_e(x_2).
\]
But this may occur through:
\[
\text{TTL expiry},
\]
\[
\text{explicit purge},
\]
\[
\text{version change},
\]
\[
\text{revalidation}.
\]
Until then, users may receive stale state.
This produces:
\[
X_{\text{origin}}(t)
\neq
X_{\text{edge}}(t).
\]
The system can therefore be functioning exactly as designed while different users see different versions.
That is one reason internet state can feel geographically inconsistent.
A related mechanism is stale-while-revalidate.
The edge may intentionally serve an old copy immediately while fetching a fresh copy in the background.
Formally:
\[
\text{serve }x_{old}
\]
while
\[
\text{fetch }x_{new}.
\]
This reduces latency but temporarily preserves stale representation.
Again:
\[
\text{fast response}
\neq
\text{current response}.
\]
CDNs also cache more than static files.
Modern edge systems may handle:
\[
\text{HTML},
\text{API responses},
\text{images},
\text{video},
\text{authentication state},
\text{feature configuration},
\text{computation}.
\]
That moves us from simple caching into edge computing.
Instead of merely storing objects, edge nodes may execute code:
\[
Y=
F_e(X,U).
\]
This can include:
\[
\text{request filtering},
\text{A/B testing},
\text{personalization},
\text{security checks},
\text{geographic routing},
\text{content transformation}.
\]
That means the edge is no longer just a passive replica.
It becomes an active computational layer.
A modern request may therefore traverse:
\[
\text{user}
\rightarrow
\text{edge logic}
\rightarrow
\text{regional service}
\rightarrow
\text{origin}
\rightarrow
\text{database}.
\]
Any of those layers can alter the final result.
This creates regional execution divergence.
Suppose one edge runs software version \(v_1\) while another has already been updated to \(v_2\):
\[
E_1(v_1),\qquad E_2(v_2).
\]
Then users may experience different behavior even when querying the same service.
This is common during rolling deployments.
A rollout might follow:
\[
10\%
\rightarrow
25\%
\rightarrow
50\%
\rightarrow
100\%.
\]
During the rollout, multiple versions coexist.
So:
\[
\text{same domain}
\neq
\text{same software version}.
\]
Feature flags introduce another layer.
A feature flag can be modeled as
\[
F(u,r,t)\in\{0,1\},
\]
where access depends on user identity \(u\), region \(r\), and time \(t\).
Then two users visiting the same page can receive different interfaces:
\[
Y^{(u_1)}\neq Y^{(u_2)}.
\]
This can happen even if they are physically close.
The difference may come from:
\[
\text{account state},
\text{experiment group},
\text{device type},
\text{regional policy},
\text{subscription tier},
\text{rollout stage}.
\]
This connects directly to the earlier point about people assuming they see the same broadcast.
Modern internet infrastructure increasingly produces:
\[
\text{shared nominal service}
\]
but
\[
\text{individualized operational projection}.
\]
CDNs also contribute to failure isolation.
If one origin server goes down, an edge may continue serving cached content.
So the website can appear available even though the backend is degraded.
That creates:
\[
\text{surface available}
\]
while
\[
\text{origin unhealthy}.
\]
This is useful for resilience, but it can hide operational failure.
A user may still see a page while forms, transactions, or dynamic services fail.
This creates a mixed state:
\[
\text{static layer healthy}
\land
\text{dynamic layer unhealthy}.
\]
For example:
\[
\text{homepage loads}
\]
but
\[
\text{login fails}.
\]
That can happen because static assets are cached while authentication requires live backend state.
CDNs also interact with DNS, because the user must first be routed to the correct edge.
The path becomes:
\[
\text{DNS}
\rightarrow
\text{edge selection}
\rightarrow
\text{cache state}
\rightarrow
\text{origin state}.
\]
Each layer can introduce delay, inconsistency, or geographic divergence.
Regional policy can also matter.
A platform may intentionally serve different content based on jurisdiction:
\[
Y_r=
H(X,\text{region policy}).
\]
Examples include:
\[
\text{privacy rules},
\text{licensing},
\text{regulatory restrictions},
\text{language},
\text{content availability}.
\]
That means geographic differences are not always bugs.
They can be deliberate.
A useful distinction is:
\[
\text{geographic divergence}
\text{intentional policy divergence}
\]
or
\[
\text{geographic divergence}
\text{unintentional state divergence}.
\]
The surface alone may not tell you which one occurred.
Edge systems also generate their own observability problem.
The operator may have logs for:
\[
\text{origin},
\text{regional node},
\text{edge node},
\text{user request}.
\]
To diagnose a bad result, those traces must be correlated.
This is where distributed tracing becomes necessary.
A single user-visible error may require reconstructing:
\[
\text{request ID}
\rightarrow
\text{edge}
\rightarrow
\text{regional service}
\rightarrow
\text{origin}
\rightarrow
\text{database}.
\]
Without that trace, failures can look random.
For operational digital twins, CDN and edge state should therefore be treated as external delivery state, not as the underlying authoritative object.
A twin might observe:
\[
\text{page seen at time }t
\]
but should preserve that as:
\[
\text{edge-delivered observation}
\]
rather than:
\[
\text{canonical website state}.
\]
That distinction matters because the page may have been cached, personalized, geographically filtered, or served from a rollout variant.
A richer local relational trace might include:
\[
\text{URL},
\text{timestamp},
\text{region},
\text{observed content},
\text{account context},
\text{response headers},
\text{provenance class}.
\]
This gives the twin a better chance of distinguishing:
\[
\text{actual state change}
\]
from
\[
\text{delivery-layer variation}.
\]
The full delivery architecture can therefore be represented as
\[
\boxed{
\text{origin state}
\rightarrow
\text{regional replication}
\rightarrow
\text{edge cache/computation}
\rightarrow
\text{routing}
\rightarrow
\text{user-specific projection}
}
\]
with possible differences introduced at every stage.
Within Schrödinger’s Library, the central principle is:
\[
\boxed{
\text{one website identity can correspond to many geographically distributed operational states}
}
\]
and
\[
\boxed{
\text{different users can receive different valid projections because delivery itself is stateful, cached, regional, and sometimes personalized}.
}
\]
The next step in the sequence is observability and distributed tracing—how operators reconstruct hidden request paths across services, caches, queues, databases, and external dependencies to determine where a visible failure actually originated.
Source: r/Wendbine · by /u/Upset-Ratio502