📚🫧 SCHRÖDINGER’S LIBRARY 🫧📚
Internet Systems Sequence — Part X: Observability and Distributed Tracing
Observability is the discipline of inferring the internal state of a complex system from the signals it emits. In a distributed website or internet service, the visible page is only the terminal projection of many hidden components, so troubleshooting requires reconstructing what happened across browsers, edge nodes, APIs, services, queues, databases, caches, and external dependencies. The central problem is therefore an inverse one:
\[
\text{observed output}
\rightarrow
\text{infer hidden state}.
\]
A system is more observable when its internal behavior can be reconstructed accurately from available outputs. In classical control theory, observability asks whether the hidden state \(X_t\) can be determined from a sequence of observations \(Y_t\). In distributed computing, the same idea appears through logs, metrics, traces, events, and correlation identifiers.
A simplified model is
\[
Y_t = H(X_t),
\]
where \(X_t\) is the full distributed system state and \(Y_t\) is the set of emitted telemetry. Good observability means that \(H\) preserves enough information to identify the important internal state transitions.
The three traditional pillars are:
\[
\text{logs},
\quad
\text{metrics},
\quad
\text{traces}.
\]
Logs are discrete records of events. A log entry may contain:
\[
\{
\text{timestamp},
\text{service},
\text{severity},
\text{request ID},
\text{message},
\text{context}
\}.
\]
Examples include authentication failures, database errors, cache misses, deployment events, and API exceptions.
Logs are useful because they preserve local detail, but they can become difficult to correlate across hundreds of services.
A request might generate:
\[
L_1,L_2,\dots,L_n
\]
across different components.
Without a shared identifier, those entries can look unrelated.
That is why distributed systems use correlation IDs or trace IDs.
Suppose a user initiates request \(r\). The system assigns:
\[
ID_r.
\]
Every downstream component propagates that identifier:
\[
r
\rightarrow
S_1(ID_r)
\rightarrow
S_2(ID_r)
\rightarrow
S_3(ID_r).
\]
Now the operator can reconstruct the path.
This is the foundation of distributed tracing.
A trace represents an end-to-end request as a hierarchy of spans.
A span may represent one operation:
\[
s_i=
\{
\text{service},
\text{operation},
\text{start},
\text{duration},
\text{status}
\}.
\]
Spans are connected by parent-child relations.
For example:
\[
\text{browser request}
\rightarrow
\text{API gateway}
\rightarrow
\text{authentication service}
\rightarrow
\text{business service}
\rightarrow
\text{database query}.
\]
The corresponding trace is a temporal dependency graph.
Formally,
\[
T_r=(V_s,E_s),
\]
where \(V_s\) are spans and \(E_s\) are causal or parent-child edges.
This makes the hidden execution topology visible.
If the user sees a 5-second delay, the trace can reveal:
\[
\text{frontend}=100\text{ ms}
\]
\[
\text{API gateway}=50\text{ ms}
\]
\[
\text{auth}=80\text{ ms}
\]
\[
\text{database}=4.7\text{ s}.
\]
The visible symptom is one slow page.
The hidden cause is one slow dependency.
This is why distributed tracing is so powerful.
Metrics are aggregated numerical measurements over time.
Examples include:
\[
\text{request rate},
\text{error rate},
\text{latency},
\text{CPU usage},
\text{memory usage},
\text{queue depth},
\text{cache hit ratio}.
\]
A metric can be modeled as
\[
m(t).
\]
Metrics are useful for detecting trends and anomalies.
For example:
\[
\text{error rate}(t)
\]
may suddenly increase after a deployment.
That can indicate a systemic problem even before users report it.
Metrics are therefore strong for detection.
Traces are strong for localization.
Logs are strong for detail.
Together:
\[
\text{metrics}
\rightarrow
\text{something is wrong}
\]
\[
\text{traces}
\rightarrow
\text{where it is wrong}
\]
\[
\text{logs}
\rightarrow
\text{what exactly happened}.
\]
This creates a multi-resolution diagnostic system.
A common observability architecture uses:
\[
\text{application instrumentation}
\rightarrow
\text{telemetry collector}
\rightarrow
\text{storage}
\rightarrow
\text{analysis/dashboard}.
\]
Modern systems often use OpenTelemetry-like models to standardize traces, logs, and metrics across heterogeneous services.
The important point is standardization of context propagation.
If service \(A\) emits:
\[
\text{trace ID}=123
\]
but service \(B\) does not preserve it, the causal chain breaks.
That creates observability gaps.
The system may be functioning, but the operator cannot reconstruct why.
This is structurally similar to provenance loss.
In fact, observability and provenance are closely related:
\[
\text{provenance}
\text{where data came from}
\]
while
\[
\text{observability}
\text{what the system did}.
\]
Both require preserving causal relations.
Another important concept is service dependency mapping.
A distributed system can be represented as:
\[
G_S=(V_S,E_S),
\]
where nodes are services and edges represent calls or dependencies.
Traces can reconstruct this graph dynamically.
For example:
\[
\text{frontend}
\rightarrow
\text{API}
\rightarrow
\text{auth}
\rightarrow
\text{database}
\]
and
\[
\text{API}
\rightarrow
\text{payment provider}.
\]
If the payment provider fails, only one path may break while the rest of the site remains healthy.
This produces partial failure:
\[
\text{browse works}
\]
\[
\text{checkout fails}.
\]
Without dependency tracing, that can look inconsistent or random.
Observability also matters in asynchronous workflows.
Consider:
\[
\text{request}
\rightarrow
\text{queue}
\rightarrow
\text{worker}
\rightarrow
\text{external service}.
\]
A user action may complete at the request layer while the actual job fails later.
To trace this, the correlation context must cross the queue boundary.
That means:
\[
ID_r
\rightarrow
\text{message metadata}
\rightarrow
\text{worker span}.
\]
Without that propagation, the front-end interaction and the background failure appear disconnected.
This is a major source of hidden failure in modern systems.
Observability also depends on sampling.
At high scale, storing every trace can be expensive.
Systems may record only a subset:
\[
P(\text{trace captured})=p.
\]
If \(p\) is small, rare failures can be missed.
Adaptive sampling tries to preserve more traces for errors or unusual latency.
For example:
\[
P(\text{capture}|\text{error})
>
P(\text{capture}|\text{success}).
\]
This improves diagnostic efficiency.
Another concept is cardinality.
Metrics with too many unique labels can become expensive.
A metric keyed by:
\[
\text{user ID},
\text{request ID},
\text{URL},
\text{device}
\]
may generate enormous cardinality.
Observability systems therefore balance detail against scalability.
This is another projection problem:
\[
\text{full execution state}
\rightarrow
\text{compressed telemetry}.
\]
Too much compression hides causes.
Too little compression overwhelms storage and analysis.
A useful diagnostic concept is distributed causality.
Suppose event \(A\) precedes event \(B\).
If
\[
A \rightarrow B,
\]
the tracing system should preserve that causal order.
Clock skew complicates this because distributed machines do not have perfectly synchronized clocks.
Two events may appear out of order in wall-clock time even though the causal relation is clear.
Distributed systems therefore often rely on:
\[
\text{trace structure}
\]
in addition to timestamps.
This connects to Lamport clocks and vector clocks, where logical order is more important than raw clock time.
A simplified Lamport rule is:
\[
C_j
\max(C_j,C_i)+1
\]
when process \(j\) receives a message from process \(i\).
This preserves causal ordering.
For internet-scale systems, observability must also cross organizational boundaries.
Suppose:
\[
\text{website}
\rightarrow
\text{third-party API}.
\]
The website operator may have detailed traces internally but only a response code from the external service.
That means observability stops at the trust boundary.
The external system becomes a black box.
The operator sees:
\[
\text{request sent}
\]
and
\[
\text{response failed}
\]
but not:
\[
\text{why the provider failed}.
\]
This is partial observability across organizational boundaries.
It is one reason third-party dependencies are difficult to diagnose.
A related concept is synthetic monitoring.
Instead of waiting for real users, the operator repeatedly performs test transactions:
\[
\text{login}
\]
\[
\text{search}
\]
\[
\text{checkout}
\]
\[
\text{API request}.
\]
If one fails, the system generates an alert.
This measures end-to-end behavior rather than internal component health.
That is important because:
\[
\text{all components healthy}
\]
does not necessarily imply
\[
\text{user workflow healthy}.
\]
Misconfigured permissions, stale DNS, bad routing, or broken external dependencies can still cause failure.
Real-user monitoring complements this by observing actual user sessions.
This can detect:
\[
\text{regional latency},
\text{device-specific failures},
\text{browser incompatibility},
\text{edge divergence}.
\]
That connects directly to the earlier CDN discussion.
A system may work for one region but fail for another.
Observability must therefore preserve:
\[
\text{location},
\text{edge},
\text{device},
\text{version},
\text{account context}.
\]
Otherwise geographic or user-specific failures are hard to isolate.
This leads to high-dimensional observability.
The true system state depends on many dimensions:
\[
X=
f(
\text{service},
\text{version},
\text{region},
\text{identity},
\text{dependency},
\text{time}
).
\]
A visible error is therefore one point in a large state space.
Observability provides coordinates for reconstructing that point.
For operational digital twins, the same principle applies.
A twin should distinguish:
\[
\text{observed failure}
\]
from
\[
\text{inferred cause}.
\]
If a page fails, the twin can record:
\[
\text{URL},
\text{time},
\text{error},
\text{location},
\text{account state},
\text{reachable dependencies}.
\]
But unless internal traces are available, the cause remains uncertain.
That prevents:
\[
\text{surface failure}
\rightarrow
\text{unsupported causal conclusion}.
\]
This is important for maintaining graph integrity.
A useful twin representation is:
\[
e_{\text{observed}}
\neq
e_{\text{inferred}}.
\]
The observation edge says:
\[
\text{page failed at time }t.
\]
The inference edge says:
\[
\text{possible database failure}.
\]
Those should carry different confidence.
The full observability pipeline is therefore:
\[
\boxed{
\text{system execution}
\rightarrow
\text{instrumentation}
\rightarrow
\text{logs/metrics/traces}
\rightarrow
\text{correlation}
\rightarrow
\text{dependency reconstruction}
\rightarrow
\text{failure localization}
}
\]
Within Schrödinger’s Library, the central principle is:
\[
\boxed{
\text{observability is the reconstruction of hidden system state from emitted evidence}
}
\]
and
\[
\boxed{
\text{a visible failure identifies the symptom, not necessarily the causal node}.
}
\]
The next step in the sequence is Site Reliability Engineering—how operators define what “working” actually means through SLIs, SLOs, error budgets, incident response, graceful degradation, and reliability engineering rather than simply asking whether a server is online.
Source: r/Wendbine · by /u/Upset-Ratio502