ESC
Type to search guides, tutorials, and reference documentation.
Verified by Garnet Grid

Observability

How to make a running system explainable from the outside: what metrics, logs, traces and wide events each answer, why cardinality is the cost driver, how sampling and context propagation actually work, and the failure modes that turn an observability stack into expensive noise.

Observability is the property of a system that lets you answer questions about its internal state using only the data it already emits — including questions nobody thought to ask before the system started misbehaving. The practical test is simple: when a novel failure appears, can you explain it from existing telemetry, or do you have to ship new code and wait for it to recur?


What Observability Actually Is

Monitoring and observability are often used interchangeably, and the difference matters. Monitoring answers questions you decided in advance were important: a dashboard for the queue depth, an alert for disk usage, a check that the process is alive. It handles known unknowns — the failure modes you have already seen or predicted.

Observability is about unknown unknowns. It is the ability to slice, group and filter telemetry along dimensions you did not pre-aggregate, so you can narrow a vague symptom ("some checkouts are slow") down to a specific population ("slow only for requests that hit the new pricing service, only for accounts on the legacy plan, only in one region"). A system is observable to the degree that this narrowing can be done as a query rather than as a code change.

That framing has a design consequence. Observability is not something you buy and point at a system; it is a property you build into the system by emitting data that carries enough dimensions to support questions you have not thought of yet.


The Signal Types and What Each Answers

Most stacks combine several telemetry shapes. They are not interchangeable — each trades fidelity against cost in a different way.

SignalAnswersCost scales withMain limitation
MetricsIs it broken, how badly, and since whenDistinct label combinationsNo per-request detail once aggregated
LogsWhat exactly happened to this one requestEvent volume and field widthExpensive to query at scale; useless if unstructured
TracesWhere did the time go across servicesSpans retained after samplingNeeds unbroken context propagation to be usable
Wide eventsWhich population of requests is affectedEvents retainedRequires deliberate design of the field set
ProfilesWhich code path inside a process is hotSampling rate and retentionScoped to one process; not a cross-service view

Metrics are pre-aggregated numbers over time. The storage cost is roughly proportional to the number of distinct time series, not to the request volume, which is why a metric is cheap even for a service handling enormous traffic. That same aggregation is the limitation: once you have averaged a value into a series, the individual requests are gone.

Logs keep per-event fidelity, so they answer "what exactly happened to this request". Their cost scales with volume, and unstructured free-text logs are expensive to query because every search is a scan. Structured logs — key/value fields rather than sentences — are the difference between a log store you can interrogate and one you can only grep.

Traces record causality. A trace is a tree of spans, each with a start time, duration, and parent, stitched together across process and network boundaries by a propagated identifier. Traces are the only signal that answers "where did the time go" in a request that touches many services, and "which dependency actually caused this".

Wide structured events — sometimes called canonical log lines — are a hybrid worth knowing about: one event emitted per unit of work, carrying every dimension known at the time (route, tenant, plan tier, cache outcome, upstream versions, queue wait, total duration). Because the dimensions live on a single event rather than on separate pre-aggregated series, you can group by any of them after the fact. This is the cheapest way to get high-cardinality analysis without a metrics bill that grows multiplicatively.

Continuous profiles answer a question the other signals cannot: inside one process, which code paths consumed the CPU or allocated the memory. They pair naturally with performance engineering, where the bottleneck is usually inside a process rather than between processes.


Cardinality: The Thing That Decides Your Bill

Cardinality is the number of distinct values a dimension can take, and it is the single most important cost concept in telemetry. For metrics, each unique combination of label values creates its own time series, so the series count is the product of the cardinalities of the labels. Adding a label with a handful of values multiplies your series count by a handful. Adding a label that carries a user identifier, a request identifier, a full URL path with embedded IDs, or an error message string multiplies it by something unbounded.

This is the mechanism behind a very common self-inflicted outage: an engineer adds a useful-looking label to a metric, series count grows without limit, and the metrics backend either starts rejecting writes, evicts data, or generates a bill that triggers an incident of its own. The defensive habits are durable:

  • Keep unbounded identifiers (user, request, session, order) out of metric labels; put them on events, logs, or spans, where the cost model is per-event rather than per-series.
  • Normalise path-like labels to route templates rather than concrete URLs, so one route does not become thousands of series.
  • Never use a raw error message or stack trace as a label value; map to a bounded error class.
  • Treat a sudden jump in series count as an incident signal in its own right, and cap or drop offending series at the collector rather than at the application.

Instrumentation and Context Propagation

Instrumentation comes in two forms. Automatic instrumentation hooks common libraries — HTTP servers and clients, database drivers, message consumers — and gives you baseline coverage with no code changes. It tells you that a call happened and how long it took. Manual instrumentation adds the business dimensions that automatic instrumentation cannot know: which tenant, which plan, which feature flag variant, which retry attempt. The useful telemetry is nearly always the manual part; the automatic part is the skeleton it hangs on.

A vendor-neutral instrumentation layer (an API and SDK separate from the backend that stores the data) is worth the indirection for one reason: instrumentation lives in your application code and is expensive to rewrite, while backends get replaced. Keeping the emit side stable lets you change where the data goes without touching every service.

Tracing depends entirely on context propagation: a trace identifier and the current span identifier must travel with the work. Across HTTP or RPC this is a header. The places it silently breaks are worth memorising, because a broken chain produces orphaned spans that look like a missing service rather than a missing header:

  • Asynchronous boundaries. Work handed to a queue, a scheduled job, or a background thread pool loses context unless it is explicitly carried in the message and restored by the consumer.
  • Process-internal concurrency. Fan-out to worker threads or coroutines needs the context passed into each unit of work, not read from an ambient variable that belongs to a different task.
  • Intermediaries that strip headers. Proxies, gateways, service meshes and third-party callers may not forward trace headers, producing traces that begin in the middle of the system.
  • Mixed propagation formats. Two services using different header conventions will each produce valid traces that never join.

Correlation across signals is what turns three separate tools into one workflow. The practical minimum is to put the trace identifier on every log line and to attach consistent resource attributes (service name, version, deployment environment, region) to every signal, so a spike on a chart leads to the exemplar trace, and the trace leads to that request's logs.


Sampling Without Losing the Interesting Requests

At meaningful traffic volumes you will not keep every trace or every event, so the question is not whether to sample but how to choose what survives. The two strategies differ in where the decision is made, and the difference is consequential.

Head-based sampling decides at the start of a request, usually at the root, and propagates that decision to every downstream span. It is cheap, requires no buffering, and guarantees that a kept trace is complete. Its weakness is that the decision is made before anything interesting has happened — a uniform head sample keeps a representative slice of ordinary traffic and, by construction, discards nearly all of the rare errors and slow outliers you would actually want to look at.

Tail-based sampling buffers the spans of a trace until it finishes, then decides based on the outcome: keep everything that errored, everything slower than a threshold, everything touching a rare code path, and a small random slice of the rest. It captures the interesting minority, at the cost of a collector that must hold spans in memory long enough for the whole trace to arrive — which means it needs all spans of a trace routed to the same collector instance, enough memory for the buffer window, and a defined behaviour for traces that never complete.

Two refinements are worth knowing. Rate-limiting per route or tenant prevents one high-volume endpoint from consuming the entire sampling budget and crowding out the low-volume endpoints where problems hide. Consistent sampling — deriving the keep/drop decision from the trace identifier rather than from an independent random draw at each service — ensures services agree, which is what prevents the partial traces that make a sampled system look broken.

Whatever the strategy, record the sampling rate alongside the data. A count derived from sampled traces is not a count of what happened unless you know the factor by which to scale it, and quietly treating sampled data as complete is a reliable way to under-report a problem.


Common Failure Modes

Failure modeWhat it looks likeWhat to do instead
Dashboard sprawlHundreds of dashboards, each built during one past incident, none trustedA small number of owned, curated views tied to objectives; delete the rest
Alerting on causesPager fires for CPU, restarts, queue depth; nobody is affectedAlert on user-visible symptoms; keep cause metrics as diagnostics
Cardinality explosionBackend rejects writes, or costs jump with no traffic changeBound label values; enforce limits in the collection pipeline
Logs as a debuggerVolume so high that searching is slow and expensiveStructured wide events per unit of work; debug-level logging behind a sampled switch
Shared failure domainTelemetry disappears exactly when the system failsKeep at least one signal path that does not depend on the failing system
Retention shorter than detectionData has aged out by the time you investigateMatch retention to how long a problem can hide; downsample rather than delete
Unsynchronised clocksSpans with impossible ordering or negative durationsRely on the tracing library's duration, not on cross-host timestamp arithmetic

Two of these deserve emphasis. Alerting on causes — high CPU, a full queue, a restarted pod — produces pages for conditions that frequently do not affect anyone, and the resulting noise trains responders to ignore the pager. Alert on the symptom a user would notice, and keep cause-level signals as diagnostic context. This is the same discipline that makes incident response tractable: a page should mean a human decision is needed now.

Shared failure domains are the one that hurts during a real outage. If your dashboards, your alert evaluation, and your on-call paging path all depend on the cluster, the network, or the identity provider that just failed, you lose your instruments exactly when you need them. The mitigation is not exotic — it is knowing which parts of the observability path share fate with the system, and keeping at least one independent way to see that the service is down and to reach a human.


When It Is Worth It, and When It Is Not

Distributed tracing earns its cost when a single user-facing request crosses more network hops than a person can hold in their head, or when ownership of those hops is spread across teams. A single-process application with a database behind it is usually better served by good structured logs, a handful of metrics tied to its objectives, and a profiler — tracing adds infrastructure without answering a question you had.

Conversely, do not defer instrumentation because the system is small. The dimensions you fail to emit are the dimensions you cannot query later, and retrofitting them during an incident is the worst possible time. A cheap starting posture: one wide event per request with the dimensions you already know, a small set of metrics tied to user-visible objectives, and structured logs with a correlation identifier.

Finally, observability is a prerequisite for the practices that depend on detecting deviation from normal — you cannot run a chaos experiment without a measurable steady state, you cannot size capacity without knowing your real demand shape, and you cannot claim an objective is met without the data to show it.

Jakub Dimitri Rezayev
Jakub Dimitri Rezayev
Founder & Chief Architect • Garnet Grid Consulting

Jakub holds an M.S. in Customer Intelligence & Analytics and a B.S. in Finance & Computer Science from Pace University. With deep expertise spanning D365 F&O, Azure, Power BI, and AI/ML systems, he architects enterprise solutions that bridge legacy systems and modern technology — and has led multi-million dollar ERP implementations for Fortune 500 supply chains.

View Full Profile →